DBScholar

Back to papers

HoloDetect: Few-Shot Learning for Error Detection

Summary: HoloDetect: few-shot error detection with a two-part model; rich representations and a data-augmentation policy learner. Augmenting a small seed of clean data yields ~94% precision, ~93% recall, ~20 F1 gains, and ~3x fewer labels than ML baselines. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h8869392bcf4ef6bb
Venue
SIGMOD
Year
2019
Pagerank
0.00013263269
Overall Rank
884 | 94.07%
DOI
10.1145/3299869.3319888

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{heidari_sigmod19,
        title = {{HoloDetect: Few-Shot Learning for Error Detection}},
        author = {Heidari, Alireza and McGrath, Joshua and Ilyas, Ihab F. and Rekatsinas, Theodoros},
        series = {{SIGMOD} '19},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3299869.3319888},
        url = {https://dl.acm.org/doi/10.1145/3299869.3319888},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 45 of 45 citing papers.

Rank Citing Paper Year Venue Pagerank
144 Neo: A Learned Query Optimizer 2019 VLDB 0.00029090793
329 Can Foundation Models Wrangle Your Data? 2023 VLDB 0.00020867521
1,342 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010963254
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
2,372 SCODED: Statistical Constraint Oriented Data Error Detection 2020 SIGMOD 8.5575299e-05
2,445 Approximate Denial Constraints 2020 VLDB 8.4546614e-05
2,600 Complaint-driven Training Data Debugging for Query 2.0 2020 SIGMOD 8.2346824e-05
2,642 Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks 2020 SIGMOD 8.1751637e-05
3,015 Saga: A Platform for Continuous Construction and Serving of Knowledge At Scale 2022 SIGMOD 7.7504975e-05
3,194 A Statistical Perspective on Discovering Functional Dependencies in Noisy Data 2020 SIGMOD 7.5469742e-05
3,264 Automatic Data Repair: Are We Ready to Deploy? 2024 VLDB 7.4774357e-05
4,197 PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models 2020 SIGMOD 6.7390419e-05
4,426 Auto-Transform: Learning-to-Transform by Patterns 2020 VLDB 6.601896e-05
5,078 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2832055e-05
5,251 Enabling SQL-based Training Data Debugging for Federated Learning 2022 VLDB 6.2092359e-05
5,574 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0773771e-05
5,956 Semi-Supervised Data Cleaning with Raha and Baran 2021 CIDR 5.9329823e-05
6,232 The Fast and the Private: Task-based Dataset Search 2024 CIDR 5.8417106e-05
6,528 Parallel Discrepancy Detection and Incremental Detection 2021 VLDB 5.7540798e-05
6,554 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.7473435e-05
7,531 Auto-Prep: Holistic Prediction of Data Preparation Steps for Self-Service Business Intelligence 2025 VLDB 5.4994676e-05
7,544 Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines 2023 SIGMOD 5.496984e-05
7,654 MisDetect: Iterative Mislabel Detection using Early Loss 2024 VLDB 5.4746904e-05
7,742 Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes 2021 SIGMOD 5.4599258e-05
7,843 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4406331e-05
7,876 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.4336606e-05
8,764 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2815275e-05
9,320 VerifAI: Verified Generative AI 2024 CIDR 5.1941278e-05
9,616 Discovering Top-k Rules using Subjective and Objective Criteria 2023 SIGMOD 5.1503279e-05
9,623 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1486163e-05
9,727 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1325223e-05
9,813 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.1233734e-05
9,881 MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series Data 2024 VLDB 5.115241e-05
10,116 Data Augmentation for ML-driven Data Preparation and Integration 2021 VLDB 5.0765311e-05
10,191 Reptile: Aggregation-level Explanations for Hierarchical Data 2022 SIGMOD 5.0628015e-05
10,307 DobLIX: A Dual-Objective Learned Index for Log-Structured Merge Trees 2025 VLDB 5.0392037e-05
10,346 Parallel Rule Discovery from Large Datasets by Sampling 2022 SIGMOD 5.0171283e-05
10,364 Towards Scalable Visual Data Wrangling via Direct Manipulation 2026 CIDR 4.9769913e-05
10,540 Minimum Change ≠ Best Cleaning: Parallel and Incremental Error Detection under Integrity Constraints 2026 SIGMOD 4.9769913e-05
10,805 Measuring Database Unfairness via Dependency Quantification Under Differential Privacy 2026 VLDB 4.9769913e-05
10,864 PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines 2026 VLDB 4.9769913e-05
11,360 UniClean: A Scalable Data Cleaning Solution for Mixed Errors based on Unified Cleaners and Optimized Cleaning Workflow 2025 VLDB 4.9769913e-05
11,417 Demonstrating Matelda for Multi-Table Error Detection 2025 VLDB 4.9769913e-05
11,667 Generalizable Data Cleaning of Tabular Data in Latent Space 2024 VLDB 4.9769913e-05
11,882 PGE: Robust Product Graph Embedding Learning for Error Detection 2022 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers