DBScholar

Back to papers

HoloClean: Holistic Data Repairs with Probabilistic Inference

Summary: HoloClean unifies constraint-/external-source-based and statistical data repair through automatically generated probabilistic programs. Scalable inference optimizations handle millions of tuples, achieving ~90% precision, >76% recall, and >2× F1 over prior methods. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
11592
Venue
VLDB
Year
2017
Pagerank
0.00032801121
Overall Rank
112 | 99.24%
DOI
10.14778/3137628.3137663

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{rekatsinas_vldb17,
        title = {{HoloClean: Holistic Data Repairs with Probabilistic Inference}},
        author = {Rekatsinas, Theodoros and Chu, Xu and Ilyas, Ihab F. and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '17},
        volume = {10},
        number = {11},
        pages = {1190--1203},
        doi = {10.14778/3137628.3137663},
        url = {https://doi.org/10.14778/3137628.3137663},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 137 citing papers.

Rank Citing Paper Year Venue Pagerank
7,200 Intermittent Query Processing 2019 VLDB 5.6756294e-05
7,232 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 5.6659017e-05
7,387 PIClean: A Probabilistic and Interactive Data Cleaning System 2019 SIGMOD 5.6270554e-05
7,396 Fast Detection of Denial Constraint Violations 2022 VLDB 5.6257228e-05
7,424 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.6214566e-05
7,556 ReStore - Neural Data Completion for Relational Databases 2021 SIGMOD 5.5997742e-05
7,687 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.5671645e-05
7,880 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 5.5244204e-05
7,982 Fast Approximate Denial Constraint Discovery 2023 VLDB 5.5138609e-05
8,058 Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness 2024 VLDB 5.4981306e-05
8,103 The Computation of Optimal Subset Repairs 2020 VLDB 5.4867244e-05
8,133 Evaluating Top-k Queries with Inconsistency Degrees 2020 VLDB 5.4813895e-05
8,160 ICARUS: Minimizing Human Effort in Iterative Data Completion 2018 VLDB 5.4754476e-05
8,225 Deducing Certain Fixes to Graphs 2019 VLDB 5.4624971e-05
8,242 Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics 2019 VLDB 5.459288e-05
8,244 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.4587712e-05
8,373 Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? 2021 SIGMOD 5.4399999e-05
8,594 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.4053095e-05
8,686 Rapidash: Efficient Detection of Constraint Violations 2024 VLDB 5.3849387e-05
8,716 nsDB: Architecting the Next Generation Database by Integrating Neural and Symbolic Systems 2024 VLDB 5.3776746e-05
8,722 Machine Learning Meets Big Spatial Data 2019 VLDB 5.3769871e-05
8,724 FastPDB: Towards Bag-Probabilistic Queries at Interactive Speeds 2025 SIGMOD 5.3766157e-05
8,838 The Cost of Representation by Subset Repairs 2025 VLDB 5.3592132e-05
9,129 DataDiff: User-Interpretable Data Transformation Summaries for Collaborative Data Analysis 2018 SIGMOD 5.318789e-05
9,177 VerifAI: Verified Generative AI 2024 CIDR 5.3078984e-05
9,193 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 5.3058708e-05
9,204 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 5.3058708e-05
9,245 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 5.2992628e-05
9,285 ZIP: Lazy Imputation during Query Processing 2024 VLDB 5.292293e-05
9,290 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 5.2910774e-05
9,427 Discovering Top-k Rules using Subjective and Objective Criteria 2023 SIGMOD 5.271035e-05
9,437 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.2687567e-05
9,539 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.2528121e-05
9,550 Data Imputation with Limited Data Redundancy Using Data Lakes 2025 VLDB 5.2528121e-05
9,614 Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation 2021 CIDR 5.2444447e-05
9,623 Incremental Detection of Denial Constraint Violations 2025 VLDB 5.2434488e-05
9,627 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.2434488e-05
9,631 Lingua Manga: A Generic Large Language Model Centric System for Data Curation 2023 VLDB 5.2434488e-05
9,647 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.2430158e-05
9,653 CoClean: Collaborative Data Cleaning 2020 SIGMOD 5.2425585e-05
9,699 MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series Data 2024 VLDB 5.2351259e-05
9,838 Don’t Be a Tattle-Tale: Preventing Leakages through Data Dependencies on Access Control Protected Data 2022 VLDB 5.2103651e-05
9,923 Efficient Differential Dependency Discovery 2024 VLDB 5.1955087e-05
9,928 EasyDR: A Human-in-the-loop Error Detection&Repair Platform for Holistic Table Cleaning 2022 VLDB 5.1955087e-05
9,993 In-Database Data Imputation 2024 SIGMOD 5.1815618e-05
9,999 Discovering Top-k Relevant and Diversified Rules 2024 SIGMOD 5.1814573e-05
10,001 Reptile: Aggregation-level Explanations for Hierarchical Data 2022 SIGMOD 5.1814573e-05
10,042 Scalable and Usable Relational Learning With Automatic Language Bias 2021 SIGMOD 5.1708123e-05
10,049 Towards Interpretable and Learnable Risk Analysis for Entity Resolution 2020 SIGMOD 5.1685424e-05
10,065 On Saving Outliers for Better Clustering over Noisy Data 2021 SIGMOD 5.1648805e-05
Previous Page 2 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
185 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00026231189
254 Record Linkage: Similarity Measures and Algorithms 2006 SIGMOD 0.00023199211
376 Discovering Denial Constraints 2013 VLDB 0.00019677674
494 Dependencies Revisited for Improving Data Quality 2008 PODS 0.00017526549
516 Data Curation at Scale: The Data Tamer System 2013 CIDR 0.00017171198
533 Improving Data Quality: Consistency and Accuracy 2007 VLDB 0.0001705859
549 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016692839
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016217563
619 Reasoning about Record Matching Rules 2009 VLDB 0.00015707247
626 Entity Resolution: Theory, Practice & Open Challenges 2012 VLDB 0.00015656958
661 Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes 2013 SIGMOD 0.0001519162
725 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014617251
800 The LLUNATIC Data-Cleaning Framework 2013 VLDB 0.00013911223
955 Truth Finding on the Deep Web: Is the Problem Solved? 2013 VLDB 0.00012996675
998 Towards Certain Fixes with Editing Rules and Master Data 2010 VLDB 0.0001275238
1,043 Tuffy: Scaling up Statistical Inference in Markov Logic Networks using an RDBMS 2011 VLDB 0.0001244565
1,101 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00012168934
1,150 DimmWitted: A Study of Main-Memory Statistical Analytics 2014 VLDB 0.00011943462
1,332 Sampling the Repairs of Functional Dependency Violations under Hard Constraints 2010 VLDB 0.00011129529
1,351 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00011064851
2,746 Dichotomies in the Complexity of Preferred Repairs 2015 PODS 8.1725014e-05
2,981 Towards Dependable Data Repairing with Fixing Rules 2014 SIGMOD 7.8960114e-05
3,715 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.1763559e-05
Previous Page 1 / 1 Next

Semantically Similar Papers