DBScholar

Back to papers

HoloClean: Holistic Data Repairs with Probabilistic Inference

Summary: HoloClean unifies constraint-/external-source-based and statistical data repair through automatically generated probabilistic programs. Scalable inference optimizations handle millions of tuples, achieving ~90% precision, >76% recall, and >2× F1 over prior methods. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hfd89a18c795e9f81
Venue
VLDB
Year
2017
Pagerank
0.00033690989
Overall Rank
104 | 99.31%
DOI
10.14778/3137628.3137663

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{rekatsinas_vldb17,
        title = {{HoloClean: Holistic Data Repairs with Probabilistic Inference}},
        author = {Rekatsinas, Theodoros and Chu, Xu and Ilyas, Ihab F. and Ré, Christopher},
        journal = {PVLDB},
        series = {{VLDB} '17},
        volume = {10},
        number = {11},
        pages = {1190--1203},
        doi = {10.14778/3137628.3137663},
        url = {https://doi.org/10.14778/3137628.3137663},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 143 citing papers.

Rank Citing Paper Year Venue Pagerank
7,068 How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses 2024 VLDB 5.6083188e-05
7,183 DataPrism: Exposing Disconnect between Data and Systems 2022 SIGMOD 5.5917354e-05
7,250 Akane: Perplexity-Guided Time Series Data Cleaning 2024 SIGMOD 5.574104e-05
7,291 Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics 2019 VLDB 5.5642829e-05
7,415 Data Integration and Machine Learning: A Natural Synergy 2018 VLDB 5.5332662e-05
7,471 Fast Detection of Denial Constraint Violations 2022 VLDB 5.5176505e-05
7,532 PIClean: A Probabilistic and Interactive Data Cleaning System 2019 SIGMOD 5.5007996e-05
7,583 Data Imputation with Limited Data Redundancy Using Data Lakes 2025 VLDB 5.4906208e-05
7,648 MisDetect: Iterative Mislabel Detection using Early Loss 2024 VLDB 5.4772833e-05
7,704 ReStore - Neural Data Completion for Relational Databases 2021 SIGMOD 5.4741304e-05
7,839 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4432099e-05
7,850 CoClean: Collaborative Data Cleaning 2020 SIGMOD 5.4405793e-05
7,872 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 5.4362062e-05
7,875 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 5.4355826e-05
8,062 Evaluating Top-k Queries with Inconsistency Degrees 2020 VLDB 5.3955302e-05
8,224 Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness 2024 VLDB 5.3747673e-05
8,277 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 5.3639084e-05
8,280 The Computation of Optimal Subset Repairs 2020 VLDB 5.363617e-05
8,335 ICARUS: Minimizing Human Effort in Iterative Data Completion 2018 VLDB 5.3525933e-05
8,398 Deducing Certain Fixes to Graphs 2019 VLDB 5.3399392e-05
8,411 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.336291e-05
8,544 Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? 2021 SIGMOD 5.3186267e-05
8,695 From Papers to Practice: The openclean Open-Source Data Cleaning Library 2021 VLDB 5.2905577e-05
8,756 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 5.2840289e-05
8,847 Machine Learning Meets Big Spatial Data 2019 VLDB 5.2646682e-05
8,852 Rapidash: Efficient Detection of Constraint Violations 2024 VLDB 5.2641151e-05
8,886 FastPDB: Towards Bag-Probabilistic Queries at Interactive Speeds 2025 SIGMOD 5.2559789e-05
8,984 nsDB: Architecting the Next Generation Database by Integrating Neural and Symbolic Systems 2024 VLDB 5.2433452e-05
9,004 The Cost of Representation by Subset Repairs 2025 VLDB 5.2389669e-05
9,229 Auto-Test: Learning Semantic-Domain Constraints for Unsupervised Error Detection in Tables 2025 SIGMOD 5.2056825e-05
9,294 DataDiff: User-Interpretable Data Transformation Summaries for Collaborative Data Analysis 2018 SIGMOD 5.1995938e-05
9,311 VerifAI: Verified Generative AI 2024 CIDR 5.1965878e-05
9,373 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 5.1868213e-05
9,384 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 5.1868213e-05
9,417 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 5.1803615e-05
9,455 ZIP: Lazy Imputation during Query Processing 2024 VLDB 5.1735482e-05
9,608 Discovering Top-k Rules using Subjective and Objective Criteria 2023 SIGMOD 5.1527671e-05
9,616 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1510548e-05
9,722 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1349531e-05
9,766 Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation 2021 CIDR 5.1344318e-05
9,801 Incremental Detection of Denial Constraint Violations 2025 VLDB 5.1257999e-05
9,806 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 5.1257999e-05
9,808 Lingua Manga: A Generic Large Language Model Centric System for Data Curation 2023 VLDB 5.1257999e-05
9,874 MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series Data 2024 VLDB 5.1176637e-05
10,025 Don’t Be a Tattle-Tale: Preventing Leakages through Data Dependencies on Access Control Protected Data 2022 VLDB 5.0934585e-05
10,106 Efficient Differential Dependency Discovery 2024 VLDB 5.0789354e-05
10,110 EasyDR: A Human-in-the-loop Error Detection&Repair Platform for Holistic Table Cleaning 2022 VLDB 5.0789354e-05
10,181 In-Database Data Imputation 2024 SIGMOD 5.0653015e-05
10,186 Discovering Top-k Relevant and Diversified Rules 2024 SIGMOD 5.0651993e-05
10,188 Reptile: Aggregation-level Explanations for Hierarchical Data 2022 SIGMOD 5.0651993e-05
Previous Page 2 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
188 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00025872962
262 Record Linkage: Similarity Measures and Algorithms 2006 SIGMOD 0.00022821392
350 Discovering Denial Constraints 2013 VLDB 0.00020253521
500 Dependencies Revisited for Improving Data Quality 2008 PODS 0.00017280722
514 Data Curation at Scale: The Data Tamer System 2013 CIDR 0.00017006745
526 Improving Data Quality: Consistency and Accuracy 2007 VLDB 0.00016886621
547 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016578131
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016086569
618 Reasoning about Record Matching Rules 2009 VLDB 0.00015536912
623 Entity Resolution: Theory, Practice & Open Challenges 2012 VLDB 0.00015483844
652 Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes 2013 SIGMOD 0.00015121325
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014694048
767 The LLUNATIC Data-Cleaning Framework 2013 VLDB 0.00014114806
938 Truth Finding on the Deep Web: Is the Problem Solved? 2013 VLDB 0.00012973266
989 Towards Certain Fixes with Editing Rules and Master Data 2010 VLDB 0.0001265344
1,062 Tuffy: Scaling up Statistical Inference in Markov Logic Networks using an RDBMS 2011 VLDB 0.0001221201
1,099 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00012037058
1,167 DimmWitted: A Study of Main-Memory Statistical Analytics 2014 VLDB 0.00011729888
1,344 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00010956518
1,348 Sampling the Repairs of Functional Dependency Violations under Hard Constraints 2010 VLDB 0.0001094733
2,762 Dichotomies in the Complexity of Preferred Repairs 2015 PODS 8.0441688e-05
2,964 Towards Dependable Data Repairing with Fixing Rules 2014 SIGMOD 7.8052551e-05
3,791 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 7.0189865e-05
Previous Page 1 / 1 Next

Semantically Similar Papers