Database Paper Browser

Back to papers

HoloClean: Holistic Data Repairs with Probabilistic Inference

Summary: HoloClean couples constraint-driven and statistical data repair via automatic probabilistic-program generation from dirty data. Scalable inference over millions of tuples; precision ~90%, recall ~76%, F1 >2x vs state-of-the-art. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
11405
Venue
VLDB
Year
2017
Pagerank
0.00035692958
Overall Rank
192 | 98.67%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 133 citing papers.

Rank Citing Paper Year Venue Pagerank
7,397 Intermittent Query Processing 2019 VLDB 4.7367491e-05
7,565 PIClean: A Probabilistic and Interactive Data Cleaning System 2019 SIGMOD 4.7048523e-05
7,634 ReStore - Neural Data Completion for Relational Databases 2021 SIGMOD 4.6866388e-05
7,666 Fast Detection of Denial Constraint Violations 2022 VLDB 4.6792751e-05
7,702 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 4.6689015e-05
7,769 ICARUS: Minimizing Human Effort in Iterative Data Completion 2018 VLDB 4.6520279e-05
7,868 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 4.6276013e-05
7,992 Data Civilizer 2.0: A Holistic Framework for Data Preparation and Analytics 2019 VLDB 4.6082565e-05
8,014 The Computation of Optimal Subset Repairs 2020 VLDB 4.6020746e-05
8,096 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 4.583522e-05
8,125 Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? 2021 SIGMOD 4.5765541e-05
8,148 Evaluating Top-k Queries with Inconsistency Degrees 2020 VLDB 4.5717374e-05
8,183 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 4.5615358e-05
8,363 Deducing Certain Fixes to Graphs 2019 VLDB 4.5311185e-05
8,470 Rapidash: Efficient Detection of Constraint Violations 2024 VLDB 4.4993203e-05
8,588 Exploratory Training: When Annotators Learn About Data 2023 SIGMOD 4.4853244e-05
8,706 nsDB: Architecting the Next Generation Database by Integrating Neural and Symbolic Systems 2024 VLDB 4.4587004e-05
8,741 Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness 2024 VLDB 4.4520434e-05
8,784 Machine Learning Meets Big Spatial Data 2019 VLDB 4.4473392e-05
8,836 Fast Approximate Denial Constraint Discovery 2023 VLDB 4.4350633e-05
8,839 The Cost of Representation by Subset Repairs 2025 VLDB 4.4346105e-05
9,043 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 4.3997447e-05
9,053 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 4.3997447e-05
9,072 DataDiff: User-Interpretable Data Transformation Summaries for Collaborative Data Analysis 2018 SIGMOD 4.3975844e-05
9,073 VerifAI: Verified Generative AI 2024 CIDR 4.396857e-05
9,116 Towards Observability for Production Machine Learning Pipelines 2022 VLDB 4.3886184e-05
9,196 Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale 2022 VLDB 4.3723457e-05
9,247 ZIP: Lazy Imputation during Query Processing 2024 VLDB 4.3648789e-05
9,354 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 4.3484715e-05
9,362 Discovering Top-k Rules using Subjective and Objective Criteria 2023 SIGMOD 4.3472627e-05
9,395 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 4.3399748e-05
9,439 Rock: Cleaning Data by Embedding ML in Logic Rules 2024 SIGMOD 4.3389137e-05
9,443 Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation 2021 CIDR 4.3383453e-05
9,480 Incremental Detection of Denial Constraint Violations 2025 VLDB 4.3300131e-05
9,481 Data Imputation with Limited Data Redundancy Using Data Lakes 2025 VLDB 4.3300131e-05
9,489 Making It Tractable to Catch Duplicates and Conflicts in Graphs 2023 SIGMOD 4.3300131e-05
9,493 Lingua Manga : A Generic Large Language Model Centric System for Data Curation 2023 VLDB 4.3300131e-05
9,560 MTSClean: Efficient Constraint-based Cleaning for Multi-Dimensional Time Series Data 2024 VLDB 4.3212967e-05
9,576 CoClean: Collaborative Data Cleaning 2020 SIGMOD 4.3207e-05
9,673 Don’t Be a Tattle-Tale: Preventing Leakages through Data Dependencies on Access Control Protected Data 2022 VLDB 4.3014213e-05
9,748 Efficient Differential Dependency Discovery 2024 VLDB 4.2856385e-05
9,773 EasyDR: A Human-in-the-loop Error Detection and Repair Platform for Holistic Table Cleaning 2022 VLDB 4.2815042e-05
9,847 Discovering Top-k Relevant and Diversified Rules 2024 SIGMOD 4.2680295e-05
9,849 Reptile: Aggregation-level Explanations for Hierarchical Data 2022 SIGMOD 4.2680295e-05
9,855 In-Database Data Imputation 2024 SIGMOD 4.2652623e-05
9,885 Scalable and Usable Relational Learning With Automatic Language Bias 2021 SIGMOD 4.2580321e-05
9,895 Towards Interpretable and Learnable Risk Analysis for Entity Resolution 2020 SIGMOD 4.2559233e-05
9,923 On Saving Outliers for Better Clustering over Noisy Data 2021 SIGMOD 4.2503475e-05
9,962 Parallel Rule Discovery from Large Datasets by Sampling 2022 SIGMOD 4.2254157e-05
9,983 Towards Scalable Visual Data Wrangling via Direct Manipulation 2026 CIDR 4.1905499e-05
Previous Page 2 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
268 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00029739054
321 Record Linkage: Similarity Measures and Algorithms 2006 SIGMOD 0.00027524716
488 Data Curation at Scale: The Data Tamer System 2013 CIDR 0.00022029993
556 Discovering Denial Constraints 2013 VLDB 0.00020214701
560 Dependencies Revisited for Improving Data Quality 2008 PODS 0.0002012328
621 Improving Data Quality: Consistency and Accuracy 2007 VLDB 0.00018978331
656 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00018590675
668 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00018428925
700 Reasoning about Record Matching Rules 2009 VLDB 0.00017927576
799 Entity Resolution: Theory, Practice & Open Challenges 2012 VLDB 0.00016479804
879 Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes 2013 SIGMOD 0.00015649604
1,012 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014638349
1,015 Tuffy: Scaling up Statistical Inference in Markov Logic Networks using an RDBMS 2011 VLDB 0.00014630577
1,044 DimmWitted: A Study of Main-Memory Statistical Analytics 2014 VLDB 0.00014465007
1,160 Towards Certain Fixes with Editing Rules and Master Data 2010 VLDB 0.0001358129
1,197 The LLUNATIC Data-Cleaning Framework 2013 VLDB 0.00013373177
1,214 Truth Finding on the Deep Web: Is the Problem Solved? 2013 VLDB 0.00013246179
1,403 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00012180046
1,544 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00011438274
1,625 Sampling the Repairs of Functional Dependency Violations under Hard Constraints 2010 VLDB 0.0001109361
3,045 Dichotomies in the Complexity of Preferred Repairs 2015 PODS 7.6618214e-05
3,198 Towards Dependable Data Repairing with Fixing Rules 2014 SIGMOD 7.4029546e-05
3,900 SLiMFast: Guaranteed Results for Data Fusion and Source Reliability 2017 SIGMOD 6.649432e-05
Previous Page 1 / 1 Next

Semantically Similar Papers