DBScholar

Back to papers

Data Cleaning: Overview and Emerging Challenges

Summary: Presents a taxonomy of data cleaning, focusing on constraint- and pattern-based detection and repair for data quality. Links qualitative cleaning to ML and statistics, addressing scalability for big data and its impact on analytics, with a statistical view on inference. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hf54e72c502dcf46e
Venue
SIGMOD
Year
2016
Pagerank
0.00012335114
Overall Rank
1,043 | 92.99%
DOI
10.1145/2882903.2912574

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{chu_sigmod16,
        title = {{Data Cleaning: Overview and Emerging Challenges}},
        author = {Chu, Xu and Ilyas, Ihab F. and Krishnan, Sanjay and Wang, Jiannan},
        series = {{SIGMOD} '16},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2882903.2912574},
        url = {https://dl.acm.org/doi/10.1145/2882903.2912574},
        year = {2016}
}

Incoming Citations (Sorted by Pagerank)

Showing 37 of 37 citing papers.

Rank Citing Paper Year Venue Pagerank
1,308 Automating Large-Scale Data Quality Verification 2018 VLDB 0.0001107886
1,978 Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks 2024 SIGMOD 9.2730152e-05
2,371 SCODED: Statistical Constraint Oriented Data Error Detection 2020 SIGMOD 8.5615698e-05
3,263 Automatic Data Repair: Are We Ready to Deploy? 2024 VLDB 7.4809771e-05
3,307 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4414303e-05
3,706 TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines 2020 SIGMOD 7.0829596e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7608137e-05
4,195 PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models 2020 SIGMOD 6.7422305e-05
4,524 DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data 2023 SIGMOD 6.5621504e-05
5,379 Efficient Knowledge Graph Accuracy Evaluation 2019 VLDB 6.1558829e-05
6,046 Your notebook is not crumby enough, REPLace it 2020 CIDR 5.9052121e-05
6,183 Causal Data Integration 2023 VLDB 5.8587542e-05
6,552 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.750057e-05
6,921 CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning 2024 SIGMOD 5.6432616e-05
7,068 How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses 2024 VLDB 5.6083188e-05
7,471 Fast Detection of Denial Constraint Violations 2022 VLDB 5.5176505e-05
7,532 PIClean: A Probabilistic and Interactive Data Cleaning System 2019 SIGMOD 5.5007996e-05
7,704 ReStore - Neural Data Completion for Relational Databases 2021 SIGMOD 5.4741304e-05
7,875 Learning Over Dirty Data Without Cleaning 2020 SIGMOD 5.4355826e-05
8,166 Fast and Reliable Missing Data Contingency Analysis with Predicate-Constraints 2020 SIGMOD 5.3849926e-05
8,280 The Computation of Optimal Subset Repairs 2020 VLDB 5.363617e-05
9,385 A Data Quality Metric (DQM): How to Estimate the Number of Undetected Errors in Data Sets 2017 VLDB 5.1868213e-05
9,722 DataVinci: Learning Syntactic and Semantic String Repairs 2025 SIGMOD 5.1349531e-05
10,181 In-Database Data Imputation 2024 SIGMOD 5.0653015e-05
10,188 Reptile: Aggregation-level Explanations for Hierarchical Data 2022 SIGMOD 5.0651993e-05
10,499 Shape-Agnostic Table Overlap Discovery: A Maximum Common Subhypergraph Approach 2026 SIGMOD 4.9793485e-05
10,515 Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis] 2026 SIGMOD 4.9793485e-05
10,607 WaveStitch: Flexible and Fast Conditional Time Series Generation With Diffusion Models 2026 SIGMOD 4.9793485e-05
10,707 Repairing Property Graphs under PG-Constraints 2026 VLDB 4.9793485e-05
10,804 Document-to-Database: Extraction Meets Relational Semantics 2026 VLDB 4.9793485e-05
10,853 Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models 2026 VLDB 4.9793485e-05
10,970 PGShield: Maintaining Path Constraints in Property Graphs 2026 VLDB 4.9793485e-05
11,572 Efficient and Reliable Estimation of Knowledge Graph Accuracy 2024 VLDB 4.9793485e-05
11,699 LinCQA: Faster Consistent Query Answering with Linear Time Guarantees 2023 SIGMOD 4.9793485e-05
11,731 Demystifying the QoS and QoE of Edge-hosted Video Streaming Applications in the Wild with SNESet 2023 SIGMOD 4.9793485e-05
12,036 LOCATER: Cleaning WiFi Connectivity Datasets for Semantic Localization 2021 VLDB 4.9793485e-05
12,177 IHCS: An Integrated Hybrid Cleaning System 2019 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 45 of 45 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
95 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034382643
188 A Cost-Based Model and Effective Heuristic for Repairing Constraints by Value Modification 2005 SIGMOD 0.00025872962
189 Scorpion: Explaining Away Outliers in Aggregate Queries 2013 VLDB 0.00025840026
198 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025555196
259 Answering Queries using Humans, Algorithms and Databases 2011 CIDR 0.00022923243
307 Eliminating Fuzzy Duplicates in Data Warehouses 2002 VLDB 0.00021499031
343 Model-Driven Data Acquisition in Sensor Networks 2004 VLDB 0.00020519525
350 Discovering Denial Constraints 2013 VLDB 0.00020253521
433 Corleone: Hands-Off Crowdsourcing for Entity Matching 2014 SIGMOD 0.00018332741
514 Data Curation at Scale: The Data Tamer System 2013 CIDR 0.00017006745
526 Improving Data Quality: Consistency and Accuracy 2007 VLDB 0.00016886621
547 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016578131
652 Don’t be SCAREd: Use SCalable Automatic REpairing with Maximal Likelihood and Bounded Changes 2013 SIGMOD 0.00015121325
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014694048
716 Guided Data Repair 2011 VLDB 0.00014553463
767 The LLUNATIC Data-Cleaning Framework 2013 VLDB 0.00014114806
871 Leveraging Transitive Relations for Crowdsourced Joins 2013 SIGMOD 0.00013343705
933 Question Selection for Crowd Entity Resolution 2013 VLDB 0.00013004422
987 CrowdScreen: Algorithms for Filtering Data with Humans 2012 SIGMOD 0.00012660627
989 Towards Certain Fixes with Editing Rules and Master Data 2010 VLDB 0.0001265344
1,031 On Generating Near-Optimal Tableaux for Conditional Functional Dependencies 2008 VLDB 0.00012405065
1,099 KATARA: A Data Cleaning System Powered by Knowledge Bases and Crowdsourcing 2015 SIGMOD 0.00012037058
1,348 Sampling the Repairs of Functional Dependency Violations under Hard Constraints 2010 VLDB 0.0001094733
1,360 Adaptive Cleaning for RFID Data Streams 2006 VLDB 0.00010918075
1,509 Data Quality and Data Cleaning: An Overview 2003 SIGMOD 0.00010439963
1,720 A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data 2014 SIGMOD 9.7965659e-05
2,070 Dedoop: Efficient Deduplication with Hadoop 2012 VLDB 9.0876046e-05
2,359 Online Outlier Detection in Sensor Data Using Non-Parametric Models 2006 VLDB 8.5813849e-05
2,423 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.4877894e-05
2,458 Tracing Data Errors with View-Conditioned Causality 2011 SIGMOD 8.4371456e-05
2,504 Query-Oriented Data Cleaning with Oracles 2015 SIGMOD 8.3782213e-05
2,642 Scaling Up Crowd-Sourcing to Very Large Datasets: A Case for Active Learning 2015 VLDB 8.1778168e-05
2,645 Progressive Approach to Relational Entity Resolution 2014 VLDB 8.177221e-05
2,653 CrowdFill: Collecting Structured Data from the Crowd 2014 SIGMOD 8.1732164e-05
2,784 Interaction between Record Matching and Data Repairing 2011 SIGMOD 8.018482e-05
2,964 Towards Dependable Data Repairing with Fixing Rules 2014 SIGMOD 7.8052551e-05
3,426 Modeling and Querying Possible Repairs in Duplicate Detection 2009 VLDB 7.3087723e-05
3,946 Continuous Outlier Detection in Data Streams: An Extensible Framework and State-Of-The-Art Algorithms 2013 SIGMOD 6.9083867e-05
4,024 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.8472018e-05
4,998 Descriptive and Prescriptive Data Cleaning 2014 SIGMOD 6.3207886e-05
5,198 QuERy: A Framework for Integrating Entity Resolution with Query Processing 2016 VLDB 6.2338998e-05
6,880 Estimating the Impact of Unknown Unknowns on Aggregate Query Results 2016 SIGMOD 5.6573214e-05
8,244 When Speed Has a Price: Fast Information Extraction Using Approximate Algorithms 2013 VLDB 5.3698153e-05
8,777 Wisteria: Nurturing Scalable Data Cleaning Infrastructure 2015 VLDB 5.2792451e-05
8,870 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.2601766e-05
Previous Page 1 / 1 Next

Semantically Similar Papers