DBScholar

Back to papers

A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data

Summary: Sample-and-Clean blends SAQP with selective cleaning on a small subset to reduce dirty-data bias. Derives confidence intervals by sample size and shows accuracy gains with speedups on noisy TPC-H, Microsoft Academic, and sensor data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
4938
Venue
SIGMOD
Year
2014
Pagerank
9.8984415e-05
Overall Rank
1,736 | 88.10%
DOI
10.1145/2588555.2610505

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{wang_sigmod14,
        title = {{A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data}},
        author = {Wang, Jiannan and Krishnan, Sanjay and Franklin, Michael J. and Goldberg, Ken and Kraska, Tim and Milo, Tova},
        series = {{SIGMOD} '14},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2588555.2610505},
        url = {https://dl.acm.org/doi/10.1145/2588555.2610505},
        year = {2014}
}

Incoming Citations (Sorted by Pagerank)

Showing 38 of 38 citing papers.

Rank Citing Paper Year Venue Pagerank
582 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00016148948
1,323 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00011152602
1,392 Northstar: An Interactive Data Science System 2018 VLDB 0.00010936065
1,401 Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems 2014 SIGMOD 0.00010889902
1,476 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010659277
1,768 Tuplex: Data Science in Python at Native Code Speed 2021 SIGMOD 9.8041636e-05
2,147 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.0831495e-05
2,398 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.631172e-05
2,570 Query-Oriented Data Cleaning with Oracles 2015 SIGMOD 8.4061164e-05
2,840 Towards Sustainable Insights or why polygamy is bad for you 2017 CIDR 8.0650295e-05
3,281 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.5706653e-05
3,361 Cleaning Denial Constraint Violations through Relaxation 2020 SIGMOD 7.4872032e-05
3,366 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.4748604e-05
3,904 QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications 2015 SIGMOD 7.0304212e-05
3,971 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.9835263e-05
4,205 Sample Debiasing in the Themis Open World Database System 2020 SIGMOD 6.8337021e-05
4,488 Horizon: Scalable Dependency-driven Data Cleaning 2021 VLDB 6.668457e-05
4,837 PrivateClean: Data Cleaning and Differential Privacy 2016 SIGMOD 6.4845444e-05
5,204 QuERy: A Framework for Integrating Entity Resolution with Query Processing 2016 VLDB 6.3187983e-05
5,794 ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning 2016 SIGMOD 6.0863045e-05
6,206 Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing 2021 SIGMOD 5.9443409e-05
6,392 Qualitative Data Cleaning 2016 VLDB 5.8883971e-05
6,471 QOCO: A Query Oriented Data Cleaning System with Oracles 2015 VLDB 5.8702801e-05
6,581 Efficient Knowledge Graph Accuracy Evaluation 2019 VLDB 5.8364579e-05
7,048 Learning to Sample: Counting with Complex Queries 2020 VLDB 5.7178054e-05
7,179 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 5.6802479e-05
7,357 Crowdsourced Data Management: Overview and Challenges 2017 SIGMOD 5.6346837e-05
7,556 ReStore - Neural Data Completion for Relational Databases 2021 SIGMOD 5.5997742e-05
8,160 ICARUS: Minimizing Human Effort in Iterative Data Completion 2018 VLDB 5.4754476e-05
8,620 Wisteria: Nurturing Scalable Data Cleaning Infrastructure 2015 VLDB 5.3991092e-05
8,714 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.3778009e-05
9,193 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 5.3058708e-05
9,204 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 5.3058708e-05
9,206 A Data Quality Metric (DQM): How to Estimate the Number of Undetected Errors in Data Sets 2017 VLDB 5.3058708e-05
9,381 Deduplicated Sampling On-Demand 2025 VLDB 5.2755515e-05
9,437 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.2687567e-05
10,303 Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis] 2026 SIGMOD 5.093636e-05
11,239 Efficient and Reliable Estimation of Knowledge Graph Accuracy 2024 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 21 of 21 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
9 Online Aggregation 1997 SIGMOD 0.00077458002
75 Sampling-Based Estimation of the Number of Distinct Values of an Attribute 1995 VLDB 0.00037277061
196 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025780596
248 Evaluation of entity resolution approaches on real-world match problems 2010 VLDB 0.00023278354
288 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00022296371
327 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00021091539
363 Approximate Query Processing: Taming the TeraBytes! A Tutorial 2001 VLDB 0.0002005475
427 Big Data Integration 2013 VLDB 0.00018661543
553 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016590619
689 On Synopses for Distinct-Value Estimation Under Multiset Operations 2007 SIGMOD 0.00014940023
714 Guided Data Repair 2011 VLDB 0.00014662041
725 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014617251
852 Leveraging Transitive Relations for Crowdsourced Joins 2013 SIGMOD 0.00013604253
866 Pay-as-you-go User Feedback for Dataspace Systems 2008 SIGMOD 0.00013519288
909 Dynamic Sample Selection for Approximate Query Processing 2003 SIGMOD 0.00013291205
998 Towards Certain Fixes with Editing Rules and Master Data 2010 VLDB 0.0001275238
1,009 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00012684342
1,582 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.00010295367
2,312 Online Aggregation and Continuous Query support in MapReduce 2010 SIGMOD 8.7642158e-05
2,636 CrowdFill: Collecting Structured Data from the Crowd 2014 SIGMOD 8.31868e-05
3,844 Distributed Online Aggregations 2009 VLDB 7.0782059e-05
Previous Page 1 / 1 Next

Semantically Similar Papers