DBScholar

Back to papers

A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data

Summary: Sample-and-Clean blends SAQP with selective cleaning on a small subset to reduce dirty-data bias. Derives confidence intervals by sample size and shows accuracy gains with speedups on noisy TPC-H, Microsoft Academic, and sensor data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h55639aa22d43d3c3
Venue
SIGMOD
Year
2014
Pagerank
9.7965659e-05
Overall Rank
1,720 | 88.44%
DOI
10.1145/2588555.2610505

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{wang_sigmod14,
        title = {{A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data}},
        author = {Wang, Jiannan and Krishnan, Sanjay and Franklin, Michael J. and Goldberg, Ken and Kraska, Tim and Milo, Tova},
        series = {{SIGMOD} '14},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2588555.2610505},
        url = {https://dl.acm.org/doi/10.1145/2588555.2610505},
        year = {2014}
}

Incoming Citations (Sorted by Pagerank)

Showing 38 of 38 citing papers.

Rank Citing Paper Year Venue Pagerank
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017590977
1,043 Data Cleaning: Overview and Emerging Challenges 2016 SIGMOD 0.00012335114
1,342 Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning 2020 VLDB 0.00010968223
1,409 Northstar: An Interactive Data Science System 2018 VLDB 0.00010743451
1,428 Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems 2014 SIGMOD 0.00010693831
1,803 Tuplex: Data Science in Python at Native Code Speed 2021 SIGMOD 9.6068397e-05
1,847 Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions 2021 VLDB 9.5120573e-05
2,423 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.4877894e-05
2,504 Query-Oriented Data Cleaning with Oracles 2015 SIGMOD 8.3782213e-05
2,838 Towards Sustainable Insights or why polygamy is bad for you 2017 CIDR 7.9527257e-05
3,307 Cleaning Crowdsourced Labels Using Oracles for Statistical Classification 2019 VLDB 7.4414303e-05
3,360 Cleaning Denial Constraint Violations through Relaxation 2020 SIGMOD 7.3767115e-05
3,424 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.3117029e-05
3,839 QASCA: A Quality-Aware Task Assignment System for Crowdsourcing Applications 2015 SIGMOD 6.9921604e-05
4,024 CLAMShell: Speeding up Crowds for Low-latency Data Labeling 2016 VLDB 6.8472018e-05
4,253 Sample Debiasing in the Themis Open World Database System 2020 SIGMOD 6.7015395e-05
4,347 Horizon: Scalable Dependency-driven Data Cleaning 2021 VLDB 6.6469984e-05
4,960 PrivateClean: Data Cleaning and Differential Privacy 2016 SIGMOD 6.3390487e-05
5,198 QuERy: A Framework for Integrating Entity Resolution with Query Processing 2016 VLDB 6.2338998e-05
5,379 Efficient Knowledge Graph Accuracy Evaluation 2019 VLDB 6.1558829e-05
5,881 ActiveClean: An Interactive Data Cleaning Framework For Modern Machine Learning 2016 SIGMOD 5.9588636e-05
6,221 Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing 2021 SIGMOD 5.8463347e-05
6,381 Qualitative Data Cleaning 2016 VLDB 5.8066959e-05
6,537 QOCO: A Query Oriented Data Cleaning System with Oracles 2015 VLDB 5.7535125e-05
7,166 Learning to Sample: Counting with Complex Queries 2020 VLDB 5.5949741e-05
7,298 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 5.561054e-05
7,500 Crowdsourced Data Management: Overview and Challenges 2017 SIGMOD 5.5083793e-05
7,704 ReStore - Neural Data Completion for Relational Databases 2021 SIGMOD 5.4741304e-05
8,335 ICARUS: Minimizing Human Effort in Iterative Data Completion 2018 VLDB 5.3525933e-05
8,777 Wisteria: Nurturing Scalable Data Cleaning Infrastructure 2015 VLDB 5.2792451e-05
8,870 Stale View Cleaning: Getting Fresh Answers from Stale Materialized Views 2015 VLDB 5.2601766e-05
9,373 Query-Guided Resolution in Uncertain Databases 2023 SIGMOD 5.1868213e-05
9,384 Selecting Data to Clean for Fact Checking: Minimizing Uncertainty vs. Maximizing Surprise 2019 VLDB 5.1868213e-05
9,385 A Data Quality Metric (DQM): How to Estimate the Number of Undetected Errors in Data Sets 2017 VLDB 5.1868213e-05
9,565 Deduplicated Sampling On-Demand 2025 VLDB 5.1571823e-05
9,616 GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language Models 2024 SIGMOD 5.1510548e-05
10,515 Understanding the Impact of Data Noise in Federated Learning: [Experiments & Analysis] 2026 SIGMOD 4.9793485e-05
11,572 Efficient and Reliable Estimation of Knowledge Graph Accuracy 2024 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 21 of 21 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
9 Online Aggregation 1997 SIGMOD 0.00076195956
77 Sampling-Based Estimation of the Number of Distinct Values of an Attribute 1995 VLDB 0.00036828234
198 CrowdER: Crowdsourcing Entity Resolution 2012 VLDB 0.00025555196
244 Evaluation of entity resolution approaches on real-world match problems 2010 VLDB 0.00023314591
295 Towards Estimation Error Guarantees for Distinct Values 2000 PODS 0.00021914399
336 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00020657819
372 Approximate Query Processing: Taming the TeraBytes! A Tutorial 2001 VLDB 0.00019720059
427 Big Data Integration 2013 VLDB 0.00018465558
564 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016296665
697 NADEEF: A Commodity Data Cleaning System 2013 SIGMOD 0.00014694048
707 On Synopses for Distinct-Value Estimation Under Multiset Operations 2007 SIGMOD 0.00014640173
716 Guided Data Repair 2011 VLDB 0.00014553463
871 Leveraging Transitive Relations for Crowdsourced Joins 2013 SIGMOD 0.00013343705
888 Pay-as-you-go User Feedback for Dataspace Systems 2008 SIGMOD 0.00013255989
931 Dynamic Sample Selection for Approximate Query Processing 2003 SIGMOD 0.00013011667
989 Towards Certain Fixes with Editing Rules and Master Data 2010 VLDB 0.0001265344
1,022 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00012438826
1,607 SciBORQ: Scientific data management with Bounds On Runtime and Quality 2011 CIDR 0.0001008742
2,361 Online Aggregation and Continuous Query support in MapReduce 2010 SIGMOD 8.5761274e-05
2,653 CrowdFill: Collecting Structured Data from the Crowd 2014 SIGMOD 8.1732164e-05
3,905 Distributed Online Aggregations 2009 VLDB 6.9335334e-05
Previous Page 1 / 1 Next

Semantically Similar Papers