A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data
Summary: Sample-and-Clean blends SAQP with selective cleaning on a small subset to reduce dirty-data bias. Derives confidence intervals by sample size and shows accuracy gains with speedups on noisy TPC-H, Microsoft Academic, and sensor data. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Jiannan Wang (University of California Berkeley)
- 2. Sanjay Krishnan (University of California Berkeley)
- 3. Michael J. Franklin (University of California Berkeley)
- 4. Ken Goldberg (University of California Berkeley)
- 5. Tim Kraska (Brown University)
- 6. Tova Milo (Tel Aviv University)
BibTeX Citation
@inproceedings{wang_sigmod14,
title = {{A Sample-and-Clean Framework for Fast and Accurate Query Processing on Dirty Data}},
author = {Wang, Jiannan and Krishnan, Sanjay and Franklin, Michael J. and Goldberg, Ken and Kraska, Tim and Milo, Tova},
series = {{SIGMOD} '14},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2588555.2610505},
url = {https://dl.acm.org/doi/10.1145/2588555.2610505},
year = {2014}
}
Incoming Citations (Sorted by Pagerank)
Showing 38 of 38 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 21 of 21 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 149 | New Sampling-Based Summary Statistics for Improving Approximate Query Answers | 1998 | SIGMOD |
| 2 | 6,206 | Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing | 2021 | SIGMOD |
| 3 | 1,872 | The Analytical Bootstrap: a New Method for Fast Error Estimation in Approximate Query Processing | 2014 | SIGMOD |
| 4 | 5,318 | Cleaning Uncertain Data with Quality Guarantees | 2008 | VLDB |
| 5 | 1,962 | Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee | 2016 | SIGMOD |
| 6 | 10,634 | Efficient Approximate Query Processing with Block Sampling | 2025 | CIDR |
| 7 | 8,108 | Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters | 2019 | VLDB |
| 8 | 4,696 | Error-bounded Sampling for Analytics on Big Sparse Data | 2014 | VLDB |
| 9 | 6,827 | Sampling Dirty Data for Matching Attributes | 2010 | SIGMOD |
| 10 | 1,401 | Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems | 2014 | SIGMOD |