Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters
Summary: An empirical study of sampling-based query approximation deployed in Microsoft’s production big-data clusters. Details implementation choices, workload use cases, and evidence on when sampling delivers useful answers at scale. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Srikanth Kandula (Microsoft)
- 2. Kukjin Lee (Microsoft)
- 3. Surajit Chaudhuri (Microsoft)
- 4. Marc Friedman (Microsoft)
BibTeX Citation
@article{kandula_vldb19,
title = {{Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters}},
author = {Kandula, Srikanth and Lee, Kukjin and Chaudhuri, Surajit and Friedman, Marc},
journal = {PVLDB},
series = {{VLDB} '19},
volume = {12},
number = {12},
pages = {2131--2142},
doi = {10.14778/3352063.3352130},
url = {https://doi.org/10.14778/3352063.3352130},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,121 | The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward | 2021 | VLDB | 5.9688569e-05 |
| 6,206 | Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing | 2021 | SIGMOD | 5.9443409e-05 |
| 8,161 | LAQy: Efficient and Reusable Query Approximations via Lazy Sampling | 2023 | SIGMOD | 5.4752972e-05 |
| 8,608 | One Size Does Not Fit All: A Bandit-Based Sampler Combination Framework with Theoretical Guarantees | 2022 | SIGMOD | 5.4024561e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,379 | Sampling Big Ideas in Query Optimization | 2023 | PODS |
| 2 | 76 | Practical Selectivity Estimation through Adaptive Sampling | 1990 | SIGMOD |
| 3 | 909 | Dynamic Sample Selection for Approximate Query Processing | 2003 | SIGMOD |
| 4 | 435 | Histogram-Based Approximation of Set-Valued Query Answers | 1999 | VLDB |
| 5 | 10,634 | Efficient Approximate Query Processing with Block Sampling | 2025 | CIDR |
| 6 | 1,962 | Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee | 2016 | SIGMOD |
| 7 | 2,608 | A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries | 2001 | SIGMOD |
| 8 | 1,401 | Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems | 2014 | SIGMOD |
| 9 | 149 | New Sampling-Based Summary Statistics for Improving Approximate Query Answers | 1998 | SIGMOD |
| 10 | 4,696 | Error-bounded Sampling for Analytics on Big Sparse Data | 2014 | VLDB |