Error-bounded Sampling for Analytics on Big Sparse Data
Summary: Introduces error-bounded stratified sampling for aggregation over massive, wide-range sparse data, preserving user-specified error guarantees. Distribution-aware strata reduce sample size by up to 99% versus uniform sampling, with robust performance in Microsoft’s search-query platform. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Ying Yan (Microsoft)
- 2. Liang Jeff Chen (Microsoft)
- 3. Zheng Zhang (Microsoft)
BibTeX Citation
@article{yan_vldb14,
title = {{Error-bounded Sampling for Analytics on Big Sparse Data}},
author = {Yan, Ying and Chen, Liang Jeff and Zhang, Zheng},
journal = {PVLDB},
series = {{VLDB} '14},
volume = {7},
number = {13},
pages = {1508--1519},
doi = {10.14778/2733004.2733012},
url = {https://doi.org/10.14778/2733004.2733012},
year = {2014}
}
Incoming Citations (Sorted by Pagerank)
Showing 11 of 11 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9 | Online Aggregation | 1997 | SIGMOD | 0.00077458002 |
| 30 | SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets | 2008 | VLDB | 0.00051174276 |
| 553 | Congressional Samples for Approximate Answering of Group-By Queries | 2000 | SIGMOD | 0.00016590619 |
| 909 | Dynamic Sample Selection for Approximate Query Processing | 2003 | SIGMOD | 0.00013291205 |
| 1,009 | Online Aggregation for Large MapReduce Jobs | 2011 | VLDB | 0.00012684342 |
| 1,166 | ICICLES: Self-tuning Samples for Approximate Query Answering | 2000 | VLDB | 0.00011850439 |
| 1,582 | SciBORQ: Scientific data management with Bounds On Runtime and Quality | 2011 | CIDR | 0.00010295367 |
| 2,312 | Online Aggregation and Continuous Query support in MapReduce | 2010 | SIGMOD | 8.7642158e-05 |
| 3,042 | Continuous Sampling for Online Aggregation Over Multiple Queries | 2010 | SIGMOD | 7.8231049e-05 |
| 3,096 | Early Accurate Results for Advanced Analytics on MapReduce | 2012 | VLDB | 7.7629371e-05 |
| 3,844 | Distributed Online Aggregations | 2009 | VLDB | 7.0782059e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 173 | Simple Random Sampling from Relational Databases | 1986 | VLDB |
| 2 | 6,206 | Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing | 2021 | SIGMOD |
| 3 | 149 | New Sampling-Based Summary Statistics for Improving Approximate Query Answers | 1998 | SIGMOD |
| 4 | 508 | Random Sampling for Histogram Construction: How much is enough? | 1998 | SIGMOD |
| 5 | 7,419 | Structure-Aware Sampling: Flexible and Accurate Summarization | 2011 | VLDB |
| 6 | 8,379 | Sampling Big Ideas in Query Optimization | 2023 | PODS |
| 7 | 1,401 | Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems | 2014 | SIGMOD |
| 8 | 2,608 | A Robust, Optimization-Based Approach for Approximate Answering of Aggregate Queries | 2001 | SIGMOD |
| 9 | 1,962 | Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee | 2016 | SIGMOD |
| 10 | 8,108 | Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters | 2019 | VLDB |