Sapprox: Enabling Efficient and Accurate Approximations on Sub-datasets with Distribution-aware Online Sampling
Summary: Sapprox enables efficient, accurate approximations on arbitrary sub-datasets via distribution-aware sampling. Uses a probabilistic map to flatten subsets, applies unequal-probability sampling, and optimizes unit size, yielding 20x speedups in Hadoop. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Xuhong Zhang (University of Central Florida)
- 2. Jun Wang (University of Central Florida)
- 3. Jiangling Yin (University of Central Florida)
BibTeX Citation
@article{zhang_vldb17,
title = {{Sapprox: Enabling Efficient and Accurate Approximations on Sub-datasets with Distribution-aware Online Sampling}},
author = {Zhang, Xuhong and Wang, Jun and Yin, Jiangling},
journal = {PVLDB},
series = {{VLDB} '17},
volume = {10},
number = {3},
pages = {109--120},
doi = {10.14778/3021924.3021925},
url = {https://doi.org/10.14778/3021924.3021925},
year = {2017}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,129 | Marviq: Quality-Aware Geospatial Visualization of Range-Selection Queries Using Materialization | 2020 | SIGMOD | 5.694968e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB | 0.00050111008 |
| 819 | Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters | 2016 | SIGMOD | 0.00013815639 |
| 1,009 | Online Aggregation for Large MapReduce Jobs | 2011 | VLDB | 0.00012684342 |
| 2,312 | Online Aggregation and Continuous Query support in MapReduce | 2010 | SIGMOD | 8.7642158e-05 |
| 2,633 | Relational Confidence Bounds Are Easy With The Bootstrap* | 2005 | SIGMOD | 8.3224527e-05 |
| 3,096 | Early Accurate Results for Advanced Analytics on MapReduce | 2012 | VLDB | 7.7629371e-05 |
| 3,803 | A Bi-Level Bernoulli Scheme for Database Sampling | 2004 | SIGMOD | 7.1114677e-05 |
| 4,696 | Error-bounded Sampling for Analytics on Big Sparse Data | 2014 | VLDB | 6.557612e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,634 | Efficient Approximate Query Processing with Block Sampling | 2025 | CIDR |
| 2 | 909 | Dynamic Sample Selection for Approximate Query Processing | 2003 | SIGMOD |
| 3 | 149 | New Sampling-Based Summary Statistics for Improving Approximate Query Answers | 1998 | SIGMOD |
| 4 | 11,418 | Efficient Approximation Framework for Attribute Recommendation | 2023 | SIGMOD |
| 5 | 261 | Approximate Medians and other Quantiles in One Pass and with Limited Memory | 1998 | SIGMOD |
| 6 | 5,193 | Sampling Algorithms in a Stream Operator | 2005 | SIGMOD |
| 7 | 6,206 | Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing | 2021 | SIGMOD |
| 8 | 1,401 | Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems | 2014 | SIGMOD |
| 9 | 8,108 | Experiences with Approximating Queries in Microsoft’s Production Big-Data Clusters | 2019 | VLDB |
| 10 | 1,962 | Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee | 2016 | SIGMOD |