DBScholar

Back to papers

ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data Generation

Summary: ShadowAQP allocates sample sizes per group-by/join attribute-value combinations and synthesizes table rows via a conditional VAE (with automatic encoding and model updates) to avoid costly raw sampling. With parallel multi-round aggregation, outlier-aware sampling and dimensionality reduction it yields up to 12.8x speedups and ~74% average error reduction versus SOTA. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
hf8c79ac12ff3b8e9
Venue
VLDB
Year
2023
Pagerank
5.2930951e-05
Overall Rank
8,659 | 41.79%
DOI
10.14778/3625054.3625059

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{gu_vldb23,
        title = {{ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data Generation}},
        author = {Gu, Rong and Li, Han and Dai, Haipeng and Huang, Wenjie and Xue, Jie and Li, Meng and Zheng, Jiaqi and Cai, Haoran and Huang, Yihua and Chen, Guihai},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {13},
        pages = {4216--4229},
        doi = {10.14778/3625054.3625059},
        url = {https://doi.org/10.14778/3625054.3625059},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 36 of 36 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
36 Accurate Estimation Of The Number Of Tuples Satisfying A Condition 1984 SIGMOD 0.00047863192
57 On Random Sampling over Joins 1999 SIGMOD 0.00040108301
85 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035864347
135 Ripple Joins for Online Aggregation 1999 SIGMOD 0.00029866033
138 Join Synopses for Approximate Query Answering 1999 SIGMOD 0.00029627449
309 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021384073
314 An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement Learning 2019 SIGMOD 0.00021282642
318 DeepDB: Learn from Data, not from Queries! 2020 VLDB 0.00021167555
437 QTune: A Query-Aware Database Tuning System with Deep Reinforcement Learning 2019 VLDB 0.00018315867
461 An End-to-End Learning-based Cost Estimator 2020 VLDB 0.00017829982
564 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016296665
596 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00015785583
784 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.00014012614
795 Random Sampling over Joins Revisited 2018 SIGMOD 0.00013938779
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
948 Aqua: A Fast Decision Support System Using Approximate Query Answers 1999 VLDB 0.00012914559
1,064 Are We Ready For Learned Cardinality Estimation? 2021 VLDB 0.00012202282
1,678 Two-Level Sampling for Join Size Estimation 2017 SIGMOD 9.9088372e-05
1,797 IDEBench: A Benchmark for Interactive Data Exploration 2020 SIGMOD 9.6187325e-05
1,829 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.5510333e-05
2,000 Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee 2016 SIGMOD 9.2112617e-05
2,522 Learned Cardinality Estimation: A Design Space Exploration and A Comparative Evaluation 2022 VLDB 8.3477168e-05
2,846 FactorJoin: A New Cardinality Estimation Framework for Join Queries 2023 SIGMOD 7.9453616e-05
2,873 SnappyData: A Unified Cluster for Streaming, Transactions, and Interactive Analytics 2017 CIDR 7.9225385e-05
3,052 A Unified Deep Model of Learning from both Data and Queries for Cardinality Estimation 2021 SIGMOD 7.7052471e-05
3,341 I've Seen "Enough": Incrementally Improving Visualizations to Support Rapid Decision Making 2017 VLDB 7.4063139e-05
3,424 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.3117029e-05
4,688 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.4697463e-05
4,781 Learned Approximate Query Processing: Make it Light, Accurate and Fast 2021 CIDR 6.4162085e-05
5,473 ABS: a System for Scalable Approximate Queries with Accuracy Guarantees 2014 SIGMOD 6.1178467e-05
7,561 Selectivity Estimation on Streaming Spatio-Textual Data Using Local Correlations 2015 VLDB 5.4958573e-05
7,965 Efficient Approximate Algorithms for Empirical Entropy and Mutual Information 2021 SIGMOD 5.4161136e-05
8,276 Consistent and Flexible Selectivity Estimation for High-Dimensional Data 2021 SIGMOD 5.3641556e-05
8,846 Data Driven Approximation with Bounded Resources 2017 VLDB 5.2649328e-05
10,051 Efficient Insights Discovery through Conditional Generative Model based Query Approximation 2022 SIGMOD 5.0891505e-05
10,269 On Saving Outliers for Better Clustering over Noisy Data 2021 SIGMOD 5.0489944e-05
Previous Page 1 / 1 Next

Semantically Similar Papers