DBScholar

Back to papers

ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data Generation

Summary: ShadowAQP allocates sample sizes per group-by/join attribute-value combinations and synthesizes table rows via a conditional VAE (with automatic encoding and model updates) to avoid costly raw sampling. With parallel multi-round aggregation, outlier-aware sampling and dimensionality reduction it yields up to 12.8x speedups and ~74% average error reduction versus SOTA. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13479
Venue
VLDB
Year
2023
Pagerank
5.4145838e-05
Overall Rank
8,492 | 41.74%
DOI
10.14778/3625054.3625059

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{gu_vldb23,
        title = {{ShadowAQP: Efficient Approximate Group-by and Join Query via Attribute-oriented Sample Size Allocation and Data Generation}},
        author = {Gu, Rong and Li, Han and Dai, Haipeng and Huang, Wenjie and Xue, Jie and Li, Meng and Zheng, Jiaqi and Cai, Haoran and Huang, Yihua and Chen, Guihai},
        journal = {PVLDB},
        series = {{VLDB} '23},
        volume = {16},
        number = {13},
        pages = {4216--4229},
        doi = {10.14778/3625054.3625059},
        url = {https://doi.org/10.14778/3625054.3625059},
        year = {2023}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 36 of 36 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
36 Accurate Estimation Of The Number Of Tuples Satisfying A Condition 1984 SIGMOD 0.00048351457
54 On Random Sampling over Joins 1999 SIGMOD 0.00040810225
84 Learned Cardinalities: Estimating Correlated Joins with Deep Learning 2019 CIDR 0.00035838391
131 Ripple Joins for Online Aggregation 1999 SIGMOD 0.00030424509
136 Join Synopses for Approximate Query Answering 1999 SIGMOD 0.00030123303
307 Approximate Query Processing Using Wavelets 2000 VLDB 0.00021792475
323 DeepDB: Learn from Data, not from Queries! 2020 VLDB 0.00021264788
334 An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement Learning 2019 SIGMOD 0.00020875082
465 An End-to-End Learning-based Cost Estimator 2020 VLDB 0.0001803934
498 QTune: A Query-Aware Database Tuning System with Deep Reinforcement Learning 2019 VLDB 0.00017440583
553 Congressional Samples for Approximate Answering of Group-By Queries 2000 SIGMOD 0.00016590619
593 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00016027871
772 VerdictDB: Universalizing Approximate Query Processing 2018 SIGMOD 0.00014147905
802 Random Sampling over Joins Revisited 2018 SIGMOD 0.00013907725
819 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013815639
931 Aqua: A Fast Decision Support System Using Approximate Query Answers 1999 VLDB 0.00013125812
1,061 Are We Ready For Learned Cardinality Estimation? 2021 VLDB 0.00012369764
1,664 Two-Level Sampling for Join Size Estimation 2017 SIGMOD 0.00010070362
1,767 IDEBench: A Benchmark for Interactive Data Exploration 2020 SIGMOD 9.805856e-05
1,799 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.7326398e-05
1,962 Sample + Seek: Approximating Aggregates with Distribution Precision Guarantee 2016 SIGMOD 9.3978414e-05
2,723 Learned Cardinality Estimation: A Design Space Exploration and A Comparative Evaluation 2022 VLDB 8.2049453e-05
2,868 SnappyData: A Unified Cluster for Streaming, Transactions, and Interactive Analytics 2017 CIDR 8.0156936e-05
2,991 FactorJoin: A New Cardinality Estimation Framework for Join Queries 2023 SIGMOD 7.8880723e-05
3,086 A Unified Deep Model of Learning from both Data and Queries for Cardinality Estimation 2021 SIGMOD 7.7708642e-05
3,288 I've Seen "Enough": Incrementally Improving Visualizations to Support Rapid Decision Making 2017 VLDB 7.5578177e-05
3,366 AQP++: Connecting Approximate Query Processing With Aggregate Precomputation for Interactive Analytics 2018 SIGMOD 7.4748604e-05
4,617 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.604437e-05
4,789 Learned Approximate Query Processing: Make it Light, Accurate and Fast 2021 CIDR 6.5072039e-05
5,358 ABS: a System for Scalable Approximate Queries with Accuracy Guarantees 2014 SIGMOD 6.2492955e-05
7,433 Selectivity Estimation on Streaming Spatio-Textual Data Using Local Correlations 2015 VLDB 5.6197531e-05
7,807 Efficient Approximate Algorithms for Empirical Entropy and Mutual Information 2021 SIGMOD 5.5404259e-05
8,124 Consistent and Flexible Selectivity Estimation for High-Dimensional Data 2021 SIGMOD 5.4829513e-05
8,701 Data Driven Approximation with Bounded Resources 2017 VLDB 5.3828806e-05
9,860 Efficient Insights Discovery through Conditional Generative Model based Query Approximation 2022 SIGMOD 5.2059582e-05
10,065 On Saving Outliers for Better Clustering over Noisy Data 2021 SIGMOD 5.1648805e-05
Previous Page 1 / 1 Next

Semantically Similar Papers