Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query Workload
Summary: Prestroid uses tree-convolution to predict SQL query resource usage from traces, reducing encoding/padding waste in large-scale DL training. On 19k Presto queries over 20PB, it outperforms baselines and cuts memory 13.5x and epoch time 3.45x, with up to 13.2x Azure savings. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Johan Kok Zhi Kang (National University of Singapore)
- 2. Gaurav (GrabTaxi Holdings)
- 3. Sien Yi Tan (GrabTaxi Holdings)
- 4. Feng Cheng (GrabTaxi Holdings)
- 5. Shixuan Sun (National University of Singapore)
- 6. Bingsheng He (National University of Singapore)
BibTeX Citation
@inproceedings{kang_sigmod21,
title = {{Efficient Deep Learning Pipelines for Accurate Cost Estimations Over Large Scale Query Workload}},
author = {Kang, Johan Kok Zhi and Gaurav and Tan, Sien Yi and Cheng, Feng and Sun, Shixuan and He, Bingsheng},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457546},
url = {https://dl.acm.org/doi/10.1145/3448016.3457546},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 16 of 16 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 84 | Learned Cardinalities: Estimating Correlated Joins with Deep Learning | 2019 | CIDR | 0.00035838391 |
| 154 | Neo: A Learned Query Optimizer | 2019 | VLDB | 0.00028726181 |
| 465 | An End-to-End Learning-based Cost Estimator | 2020 | VLDB | 0.0001803934 |
| 1,170 | QuickSel: Quick Selectivity Learning with Mixture Models | 2020 | SIGMOD | 0.00011827259 |
| 2,177 | The Case for Predictive Database Systems: Opportunities and Challenges | 2011 | CIDR | 9.0159844e-05 |
| 3,051 | Towards a Hands-Free Query Optimizer through Deep Learning | 2019 | CIDR | 7.8121919e-05 |
| 4,816 | A Top-Down Approach to Achieving Performance Predictability in Database Systems | 2017 | SIGMOD | 6.4949878e-05 |
| 5,132 | Facilitating SQL Query Composition and Analysis | 2020 | SIGMOD | 6.3534526e-05 |
| 7,619 | AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft | 2020 | VLDB | 5.5810604e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,799 | DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models | 2019 | SIGMOD |
| 2 | 5,073 | Database Workload Characterization with Query Plan Encoders | 2022 | VLDB |
| 3 | 9,353 | On Efficient Approximate Queries over Machine Learning Models | 2023 | VLDB |
| 4 | 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD |
| 5 | 465 | An End-to-End Learning-based Cost Estimator | 2020 | VLDB |
| 6 | 11,845 | Query-Driven Learning for Next Generation Predictive Modeling & Analytics | 2019 | SIGMOD |
| 7 | 563 | Plan-Structured Deep Neural Network Models for Query Performance Prediction | 2019 | VLDB |
| 8 | 9,215 | Deep Query Optimization | 2019 | SIGMOD |
| 9 | 2,822 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings | 2020 | SIGMOD |
| 10 | 2,844 | Zero-Shot Cost Models for Out-of-the-box Learned Cost Prediction | 2022 | VLDB |