KEA: Tuning an Exabyte-Scale Data Infrastructure
Summary: KEA automates tuning of exabyte-scale data infra with ML models from telemetry, using observational tuning and cautious production flighting. First study addressing exabyte-scale data-management tuning, with potential tens of millions in annual savings. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yiwen Zhu (Microsoft)
- 2. Subru Krishnan (Microsoft)
- 3. Konstantinos Karanasos (Microsoft)
- 4. Isha Tarte (Microsoft)
- 5. Conor Power (Microsoft)
- 6. Abhishek Modi (Microsoft)
- 7. Manoj Kumar (Microsoft)
- 8. Deli Zhang (Microsoft)
- 9. Kartheek Muthyala (Microsoft)
- 10. Nick Jurgens (Microsoft)
- 11. Sarvesh Sakalanaga (Salesforce)
- 12. Sudhir Darbha (Microsoft)
- 13. Minu Iyer (Microsoft)
- 14. Ankita Agarwal (Microsoft)
- 15. Carlo Curino (Microsoft)
BibTeX Citation
@inproceedings{zhu_sigmod21,
title = {{KEA: Tuning an Exabyte-Scale Data Infrastructure}},
author = {Zhu, Yiwen and Krishnan, Subru and Karanasos, Konstantinos and Tarte, Isha and Power, Conor and Modi, Abhishek and Kumar, Manoj and Zhang, Deli and Muthyala, Kartheek and Jurgens, Nick and Sakalanaga, Sarvesh and Darbha, Sudhir and Iyer, Minu and Agarwal, Ankita and Curino, Carlo},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457569},
url = {https://dl.acm.org/doi/10.1145/3448016.3457569},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 14 of 14 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,008 | Making Data Clouds Smarter at Keebo: Automated Warehouse Optimization using Data Learning | 2023 | SIGMOD |
| 2 | 6,121 | The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward | 2021 | VLDB |
| 3 | 3,343 | LlamaTune: Sample-Efficient DBMS Configuration Tuning | 2022 | VLDB |
| 4 | 5,978 | From Auto-tuning One Size Fits All to Self-designed and Learned Data-intensive Systems | 2019 | SIGMOD |
| 5 | 5,865 | Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems | 2019 | VLDB |
| 6 | 7,762 | Runtime Variation in Big Data Analytics | 2023 | SIGMOD |
| 7 | 4,742 | Continuous Cloud-Scale Query Optimization and Processing | 2013 | VLDB |
| 8 | 7,619 | AutoToken: Predicting Peak Parallelism for Big Data Analytics at Microsoft | 2020 | VLDB |
| 9 | 2,822 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings | 2020 | SIGMOD |
| 10 | 5,059 | Steering Query Optimizers: A Practical Take on Big Data Workloads | 2021 | SIGMOD |