Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance
Summary: ByteDance’s governance framework addresses Spark I/O stalls, coarse resource controls, and misconfiguration across 1M+ jobs and 500PB daily shuffle. Push-based CSS, fine-grained controls, and two-stage autotuning deliver 22% higher CPU utilization and substantial resource savings. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Yixin Wu (ByteDance)
- 2. Xiuqi Huang (Shanghai Jiao Tong University)
- 3. Zhongjia Wei (ByteDance)
- 4. Hang Cheng (ByteDance)
- 5. Chaohui Xin (ByteDance)
- 6. Zuzhi Chen (ByteDance)
- 7. Binbin Chen (ByteDance)
- 8. Yufei Wu (ByteDance)
- 9. Hao Wang (ByteDance)
- 10. Tieying Zhang (ByteDance)
- 11. Rui Shi (ByteDance)
- 12. Xiaofeng Gao (Shanghai Jiao Tong University)
- 13. Yuming Liang (ByteDance)
- 14. Pengwei Zhao (ByteDance)
- 15. Guihai Chen (Shanghai Jiao Tong University)
BibTeX Citation
@article{wu_vldb24,
title = {{Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance}},
author = {Wu, Yixin and Huang, Xiuqi and Wei, Zhongjia and Cheng, Hang and Xin, Chaohui and Chen, Zuzhi and Chen, Binbin and Wu, Yufei and Wang, Hao and Zhang, Tieying and Shi, Rui and Gao, Xiaofeng and Liang, Yuming and Zhao, Pengwei and Chen, Guihai},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {12},
pages = {3759--3771},
doi = {10.14778/3685800.3685804},
url = {https://doi.org/10.14778/3685800.3685804},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,771 | Rockhopper: A Robust Optimizer for Spark Configuration Tuning in Production Environment | 2025 | SIGMOD | 5.2209769e-05 |
| 10,565 | LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration | 2026 | VLDB | 5.093636e-05 |
| 11,006 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning Workloads | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,116 | ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud Databases | 2021 | SIGMOD | 7.7390737e-05 |
| 3,173 | Why You Should Run TPC-DS:A Workload Analysis | 2007 | VLDB | 7.6664516e-05 |
| 3,219 | iBTune: Individualized Buffer Tuning for Large-scale Cloud Databases | 2019 | VLDB | 7.6283153e-05 |
| 4,854 | LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications | 2022 | SIGMOD | 6.4779623e-05 |
| 5,390 | Magnet: Push-based Shuffle Service for Large-scale Data Processing | 2020 | VLDB | 6.2346971e-05 |
| 6,102 | AutoExecutor: Predictive Parallelism for Spark SQL Queries | 2021 | VLDB | 5.976708e-05 |
| 6,344 | Towards General and Efficient Online Tuning for Spark | 2023 | VLDB | 5.9060457e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,854 | SparkCruise: Handsfree Computation Reuse in Spark | 2019 | VLDB |
| 2 | 5,821 | StreamOps: Cloud-Native Runtime Management for Streaming Services in ByteDance | 2023 | VLDB |
| 3 | 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB |
| 4 | 5,865 | Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems | 2019 | VLDB |
| 5 | 11,284 | ResLake: Towards Minimum Job Latency and Balanced Resource Utilization in Geo-distributed Job Scheduling | 2024 | VLDB |
| 6 | 8,175 | SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft | 2021 | VLDB |
| 7 | 11,222 | Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service | 2024 | VLDB |
| 8 | 8,615 | A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning | 2024 | VLDB |
| 9 | 5,390 | Magnet: Push-based Shuffle Service for Large-scale Data Processing | 2020 | VLDB |
| 10 | 6,344 | Towards General and Efficient Online Tuning for Spark | 2023 | VLDB |