Back to papers
SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft
Summary: SparkCruise injects a workload-driven feedback loop into the Spark SQL optimizer to optimize large workloads without accessing user data. Analysis of production Spark SQL workloads vs. TPC-DS demonstrates online learning and a computation-reuse optimization.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 12518
- Venue
- VLDB
- Year
- 2021
- Pagerank
- 4.5568952e-05
- Overall Rank
- 8,196 | 43.04%
- DOI
-
10.14778/3476311.3476388
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 22 |
SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets |
2008 |
VLDB |
0.00084679526 |
| 66 |
Spark SQL: Relational Data Processing in Spark |
2015 |
SIGMOD |
0.00061707583 |
| 304 |
Generic Schema Matching with Cupid |
2001 |
VLDB |
0.00028282278 |
| 539 |
Shark: SQL and Rich Analytics at Scale |
2013 |
SIGMOD |
0.00020615453 |
| 796 |
SageDB: A Learned Database System |
2019 |
CIDR |
0.00016541749 |
| 1,921 |
Selecting Subexpressions to Materialize at Datacenter Scale |
2018 |
VLDB |
0.00010085899 |
| 2,080 |
Towards a Learning Optimizer for Shared Clouds |
2019 |
VLDB |
9.5954034e-05 |
| 2,126 |
IDEBench: A Benchmark for Interactive Data Exploration |
2020 |
SIGMOD |
9.4814404e-05 |
| 3,623 |
Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings |
2020 |
SIGMOD |
6.9017341e-05 |
| 3,750 |
DIAMetrics: Benchmarking Query Engines at Scale |
2020 |
VLDB |
6.7864305e-05 |
| 4,171 |
Computation Reuse in Analytics Job Service at Microsoft |
2018 |
SIGMOD |
6.3800823e-05 |
| 5,994 |
Steering Query Optimizers: A Practical Take on Big Data Workloads |
2021 |
SIGMOD |
5.2367998e-05 |
| 9,734 |
SparkCruise: Handsfree Computation Reuse in Spark |
2019 |
VLDB |
4.2901665e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 11,014 |
Intelligent Pooling: Proactive Resource Provisioning in Large-scale Cloud Service |
2024 |
VLDB |
4.1905499e-05 |
| 9,643 |
Rockhopper: A Robust Optimizer for Spark Configuration Tuning in Production Environment |
2025 |
SIGMOD |
4.3067693e-05 |
| 8,454 |
Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance |
2024 |
VLDB |
4.5022073e-05 |
| 3,536 |
Scaling Spark in the Real World: Performance and Usability |
2015 |
VLDB |
6.9938207e-05 |
| 9,122 |
Dynamic Speculative Optimizations for SQL Compilation in Apache Spark |
2020 |
VLDB |
4.3877539e-05 |
| 6,489 |
Towards General and Efficient Online Tuning for Spark |
2023 |
VLDB |
5.0373773e-05 |
| 5,309 |
Continuous Cloud-Scale Query Optimization and Processing |
2013 |
VLDB |
5.5714729e-05 |
| 8,502 |
New Query Optimization Techniques in the Spark Engine of Azure Synapse |
2022 |
VLDB |
4.491819e-05 |
| 8,585 |
A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning |
2024 |
VLDB |
4.4856045e-05 |
| 9,734 |
SparkCruise: Handsfree Computation Reuse in Spark |
2019 |
VLDB |
4.2901665e-05 |