Blink Twice - Automatic Workload Pinning and Regression Detection for Versionless Apache Spark using Retries
Summary: Decouples Spark client and engine via Spark Connect, enabling versionless Spark and upgrades with zero code changes. Proposes automatic workload pinning and regression detection via retries to remediate failures, keeping workloads always-on. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Justin Breese (Databricks)
- 2. Vijayan Prabhakaran (Databricks)
- 3. Martin Grund (Databricks)
- 4. Stefania Leone (Databricks)
- 5. Amit Shukla (Databricks)
- 6. Michael Armbrust (Databricks)
- 7. Reynold Xin (Databricks)
- 8. Matei Zaharia (Databricks)
- 9. Lennart Kats (Databricks)
- 10. Sung Chiu (Databricks)
- 11. Tatiana Romanova (Databricks)
- 12. Philip Nord (Databricks)
- 13. Mitchell Webster (Databricks)
- 14. Chris Munson (Databricks)
- 15. Bo Pang (Databricks)
- 16. David Ma (Databricks)
BibTeX Citation
@inproceedings{breese_sigmod25,
title = {{Blink Twice - Automatic Workload Pinning and Regression Detection for Versionless Apache Spark using Retries}},
author = {Breese, Justin and Prabhakaran, Vijayan and Grund, Martin and Leone, Stefania and Shukla, Amit and Armbrust, Michael and Xin, Reynold and Zaharia, Matei and Kats, Lennart and Chiu, Sung and Romanova, Tatiana and Nord, Philip and Webster, Mitchell and Munson, Chris and Pang, Bo and Ma, David},
series = {{SIGMOD} '25},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3722212.3725084},
url = {https://dl.acm.org/doi/10.1145/3722212.3725084},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 5,983 | Adaptive and Robust Query Execution for Lakehouses at Scale | 2024 | VLDB | 6.0206841e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,771 | Rockhopper: A Robust Optimizer for Spark Configuration Tuning in Production Environment | 2025 | SIGMOD |
| 2 | 2,594 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD |
| 3 | 6,850 | Bridging the Gap Between HPC and Big Data Frameworks | 2017 | VLDB |
| 4 | 8,175 | SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft | 2021 | VLDB |
| 5 | 11,112 | Off-the-shelf Data Analytics on Serverless | 2024 | CIDR |
| 6 | 9,127 | Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads | 2025 | SIGMOD |
| 7 | 1,350 | Automating Large-Scale Data Quality Verification | 2018 | VLDB |
| 8 | 7,344 | Building the Enterprise Fabric for Big Data with Vertica and Spark Integration | 2016 | SIGMOD |
| 9 | 1,190 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark | 2018 | SIGMOD |
| 10 | 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB |