New Query Optimization Techniques in the Spark Engine of Azure Synapse
Summary: Azure Synapse Spark introduces exchange placement that jointly minimizes exchanges and maximizes multi-consumer reuse, alongside aggressive partial pushdowns for aggregates, joins, and intersections. Stateful-operator specialization delivers a 1.8× TPC-DS speedup over Apache Spark 3.0.1. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Abhishek Modi (Microsoft)
- 2. Kaushik Rajan (Microsoft)
- 3. Srinivas Thimmaiah (Microsoft)
- 4. Prakhar Jain (Databricks)
- 5. Swinky Mann (Microsoft)
- 6. Ayushi Agarwal (Microsoft)
- 7. Ajith Shetty (Microsoft)
- 8. Shahid K I (Microsoft)
- 9. Ashit Gosalia (Microsoft)
- 10. Partho Sarthi (University of Wisconsin)
BibTeX Citation
@article{modi_vldb22,
title = {{New Query Optimization Techniques in the Spark Engine of Azure Synapse}},
author = {Modi, Abhishek and Rajan, Kaushik and Thimmaiah, Srinivas and Jain, Prakhar and Mann, Swinky and Agarwal, Ayushi and Shetty, Ajith and I, Shahid K and Gosalia, Ashit and Sarthi, Partho},
journal = {PVLDB},
series = {{VLDB} '22},
volume = {15},
number = {4},
pages = {936--948},
doi = {10.14778/3503585.3503601},
url = {https://doi.org/10.14778/3503585.3503601},
year = {2022}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,363 | GenRewrite: Query Rewriting via Large Language Models | 2026 | SIGMOD | 6.7423909e-05 |
| 10,409 | TQEx: Tensor-based Query Engine Enhanced by Bridging the Gap | 2026 | SIGMOD | 5.093636e-05 |
| 10,985 | Scaling GPU-Accelerated Databases beyond GPU Memory Size | 2025 | VLDB | 5.093636e-05 |
| 11,466 | Anser: Adaptive Information Sharing Framework of AnalyticDB | 2023 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,167 | Dynamic Speculative Optimizations for SQL Compilation in Apache Spark | 2020 | VLDB |
| 2 | 6,102 | AutoExecutor: Predictive Parallelism for Spark SQL Queries | 2021 | VLDB |
| 3 | 11,399 | QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark | 2023 | SIGMOD |
| 4 | 2,078 | Query Optimization in Microsoft SQL Server PDW | 2012 | SIGMOD |
| 5 | 4,481 | Dynamically Optimizing Queries over Large Scale Data Platforms | 2014 | SIGMOD |
| 6 | 8,175 | SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft | 2021 | VLDB |
| 7 | 6,194 | Incorporating Super-Operators in Big-Data Query Optimizers | 2020 | VLDB |
| 8 | 9,446 | Parallelizing Query Optimization on Shared-Nothing Architectures | 2016 | VLDB |
| 9 | 4,742 | Continuous Cloud-Scale Query Optimization and Processing | 2013 | VLDB |
| 10 | 8,615 | A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning | 2024 | VLDB |