Database Paper Browser

Back to papers

A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning

Summary: Controls all per-query tunable parameters in Spark’s Adaptive Query Execution via a hybrid compile-time/runtime, multi-granularity scheme to handle diverse, correlated parameters. Poses tuning as multi-objective (latency vs cost), with models/solvers achieving sub-second solves and much better latency/cost tradeoffs (63–65% latency reduction vs 18–25% for prior MOO methods) and superior adaptability to user preferences. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13565
Venue
VLDB
Year
2024
Pagerank
4.4856045e-05
Overall Rank
8,585 | 40.34%
DOI
10.14778/3681954.3682021

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 39 of 39 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
66 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00061707583
70 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00059744625
183 Automatic Database Management System Tuning Through Large-scale Machine Learning 2017 SIGMOD 0.00036859633
510 An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement Learning 2019 SIGMOD 0.00021420477
539 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00020615453
776 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00016765694
779 QTune: A Query-Aware Database Tuning System with Deep Reinforcement Learning 2019 VLDB 0.00016719473
1,644 Parametric Query Optimization for Linear and Piecewise Linear Cost Functions 2002 VLDB 0.0001102889
1,678 Fuxi: a Fault-Tolerant Resource Management and Job Scheduling System at Internet Scale 2014 VLDB 0.00010928011
1,892 Black or White? How to Develop an AutoTuner for Memory-based Analytics 2020 SIGMOD 0.00010176219
2,364 Deep Learning Models for Selectivity Estimation of Multi-Attribute Queries 2020 SIGMOD 8.955077e-05
2,553 Towards Cost-Optimal Query Processing in the Cloud 2021 VLDB 8.5522099e-05
2,652 Multi-Objective Parametric Query Optimization 2015 VLDB 8.3662031e-05
2,693 Greenplum: A Hybrid Database for Transactional and Analytical Workloads 2021 SIGMOD 8.2845883e-05
2,769 FLAT: Fast, Lightweight and Accurate Method for Cardinality Estimation 2021 VLDB 8.1512848e-05
2,781 Flow-Loss: Learning Cardinality Estimates That Matter 2021 VLDB 8.1282042e-05
3,222 WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases 2016 VLDB 7.3531422e-05
3,492 Fauce: Fast and Accurate Deep Ensembles with Uncertainty for Cardinality Estimation 2021 VLDB 7.0435484e-05
3,924 A Unified Deep Model of Learning from both Data and Queries for Cardinality Estimation 2021 SIGMOD 6.6227223e-05
3,995 ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud Databases 2021 SIGMOD 6.5475871e-05
4,543 FACE: A Normalizing Flow based Cardinality Estimator 2022 VLDB 6.0953507e-05
4,698 Schedule Optimization for Data Processing Flows on the Cloud 2011 SIGMOD 5.9835195e-05
4,730 UDO: Universal Database Optimization using Reinforcement Learning 2021 VLDB 5.9604983e-05
4,799 Towards Dynamic and Safe Configuration Tuning for Cloud Databases 2022 SIGMOD 5.9082876e-05
4,876 Approximation Schemes for Many-Objective Query Optimization 2014 SIGMOD 5.8544467e-05
5,073 An Incremental Anytime Algorithm for Multi-Objective Query Optimization 2015 SIGMOD 5.7118738e-05
5,318 LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications 2022 SIGMOD 5.5685434e-05
5,373 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing 2022 VLDB 5.5410059e-05
5,477 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 5.4856699e-05
5,537 Presto: A Decade of SQL Analytics at Meta 2023 SIGMOD 5.453017e-05
5,643 Intelligent Scaling in Amazon Redshift 2024 SIGMOD 5.3949759e-05
6,365 Pre-training Summarization Models of Structured Datasets for Cardinality Estimation 2022 VLDB 5.0892829e-05
6,489 Towards General and Efficient Online Tuning for Spark 2023 VLDB 5.0373773e-05
7,340 Weighted Distinct Sampling: Cardinality Estimation for SPJ Queries 2021 SIGMOD 4.7526052e-05
8,573 PostCENN: PostgreSQL with Machine Learning Models for Cardinality Estimation 2021 VLDB 4.4885745e-05
9,064 Tempo: Robust and Self-Tuning Resource Management in Multi-tenant Parallel Databases 2016 VLDB 4.3994049e-05
9,547 Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms 2021 VLDB 4.3219254e-05
9,548 Optimistic Recovery for Iterative Dataflows in Action 2015 SIGMOD 4.3219254e-05
9,735 UDAO: A Next-Generation Unified Data Analytics Optimizer 2019 VLDB 4.2901665e-05
Previous Page 1 / 1 Next

Semantically Similar Papers