DBScholar

Back to papers

A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning

Summary: Controls all per-query tunable parameters in Spark’s Adaptive Query Execution via a hybrid compile-time/runtime, multi-granularity scheme to handle diverse, correlated parameters. Poses tuning as multi-objective (latency vs cost), with models/solvers achieving sub-second solves and much better latency/cost tradeoffs (63–65% latency reduction vs 18–25% for prior MOO methods) and superior adaptability to user preferences. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13752
Venue
VLDB
Year
2024
Pagerank
5.4005602e-05
Overall Rank
8,615 | 40.90%
DOI
10.14778/3681954.3682021

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{lyu_vldb24,
        title = {{A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning}},
        author = {Lyu, Chenghao and Fan, Qi and Guyard, Philippe and Diao, Yanlei},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {11},
        pages = {3565--3579},
        doi = {10.14778/3681954.3682021},
        url = {https://doi.org/10.14778/3681954.3682021},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 39 of 39 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
86 Automatic Database Management System Tuning Through Large-scale Machine Learning 2017 SIGMOD 0.00035316107
334 An End-to-End Automatic Cloud Database Tuning System Using Deep Reinforcement Learning 2019 SIGMOD 0.00020875082
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
498 QTune: A Query-Aware Database Tuning System with Deep Reinforcement Learning 2019 VLDB 0.00017440583
642 Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience 2009 VLDB 0.00015395331
1,321 Parametric Query Optimization for Linear and Piecewise Linear Cost Functions 2002 VLDB 0.00011162369
1,573 Deep Learning Models for Selectivity Estimation of Multi-Attribute Queries 2020 SIGMOD 0.00010328171
1,686 Black or White? How to Develop an AutoTuner for Memory-based Analytics 2020 SIGMOD 0.00010008686
1,876 Flow-Loss: Learning Cardinality Estimates That Matter 2021 VLDB 9.5717543e-05
1,948 Greenplum: A Hybrid Database for Transactional and Analytical Workloads 2021 SIGMOD 9.432395e-05
1,988 FLAT: Fast, Lightweight and Accurate Method for Cardinality Estimation 2021 VLDB 9.3501502e-05
2,331 Fuxi: a Fault-Tolerant Resource Management and Job Scheduling System at Internet Scale 2014 VLDB 8.7430753e-05
2,344 Towards Cost-Optimal Query Processing in the Cloud 2021 VLDB 8.7199754e-05
2,540 Multi-Objective Parametric Query Optimization 2015 VLDB 8.45187e-05
2,620 Fauce: Fast and Accurate Deep Ensembles with Uncertainty for Cardinality Estimation 2021 VLDB 8.3363963e-05
2,958 WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases 2016 VLDB 7.9197796e-05
3,086 A Unified Deep Model of Learning from both Data and Queries for Cardinality Estimation 2021 SIGMOD 7.7708642e-05
3,116 ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud Databases 2021 SIGMOD 7.7390737e-05
3,688 FACE: A Normalizing Flow based Cardinality Estimator 2022 VLDB 7.201795e-05
3,926 UDO: Universal Database Optimization using Reinforcement Learning 2021 VLDB 7.0128068e-05
4,011 Towards Dynamic and Safe Configuration Tuning for Cloud Databases 2022 SIGMOD 6.959982e-05
4,117 Schedule Optimization for Data Processing Flows on the Cloud 2011 SIGMOD 6.8919486e-05
4,617 Learned Cardinality Estimation for Similarity Queries 2021 SIGMOD 6.604437e-05
4,705 Presto: A Decade of SQL Analytics at Meta 2023 SIGMOD 6.5529421e-05
4,854 LOCAT: Low-Overhead Online Configuration Auto-Tuning of Spark SQL Applications 2022 SIGMOD 6.4779623e-05
4,880 An Incremental Anytime Algorithm for Multi-Objective Query Optimization 2015 SIGMOD 6.4656221e-05
5,124 Approximation Schemes for Many-Objective Query Optimization 2014 SIGMOD 6.3567981e-05
5,369 Intelligent Scaling in Amazon Redshift 2024 SIGMOD 6.2437078e-05
5,388 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing 2022 VLDB 6.2362811e-05
5,792 Pre-training Summarization Models of Structured Datasets for Cardinality Estimation 2022 VLDB 6.0871213e-05
6,344 Towards General and Efficient Online Tuning for Spark 2023 VLDB 5.9060457e-05
7,256 Weighted Distinct Sampling: Cardinality Estimation for SPJ Queries 2021 SIGMOD 5.6625146e-05
8,581 PostCENN: PostgreSQL with Machine Learning Models for Cardinality Estimation 2021 VLDB 5.4082749e-05
9,039 Tempo: Robust and Self-Tuning Resource Management in Multi-tenant Parallel Databases 2016 VLDB 5.3267474e-05
9,669 Trident: Task Scheduling over Tiered Storage Systems in Big Data Platforms 2021 VLDB 5.2389261e-05
9,670 Optimistic Recovery for Iterative Dataflows in Action 2015 SIGMOD 5.2389261e-05
9,902 UDAO: A Next-Generation Unified Data Analytics Optimizer 2019 VLDB 5.1997534e-05
Previous Page 1 / 1 Next

Semantically Similar Papers