Back to papers
Adaptive and Robust Query Execution for Lakehouses at Scale
Summary: AQE for lakehouses: use pipeline breakers to collect runtime statistics and reoptimize plans, mitigating missing/incorrect table/column stats and bad cardinality/UDF estimates. Up to 25× TPC‑DS speedup; deployed at Databricks for exabyte‑scale workloads to reduce data movement, spills, and memory pressure.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13597
- Venue
- VLDB
- Year
- 2024
- Pagerank
- 4.9593505e-05
- Overall Rank
- 6,683 | 53.56%
- DOI
-
10.14778/3685800.3685818
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 10 of 10 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 6,898 |
Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine |
2025 |
VLDB |
4.8878659e-05 |
| 9,090 |
Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads |
2025 |
SIGMOD |
4.3939337e-05 |
| 9,746 |
Still Asking: How Good Are Query Optimizers, Really? |
2025 |
VLDB |
4.2856385e-05 |
| 9,765 |
The UDFBench Benchmark for General-purpose UDF Queries |
2025 |
VLDB |
4.2815042e-05 |
| 10,241 |
Robust Predicate Transfer with Dynamic Execution |
2026 |
VLDB |
4.1905499e-05 |
| 10,278 |
LakeHelm: Zero-Shot Lakehouse Advisor for Joint Engine-Format Selection and Configuration |
2026 |
VLDB |
4.1905499e-05 |
| 10,773 |
The HANA Native Query Engine for Lakehouse Systems |
2025 |
VLDB |
4.1905499e-05 |
| 10,778 |
veDB-HTAP: a Highly Integrated, Efficient and Adaptive HTAP System |
2025 |
VLDB |
4.1905499e-05 |
| 10,863 |
Graph Transformers for Query Plan Representation: Potentials and Challenges |
2025 |
VLDB |
4.1905499e-05 |
| 13,110 |
Blink Twice - Automatic Workload Pinning and Regression Detection for Versionless Apache Spark using Retries |
2025 |
SIGMOD |
- |
Outgoing Citations (Sorted by Pagerank)
Showing 24 of 24 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 1 |
Access Path Selection in a Relational Database Management System |
1979 |
SIGMOD |
0.0040465394 |
| 66 |
Spark SQL: Relational Data Processing in Spark |
2015 |
SIGMOD |
0.00061707583 |
| 116 |
Eddies: Continuously Adaptive Query Processing |
2000 |
SIGMOD |
0.00046191288 |
| 167 |
The Snowflake Elastic Data Warehouse |
2016 |
SIGMOD |
0.00039408116 |
| 181 |
LEO - DB2's LEarning Optimizer |
2001 |
VLDB |
0.00036970794 |
| 221 |
Efficient Mid-Query Re-Optimization of Sub-Optimal Query Execution Plans |
1998 |
SIGMOD |
0.00033182072 |
| 310 |
The Vertica Analytic Database: C-Store 7 Years Later |
2012 |
VLDB |
0.0002815547 |
| 424 |
Amazon Redshift and the Case for Simpler Data Warehouses |
2015 |
SIGMOD |
0.00023604384 |
| 476 |
Impala: A Modern, Open-Source SQL Engine for Hadoop |
2015 |
CIDR |
0.00022216002 |
| 509 |
Dynamic Query Evaluation Plans |
1989 |
SIGMOD |
0.00021463676 |
| 518 |
An Overview of The System Software of A Parallel Relational Database Machine GRACE |
1986 |
VLDB |
0.00021161322 |
| 539 |
Shark: SQL and Rich Analytics at Scale |
2013 |
SIGMOD |
0.00020615453 |
| 650 |
Robust Query Processing through Progressive Optimization |
2004 |
SIGMOD |
0.0001865144 |
| 739 |
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores |
2020 |
VLDB |
0.00017365933 |
| 1,356 |
Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics |
2021 |
CIDR |
0.00012409986 |
| 1,574 |
Approximate Query Processing: No Silver Bullet |
2017 |
SIGMOD |
0.00011289028 |
| 1,790 |
Effective Use of Block-Level Sampling in Statistics Estimation |
2004 |
SIGMOD |
0.00010529479 |
| 1,919 |
Handling Data Skew in Parallel Joins in Shared-Nothing Systems |
2008 |
SIGMOD |
0.00010097452 |
| 2,472 |
Photon: A Fast Query Engine for Lakehouse Systems |
2022 |
SIGMOD |
8.7156826e-05 |
| 2,503 |
Enhanced Subquery Optimizations in Oracle |
2009 |
VLDB |
8.6306662e-05 |
| 2,941 |
Operating System Extensions for the Teradata Parallel VLDB |
2001 |
VLDB |
7.848008e-05 |
| 3,344 |
F1 Query: Declarative Querying at Scale |
2018 |
VLDB |
7.1944106e-05 |
| 5,537 |
Presto: A Decade of SQL Analytics at Meta |
2023 |
SIGMOD |
5.453017e-05 |
| 6,384 |
Proactive Re-optimization with Rio |
2005 |
SIGMOD |
5.0828986e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 13,372 |
Robust Data Transformations |
2015 |
CIDR |
- |
| 3,307 |
Robust Query Processing in Co-Processor-accelerated Databases |
2016 |
SIGMOD |
7.2391191e-05 |
| 5,309 |
Continuous Cloud-Scale Query Optimization and Processing |
2013 |
VLDB |
5.5714729e-05 |
| 8,579 |
Towards Query Optimizer as a Service (QOaaS) in a Unified LakeHouse Ecosystem: Can One QO Rule Them All? |
2025 |
CIDR |
4.4877266e-05 |
| 2,472 |
Photon: A Fast Query Engine for Lakehouse Systems |
2022 |
SIGMOD |
8.7156826e-05 |
| 6,397 |
BigLake: BigQuery’s Evolution toward a Multi-Cloud Lakehouse |
2024 |
SIGMOD |
5.0749432e-05 |
| 8,585 |
A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning |
2024 |
VLDB |
4.4856045e-05 |
| 1,430 |
A Scalable, Predictable Join Operator for Highly Concurrent Data Warehouses |
2009 |
VLDB |
0.0001202506 |
| 10,248 |
Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability |
2026 |
VLDB |
4.1905499e-05 |
| 7,907 |
Petabyte-Scale Row-Level Operations in Data Lakehouses |
2024 |
VLDB |
4.6161532e-05 |