DBScholar

Back to papers

Why TPC Is Not Enough: An Analysis of the Amazon Redshift Fleet

Summary: Telemetry from 400 Amazon Redshift instances shows cloud data-warehouse workloads diverge from TPC-H/DS: prominent write-heavy pipelines, temporal variability in load and query types, repetitive queries, and heavy-tailed distributions of query/workload properties. Argues benchmarks must broaden beyond pure query-engine throughput to model ingestion, temporal dynamics, repetition and tail behavior, and releases a 3-month query-statistics dataset to seed more realistic benchmark design. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
h9f843dcbec00aa86
Venue
VLDB
Year
2024
Pagerank
9.5349903e-05
Overall Rank
1,834 | 87.68%
DOI
10.14778/3681954.3682031
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{renen_vldb24,
        title = {{Why TPC Is Not Enough: An Analysis of the Amazon Redshift Fleet}},
        author = {van Renen, Alexander and Horn, Dominik and Pfeil, Pascal and Vaidya, Kapil and Dong, Wenjian and Narayanaswamy, Murali and Liu, Zhengchun and Saxena, Gaurav and Kipf, Andreas and Kraska, Tim},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {11},
        pages = {3694--3706},
        doi = {10.14778/3681954.3682031},
        url = {https://doi.org/10.14778/3681954.3682031},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 37 of 37 citing papers.

Rank Citing Paper Year Venue Pagerank
3,900 SQLStorm: Taking Database Benchmarking into the LLM Era 2025 VLDB 6.9320915e-05
5,316 How Good are Learned Cost Models, Really? Insights from Query Optimization Tasks 2025 SIGMOD 6.1827415e-05
5,514 Runtime-Extensible Parsers 2025 CIDR 6.0965873e-05
6,646 Workload Insights From The Snowflake Data Cloud: What Do Production Analytic Queries Really Look Like? 2025 VLDB 5.7214206e-05
7,386 Towards Foundation Database Models 2025 CIDR 5.5360368e-05
7,546 Learned Offline Query Planning via Bayesian Optimization 2025 SIGMOD 5.4966669e-05
7,810 Parachute: Single-Pass Bi-Directional Information Passing 2025 VLDB 5.4477354e-05
7,826 Bespoke OLAP: Synthesizing Workload-Specific One-size-fits-one Database Engines 2026 VLDB 5.4435842e-05
8,043 Demonstrating SQLBarber: Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads 2025 SIGMOD 5.3999511e-05
8,643 PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking 2025 VLDB 5.2956539e-05
9,463 SQLBarber: A System Leveraging Large Language Models to Generate Customized and Realistic SQL Workloads 2026 SIGMOD 5.1712291e-05
9,559 B-Trees Are Back: Engineering Fast and Pageable Node Layouts 2025 SIGMOD 5.1577035e-05
9,635 Low Rank Learning for Offline Query Optimization 2025 SIGMOD 5.1453041e-05
9,690 The HANA Native Query Engine for Lakehouse Systems 2025 VLDB 5.1399285e-05
9,962 Graph Transformers for Query Plan Representation: Potentials and Challenges 2025 VLDB 5.1014161e-05
10,085 LiquidCache: Efficient Pushdown Caching for Cloud-Native Data Analytics 2025 VLDB 5.0806786e-05
10,129 End-to-End Declarative Data Analytics: Co-designing Engines, Interfaces, and Cloud Infrastructure 2026 CIDR 5.0727027e-05
10,152 Improving DBMS Scheduling Decisions with Accurate Performance Prediction on Concurrent Queries 2025 VLDB 5.0691578e-05
10,232 Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability 2026 VLDB 5.0547568e-05
10,248 Redbench: Workload Synthesis From Cloud Traces 2026 VLDB 5.0501823e-05
10,296 This is Going to Sound Crazy, But What If We Used Large Language Models to Boost Automatic Database Tuning Algorithms By Leveraging Prior History? We Will Find Better Configurations More Quickly Than Retraining From Scratch! 2026 SIGMOD 5.0407989e-05
10,362 Survivorship Bias in Industrial Database Workloads 2026 CIDR 4.9769913e-05
10,479 HotHash: Hotness-Aware Consistent Hashing for Cloud Databases 2026 SIGMOD 4.9769913e-05
10,495 NeurBench: A Benchmark Suite for Learned Database Components with Drift Modeling: [Experiments & Analysis] 2026 SIGMOD 4.9769913e-05
10,703 Practical Parameterized Query Optimization via Efficient Plan Reuse and List-wise Ranking 2026 SIGMOD 4.9769913e-05
10,761 Toward Drift-Aware Database Benchmarking 2026 VLDB 4.9769913e-05
10,833 Revisiting Filtered ANN Benchmarks: A Hardness-Controlled Benchmark Generator for Realistic Evaluation 2026 VLDB 4.9769913e-05
10,890 One Pass to Parse Them All: Fused Parallel CSV Processing 2026 VLDB 4.9769913e-05
10,895 Accelerating String-Heavy Queries with LLM Token Tables 2026 VLDB 4.9769913e-05
10,918 Incremental Query Optimizer Statistics in Amazon Redshift 2026 VLDB 4.9769913e-05
10,923 FastCompose: Eliminating Compilation Cold Starts in Query Execution with Composition 2026 VLDB 4.9769913e-05
10,950 Ultron: History-Based Query Optimization at Databricks 2026 VLDB 4.9769913e-05
10,982 Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents 2026 VLDB 4.9769913e-05
11,136 Flux: Unifying Heterogeneous Infrastructure for Alibaba AnalyticDB 2025 SIGMOD 4.9769913e-05
11,435 Sampling-based Predictive Database Buffer Management 2025 VLDB 4.9769913e-05
11,437 Mayura: Exploiting Similarities in Motifs for Temporal Co-Mining 2025 VLDB 4.9769913e-05
11,439 CloudGlide: Deconstructing the Landscape of Cloud-Based Analytics 2025 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 13 of 13 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers