TPCx-AI under the Microscope: A Benchmarking Debt Analysis
Summary: Dissects TPCx-AI for "benchmarking debt": kit/spec divergences, data errors, weak metrics, and workload artifacts that distort what is actually being measured. Shows these issues can skew training/serving by 350x/800x and that fixing them yields up to 3.8x higher end-to-end throughput. (summarized by gpt-5.4-mini on Apr 12 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Ilin Tolovski
- 2. Philipp Hildebrandt
- 3. Khuzaima Daudjee
- 4. Tilmann Rabl
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,914 | Quantifying TPC-H Choke Points and Their Optimizations | 2020 | VLDB | 7.9197583e-05 |
| 5,575 | Optimizing Data Pipelines for Machine Learning in Feature Stores | 2023 | VLDB | 5.4253204e-05 |
| 5,615 | TPCx-AI - An Industry Standard Benchmark for Artificial Intelligence and Machine Learning Systems | 2023 | VLDB | 5.409e-05 |
| 6,377 | Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine | 2025 | SIGMOD | 5.0860948e-05 |
| 6,898 | Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine | 2025 | VLDB | 4.8878659e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,645 | PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking | 2025 | VLDB | 4.3067693e-05 |
| 339 | OLTP-Bench: An Extensible Testbed for Benchmarking Relational Databases | 2014 | VLDB | 0.00026895683 |
| 14,267 | Database System Performance Measurement | 1986 | SIGMOD | - |
| 4,441 | Generating Databases for Query Workloads | 2010 | VLDB | 6.1795515e-05 |
| 3,800 | Why You Should Run TPC-DS: A Workload Analysis | 2007 | VLDB | 6.7529849e-05 |
| 9,372 | FEBench: A Benchmark for Real-Time Relational Data Feature Extraction | 2023 | VLDB | 4.3461481e-05 |
| 3,260 | Query Processing on Tensor Computation Runtimes | 2022 | VLDB | 7.3091312e-05 |
| 3,131 | Why TPC Is Not Enough: An Analysis of the Amazon Redshift Fleet | 2024 | VLDB | 7.5054309e-05 |
| 8,621 | A Study of Database Performance Sensitivity to Experiment Settings | 2022 | VLDB | 4.4787515e-05 |
| 5,615 | TPCx-AI - An Industry Standard Benchmark for Artificial Intelligence and Machine Learning Systems | 2023 | VLDB | 5.409e-05 |