Back to papers
Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs
Summary: Jointly scales hybrid CPU–GPU DBMSs across multiple GPUs by optimizing data placement and distributed query execution. Introduces cache-aware replication that accounts for shuffle cost and coordinates caching/replication; Lancelot implements distributed hybrid execution and achieves 2–12× speedups on SSB and TPC‑H.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13716
- Venue
- VLDB
- Year
- 2024
- Pagerank
- 4.432948e-05
- Overall Rank
- 8,846 | 38.53%
- DOI
-
10.14778/3704965.3704977
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 29 of 29 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 145 |
Quickly Generating Billion-Record Synthetic Databases |
1994 |
SIGMOD |
0.00041403894 |
| 185 |
DuckDB: an Embeddable Analytical Database |
2019 |
SIGMOD |
0.00036529607 |
| 771 |
Relational Joins on Graphics Processors |
2008 |
SIGMOD |
0.00016813054 |
| 1,271 |
The Yin and Yang of Processing Data Warehousing Queries on GPU Devices |
2013 |
VLDB |
0.00012900735 |
| 1,285 |
Hardware-Oblivious Parallelism for In-Memory Column-Stores |
2013 |
VLDB |
0.00012809552 |
| 2,044 |
A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics |
2020 |
SIGMOD |
9.6963999e-05 |
| 2,072 |
HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics |
2016 |
VLDB |
9.6300019e-05 |
| 2,292 |
Pipelined Query Processing in Coprocessor Environments |
2018 |
SIGMOD |
9.0884645e-05 |
| 2,523 |
Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture |
2013 |
VLDB |
8.599693e-05 |
| 2,659 |
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines |
2019 |
VLDB |
8.3615158e-05 |
| 3,161 |
A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs |
2017 |
SIGMOD |
7.4648665e-05 |
| 3,260 |
Query Processing on Tensor Computation Runtimes |
2022 |
VLDB |
7.3091312e-05 |
| 3,307 |
Robust Query Processing in Co-Processor-accelerated Databases |
2016 |
SIGMOD |
7.2391191e-05 |
| 3,328 |
Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects |
2020 |
SIGMOD |
7.2136181e-05 |
| 3,698 |
Why it is time for a HyPE: A Hybrid Query Processing Engine for Efficient GPU Coprocessing in DBMS |
2013 |
VLDB |
6.8278944e-05 |
| 3,899 |
Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment |
2021 |
VLDB |
6.6513982e-05 |
| 4,000 |
MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures |
2021 |
SIGMOD |
6.5419402e-05 |
| 4,359 |
Hardware-conscious Query Processing in GPU-accelerated Analytical Engines |
2019 |
CIDR |
6.2493951e-05 |
| 4,678 |
OmniDB: Towards Portable and Efficient Query Processing on Parallel CPU/GPU Architectures |
2013 |
VLDB |
5.9988623e-05 |
| 4,996 |
Adaptive Work Placement for Query Processing on Heterogeneous Computing Resources |
2017 |
VLDB |
5.7697228e-05 |
| 5,018 |
Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS |
2022 |
VLDB |
5.7503878e-05 |
| 5,039 |
Tile-based Lightweight Integer Compression in GPU |
2022 |
SIGMOD |
5.7369993e-05 |
| 5,199 |
Data-Parallel Query Processing on Non-Uniform Data |
2020 |
VLDB |
5.6294232e-05 |
| 5,251 |
Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects |
2022 |
SIGMOD |
5.6003972e-05 |
| 5,826 |
Towards a Hybrid Design for Fast Query Processing in DB2 with BLU Acceleration Using Graphical Processing Units: A Technology Demonstration |
2016 |
SIGMOD |
5.3116062e-05 |
| 6,068 |
GPU Database Systems Characterization and Optimization |
2024 |
VLDB |
5.2240241e-05 |
| 6,220 |
Distributed GPU Joins on Fast RDMA-capable Networks |
2023 |
SIGMOD |
5.1446966e-05 |
| 6,367 |
Improving Execution Efficiency of Just-in-time Compilation based Query Processing on GPUs |
2021 |
VLDB |
5.0887599e-05 |
| 6,861 |
HetCache: Synergising NVMe Storage and GPU acceleration for Memory-Efficient Analytics |
2023 |
CIDR |
4.9005561e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 3,899 |
Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment |
2021 |
VLDB |
6.6513982e-05 |
| 4,359 |
Hardware-conscious Query Processing in GPU-accelerated Analytical Engines |
2019 |
CIDR |
6.2493951e-05 |
| 2,072 |
HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics |
2016 |
VLDB |
9.6300019e-05 |
| 3,698 |
Why it is time for a HyPE: A Hybrid Query Processing Engine for Efficient GPU Coprocessing in DBMS |
2013 |
VLDB |
6.8278944e-05 |
| 7,917 |
Terabyte-Scale Analytics in the Blink of an Eye |
2026 |
VLDB |
4.6129625e-05 |
| 3,307 |
Robust Query Processing in Co-Processor-accelerated Databases |
2016 |
SIGMOD |
7.2391191e-05 |
| 2,336 |
Concurrent Analytical Query Processing with GPUs |
2014 |
VLDB |
9.0106308e-05 |
| 6,068 |
GPU Database Systems Characterization and Optimization |
2024 |
VLDB |
5.2240241e-05 |
| 5,018 |
Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS |
2022 |
VLDB |
5.7503878e-05 |
| 10,755 |
Scaling GPU-Accelerated Databases beyond GPU Memory Size |
2025 |
VLDB |
4.1905499e-05 |