Back to papers
Distributed GPU Joins on Fast RDMA-capable Networks
Summary: Pipelined distributed GPU joins on fast RDMA networks overlap shuffling with build/probe to hide GPU idle time. RDMA/GPUDirect-based algorithms scale to arbitrarily large tables and show up to 6x faster full queries versus CPU-only joins.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6533
- Venue
- SIGMOD
- Year
- 2023
- Pagerank
- 5.1446966e-05
- Overall Rank
- 6,220 | 56.78%
- DOI
-
10.1145/3588709
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 7,570 |
Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs |
2025 |
VLDB |
4.7039167e-05 |
| 7,752 |
Efficiently Processing Joins and Grouped Aggregations on GPUs |
2025 |
SIGMOD |
4.6558737e-05 |
| 7,917 |
Terabyte-Scale Analytics in the Blink of an Eye |
2026 |
VLDB |
4.6129625e-05 |
| 8,647 |
Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs |
2024 |
SIGMOD |
4.4720005e-05 |
| 8,846 |
Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs |
2024 |
VLDB |
4.432948e-05 |
| 9,458 |
DPDPU: Data Processing with DPUs |
2025 |
CIDR |
4.3344018e-05 |
| 9,837 |
Efficiently Joining Large Relations on Multi-GPU Systems |
2025 |
VLDB |
4.269939e-05 |
| 10,143 |
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,253 |
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization |
2026 |
VLDB |
4.1905499e-05 |
| 10,755 |
Scaling GPU-Accelerated Databases beyond GPU Memory Size |
2025 |
VLDB |
4.1905499e-05 |
| 10,860 |
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs |
2025 |
VLDB |
4.1905499e-05 |
| 10,984 |
Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality |
2024 |
SIGMOD |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 403 |
Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited |
2014 |
VLDB |
0.00024176677 |
| 771 |
Relational Joins on Graphics Processors |
2008 |
SIGMOD |
0.00016813054 |
| 932 |
Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort |
2010 |
SIGMOD |
0.00015227954 |
| 1,196 |
Rack-Scale In-Memory Join Processing using RDMA |
2015 |
SIGMOD |
0.00013378379 |
| 1,271 |
The Yin and Yang of Processing Data Warehousing Queries on GPU Devices |
2013 |
VLDB |
0.00012900735 |
| 1,351 |
The End of Slow Networks: It's Time for a Redesign |
2016 |
VLDB |
0.00012439556 |
| 1,813 |
The End of a Myth: Distributed Transactions Can Scale |
2017 |
VLDB |
0.00010452333 |
| 1,834 |
Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks |
2019 |
SIGMOD |
0.00010364718 |
| 2,044 |
A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics |
2020 |
SIGMOD |
9.6963999e-05 |
| 2,292 |
Pipelined Query Processing in Coprocessor Environments |
2018 |
SIGMOD |
9.0884645e-05 |
| 2,523 |
Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture |
2013 |
VLDB |
8.599693e-05 |
| 2,914 |
Quantifying TPC-H Choke Points and Their Optimizations |
2020 |
VLDB |
7.9197583e-05 |
| 3,307 |
Robust Query Processing in Co-Processor-accelerated Databases |
2016 |
SIGMOD |
7.2391191e-05 |
| 3,328 |
Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects |
2020 |
SIGMOD |
7.2136181e-05 |
| 3,428 |
Distributed Join Algorithms on Thousands of Cores |
2017 |
VLDB |
7.1002401e-05 |
| 3,698 |
Why it is time for a HyPE: A Hybrid Query Processing Engine for Efficient GPU Coprocessing in DBMS |
2013 |
VLDB |
6.8278944e-05 |
| 3,899 |
Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment |
2021 |
VLDB |
6.6513982e-05 |
| 4,000 |
MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures |
2021 |
SIGMOD |
6.5419402e-05 |
| 4,449 |
DFI: The Data Flow Interface for High-Speed Networks |
2021 |
SIGMOD |
6.1740574e-05 |
| 5,251 |
Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects |
2022 |
SIGMOD |
5.6003972e-05 |
Semantically Similar Papers