Distributed GPU Joins on Fast RDMA-capable Networks
Summary: Pipelined distributed GPU joins on fast RDMA networks overlap shuffling with build/probe to hide GPU idle time. RDMA/GPUDirect-based algorithms scale to arbitrarily large tables and show up to 6x faster full queries versus CPU-only joins. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Lasse Thostrup (Technical University of Darmstadt)
- 2. Gloria Doci (Snowflake)
- 3. Nils Boeschen (Technical University of Darmstadt)
- 4. Manisha Luthra (German National Research Center for Information Technology; Technical University of Darmstadt)
- 5. Carsten Binnig (German National Research Center for Information Technology; Technical University of Darmstadt)
BibTeX Citation
@inproceedings{thostrup_sigmod23,
title = {{Distributed GPU Joins on Fast RDMA-capable Networks}},
author = {Thostrup, Lasse and Doci, Gloria and Boeschen, Nils and Luthra, Manisha and Binnig, Carsten},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3588709},
url = {https://dl.acm.org/doi/10.1145/3588709},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,011 | Design and Evaluation of Parallel Pipelined Join Algorithms | 1987 | SIGMOD |
| 2 | 2,140 | Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture | 2013 | VLDB |
| 3 | 6,835 | GPU-accelerated data management under the test of time | 2020 | CIDR |
| 4 | 2,566 | Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects | 2020 | SIGMOD |
| 5 | 4,535 | Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects | 2022 | SIGMOD |
| 6 | 631 | Relational Joins on Graphics Processors | 2008 | SIGMOD |
| 7 | 892 | Rack-Scale In-Memory Join Processing using RDMA | 2015 | SIGMOD |
| 8 | 7,591 | Efficiently Processing Joins and Grouped Aggregations on GPUs | 2025 | SIGMOD |
| 9 | 9,333 | Efficiently Joining Large Relations on Multi-GPU Systems | 2025 | VLDB |
| 10 | 3,134 | Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment | 2021 | VLDB |