Back to papers
Distributed GPU Joins on Fast RDMA-capable Networks
Summary: Pipelined distributed GPU joins on fast RDMA networks overlap shuffling with build/probe to hide GPU idle time. RDMA/GPUDirect-based algorithms scale to arbitrarily large tables and show up to 6x faster full queries versus CPU-only joins.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h7c3e52b7b5656740
Venue
SIGMOD
Year
2023
Pagerank
6.1584802e-05
Overall Rank
5,371 | 63.89%
DOI
10.1145/3588709
Incoming Non-self Citations Over Time
Authors
1.
Lasse Thostrup
(Technical University of Darmstadt)
2.
Gloria Doci
(Snowflake)
3.
Nils Boeschen
(Technical University of Darmstadt)
4.
Manisha Luthra
(German National Research Center for Information Technology; Technical University of Darmstadt)
5.
Carsten Binnig
(German National Research Center for Information Technology; Technical University of Darmstadt)
BibTeX Citation
Copy BibTeX
@inproceedings{thostrup_sigmod23,
title = {{Distributed GPU Joins on Fast RDMA-capable Networks}},
author = {Thostrup, Lasse and Doci, Gloria and Boeschen, Nils and Luthra, Manisha and Binnig, Carsten},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3588709},
url = {https://dl.acm.org/doi/10.1145/3588709},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 14 of 14 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
5,991
Terabyte-Scale Analytics in the Blink of an Eye
2026
VLDB
5.9247663e-05
6,151
Efficiently Processing Joins and Grouped Aggregations on GPUs
2025
SIGMOD
5.8697699e-05
6,667
Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs
2025
VLDB
5.7152948e-05
8,236
Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs
2024
VLDB
5.3710413e-05
8,274
Scaling GPU-Accelerated Databases beyond GPU Memory Size
2025
VLDB
5.3642256e-05
8,752
Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs
2024
SIGMOD
5.2847178e-05
9,429
Efficiently Joining Large Relations on Multi-GPU Systems
2025
VLDB
5.1786456e-05
9,687
DPDPU: Data Processing with DPUs
2025
CIDR
5.1408809e-05
10,621
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
2026
SIGMOD
4.9793485e-05
10,723
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization
2026
VLDB
4.9793485e-05
10,785
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
2026
VLDB
4.9793485e-05
10,801
PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
2026
VLDB
4.9793485e-05
11,436
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs
2025
VLDB
4.9793485e-05
11,536
Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality
2024
SIGMOD
4.9793485e-05
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
251
Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited
2014
VLDB
0.00023143736
616
Relational Joins on Graphics Processors
2008
SIGMOD
0.00015561564
661
Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort
2010
SIGMOD
0.00015003815
771
The Yin and Yang of Processing Data Warehousing Queries on GPU Devices
2013
VLDB
0.00014085862
889
Rack-Scale In-Memory Join Processing using RDMA
2015
SIGMOD
0.00013247362
945
The End of Slow Networks: It's Time for a Redesign
2016
VLDB
0.00012939225
1,268
A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics
2020
SIGMOD
0.0001126007
1,339
The End of a Myth: Distributed Transactions Can Scale
2017
VLDB
0.00010975458
1,366
Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks
2019
SIGMOD
0.00010912114
1,533
Pipelined Query Processing in Coprocessor Environments
2018
SIGMOD
0.00010332035
1,952
Quantifying TPC-H Choke Points and Their Optimizations
2020
VLDB
9.3189525e-05
2,122
Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture
2013
VLDB
9.0084047e-05
2,284
Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects
2020
SIGMOD
8.6954168e-05
2,600
Robust Query Processing in Co-Processor-accelerated Databases
2016
SIGMOD
8.2363864e-05
2,802
Distributed Join Algorithms on Thousands of Cores
2017
VLDB
7.9903139e-05
2,886
Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment
2021
VLDB
7.9081605e-05
3,053
MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures
2021
SIGMOD
7.7019663e-05
3,200
Why it is time for a HyPE: A Hybrid Query Processing Engine for Efficient GPU Coprocessing in DBMS
2013
VLDB
7.5437496e-05
3,608
DFI: The Data Flow Interface for High-Speed Networks
2021
SIGMOD
7.1650501e-05
3,626
Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects
2022
SIGMOD
7.1524537e-05
Semantically Similar Papers