Back to papers
Distributed GPU Joins on Fast RDMA-capable Networks
Summary: Pipelined distributed GPU joins on fast RDMA networks overlap shuffling with build/probe to hide GPU idle time. RDMA/GPUDirect-based algorithms scale to arbitrarily large tables and show up to 6x faster full queries versus CPU-only joins.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h7c3e52b7b5656740
Venue
SIGMOD
Year
2023
Pagerank
6.1555648e-05
Overall Rank
5,376 | 63.87%
DOI
10.1145/3588709
Incoming Non-self Citations Over Time
Authors
1.
Lasse Thostrup
(Technical University of Darmstadt)
2.
Gloria Doci
(Snowflake)
3.
Nils Boeschen
(Technical University of Darmstadt)
4.
Manisha Luthra
(German National Research Center for Information Technology; Technical University of Darmstadt)
5.
Carsten Binnig
(German National Research Center for Information Technology; Technical University of Darmstadt)
BibTeX Citation
Copy BibTeX
@inproceedings{thostrup_sigmod23,
title = {{Distributed GPU Joins on Fast RDMA-capable Networks}},
author = {Thostrup, Lasse and Doci, Gloria and Boeschen, Nils and Luthra, Manisha and Binnig, Carsten},
series = {{SIGMOD} '23},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3588709},
url = {https://dl.acm.org/doi/10.1145/3588709},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 14 of 14 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
5,991
Terabyte-Scale Analytics in the Blink of an Eye
2026
VLDB
5.9219616e-05
6,153
Efficiently Processing Joins and Grouped Aggregations on GPUs
2025
SIGMOD
5.8669913e-05
6,671
Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs
2025
VLDB
5.7125893e-05
8,242
Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs
2024
VLDB
5.3684987e-05
8,280
Scaling GPU-Accelerated Databases beyond GPU Memory Size
2025
VLDB
5.3616863e-05
8,760
Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs
2024
SIGMOD
5.2822161e-05
9,438
Efficiently Joining Large Relations on Multi-GPU Systems
2025
VLDB
5.1761941e-05
9,693
DPDPU: Data Processing with DPUs
2025
CIDR
5.1384473e-05
10,632
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management
2026
SIGMOD
4.9769913e-05
10,733
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization
2026
VLDB
4.9769913e-05
10,795
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
2026
VLDB
4.9769913e-05
10,811
PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
2026
VLDB
4.9769913e-05
11,442
Analyzing Near-Network Hardware Acceleration with Co-Processing on DPUs
2025
VLDB
4.9769913e-05
11,542
Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality
2024
SIGMOD
4.9769913e-05
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
251
Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited
2014
VLDB
0.00023136934
616
Relational Joins on Graphics Processors
2008
SIGMOD
0.00015554627
663
Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort
2010
SIGMOD
0.00014997516
775
The Yin and Yang of Processing Data Warehousing Queries on GPU Devices
2013
VLDB
0.0001407924
890
Rack-Scale In-Memory Join Processing using RDMA
2015
SIGMOD
0.00013241413
945
The End of Slow Networks: It's Time for a Redesign
2016
VLDB
0.00012933247
1,269
A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics
2020
SIGMOD
0.00011254742
1,339
The End of a Myth: Distributed Transactions Can Scale
2017
VLDB
0.0001097043
1,366
Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks
2019
SIGMOD
0.00010908555
1,534
Pipelined Query Processing in Coprocessor Environments
2018
SIGMOD
0.00010327147
1,949
Quantifying TPC-H Choke Points and Their Optimizations
2020
VLDB
9.3172855e-05
2,124
Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture
2013
VLDB
9.0041425e-05
2,287
Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects
2020
SIGMOD
8.691301e-05
2,601
Robust Query Processing in Co-Processor-accelerated Databases
2016
SIGMOD
8.2325291e-05
2,802
Distributed Join Algorithms on Thousands of Cores
2017
VLDB
7.9866934e-05
2,887
Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment
2021
VLDB
7.9044174e-05
3,055
MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures
2021
SIGMOD
7.6983202e-05
3,202
Why it is time for a HyPE: A Hybrid Query Processing Engine for Efficient GPU Coprocessing in DBMS
2013
VLDB
7.5401819e-05
3,609
DFI: The Data Flow Interface for High-Speed Networks
2021
SIGMOD
7.1617145e-05
3,628
Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects
2022
SIGMOD
7.1490678e-05
Semantically Similar Papers