DBScholar

Back to papers

Efficiently Joining Large Relations on Multi-GPU Systems

Summary: Introduces a heterogeneous, out-of-core multi-GPU sort-merge join exploiting NVLink/NVSwitch peer-to-peer transfers, CPU multiway merging, and hybrid CPU/GPU joining. Scales across GPUs and outperforms CPU and non-P2P GPU joins by up to 15.2× and 8.7×, respectively. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hffc43349023f178e
Venue
VLDB
Year
2025
Pagerank
5.1786456e-05
Overall Rank
9,429 | 36.61%
DOI
10.14778/3749646.3749720

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{maltenberger_vldb25,
        title = {{Efficiently Joining Large Relations on Multi-GPU Systems}},
        author = {Maltenberger, Tobias and Tolovski, Ilin and Rabl, Tilmann},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {11},
        pages = {4653--4667},
        doi = {10.14778/3749646.3749720},
        url = {https://doi.org/10.14778/3749646.3749720},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Rank Citing Paper Year Venue Pagerank
5,991 Terabyte-Scale Analytics in the Blink of an Eye 2026 VLDB 5.9247663e-05
10,465 GraphRTX: Lighting the Way to Scalable Graph Analytics 2026 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
106 Quickly Generating Billion-Record Synthetic Databases 1994 SIGMOD 0.00033526937
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024851502
233 Amazon Redshift and the Case for Simpler Data Warehouses 2015 SIGMOD 0.00023783585
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
361 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00020006406
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018491327
616 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015561564
627 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015460957
661 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015003815
722 Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture 2008 VLDB 0.00014488003
724 Fundamental Techniques for Order Optimization 1996 SIGMOD 0.00014477566
771 The Yin and Yang of Processing Data Warehousing Queries on GPU Devices 2013 VLDB 0.00014085862
969 Memory-Efficient Hash Joins 2015 VLDB 0.0001278184
1,203 Fixed-Precision Estimation of Join Selectivity 1993 PODS 0.00011548537
1,266 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011269175
1,268 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001126007
1,467 On the Relative Cost of Sampling for Join Selectivity Estimation 1994 PODS 0.00010567959
1,533 Pipelined Query Processing in Coprocessor Environments 2018 SIGMOD 0.00010332035
1,995 Track Join: Distributed Joins with Minimal Network Traffic 2014 SIGMOD 9.2169073e-05
2,284 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.6954168e-05
2,487 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.3979719e-05
2,802 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9903139e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9739791e-05
2,886 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.9081605e-05
3,053 MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures 2021 SIGMOD 7.7019663e-05
3,626 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 7.1524537e-05
4,225 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.7207703e-05
4,931 FPGA-based Data Partitioning 2017 SIGMOD 6.348544e-05
5,371 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1584802e-05
5,628 FPGA-based Multithreading for In-Memory Hash Joins 2015 CIDR 6.060354e-05
6,151 Efficiently Processing Joins and Grouped Aggregations on GPUs 2025 SIGMOD 5.8697699e-05
6,414 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7927521e-05
Previous Page 1 / 1 Next

Semantically Similar Papers