DBScholar

Back to papers

Efficiently Joining Large Relations on Multi-GPU Systems

Summary: Introduces a heterogeneous, out-of-core multi-GPU sort-merge join exploiting NVLink/NVSwitch peer-to-peer transfers, CPU multiway merging, and hybrid CPU/GPU joining. Scales across GPUs and outperforms CPU and non-P2P GPU joins by up to 15.2× and 8.7×, respectively. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
14262
Venue
VLDB
Year
2025
Pagerank
5.2887551e-05
Overall Rank
9,333 | 35.97%
DOI
10.14778/3749646.3749720

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{maltenberger_vldb25,
        title = {{Efficiently Joining Large Relations on Multi-GPU Systems}},
        author = {Maltenberger, Tobias and Tolovski, Ilin and Rabl, Tilmann},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {11},
        pages = {4653--4667},
        doi = {10.14778/3749646.3749720},
        url = {https://doi.org/10.14778/3749646.3749720},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Rank Citing Paper Year Venue Pagerank
7,959 Terabyte-Scale Analytics in the Blink of an Eye 2026 VLDB 5.5181056e-05
10,252 GraphRTX: Lighting the Way to Scalable Graph Analytics 2026 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 32 of 32 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
105 Quickly Generating Billion-Record Synthetic Databases 1994 SIGMOD 0.00033877899
209 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024932174
237 Amazon Redshift and the Case for Simpler Data Warehouses 2015 SIGMOD 0.0002369895
252 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023242719
360 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00020182846
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018725853
631 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015591241
634 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015533814
678 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015061068
712 Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture 2008 VLDB 0.0001468812
744 Fundamental Techniques for Order Optimization 1996 SIGMOD 0.00014411295
823 The Yin and Yang of Processing Data Warehousing Queries on GPU Devices 2013 VLDB 0.00013792901
959 Memory-Efficient Hash Joins 2015 VLDB 0.00012953588
1,186 Fixed-Precision Estimation of Join Selectivity 1993 PODS 0.00011764128
1,265 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011415709
1,440 On the Relative Cost of Sampling for Join Selectivity Estimation 1994 PODS 0.00010778889
1,466 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001068941
1,574 Pipelined Query Processing in Coprocessor Environments 2018 SIGMOD 0.00010321274
1,974 Track Join: Distributed Joins with Minimal Network Traffic 2014 SIGMOD 9.3658402e-05
2,566 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.4116562e-05
2,667 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.2756346e-05
2,926 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9549783e-05
2,962 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9170451e-05
3,134 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.7231028e-05
3,227 MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures 2021 SIGMOD 7.6217889e-05
4,177 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.8499317e-05
4,535 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 6.6419266e-05
4,840 FPGA-based Data Partitioning 2017 SIGMOD 6.483442e-05
5,515 FPGA-based Multithreading for In-Memory Hash Joins 2015 CIDR 6.189963e-05
5,587 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1596139e-05
6,774 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7778738e-05
7,591 Efficiently Processing Joins and Grouped Aggregations on GPUs 2025 SIGMOD 5.5900315e-05
Previous Page 1 / 1 Next

Semantically Similar Papers