DBScholar

Back to papers

Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects

Summary: Triton Join scales large joins on GPUs by using fast interconnects (NVLink 2.0) to spill state to main memory. Delivers >100x GPU hash-join gains over non-partitioned approaches, and up to 2.5x vs CPU radix-join, enabling GPU DBMSs to exceed GPU memory. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6425
Venue
SIGMOD
Year
2022
Pagerank
6.6419266e-05
Overall Rank
4,535 | 68.89%
DOI
10.1145/3514221.3517911

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{lutz_sigmod22,
        title = {{Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects}},
        author = {Lutz, Clemens and Breß, Sebastian and Zeuch, Steffen and Rabl, Tilmann and Markl, Volker},
        series = {{SIGMOD} '22},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3514221.3517911},
        url = {https://dl.acm.org/doi/10.1145/3514221.3517911},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 15 of 15 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 41 of 41 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
7 Implementation Techniques For Main Memory Database Systems 1984 SIGMOD 0.00083340894
74 Cache Conscious Algorithms for Relational Query Processing 1994 VLDB 0.00037330605
209 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024932174
252 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023242719
631 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015591241
634 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015533814
678 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015061068
823 The Yin and Yang of Processing Data Warehousing Queries on GPU Devices 2013 VLDB 0.00013792901
892 Rack-Scale In-Memory Join Processing using RDMA 2015 SIGMOD 0.00013376761
912 Hardware-Oblivious Parallelism for In-Memory Column-Stores 2013 VLDB 0.00013269804
959 Memory-Efficient Hash Joins 2015 VLDB 0.00012953588
1,265 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011415709
1,278 A Seven-Dimensional Analysis of Hashing Methods and its Implications on Query Processing 2016 VLDB 0.00011362007
1,466 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001068941
1,574 Pipelined Query Processing in Coprocessor Environments 2018 SIGMOD 0.00010321274
1,977 HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines 2019 VLDB 9.3641101e-05
2,140 Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture 2013 VLDB 9.0991487e-05
2,566 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.4116562e-05
2,667 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.2756346e-05
2,713 Robust Query Processing in Co-Processor-accelerated Databases 2016 SIGMOD 8.2122726e-05
2,848 GPL: A GPU-based Pipelined Query Processing Engine 2016 SIGMOD 8.0538815e-05
2,926 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9549783e-05
2,962 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9170451e-05
3,134 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.7231028e-05
3,227 MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures 2021 SIGMOD 7.6217889e-05
3,406 Cache-Conscious Radix-Decluster Projections 2004 VLDB 7.4392655e-05
3,435 Improving Main Memory Hash Joins on Intel Xeon Phi Processors: An Experimental Approach 2015 VLDB 7.4172582e-05
3,471 In-Cache Query Co-Processing on Coupled CPU-GPU Architectures 2015 VLDB 7.3885861e-05
3,682 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.2035518e-05
3,791 Hardware-conscious Query Processing in GPU-accelerated Analytical Engines 2019 CIDR 7.1235328e-05
4,603 The Art of Balance: A RateupDB Experience of Building a CPU/GPU Hybrid Database Product 2021 VLDB 6.6105578e-05
4,840 FPGA-based Data Partitioning 2017 SIGMOD 6.483442e-05
5,094 On the Surprising Difficulty of Simple Things: the Case of Radix Partitioning 2015 VLDB 6.3657592e-05
5,357 Joins on Encoded and Partitioned Data 2014 VLDB 6.2497031e-05
5,480 EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal in GPUs 2021 VLDB 6.2040237e-05
5,515 FPGA-based Multithreading for In-Memory Hash Joins 2015 CIDR 6.189963e-05
5,812 MCJoin: A Memory-Constrained Join for Column-Store Main-Memory Databases. 2012 SIGMOD 6.0782357e-05
6,161 Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities 2019 CIDR 5.9537202e-05
6,774 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7778738e-05
6,835 GPU-accelerated data management under the test of time 2020 CIDR 5.7584924e-05
10,004 A four-dimensional Analysis of Partitioned Approximate Filters 2021 VLDB 5.1810297e-05
Previous Page 1 / 1 Next

Semantically Similar Papers