DBScholar

Back to papers

Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects

Summary: Triton Join scales large joins on GPUs by using fast interconnects (NVLink 2.0) to spill state to main memory. Delivers >100x GPU hash-join gains over non-partitioned approaches, and up to 2.5x vs CPU radix-join, enabling GPU DBMSs to exceed GPU memory. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h4f933eee98a5ce53
Venue
SIGMOD
Year
2022
Pagerank
7.1524537e-05
Overall Rank
3,626 | 75.63%
DOI
10.1145/3514221.3517911

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{lutz_sigmod22,
        title = {{Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects}},
        author = {Lutz, Clemens and Breß, Sebastian and Zeuch, Steffen and Rabl, Tilmann and Markl, Volker},
        series = {{SIGMOD} '22},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3514221.3517911},
        url = {https://dl.acm.org/doi/10.1145/3514221.3517911},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 20 of 20 citing papers.

Rank Citing Paper Year Venue Pagerank
4,480 GPU Database Systems Characterization and Optimization 2024 VLDB 6.5807917e-05
5,371 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1584802e-05
5,541 Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics 2025 VLDB 6.0902254e-05
5,991 Terabyte-Scale Analytics in the Blink of an Eye 2026 VLDB 5.9247663e-05
6,151 Efficiently Processing Joins and Grouped Aggregations on GPUs 2025 SIGMOD 5.8697699e-05
6,667 Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs 2025 VLDB 5.7152948e-05
7,097 Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMs 2023 SIGMOD 5.601767e-05
7,814 Analyzing Vectorized Hash Tables Across CPU Architectures 2023 VLDB 5.4482093e-05
8,236 Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs 2024 VLDB 5.3710413e-05
8,595 A Case for Graphics-driven Query Processing 2023 VLDB 5.3062403e-05
9,429 Efficiently Joining Large Relations on Multi-GPU Systems 2025 VLDB 5.1786456e-05
9,687 DPDPU: Data Processing with DPUs 2025 CIDR 5.1408809e-05
10,777 Succinct and Fast Tiny Pointer Hash Tables 2026 VLDB 4.9793485e-05
10,785 MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures 2026 VLDB 4.9793485e-05
10,801 PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage 2026 VLDB 4.9793485e-05
10,830 Bridging the Indexing Gap in Fused GPU Query Engines 2026 VLDB 4.9793485e-05
10,923 TQP++: Bridging ML Compilers and Analytical Query Processing on GPUs 2026 VLDB 4.9793485e-05
11,536 Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality 2024 SIGMOD 4.9793485e-05
11,545 SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMs 2024 SIGMOD 4.9793485e-05
11,567 Accelerating Merkle Patricia Trie with GPU 2024 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 41 of 41 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
7 Implementation Techniques For Main Memory Database Systems 1984 SIGMOD 0.00081992507
76 Cache Conscious Algorithms for Relational Query Processing 1994 VLDB 0.00036898845
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024851502
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
616 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015561564
627 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015460957
661 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015003815
771 The Yin and Yang of Processing Data Warehousing Queries on GPU Devices 2013 VLDB 0.00014085862
856 Hardware-Oblivious Parallelism for In-Memory Column-Stores 2013 VLDB 0.00013428547
889 Rack-Scale In-Memory Join Processing using RDMA 2015 SIGMOD 0.00013247362
969 Memory-Efficient Hash Joins 2015 VLDB 0.0001278184
1,266 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011269175
1,268 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001126007
1,283 A Seven-Dimensional Analysis of Hashing Methods and its Implications on Query Processing 2016 VLDB 0.00011209209
1,533 Pipelined Query Processing in Coprocessor Environments 2018 SIGMOD 0.00010332035
1,747 HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines 2019 VLDB 9.7335416e-05
2,122 Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture 2013 VLDB 9.0084047e-05
2,284 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.6954168e-05
2,487 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.3979719e-05
2,600 Robust Query Processing in Co-Processor-accelerated Databases 2016 SIGMOD 8.2363864e-05
2,802 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9903139e-05
2,815 GPL: A GPU-based Pipelined Query Processing Engine 2016 SIGMOD 7.9777435e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9739791e-05
2,886 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.9081605e-05
3,053 MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures 2021 SIGMOD 7.7019663e-05
3,432 Cache-Conscious Radix-Decluster Projections 2004 VLDB 7.3040202e-05
3,459 Improving Main Memory Hash Joins on Intel Xeon Phi Processors: An Experimental Approach 2015 VLDB 7.2848113e-05
3,491 In-Cache Query Co-Processing on Coupled CPU-GPU Architectures 2015 VLDB 7.2619913e-05
3,737 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.0628666e-05
3,773 Hardware-conscious Query Processing in GPU-accelerated Analytical Engines 2019 CIDR 7.0294475e-05
4,137 The Art of Balance: A RateupDB Experience of Building a CPU/GPU Hybrid Database Product 2021 VLDB 6.7861661e-05
4,931 FPGA-based Data Partitioning 2017 SIGMOD 6.348544e-05
5,179 On the Surprising Difficulty of Simple Things: the Case of Radix Partitioning 2015 VLDB 6.2408516e-05
5,423 Joins on Encoded and Partitioned Data 2014 VLDB 6.1360461e-05
5,521 EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal in GPUs 2021 VLDB 6.0959947e-05
5,628 FPGA-based Multithreading for In-Memory Hash Joins 2015 CIDR 6.060354e-05
5,884 MCJoin: A Memory-Constrained Join for Column-Store Main-Memory Databases. 2012 SIGMOD 5.9585932e-05
6,213 Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities 2019 CIDR 5.8479612e-05
6,414 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7927521e-05
6,516 GPU-accelerated data management under the test of time 2020 CIDR 5.7588588e-05
10,142 A four-dimensional Analysis of Partitioned Approximate Filters 2021 VLDB 5.0734597e-05
Previous Page 1 / 1 Next

Semantically Similar Papers