DBScholar

Back to papers

Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects

Summary: Triton Join scales large joins on GPUs by using fast interconnects (NVLink 2.0) to spill state to main memory. Delivers >100x GPU hash-join gains over non-partitioned approaches, and up to 2.5x vs CPU radix-join, enabling GPU DBMSs to exceed GPU memory. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h4f933eee98a5ce53
Venue
SIGMOD
Year
2022
Pagerank
7.1490678e-05
Overall Rank
3,628 | 75.62%
DOI
10.1145/3514221.3517911

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{lutz_sigmod22,
        title = {{Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects}},
        author = {Lutz, Clemens and Breß, Sebastian and Zeuch, Steffen and Rabl, Tilmann and Markl, Volker},
        series = {{SIGMOD} '22},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3514221.3517911},
        url = {https://dl.acm.org/doi/10.1145/3514221.3517911},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 20 of 20 citing papers.

Rank Citing Paper Year Venue Pagerank
4,484 GPU Database Systems Characterization and Optimization 2024 VLDB 6.5776765e-05
5,376 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1555648e-05
5,543 Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics 2025 VLDB 6.0873423e-05
5,991 Terabyte-Scale Analytics in the Blink of an Eye 2026 VLDB 5.9219616e-05
6,153 Efficiently Processing Joins and Grouped Aggregations on GPUs 2025 SIGMOD 5.8669913e-05
6,671 Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs 2025 VLDB 5.7125893e-05
7,099 Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMs 2023 SIGMOD 5.5991152e-05
7,808 Analyzing Vectorized Hash Tables Across CPU Architectures 2023 VLDB 5.448023e-05
8,242 Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs 2024 VLDB 5.3684987e-05
8,602 A Case for Graphics-driven Query Processing 2023 VLDB 5.3037284e-05
9,438 Efficiently Joining Large Relations on Multi-GPU Systems 2025 VLDB 5.1761941e-05
9,693 DPDPU: Data Processing with DPUs 2025 CIDR 5.1384473e-05
10,787 Succinct and Fast Tiny Pointer Hash Tables 2026 VLDB 4.9769913e-05
10,795 MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures 2026 VLDB 4.9769913e-05
10,811 PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage 2026 VLDB 4.9769913e-05
10,840 Bridging the Indexing Gap in Fused GPU Query Engines 2026 VLDB 4.9769913e-05
10,932 TQP++: Bridging ML Compilers and Analytical Query Processing on GPUs 2026 VLDB 4.9769913e-05
11,542 Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality 2024 SIGMOD 4.9769913e-05
11,551 SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMs 2024 SIGMOD 4.9769913e-05
11,573 Accelerating Merkle Patricia Trie with GPU 2024 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 41 of 41 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
7 Implementation Techniques For Main Memory Database Systems 1984 SIGMOD 0.00081971778
76 Cache Conscious Algorithms for Relational Query Processing 1994 VLDB 0.00036891569
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024844328
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023136934
616 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015554627
627 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015454197
663 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00014997516
775 The Yin and Yang of Processing Data Warehousing Queries on GPU Devices 2013 VLDB 0.0001407924
857 Hardware-Oblivious Parallelism for In-Memory Column-Stores 2013 VLDB 0.00013422539
890 Rack-Scale In-Memory Join Processing using RDMA 2015 SIGMOD 0.00013241413
963 Memory-Efficient Hash Joins 2015 VLDB 0.00012815832
1,212 A Seven-Dimensional Analysis of Hashing Methods and its Implications on Query Processing 2016 VLDB 0.00011521857
1,267 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011265987
1,269 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.00011254742
1,534 Pipelined Query Processing in Coprocessor Environments 2018 SIGMOD 0.00010327147
1,748 HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines 2019 VLDB 9.7289385e-05
2,124 Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture 2013 VLDB 9.0041425e-05
2,287 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.691301e-05
2,487 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.3939994e-05
2,601 Robust Query Processing in Co-Processor-accelerated Databases 2016 SIGMOD 8.2325291e-05
2,802 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9866934e-05
2,816 GPL: A GPU-based Pipelined Query Processing Engine 2016 SIGMOD 7.9740083e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9703078e-05
2,887 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.9044174e-05
3,055 MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures 2021 SIGMOD 7.6983202e-05
3,432 Cache-Conscious Radix-Decluster Projections 2004 VLDB 7.300628e-05
3,458 Improving Main Memory Hash Joins on Intel Xeon Phi Processors: An Experimental Approach 2015 VLDB 7.2815031e-05
3,491 In-Cache Query Co-Processing on Coupled CPU-GPU Architectures 2015 VLDB 7.2586387e-05
3,739 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.0595382e-05
3,775 Hardware-conscious Query Processing in GPU-accelerated Analytical Engines 2019 CIDR 7.0261504e-05
4,138 The Art of Balance: A RateupDB Experience of Building a CPU/GPU Hybrid Database Product 2021 VLDB 6.782996e-05
4,932 FPGA-based Data Partitioning 2017 SIGMOD 6.3455411e-05
5,180 On the Surprising Difficulty of Simple Things: the Case of Radix Partitioning 2015 VLDB 6.2379001e-05
5,427 Joins on Encoded and Partitioned Data 2014 VLDB 6.133151e-05
5,524 EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal in GPUs 2021 VLDB 6.0931089e-05
5,629 FPGA-based Multithreading for In-Memory Hash Joins 2015 CIDR 6.0574868e-05
5,881 MCJoin: A Memory-Constrained Join for Column-Store Main-Memory Databases. 2012 SIGMOD 5.9575926e-05
6,216 Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities 2019 CIDR 5.8451996e-05
6,417 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7900099e-05
6,518 GPU-accelerated data management under the test of time 2020 CIDR 5.7561332e-05
10,146 A four-dimensional Analysis of Partitioned Approximate Filters 2021 VLDB 5.071058e-05
Previous Page 1 / 1 Next

Semantically Similar Papers