Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment
Summary: Multi-GPU join algorithms for very large tables; three designs: nested-loop, global-sort-merge, hybrid. Addresses CPU–GPU data-transfer bottlenecks; shows scalability and up to 25× vs multi-core CPU and 2.8× vs single-GPU. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Ran Rui
- 2. Hao Li
- 3. Yi-Cheng Tu
Incoming Citations (Sorted by Pagerank)
Showing 20 of 20 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 350 | Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs | 2009 | VLDB | 0.00026368305 |
| 403 | Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited | 2014 | VLDB | 0.00024176677 |
| 538 | Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs | 2011 | SIGMOD | 0.00020632609 |
| 584 | Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems | 2012 | VLDB | 0.00019700451 |
| 771 | Relational Joins on Graphics Processors | 2008 | SIGMOD | 0.00016813054 |
| 1,016 | Memory-Efficient Hash Joins | 2015 | VLDB | 0.00014630024 |
| 1,271 | The Yin and Yang of Processing Data Warehousing Queries on GPU Devices | 2013 | VLDB | 0.00012900735 |
| 2,523 | Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture | 2013 | VLDB | 8.599693e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 350 | Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs | 2009 | VLDB | 0.00026368305 |
| 2,599 | Design and Evaluation of Parallel Pipelined Join Algorithms | 1987 | SIGMOD | 8.4655075e-05 |
| 1,938 | From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System | 2015 | SIGMOD | 0.00010025547 |
| 2,051 | Optimization of Multi-Way Join Queries for Parallel Execution | 1991 | VLDB | 9.6871984e-05 |
| 6,060 | Efficient Massively Parallel Join Optimization for Large Queries* | 2022 | SIGMOD | 5.2271244e-05 |
| 6,220 | Distributed GPU Joins on Fast RDMA-capable Networks | 2023 | SIGMOD | 5.1446966e-05 |
| 5,251 | Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects | 2022 | SIGMOD | 5.6003972e-05 |
| 7,752 | Efficiently Processing Joins and Grouped Aggregations on GPUs | 2025 | SIGMOD | 4.6558737e-05 |
| 771 | Relational Joins on Graphics Processors | 2008 | SIGMOD | 0.00016813054 |
| 9,837 | Efficiently Joining Large Relations on Multi-GPU Systems | 2025 | VLDB | 4.269939e-05 |