Back to papers
MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures
Summary: MG-Join proposes a scalable partitioned hash join for multi-GPU single-machine architectures. Adaptive multi-hop cross-GPU routing minimizes congestion, achieving up to 97% bisection-bandwidth utilization, and up to 2.5x join speedups with 4.5x TPC-H gains over Omnisci.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6146
- Venue
- SIGMOD
- Year
- 2021
- Pagerank
- 6.5419402e-05
- Overall Rank
- 4,000 | 72.21%
- DOI
-
10.1145/3448016.3457254
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 20 of 20 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 5,018 |
Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS |
2022 |
VLDB |
5.7503878e-05 |
| 5,039 |
Tile-based Lightweight Integer Compression in GPU |
2022 |
SIGMOD |
5.7369993e-05 |
| 5,251 |
Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects |
2022 |
SIGMOD |
5.6003972e-05 |
| 5,437 |
RTIndeX: Exploiting Hardware-Accelerated GPU Raytracing for Database Indexing |
2023 |
VLDB |
5.5043792e-05 |
| 6,068 |
GPU Database Systems Characterization and Optimization |
2024 |
VLDB |
5.2240241e-05 |
| 6,220 |
Distributed GPU Joins on Fast RDMA-capable Networks |
2023 |
SIGMOD |
5.1446966e-05 |
| 6,451 |
Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics |
2025 |
VLDB |
5.0522576e-05 |
| 7,155 |
Evaluating Multi-GPU Sorting with Modern Interconnects |
2022 |
SIGMOD |
4.810361e-05 |
| 7,324 |
BOSS - An Architecture for Database Kernel Composition |
2024 |
VLDB |
4.7565238e-05 |
| 7,570 |
Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs |
2025 |
VLDB |
4.7039167e-05 |
| 7,752 |
Efficiently Processing Joins and Grouped Aggregations on GPUs |
2025 |
SIGMOD |
4.6558737e-05 |
| 7,917 |
Terabyte-Scale Analytics in the Blink of an Eye |
2026 |
VLDB |
4.6129625e-05 |
| 8,846 |
Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs |
2024 |
VLDB |
4.432948e-05 |
| 9,143 |
Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMs |
2023 |
SIGMOD |
4.381112e-05 |
| 9,837 |
Efficiently Joining Large Relations on Multi-GPU Systems |
2025 |
VLDB |
4.269939e-05 |
| 10,253 |
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization |
2026 |
VLDB |
4.1905499e-05 |
| 10,755 |
Scaling GPU-Accelerated Databases beyond GPU Memory Size |
2025 |
VLDB |
4.1905499e-05 |
| 10,996 |
SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMs |
2024 |
SIGMOD |
4.1905499e-05 |
| 11,023 |
Accelerating Merkle Patricia Trie with GPU |
2024 |
VLDB |
4.1905499e-05 |
| 11,360 |
Scaling Equi-Joins |
2022 |
SIGMOD |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 771 |
Relational Joins on Graphics Processors |
2008 |
SIGMOD |
0.00016813054 |
| 1,196 |
Rack-Scale In-Memory Join Processing using RDMA |
2015 |
SIGMOD |
0.00013378379 |
| 1,271 |
The Yin and Yang of Processing Data Warehousing Queries on GPU Devices |
2013 |
VLDB |
0.00012900735 |
| 1,800 |
An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory |
2016 |
SIGMOD |
0.00010494121 |
| 2,518 |
Track Join: Distributed Joins with Minimal Network Traffic |
2014 |
SIGMOD |
8.6052941e-05 |
| 2,523 |
Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture |
2013 |
VLDB |
8.599693e-05 |
| 2,659 |
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines |
2019 |
VLDB |
8.3615158e-05 |
| 3,372 |
CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers |
2019 |
VLDB |
7.1633912e-05 |
| 3,674 |
A Distributed Multi-GPU System for Fast Graph Processing |
2018 |
VLDB |
6.8502146e-05 |
| 4,089 |
In-Cache Query Co-Processing on Coupled CPU-GPU Architectures |
2015 |
VLDB |
6.4559891e-05 |
| 4,359 |
Hardware-conscious Query Processing in GPU-accelerated Analytical Engines |
2019 |
CIDR |
6.2493951e-05 |
| 5,199 |
Data-Parallel Query Processing on Non-Uniform Data |
2020 |
VLDB |
5.6294232e-05 |
| 5,584 |
Ocelot/HyPE: Optimized Data Processing on Heterogeneous Hardware |
2014 |
VLDB |
5.4201842e-05 |
| 6,367 |
Improving Execution Efficiency of Just-in-time Compilation based Query Processing on GPUs |
2021 |
VLDB |
5.0887599e-05 |
| 7,056 |
SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning |
2017 |
VLDB |
4.8420027e-05 |
Semantically Similar Papers