Back to papers
MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures
Summary: MG-Join proposes a scalable partitioned hash join for multi-GPU single-machine architectures. Adaptive multi-hop cross-GPU routing minimizes congestion, achieving up to 97% bisection-bandwidth utilization, and up to 2.5x join speedups with 4.5x TPC-H gains over Omnisci.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h15da8b410dd3e296
Venue
SIGMOD
Year
2021
Pagerank
7.6983202e-05
Overall Rank
3,055 | 79.47%
DOI
10.1145/3448016.3457254
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{paul_sigmod21,
title = {{MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures}},
author = {Paul, Johns and Lu, Shengliang and He, Bingsheng and Lau, Chiew Tong},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457254},
url = {https://dl.acm.org/doi/10.1145/3448016.3457254},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 22 of 22 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
3,628
Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects
2022
SIGMOD
7.1490678e-05
3,657
Tile-based Lightweight Integer Compression in GPU
2022
SIGMOD
7.1227107e-05
3,691
RTIndex: Exploiting Hardware-Accelerated GPU Raytracing for Database Indexing
2023
VLDB
7.0909492e-05
3,888
Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS
2022
VLDB
6.942437e-05
4,484
GPU Database Systems Characterization and Optimization
2024
VLDB
6.5776765e-05
5,376
Distributed GPU Joins on Fast RDMA-capable Networks
2023
SIGMOD
6.1555648e-05
5,543
Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics
2025
VLDB
6.0873423e-05
5,616
BOSS - An Architecture for Database Kernel Composition
2024
VLDB
6.0623704e-05
5,991
Terabyte-Scale Analytics in the Blink of an Eye
2026
VLDB
5.9219616e-05
6,153
Efficiently Processing Joins and Grouped Aggregations on GPUs
2025
SIGMOD
5.8669913e-05
6,417
Evaluating Multi-GPU Sorting with Modern Interconnects
2022
SIGMOD
5.7900099e-05
6,671
Powerful GPUs or Fast Interconnects: Analyzing Relational Workloads on Modern GPUs
2025
VLDB
5.7125893e-05
7,099
Design and Analysis of a Processing-in-DIMM Join Algorithm: A Case Study with UPMEM DIMMs
2023
SIGMOD
5.5991152e-05
8,242
Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs
2024
VLDB
5.3684987e-05
8,280
Scaling GPU-Accelerated Databases beyond GPU Memory Size
2025
VLDB
5.3616863e-05
9,438
Efficiently Joining Large Relations on Multi-GPU Systems
2025
VLDB
5.1761941e-05
10,733
Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization
2026
VLDB
4.9769913e-05
10,795
MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures
2026
VLDB
4.9769913e-05
10,840
Bridging the Indexing Gap in Fused GPU Query Engines
2026
VLDB
4.9769913e-05
11,551
SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMs
2024
SIGMOD
4.9769913e-05
11,573
Accelerating Merkle Patricia Trie with GPU
2024
VLDB
4.9769913e-05
11,871
Scaling Equi-Joins
2022
SIGMOD
4.9769913e-05
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
616
Relational Joins on Graphics Processors
2008
SIGMOD
0.00015554627
775
The Yin and Yang of Processing Data Warehousing Queries on GPU Devices
2013
VLDB
0.0001407924
890
Rack-Scale In-Memory Join Processing using RDMA
2015
SIGMOD
0.00013241413
1,267
An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory
2016
SIGMOD
0.00011265987
1,748
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines
2019
VLDB
9.7289385e-05
1,997
Track Join: Distributed Joins with Minimal Network Traffic
2014
SIGMOD
9.2126022e-05
2,124
Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture
2013
VLDB
9.0041425e-05
3,309
A Distributed Multi-GPU System for Fast Graph Processing
2018
VLDB
7.4374573e-05
3,491
In-Cache Query Co-Processing on Coupled CPU-GPU Architectures
2015
VLDB
7.2586387e-05
3,766
CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
2019
VLDB
7.0360981e-05
3,775
Hardware-conscious Query Processing in GPU-accelerated Analytical Engines
2019
CIDR
7.0261504e-05
3,962
Data-Parallel Query Processing on Non-Uniform Data
2020
VLDB
6.8916038e-05
4,647
Improving Execution Efficiency of Just-in-time Compilation based Query Processing on GPUs
2021
VLDB
6.4850635e-05
5,135
Ocelot/HyPE: Optimized Data Processing on Heterogeneous Hardware
2014
VLDB
6.2568126e-05
6,863
SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning
2017
VLDB
5.6604842e-05
Semantically Similar Papers