DBScholar

Back to papers

Distributed Join Algorithms on Thousands of Cores

Summary: MPI-based radix-hash and sort-merge joins scale to 4,096 cores and 4.8 TB, reaching 48.7B tuples/s with SIMD, one-sided operations, and RDMA. Communication scheduling and compute/network balance dominate scalability; sort-merge nears its model peak, unlike hashing. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hada67e191ff880b0
Venue
VLDB
Year
2017
Pagerank
7.9866934e-05
Overall Rank
2,802 | 81.17%
DOI
10.14778/3055540.3055546
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{barthels_vldb17,
        title = {{Distributed Join Algorithms on Thousands of Cores}},
        author = {Barthels, Claude and Müller, Ingo and Schneider, Timo and Alonso, Gustavo and Hoefler, Torsten},
        journal = {PVLDB},
        series = {{VLDB} '17},
        volume = {10},
        number = {5},
        pages = {517--528},
        doi = {10.14778/3055540.3055546},
        url = {https://doi.org/10.14778/3055540.3055546},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 23 of 23 citing papers.

Rank Citing Paper Year Venue Pagerank
1,242 BatchDB: Efficient Isolated Execution of Hybrid OLTP+OLAP Workloads for Interactive Applications 2017 SIGMOD 0.00011378275
1,366 Designing Distributed Tree-based Index Structures for Fast RDMA-capable Networks 2019 SIGMOD 0.00010908555
1,423 Procella: Unifying serving and analytical data at YouTube 2019 VLDB 0.00010715412
1,966 Sherman: A Write-Optimized Distributed B+Tree Index on Disaggregated Memory 2022 SIGMOD 9.3014925e-05
3,229 Farview: Disaggregated Memory with Operator Off-loading for Database Engines 2022 CIDR 7.5048882e-05
3,609 DFI: The Data Flow Interface for High-Speed Networks 2021 SIGMOD 7.1617145e-05
3,628 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 7.1490678e-05
3,955 Tensors: An abstraction for general data processing 2021 VLDB 6.8994712e-05
4,416 Strong consistency is not hard to get: Two-Phase Locking and Two-Phase Commit on Thousands of Cores 2019 VLDB 6.6067888e-05
4,932 FPGA-based Data Partitioning 2017 SIGMOD 6.3455411e-05
5,376 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1555648e-05
6,216 Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities 2019 CIDR 5.8451996e-05
6,536 DPI: The Data Processing Interface for Modern Networks 2019 CIDR 5.7514775e-05
7,455 Modularis: Modular Relational Analytics over Heterogeneous Distributed Platforms 2021 VLDB 5.5210654e-05
7,735 Hardware-Oblivious SIMD Parallelism for In-Memory Column-Stores 2020 CIDR 5.4632401e-05
8,275 Accelerate Distributed Joins with Predicate Transfer 2025 SIGMOD 5.3623175e-05
8,402 The Case for Learned In-Memory Joins 2023 VLDB 5.3375308e-05
8,868 OLAP on Modern Chiplet-Based Processors 2024 VLDB 5.2607671e-05
9,438 Efficiently Joining Large Relations on Multi-GPU Systems 2025 VLDB 5.1761941e-05
9,666 Parallel Query Processing: To Separate Communication from Computation 2022 SIGMOD 5.142891e-05
10,311 Data Chunk Compaction in Vectorized Execution 2025 SIGMOD 5.0376863e-05
10,811 PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage 2026 VLDB 4.9769913e-05
11,871 Scaling Equi-Joins 2022 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 11 of 11 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers