DBScholar

Back to papers

Evaluating Multi-GPU Sorting with Modern Interconnects

Summary: Evaluates multi-GPU sorting across PCIe/NVLink/NVSwitch; proposes a P2P GPU-only sort and a heterogeneous sort, benchmarked on three modern platforms. Reports up to 35x higher P2P throughput with NVSwitch, up to 14x CPU radix-sort speedup (P2P) and 9x (HET); on fast interconnects P2P beats HET by ~1.65x, and copy/compute overlap does not hide transfer bottlenecks. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h43b80d53c55d750a
Venue
SIGMOD
Year
2022
Pagerank
5.7927521e-05
Overall Rank
6,414 | 56.88%
DOI
10.1145/3514221.3517842

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{maltenberger_sigmod22,
        title = {{Evaluating Multi-GPU Sorting with Modern Interconnects}},
        author = {Maltenberger, Tobias and Ilic, Ivan and Tolovski, Ilin and Rabl, Tilmann},
        series = {{SIGMOD} '22},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3514221.3517842},
        url = {https://dl.acm.org/doi/10.1145/3514221.3517842},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 10 of 10 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
219 SIMD-Scan: Ultra Fast in-Memory Table Scan using on-Chip Vector Processing Units 2009 VLDB 0.00024363532
233 Amazon Redshift and the Case for Simpler Data Warehouses 2015 SIGMOD 0.00023783585
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
394 One Trillion Edges: Graph Processing at Facebook-Scale 2015 VLDB 0.00019191286
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018491327
431 HYRISE—A Main Memory Hybrid Storage Engine 2011 VLDB 0.00018403783
575 Efficient Transaction Processing in SAP HANA Database – The End of a Column Store Myth 2012 SIGMOD 0.00016132182
661 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015003815
1,116 A Comprehensive Study of Main-Memory Partitioning and its Application to Large-Scale Comparison- and Radix-Sort 2014 SIGMOD 0.00011962096
1,268 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001126007
1,508 Self-Tuning, GPU-Accelerated Kernel Density Models for Multidimensional Selectivity Estimation 2015 SIGMOD 0.00010440205
2,284 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.6954168e-05
2,487 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.3979719e-05
2,886 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.9081605e-05
3,053 MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures 2021 SIGMOD 7.7019663e-05
4,065 PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort 2015 VLDB 6.8227646e-05
4,225 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.7207703e-05
6,516 GPU-accelerated data management under the test of time 2020 CIDR 5.7588588e-05
Previous Page 1 / 1 Next

Semantically Similar Papers