DBScholar

Back to papers

A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs

Summary: Memory-bandwidth-aware hybrid GPU radix sort halves transfers, delivering ~2.3× speedup on uniform data. A pipelined heterogeneous mode handles off-GPU/large inputs, enabling strong end-to-end gains over CPU radix sort on large key-value workloads. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h4da9f131cf424531
Venue
SIGMOD
Year
2017
Pagerank
8.3979719e-05
Overall Rank
2,487 | 83.29%
DOI
10.1145/3035918.3064043

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{stehle_sigmod17,
        title = {{A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs}},
        author = {Stehle, Elias and Jacobsen, Hans-Arno},
        series = {{SIGMOD} '17},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3035918.3064043},
        url = {https://dl.acm.org/doi/10.1145/3035918.3064043},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 24 of 24 citing papers.

Rank Citing Paper Year Venue Pagerank
1,268 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001126007
1,747 HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines 2019 VLDB 9.7335416e-05
2,284 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.6954168e-05
3,626 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 7.1524537e-05
3,655 Tile-based Lightweight Integer Compression in GPU 2022 SIGMOD 7.1260841e-05
3,773 Hardware-conscious Query Processing in GPU-accelerated Analytical Engines 2019 CIDR 7.0294475e-05
3,888 Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS 2022 VLDB 6.945725e-05
4,137 The Art of Balance: A RateupDB Experience of Building a CPU/GPU Hybrid Database Product 2021 VLDB 6.7861661e-05
5,615 BOSS - An Architecture for Database Kernel Composition 2024 VLDB 6.0652416e-05
6,063 ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data 2020 VLDB 5.8982522e-05
6,123 Parallel Index-based Stream Join on a Multicore CPU 2020 SIGMOD 5.8789534e-05
6,213 Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities 2019 CIDR 5.8479612e-05
6,246 ColumnML: Column-Store Machine Learning with On-The-Fly Data Transformation 2019 VLDB 5.8370793e-05
6,414 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7927521e-05
7,066 Efficient Top-K Query Processing on Massively Parallel Hardware 2018 SIGMOD 5.6087334e-05
8,236 Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs 2024 VLDB 5.3710413e-05
8,274 Scaling GPU-Accelerated Databases beyond GPU Memory Size 2025 VLDB 5.3642256e-05
8,493 SPRINTER: A Fast n-ary Join Query Processing Method for Complex OLAP Queries 2020 SIGMOD 5.3310013e-05
9,359 A Case for Ecological Efficiency in Database Server Lifecycles 2025 CIDR 5.1868213e-05
9,429 Efficiently Joining Large Relations on Multi-GPU Systems 2025 VLDB 5.1786456e-05
10,329 Distributed Stream KNN Join 2021 SIGMOD 5.0333462e-05
10,785 MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures 2026 VLDB 4.9793485e-05
10,905 ZipFlow: a Compiler-based Framework to Unleash Compressed Data Movement for Modern GPUs 2026 VLDB 4.9793485e-05
11,567 Accelerating Merkle Patricia Trie with GPU 2024 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 14 of 14 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
106 Quickly Generating Billion-Record Synthetic Databases 1994 SIGMOD 0.00033526937
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024851502
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
304 GPUTeraSort: High Performance Graphics Co-processor Sorting for Large Database Management 2006 SIGMOD 0.00021604795
394 One Trillion Edges: Graph Processing at Facebook-Scale 2015 VLDB 0.00019191286
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018491327
661 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015003815
722 Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture 2008 VLDB 0.00014488003
1,116 A Comprehensive Study of Main-Memory Partitioning and its Application to Large-Scale Comparison- and Radix-Sort 2014 SIGMOD 0.00011962096
3,375 A Hybrid B+-tree as Solution for In-Memory Indexing on CPU-GPU Heterogeneous Computing Platforms 2016 SIGMOD 7.3608441e-05
3,393 CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster 2012 SIGMOD 7.3471344e-05
4,065 PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort 2015 VLDB 6.8227646e-05
4,225 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.7207703e-05
6,194 Patience is a Virtue: Revisiting Merge and Sort on Modern Processors 2014 SIGMOD 5.8544215e-05
Previous Page 1 / 1 Next

Semantically Similar Papers