Back to papers
A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs
Summary: Memory-bandwidth-aware hybrid GPU radix sort halves transfers, delivering ~2.3× speedup on uniform data. A pipelined heterogeneous mode handles off-GPU/large inputs, enabling strong end-to-end gains over CPU radix sort on large key-value workloads.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 5420
- Venue
- SIGMOD
- Year
- 2017
- Pagerank
- 7.4648665e-05
- Overall Rank
- 3,161 | 78.04%
- DOI
-
10.1145/3035918.3064043
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 22 of 22 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 2,044 |
A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics |
2020 |
SIGMOD |
9.6963999e-05 |
| 2,659 |
HetExchange: Encapsulating heterogeneous CPU-GPU parallelism in JIT compiled engines |
2019 |
VLDB |
8.3615158e-05 |
| 3,328 |
Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects |
2020 |
SIGMOD |
7.2136181e-05 |
| 4,359 |
Hardware-conscious Query Processing in GPU-accelerated Analytical Engines |
2019 |
CIDR |
6.2493951e-05 |
| 5,018 |
Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMS |
2022 |
VLDB |
5.7503878e-05 |
| 5,039 |
Tile-based Lightweight Integer Compression in GPU |
2022 |
SIGMOD |
5.7369993e-05 |
| 5,129 |
The Art of Balance: A RateupDBTM Experience of Building a CPU/GPU Hybrid Database Product |
2021 |
VLDB |
5.6724875e-05 |
| 5,251 |
Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects |
2022 |
SIGMOD |
5.6003972e-05 |
| 6,400 |
ColumnML: Column-Store Machine Learning with On-The-Fly Data Transformation |
2019 |
VLDB |
5.0739311e-05 |
| 6,469 |
Parallel Index-based Stream Join on a Multicore CPU |
2020 |
SIGMOD |
5.0448159e-05 |
| 6,538 |
Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities |
2019 |
CIDR |
5.0173391e-05 |
| 7,155 |
Evaluating Multi-GPU Sorting with Modern Interconnects |
2022 |
SIGMOD |
4.810361e-05 |
| 7,324 |
BOSS - An Architecture for Database Kernel Composition |
2024 |
VLDB |
4.7565238e-05 |
| 7,356 |
ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw Data |
2020 |
VLDB |
4.748111e-05 |
| 7,550 |
Efficient Top-K Query Processing on Massively Parallel Hardware |
2018 |
SIGMOD |
4.7089541e-05 |
| 8,423 |
SPRINTER: A Fast n-ary Join Query Processing Method for Complex OLAP Queries |
2020 |
SIGMOD |
4.5112315e-05 |
| 8,846 |
Scaling your Hybrid CPU-GPU DBMS to Multiple GPUs |
2024 |
VLDB |
4.432948e-05 |
| 9,033 |
A Case for Ecological Efficiency in Database Server Lifecycles |
2025 |
CIDR |
4.3997447e-05 |
| 9,837 |
Efficiently Joining Large Relations on Multi-GPU Systems |
2025 |
VLDB |
4.269939e-05 |
| 9,952 |
Distributed Stream KNN Join |
2021 |
SIGMOD |
4.236537e-05 |
| 10,755 |
Scaling GPU-Accelerated Databases beyond GPU Memory Size |
2025 |
VLDB |
4.1905499e-05 |
| 11,023 |
Accelerating Merkle Patricia Trie with GPU |
2024 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers