DBScholar

Back to papers

Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort

Summary: Bandwidth-oblivious SIMD sort for CPU/GPU; competitive analysis across SIMD, radix, and merge approaches. Proposes CPU radix sort and GPU merge sort ~2× faster than prior work; radix dominates on current HW, GPU advantage narrows; merge sort favors large-key cardinalities. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h7228fa3674e94a70
Venue
SIGMOD
Year
2010
Pagerank
0.00015003815
Overall Rank
661 | 95.56%
DOI
10.1145/1807167.1807207

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{satish_sigmod10,
        title = {{Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort}},
        author = {Satish, Nadathur and Kim, Changkyu and Chhugani, Jatin and Nguyen, Anthony D. and Lee, Victor W. and Kim, Daehyun and Dubey, Pradeep},
        series = {{SIGMOD} '10},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/1807167.1807207},
        url = {https://dl.acm.org/doi/10.1145/1807167.1807207},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 38 of 38 citing papers.

Rank Citing Paper Year Venue Pagerank
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
627 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015460957
771 The Yin and Yang of Processing Data Warehousing Queries on GPU Devices 2013 VLDB 0.00014085862
856 Hardware-Oblivious Parallelism for In-Memory Column-Stores 2013 VLDB 0.00013428547
1,116 A Comprehensive Study of Main-Memory Partitioning and its Application to Large-Scale Comparison- and Radix-Sort 2014 SIGMOD 0.00011962096
1,266 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011269175
1,268 A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database Analytics 2020 SIGMOD 0.0001126007
1,855 PALM: Parallel Architecture-Friendly Latch-Free Modifications to B+ Trees on Many-Core Processors 2011 VLDB 9.4973014e-05
1,995 Track Join: Distributed Joins with Minimal Network Traffic 2014 SIGMOD 9.2169073e-05
2,362 Streaming Similarity Search over one Billion Tweets using Parallel Locality-Sensitive Hashing 2013 VLDB 8.5746436e-05
2,466 Concurrent Analytical Query Processing with GPUs 2014 VLDB 8.4222621e-05
2,487 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.3979719e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9739791e-05
3,393 CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster 2012 SIGMOD 7.3471344e-05
3,459 Improving Main Memory Hash Joins on Intel Xeon Phi Processors: An Experimental Approach 2015 VLDB 7.2848113e-05
3,491 In-Cache Query Co-Processing on Coupled CPU-GPU Architectures 2015 VLDB 7.2619913e-05
3,626 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 7.1524537e-05
3,714 The Case for a Learned Sorting Algorithm 2020 SIGMOD 7.0769061e-05
4,011 Deployment of Query Plans on Multicores 2015 VLDB 6.8573269e-05
4,065 PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort 2015 VLDB 6.8227646e-05
4,225 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.7207703e-05
4,931 FPGA-based Data Partitioning 2017 SIGMOD 6.348544e-05
5,179 On the Surprising Difficulty of Simple Things: the Case of Radix Partitioning 2015 VLDB 6.2408516e-05
5,371 Distributed GPU Joins on Fast RDMA-capable Networks 2023 SIGMOD 6.1584802e-05
5,394 Towards a Hybrid Design for Fast Query Processing in DB2 with BLU Acceleration Using Graphical Processing Units: A Technology Demonstration 2016 SIGMOD 6.1497789e-05
5,771 Database Technology for the Masses: Sub-Operators as First-Class Entities 2021 VLDB 5.9992698e-05
6,213 Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities 2019 CIDR 5.8479612e-05
6,414 Evaluating Multi-GPU Sorting with Modern Interconnects 2022 SIGMOD 5.7927521e-05
6,417 MorphStore: Analytical Query Engine with a Holistic Compression-Enabled Processing Model 2020 VLDB 5.791405e-05
7,495 Fast Multi-Column Sorting in Main-Memory Column-Stores 2016 SIGMOD 5.5091753e-05
8,628 Inferray: fast in-memory RDF inference 2016 VLDB 5.2995392e-05
8,673 Adaptive Code Generation for Data-Intensive Analytics 2021 VLDB 5.2913671e-05
8,752 Zero-sided RDMA: Network-driven Data Shuffling for Disaggregated Heterogeneous Cloud DBMSs 2024 SIGMOD 5.2847178e-05
9,429 Efficiently Joining Large Relations on Multi-GPU Systems 2025 VLDB 5.1786456e-05
10,174 Thriving in the No Man’s Land between Compilers and Databases 2019 CIDR 5.0681899e-05
10,785 MGI: A Communication Framework for Data Processing in Massive GPU Infrastructures 2026 VLDB 4.9793485e-05
11,887 Origami: A High-Performance Mergesort Framework 2022 VLDB 4.9793485e-05
12,586 Permuting Data on Random-Access Block Storage 2013 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers