DBScholar

Back to papers

Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture

Summary: SIMD-optimized MergeSort with cache-efficient multiway merging achieves 3.3× scalar speedup and sorts 64M floats in <0.5s on a 4-core CPU. Analytical and cycle-accurate studies show scalability across SIMD widths and 32+ cores. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
ha8d96555afff33ae
Venue
VLDB
Year
2008
Pagerank
0.00014488003
Overall Rank
722 | 95.15%
DOI
10.14778/1454159.1454171

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{chhugani_vldb08,
        title = {{Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture}},
        author = {Chhugani, Jatin and Nguyen, Anthony D. and Lee, Victor W. and Macy, William and Hagog, Mostafa and Chen, Yen-Kuang and Baransi, Akram and Kumar, Sanjeev and Dubey, Pradeep},
        journal = {PVLDB},
        series = {{VLDB} '08},
        volume = {1},
        number = {1},
        pages = {1313--1324},
        doi = {10.14778/1454159.1454171},
        url = {https://doi.org/10.14778/1454159.1454171},
        year = {2008}
}

Incoming Citations (Sorted by Pagerank)

Showing 38 of 38 citing papers.

Rank Citing Paper Year Venue Pagerank
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024851502
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
282 FAST: Fast Architecture Sensitive Tree Search on Modern CPUs and GPUs 2010 SIGMOD 0.00022264207
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018491327
627 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015460957
661 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015003815
713 Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan 2016 VLDB 0.00014571977
906 Data Blocks: Hybrid OLTP and OLAP on Compressed Storage using both Vectorization and Compilation 2016 SIGMOD 0.00013160654
941 Data Processing on FPGAs 2009 VLDB 0.00012957405
1,116 A Comprehensive Study of Main-Memory Partitioning and its Application to Large-Scale Comparison- and Radix-Sort 2014 SIGMOD 0.00011962096
1,335 Fast Updates on Read-Optimized Databases Using Multi-Core CPUs 2012 VLDB 0.0001099401
1,723 ByteSlice: Pushing the Envelop of Main Memory Data Processing with a New Storage Layout 2015 SIGMOD 9.7931223e-05
1,855 PALM: Parallel Architecture-Friendly Latch-Free Modifications to B+ Trees on Many-Core Processors 2011 VLDB 9.4973014e-05
2,487 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.3979719e-05
3,289 Faster Set Intersection with SIMD instructions by Reducing Branch Mispredictions 2015 VLDB 7.4528818e-05
3,393 CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster 2012 SIGMOD 7.3471344e-05
4,065 PARADIS: An Efficient Parallel Algorithm for In-place Radix Sort 2015 VLDB 6.8227646e-05
4,137 The Art of Balance: A RateupDB Experience of Building a CPU/GPU Hybrid Database Product 2021 VLDB 6.7861661e-05
4,225 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.7207703e-05
5,598 Database Processing-in-Memory: An Experimental Study 2020 VLDB 6.0713144e-05
6,154 FPGA: What's in it for a Database? 2009 SIGMOD 5.8687334e-05
6,194 Patience is a Virtue: Revisiting Merge and Sort on Modern Processors 2014 SIGMOD 5.8544215e-05
6,213 Data Partitioning for In-Memory Systems: Myths, Challenges, and Opportunities 2019 CIDR 5.8479612e-05
6,417 MorphStore: Analytical Query Engine with a Holistic Compression-Enabled Processing Model 2020 VLDB 5.791405e-05
6,731 What Is the Price for Joining Securely? Benchmarking Equi-Joins in Trusted Execution Environments 2022 VLDB 5.6939772e-05
7,066 Efficient Top-K Query Processing on Massively Parallel Hardware 2018 SIGMOD 5.6087334e-05
7,495 Fast Multi-Column Sorting in Main-Memory Column-Stores 2016 SIGMOD 5.5091753e-05
7,866 Building Advanced SQL Analytics From Low-Level Plan Operators 2021 SIGMOD 5.4365876e-05
7,964 Parallelizing Intra-Window Join on Multicores: An Experimental Study 2021 SIGMOD 5.4165494e-05
8,380 Interleaved Multi-Vectorizing 2020 VLDB 5.3429493e-05
8,673 Adaptive Code Generation for Data-Intensive Analytics 2021 VLDB 5.2913671e-05
8,821 Efficient Evaluation of Arbitrarily-Framed Holistic SQL Aggregates and Window Functions 2022 SIGMOD 5.2691478e-05
9,201 An Application-Specific Instruction Set for Accelerating Set-Oriented Database Primitives 2014 SIGMOD 5.2083769e-05
9,429 Efficiently Joining Large Relations on Multi-GPU Systems 2025 VLDB 5.1786456e-05
10,601 TQEx: Tensor-based Query Engine Enhanced by Bridging the Gap 2026 SIGMOD 4.9793485e-05
11,536 Enabling Adaptive Sampling for Intra-Window Join: Simultaneously Optimizing Quantity and Quality 2024 SIGMOD 4.9793485e-05
11,887 Origami: A High-Performance Mergesort Framework 2022 VLDB 4.9793485e-05
12,339 Efficient Query Processing on Many-core Architectures: A Case Study with Intel Xeon Phi Processor 2016 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 2 of 2 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
304 GPUTeraSort: High Performance Graphics Co-processor Sorting for Large Database Management 2006 SIGMOD 0.00021604795
2,161 CellSort: High Performance Sorting on the Cell Processor 2007 VLDB 8.9357193e-05
Previous Page 1 / 1 Next

Semantically Similar Papers