DBScholar

Back to papers

Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs

Summary: Revisits hash join versus sort-merge join with highly optimized multicore implementations, achieving record CPU throughput and robust performance under skew and varying input sizes. Models predict wider SIMD, more cores, and bandwidth limits will soon favor sort-merge joins. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
10056
Venue
VLDB
Year
2009
Pagerank
0.00024932174
Overall Rank
209 | 98.57%
DOI
10.14778/1687553.1687564

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{kim_vldb09,
        title = {{Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs}},
        author = {Kim, Changkyu and Kaldewey, Tim and Lee, Victor W. and Sedlar, Eric and Nguyen, Anthony D. and Satish, Nadathur and Chhugani, Jatin and Di Blas, Andrea and Dubey, Pradeep},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687553.1687564},
        url = {https://doi.org/10.14778/1687553.1687564},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 80 citing papers.

Rank Citing Paper Year Venue Pagerank
241 Morsel-Driven Parallelism: A NUMA-Aware Query Evaluation Framework for the Many-Core Age 2014 SIGMOD 0.00023654664
252 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023242719
278 FAST: Fast Architecture Sensitive Tree Search on Modern CPUs and GPUs 2010 SIGMOD 0.00022476841
360 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00020182846
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018725853
634 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015533814
670 SharedDB: Killing One Thousand Queries With One Stone 2012 VLDB 0.00015157572
678 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015061068
805 Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan 2016 VLDB 0.00013891999
892 Rack-Scale In-Memory Join Processing using RDMA 2015 SIGMOD 0.00013376761
912 Hardware-Oblivious Parallelism for In-Memory Column-Stores 2013 VLDB 0.00013269804
959 Memory-Efficient Hash Joins 2015 VLDB 0.00012953588
1,150 DimmWitted: A Study of Main-Memory Statistical Analytics 2014 VLDB 0.00011943462
1,177 A Comprehensive Study of Main-Memory Partitioning and its Application to Large-Scale Comparison- and Radix-Sort 2014 SIGMOD 0.00011808761
1,265 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011415709
1,315 Fast Updates on Read-Optimized Databases Using Multi-Core CPUs 2012 VLDB 0.00011181796
1,761 ByteSlice: Pushing the Envelop of Main Memory Data Processing with a New Storage Layout 2015 SIGMOD 9.8154969e-05
1,857 Joins via Geometric Resolutions: Worst-case and Beyond 2015 PODS 9.6047945e-05
1,974 Track Join: Distributed Joins with Minimal Network Traffic 2014 SIGMOD 9.3658402e-05
2,140 Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture 2013 VLDB 9.0991487e-05
2,250 Cache-Efficient Aggregation: Hashing Is Sorting 2015 SIGMOD 8.8694486e-05
2,390 Streaming Similarity Search over one Billion Tweets using Parallel Locality-Sensitive Hashing 2013 VLDB 8.6438351e-05
2,566 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.4116562e-05
2,649 Asynchronous Memory Access Chaining 2016 VLDB 8.2926258e-05
2,664 Adaptive and Big Data Scale Parallel Execution in Oracle 2013 VLDB 8.2816537e-05
2,667 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.2756346e-05
2,823 Query Processing on Tensor Computation Runtimes 2022 VLDB 8.0893814e-05
2,926 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9549783e-05
2,962 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9170451e-05
3,117 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.7382559e-05
3,134 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.7231028e-05
3,368 CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster 2012 SIGMOD 7.4713287e-05
3,435 Improving Main Memory Hash Joins on Intel Xeon Phi Processors: An Experimental Approach 2015 VLDB 7.4172582e-05
3,471 In-Cache Query Co-Processing on Coupled CPU-GPU Architectures 2015 VLDB 7.3885861e-05
4,085 Deployment of Query Plans on Multicores 2015 VLDB 6.9149518e-05
4,177 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.8499317e-05
4,478 OmniDB: Towards Portable and Efficient Query Processing on Parallel CPU/GPU Architectures 2013 VLDB 6.6789258e-05
4,535 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 6.6419266e-05
4,553 Predicate Transfer: Efficient Pre-Filtering on Multi-Join Queries 2024 CIDR 6.6346951e-05
4,603 The Art of Balance: A RateupDB Experience of Building a CPU/GPU Hybrid Database Product 2021 VLDB 6.6105578e-05
4,639 MQJoin: Efficient Shared Execution of Main-Memory Joins 2016 VLDB 6.5918791e-05
4,840 FPGA-based Data Partitioning 2017 SIGMOD 6.483442e-05
5,039 Holistic Indexing in Main-memory Column-stores 2015 SIGMOD 6.3909067e-05
5,368 Density-optimized Intersection-free Mapping and Matrix Multiplication for Join-Project Operations 2022 VLDB 6.2448114e-05
5,515 FPGA-based Multithreading for In-Memory Hash Joins 2015 CIDR 6.189963e-05
5,536 Database Processing-in-Memory: An Experimental Study 2020 VLDB 6.1822684e-05
5,715 Design and Evaluation of Storage Organizations for Read-Optimized Main Memory Databases 2013 VLDB 6.1105343e-05
5,765 Charting the Design Space of Query Execution using VOILA 2021 VLDB 6.0953705e-05
5,812 MCJoin: A Memory-Constrained Join for Column-Store Main-Memory Databases. 2012 SIGMOD 6.0782357e-05
6,003 ThunderRW: An In-Memory Graph Random Walk Engine 2021 VLDB 6.0130964e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
29 Database Architecture Optimized for the New Bottleneck: Memory Access 1999 VLDB 0.00052093615
74 Cache Conscious Algorithms for Relational Query Processing 1994 VLDB 0.00037330605
105 Quickly Generating Billion-Record Synthetic Databases 1994 SIGMOD 0.00033877899
152 Multiprocessor Hash-Based Join Algorithms 1985 VLDB 0.00029038365
215 AlphaSort: A RISC Machine Sort 1994 SIGMOD 0.00024507963
219 A Study of Index Structures for Main Memory Database Management Systems 1986 VLDB 0.00024293529
293 Implementing Database Operations Using SIMD Instructions 2002 SIGMOD 0.00022259273
305 GPUTeraSort: High Performance Graphics Co-processor Sorting for Large Database Management 2006 SIGMOD 0.00021872796
631 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015591241
632 Adaptive Aggregation on Chip Multiprocessors 2007 VLDB 0.00015575286
712 Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture 2008 VLDB 0.0001468812
987 What happens during a Join? Dissecting CPU and Memory Optimization Effects 2000 VLDB 0.00012814017
1,236 Handling Data Skew in Multiprocessor Database Computers Using Partition Tuning 1991 VLDB 0.00011548179
1,779 An Adaptive Hash Join Algorithm for Multiuser Environments 1990 VLDB 9.7764427e-05
2,470 Hash-Based Join Algorithms for Multiprocessor Computers with Shared Memory 1990 VLDB 8.5330174e-05
2,655 Executing Stream Joins on the Cell Processor 2007 VLDB 8.2888851e-05
2,679 Database Servers on Chip Multiprocessors: Limitations and Opportunities 2007 CIDR 8.2675008e-05
Previous Page 1 / 1 Next

Semantically Similar Papers