DBScholar

Back to papers

Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs

Summary: Revisits hash join versus sort-merge join with highly optimized multicore implementations, achieving record CPU throughput and robust performance under skew and varying input sizes. Models predict wider SIMD, more cores, and bandwidth limits will soon favor sort-merge joins. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hd002d9f853e65366
Venue
VLDB
Year
2009
Pagerank
0.00024851502
Overall Rank
210 | 98.59%
DOI
10.14778/1687553.1687564

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{kim_vldb09,
        title = {{Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs}},
        author = {Kim, Changkyu and Kaldewey, Tim and Lee, Victor W. and Sedlar, Eric and Nguyen, Anthony D. and Satish, Nadathur and Chhugani, Jatin and Di Blas, Andrea and Dubey, Pradeep},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687553.1687564},
        url = {https://doi.org/10.14778/1687553.1687564},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 81 citing papers.

Rank Citing Paper Year Venue Pagerank
215 Morsel-Driven Parallelism: A NUMA-Aware Query Evaluation Framework for the Many-Core Age 2014 SIGMOD 0.00024598661
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
282 FAST: Fast Architecture Sensitive Tree Search on Modern CPUs and GPUs 2010 SIGMOD 0.00022264207
361 Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs 2011 SIGMOD 0.00020006406
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018491327
627 Rethinking SIMD Vectorization for In-Memory Databases 2015 SIGMOD 0.00015460957
661 Fast Sort on CPUs and GPUs: A Case for Bandwidth Oblivious SIMD Sort 2010 SIGMOD 0.00015003815
667 SharedDB: Killing One Thousand Queries With One Stone 2012 VLDB 0.00014978213
713 Cache locality is not enough: High-Performance Nearest Neighbor Search with Product Quantization Fast Scan 2016 VLDB 0.00014571977
856 Hardware-Oblivious Parallelism for In-Memory Column-Stores 2013 VLDB 0.00013428547
889 Rack-Scale In-Memory Join Processing using RDMA 2015 SIGMOD 0.00013247362
969 Memory-Efficient Hash Joins 2015 VLDB 0.0001278184
1,116 A Comprehensive Study of Main-Memory Partitioning and its Application to Large-Scale Comparison- and Radix-Sort 2014 SIGMOD 0.00011962096
1,167 DimmWitted: A Study of Main-Memory Statistical Analytics 2014 VLDB 0.00011729888
1,266 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011269175
1,335 Fast Updates on Read-Optimized Databases Using Multi-Core CPUs 2012 VLDB 0.0001099401
1,723 ByteSlice: Pushing the Envelop of Main Memory Data Processing with a New Storage Layout 2015 SIGMOD 9.7931223e-05
1,812 Joins via Geometric Resolutions: Worst-case and Beyond 2015 PODS 9.5803973e-05
1,995 Track Join: Distributed Joins with Minimal Network Traffic 2014 SIGMOD 9.2169073e-05
2,122 Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture 2013 VLDB 9.0084047e-05
2,245 Cache-Efficient Aggregation: Hashing Is Sorting 2015 SIGMOD 8.7649358e-05
2,284 Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects 2020 SIGMOD 8.6954168e-05
2,362 Streaming Similarity Search over one Billion Tweets using Parallel Locality-Sensitive Hashing 2013 VLDB 8.5746436e-05
2,460 Query Processing on Tensor Computation Runtimes 2022 VLDB 8.4348335e-05
2,487 A Memory Bandwidth-Efficient Hybrid Radix Sort on GPUs 2017 SIGMOD 8.3979719e-05
2,530 Adaptive and Big Data Scale Parallel Execution in Oracle 2013 VLDB 8.3391175e-05
2,673 Asynchronous Memory Access Chaining 2016 VLDB 8.1482775e-05
2,802 Distributed Join Algorithms on Thousands of Cores 2017 VLDB 7.9903139e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9739791e-05
2,886 Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment 2021 VLDB 7.9081605e-05
3,033 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.7337998e-05
3,393 CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster 2012 SIGMOD 7.3471344e-05
3,459 Improving Main Memory Hash Joins on Intel Xeon Phi Processors: An Experimental Approach 2015 VLDB 7.2848113e-05
3,491 In-Cache Query Co-Processing on Coupled CPU-GPU Architectures 2015 VLDB 7.2619913e-05
3,626 Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects 2022 SIGMOD 7.1524537e-05
4,011 Deployment of Query Plans on Multicores 2015 VLDB 6.8573269e-05
4,137 The Art of Balance: A RateupDB Experience of Building a CPU/GPU Hybrid Database Product 2021 VLDB 6.7861661e-05
4,225 SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures 2015 VLDB 6.7207703e-05
4,369 Predicate Transfer: Efficient Pre-Filtering on Multi-Join Queries 2024 CIDR 6.6315141e-05
4,539 OmniDB: Towards Portable and Efficient Query Processing on Parallel CPU/GPU Architectures 2013 VLDB 6.5532239e-05
4,730 MQJoin: Efficient Shared Execution of Main-Memory Joins 2016 VLDB 6.4468007e-05
4,931 FPGA-based Data Partitioning 2017 SIGMOD 6.348544e-05
5,131 Holistic Indexing in Main-memory Column-stores 2015 SIGMOD 6.2602133e-05
5,142 The 3D Hash Join: Building On Non-Unique Join Attributes 2022 CIDR 6.2571095e-05
5,496 Density-optimized Intersection-free Mapping and Matrix Multiplication for Join-Project Operations 2022 VLDB 6.1049663e-05
5,543 Charting the Design Space of Query Execution using VOILA 2021 VLDB 6.0889645e-05
5,598 Database Processing-in-Memory: An Experimental Study 2020 VLDB 6.0713144e-05
5,628 FPGA-based Multithreading for In-Memory Hash Joins 2015 CIDR 6.060354e-05
5,771 Database Technology for the Masses: Sub-Operators as First-Class Entities 2021 VLDB 5.9992698e-05
5,827 Design and Evaluation of Storage Organizations for Read-Optimized Main Memory Databases 2013 VLDB 5.979772e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
27 Database Architecture Optimized for the New Bottleneck: Memory Access 1999 VLDB 0.0005158963
76 Cache Conscious Algorithms for Relational Query Processing 1994 VLDB 0.00036898845
106 Quickly Generating Billion-Record Synthetic Databases 1994 SIGMOD 0.00033526937
156 Multiprocessor Hash-Based Join Algorithms 1985 VLDB 0.00028522117
223 AlphaSort: A RISC Machine Sort 1994 SIGMOD 0.0002412513
229 A Study of Index Structures for Main Memory Database Management Systems 1986 VLDB 0.00023911856
287 Implementing Database Operations Using SIMD Instructions 2002 SIGMOD 0.00021970198
304 GPUTeraSort: High Performance Graphics Co-processor Sorting for Large Database Management 2006 SIGMOD 0.00021604795
616 Relational Joins on Graphics Processors 2008 SIGMOD 0.00015561564
626 Adaptive Aggregation on Chip Multiprocessors 2007 VLDB 0.00015473276
722 Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture 2008 VLDB 0.00014488003
998 What happens during a Join? Dissecting CPU and Memory Optimization Effects 2000 VLDB 0.00012630367
1,254 Handling Data Skew in Multiprocessor Database Computers Using Partition Tuning 1991 VLDB 0.00011330673
1,799 An Adaptive Hash Join Algorithm for Multiuser Environments 1990 VLDB 9.6155018e-05
2,507 Hash-Based Join Algorithms for Multiprocessor Computers with Shared Memory 1990 VLDB 8.3723695e-05
2,696 Executing Stream Joins on the Cell Processor 2007 VLDB 8.1205649e-05
2,699 Database Servers on Chip Multiprocessors: Limitations and Opportunities 2007 CIDR 8.1195744e-05
Previous Page 1 / 1 Next

Semantically Similar Papers