Faster Set Intersection with SIMD instructions by Reducing Branch Mispredictions
Summary: Presents a SIMD-accelerated set intersection for sorted arrays by extending the merge-based approach to read multiple elements per step and compare many pairs, reducing branch mispredictions. No preprocessing; can replace std::set_intersection; up to 5.2x with SIMD (2.1x without) on 32/64-bit ints, on Xeon and POWER7+. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Hiroshi Inoue
- 2. Moriyoshi Ohara
- 3. Kenjiro Taura
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 341 | EmptyHeaded: A Relational Engine for Graph Processing | 2016 | SIGMOD | 0.00026850764 |
| 1,973 | Speeding Up Set Intersections in Graph Algorithms using SIMD Instructions | 2018 | SIGMOD | 9.8834701e-05 |
| 4,554 | Distributed Subgraph Matching on Timely Dataflow | 2019 | VLDB | 6.0839934e-05 |
| 4,649 | SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures | 2015 | VLDB | 6.0171025e-05 |
| 7,693 | Processing and Optimizing Main Memory Spatial-Keyword Queries | 2016 | VLDB | 4.6714423e-05 |
| 8,160 | List Intersection for Web Search: Algorithms, Cost Models, and Optimizations | 2019 | VLDB | 4.5697316e-05 |
| 8,368 | Interleaved Multi-Vectorizing | 2020 | VLDB | 4.5295768e-05 |
| 9,177 | HERO: A Hierarchical Set Partitioning and Join Framework for Speeding up the Set Intersection Over Graphs | 2024 | SIGMOD | 4.3800806e-05 |
| 11,383 | Origami: A High-Performance Mergesort Framework | 2022 | VLDB | 4.1905499e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 343 | Implementing Database Operations Using SIMD Instructions | 2002 | SIGMOD | 0.00026756534 |
| 944 | Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture | 2008 | VLDB | 0.0001512998 |
| 1,121 | Improving the Performance of List Intersection | 2009 | VLDB | 0.00013838956 |
| 2,046 | Efficient Parallel Lists Intersection and Index Compression Algorithms using Graphics Processing Units | 2011 | VLDB | 9.6922338e-05 |
| 2,465 | Fast Set Intersection in Memory | 2011 | VLDB | 8.7344475e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,368 | Interleaved Multi-Vectorizing | 2020 | VLDB | 4.5295768e-05 |
| 8,160 | List Intersection for Web Search: Algorithms, Cost Models, and Optimizations | 2019 | VLDB | 4.5697316e-05 |
| 248 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00030888982 |
| 959 | Rethinking SIMD Vectorization for In-Memory Databases | 2015 | SIGMOD | 0.00015034808 |
| 4,649 | SIMD- and Cache-Friendly Algorithm for Sorting an Array of Structures | 2015 | VLDB | 6.0171025e-05 |
| 944 | Efficient Implementation of Sorting on Multi-Core SIMD CPU Architecture | 2008 | VLDB | 0.0001512998 |
| 1,121 | Improving the Performance of List Intersection | 2009 | VLDB | 0.00013838956 |
| 343 | Implementing Database Operations Using SIMD Instructions | 2002 | SIGMOD | 0.00026756534 |
| 2,465 | Fast Set Intersection in Memory | 2011 | VLDB | 8.7344475e-05 |
| 1,973 | Speeding Up Set Intersections in Graph Algorithms using SIMD Instructions | 2018 | SIGMOD | 9.8834701e-05 |