FPGA-based Multithreading for In-Memory Hash Joins
Summary: First end-to-end in-memory FPGA hash join using massive multithreading to hide memory latency during build/probe with hundreds of on-chip thread contexts. 2–3.4× faster than multicore at similar bandwidth (up to 76.8GB/s, 1.6B tps); benefit drops on extreme skew. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Robert J. Halstead (University of California Riverside)
- 2. Ildar Absalyamov (University of California Riverside)
- 3. Walid A. Najjar (University of California Riverside)
- 4. Vassilis J. Tsotras (University of California Riverside)
BibTeX Citation
@inproceedings{halstead_cidr15,
address = {Amsterdam, Netherlands},
series = {{CIDR} '15},
title = {{FPGA-based Multithreading for In-Memory Hash Joins}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Halstead, Robert J. and Absalyamov, Ildar and Najjar, Walid A. and Tsotras, Vassilis J.},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,626 | Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast Interconnects | 2022 | SIGMOD | 7.1524537e-05 |
| 4,931 | FPGA-based Data Partitioning | 2017 | SIGMOD | 6.348544e-05 |
| 7,964 | Parallelizing Intra-Window Join on Multicores: An Experimental Study | 2021 | SIGMOD | 5.4165494e-05 |
| 9,429 | Efficiently Joining Large Relations on Multi-GPU Systems | 2025 | VLDB | 5.1786456e-05 |
| 10,120 | Is FPGA Useful for Hash Joins? Exploring Hash Joins on Coupled CPU-FPGA Architecture | 2020 | CIDR | 5.0782186e-05 |
| 11,545 | SPID-Join: A Skew-resistant Processing-in-DIMM Join Algorithm Exploiting the Bank- and Rank-level Parallelisms of DIMMs | 2024 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 27 | Database Architecture Optimized for the New Bottleneck: Memory Access | 1999 | VLDB | 0.0005158963 |
| 76 | Cache Conscious Algorithms for Relational Query Processing | 1994 | VLDB | 0.00036898845 |
| 106 | Quickly Generating Billion-Record Synthetic Databases | 1994 | SIGMOD | 0.00033526937 |
| 210 | Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs | 2009 | VLDB | 0.00024851502 |
| 251 | Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited | 2014 | VLDB | 0.00023143736 |
| 361 | Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs | 2011 | SIGMOD | 0.00020006406 |
| 423 | Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems | 2012 | VLDB | 0.00018491327 |
| 1,387 | How Soccer Players Would do Stream Joins | 2011 | SIGMOD | 0.00010825457 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,754 | Performance Analysis of a Load Balancing Hash-Join Algorithm for a Shared Memory Multiprocessor | 1991 | VLDB |
| 2 | 251 | Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited | 2014 | VLDB |
| 3 | 4,710 | On Parallel Execution Of Multiple Pipelined Hash Joins | 1994 | SIGMOD |
| 4 | 2,122 | Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture | 2013 | VLDB |
| 5 | 156 | Multiprocessor Hash-Based Join Algorithms | 1985 | VLDB |
| 6 | 2,507 | Hash-Based Join Algorithms for Multiprocessor Computers with Shared Memory | 1990 | VLDB |
| 7 | 210 | Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs | 2009 | VLDB |
| 8 | 4,931 | FPGA-based Data Partitioning | 2017 | SIGMOD |
| 9 | 361 | Design and Evaluation of Main Memory Hash Join Algorithms for Multi-core CPUs | 2011 | SIGMOD |
| 10 | 10,120 | Is FPGA Useful for Hash Joins? Exploring Hash Joins on Coupled CPU-FPGA Architecture | 2020 | CIDR |