Back to papers
NUMA-aware algorithms: the case of data shuffling
Summary: Demonstrates that NUMA effects critically impact data shuffling on multi-socket multicore servers, with naive shuffling up to 3× slower than NUMA-aware variants. Achieves top performance using thread binding, NUMA-aware thread allocation, and relaxed global coordination, arguing such algorithmic redesign is essential as socket counts and memory heterogeneity grow.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 186
- Venue
- CIDR
- Year
- 2013
- Pagerank
- 0.00011451745
- Overall Rank
- 1,540 | 89.30%
- DOI
-
-
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 27 of 27 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 241 |
DB2 with BLU Acceleration: So Much More than Just a Column Store |
2013 |
VLDB |
0.00031314629 |
| 403 |
Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited |
2014 |
VLDB |
0.00024176677 |
| 417 |
Morsel-Driven Parallelism: A NUMA-Aware Query Evaluation Framework for the Many-Core Age |
2014 |
SIGMOD |
0.00023734582 |
| 1,016 |
Memory-Efficient Hash Joins |
2015 |
VLDB |
0.00014630024 |
| 1,044 |
DimmWitted: A Study of Main-Memory Statistical Analytics |
2014 |
VLDB |
0.00014465007 |
| 1,397 |
High-Speed Query Processing over High-Speed Networks |
2016 |
VLDB |
0.00012208385 |
| 1,610 |
A Comprehensive Study of Main-Memory Partitioning and its Application to Large-Scale Comparison- and Radix-Sort |
2014 |
SIGMOD |
0.00011155922 |
| 2,420 |
Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud Infrastructure |
2020 |
SIGMOD |
8.8474951e-05 |
| 2,523 |
Revisiting Co-Processing for Hash Joins on the Coupled CPU-GPU Architecture |
2013 |
VLDB |
8.599693e-05 |
| 3,328 |
Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects |
2020 |
SIGMOD |
7.2136181e-05 |
| 4,226 |
Hyper Dimension Shuffle: Efficient Data Repartition at Petabyte Scale in SCOPE |
2019 |
VLDB |
6.3382156e-05 |
| 4,277 |
Scaling Up Concurrent Main-Memory Column-Store Scans: Towards Adaptive NUMA-aware Data and Task Placement |
2015 |
VLDB |
6.2878362e-05 |
| 4,609 |
Deployment of Query Plans on Multicores |
2015 |
VLDB |
6.0458679e-05 |
| 5,111 |
Adaptive NUMA-aware data placement and task scheduling for analytical workloads in main-memory column-stores |
2017 |
VLDB |
5.6855393e-05 |
| 5,670 |
BriskStream: Scaling Data Stream Processing on Shared-Memory Multicore Architectures |
2019 |
SIGMOD |
5.3799054e-05 |
| 5,869 |
Low-Latency Handshake Join |
2014 |
VLDB |
5.2918426e-05 |
| 5,871 |
Taming Subgraph Isomorphism for RDF Query Processing |
2015 |
VLDB |
5.2912806e-05 |
| 6,649 |
Grizzly: Efficient Stream Processing Through Adaptive Query Compilation |
2020 |
SIGMOD |
4.9724735e-05 |
| 7,864 |
Operational Analytics Data Management Systems |
2016 |
VLDB |
4.6284093e-05 |
| 7,917 |
Terabyte-Scale Analytics in the Blink of an Eye |
2026 |
VLDB |
4.6129625e-05 |
| 8,411 |
The Case for Learned In-Memory Joins |
2023 |
VLDB |
4.5151296e-05 |
| 8,511 |
CXL Memory Performance for In-Memory Data Processing |
2025 |
VLDB |
4.4904708e-05 |
| 9,068 |
How to Stop Under-Utilization and Love Multicores |
2014 |
SIGMOD |
4.398897e-05 |
| 9,822 |
Thriving in the No Man’s Land between Compilers and Databases |
2019 |
CIDR |
4.2713516e-05 |
| 10,190 |
P-MOSS: Scheduling Main-Memory Indexes Over NUMA Servers Using Next Token Prediction |
2026 |
SIGMOD |
4.1905499e-05 |
| 11,157 |
Templating Shuffles |
2023 |
CIDR |
4.1905499e-05 |
| 12,070 |
Next Generation Data Analytics at IBM Research |
2013 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers