RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
Summary: RSR-core turns Redundant Segment Reduction for 1-/1.58-bit matrix-vector multiplication into optimized CPU/CUDA kernels, enabling practical binary/ternary LLM inference. HuggingFace integration delivers up to 62× CPU and 1.9× CUDA token-generation speedups. (summarized by gpt-5.6-luna on Aug 28 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Mohsen Dehghankar (University of Illinois Chicago)
- 2. Abolfazl Asudeh (University of Illinois Chicago)
BibTeX Citation
@article{dehghankar_vldb26,
title = {{RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication}},
author = {Dehghankar, Mohsen and Asudeh, Abolfazl},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {12},
pages = {4726--4729},
doi = {10.14778/3827998.3828107},
url = {https://doi.org/10.14778/3827998.3828107},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 4,715 | Maximum Inner Product is Query-Scaled Nearest Neighbor | 2025 | VLDB | 6.4554594e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,289 | Tensor Relational Algebra for Distributed Machine Learning System Design | 2021 | VLDB |
| 2 | 3,972 | TCUDB: Accelerating Database with Tensor Processors | 2022 | SIGMOD |
| 3 | 10,776 | Efficient Cooperation-Aware Key and Value Management for LLM Inference | 2026 | VLDB |
| 4 | 9,649 | GaussDB-Vector: A Large-Scale Persistent Real-Time Vector Database for LLM Applications | 2025 | VLDB |
| 5 | 10,711 | BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMs | 2026 | VLDB |
| 6 | 10,146 | QStore: Quantization-Aware Compressed Model Storage | 2026 | VLDB |
| 7 | 10,602 | TranSQL+: Serving Large Language Models with SQL on Low-Resource Hardware | 2026 | SIGMOD |
| 8 | 4,351 | Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity | 2024 | VLDB |
| 9 | 10,838 | Unified Static–Dynamic Pruning for Efficient LLM Inference | 2026 | VLDB |
| 10 | 6,560 | RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference | 2026 | VLDB |