Back to papers
Experimental Analysis of Large-scale Learnable Vector Storage Compression
Summary: Taxonomy and comprehensive benchmark of 14 embedding-compression methods for large-scale learnable vectors using a uniform testbed. Quantifies memory–quality trade-offs, recommends per-use-case winners, and exposes method limitations and research gaps.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13757
- Venue
- VLDB
- Year
- 2024
- Pagerank
- 4.3399748e-05
- Overall Rank
- 9,414 | 34.58%
- DOI
-
10.14778/3636218.3636234
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 740 |
Distributed Representations of Tuples for Entity Resolution |
2018 |
VLDB |
0.00017358024 |
| 973 |
Natural language to SQL: Where are we today? |
2020 |
VLDB |
0.0001488435 |
| 1,368 |
SlimDB: A Space-Efficient Key-Value Storage Engine For Semi-Sorted Data |
2017 |
VLDB |
0.0001235708 |
| 2,157 |
MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis |
2018 |
SIGMOD |
9.4153917e-05 |
| 2,264 |
Manu: A Cloud Native Vector Database Management System |
2022 |
VLDB |
9.1587362e-05 |
| 2,678 |
HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework |
2022 |
VLDB |
8.3224016e-05 |
| 3,167 |
QueryFormer: A Tree Transformer Model for Query Plan Representation |
2022 |
VLDB |
7.4561078e-05 |
| 3,492 |
Fauce: Fast and Accurate Deep Ensembles with Uncertainty for Cardinality Estimation |
2021 |
VLDB |
7.0435484e-05 |
| 3,805 |
Scaling Attributed Network Embedding to Massive Graphs |
2021 |
VLDB |
6.7485579e-05 |
| 4,464 |
LOGER: A Learned Optimizer towards Generating Efficient and Robust Query Execution Plans |
2023 |
VLDB |
6.1552798e-05 |
| 5,169 |
HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training |
2022 |
SIGMOD |
5.642415e-05 |
| 5,383 |
Parallel Training of Knowledge Graph Embedding Models: A Comparison of Techniques |
2022 |
VLDB |
5.5357645e-05 |
| 6,754 |
Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising Systems |
2021 |
SIGMOD |
4.9344367e-05 |
| 7,055 |
Serving Deep Learning Models with Deduplication from Relational Databases |
2022 |
VLDB |
4.8428976e-05 |
| 7,253 |
Effective and Efficient Retrieval of Structured Entities |
2020 |
VLDB |
4.7825469e-05 |
| 7,473 |
Cardinality Estimation of Approximate Substring Queries using Deep Learning |
2022 |
VLDB |
4.7149077e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 8,655 |
Improving Matrix-vector Multiplication via Lossless Grammar-Compressed Matrices |
2022 |
VLDB |
4.4687769e-05 |
| 4,622 |
Graph-Based Vector Search: An Experimental Evaluation of the State-of-the-Art |
2025 |
SIGMOD |
6.0356382e-05 |
| 2,821 |
An Experimental Study of Bitmap Compression vs. Inverted List Compression |
2017 |
SIGMOD |
8.0639288e-05 |
| 11,578 |
An Evaluation of Methods of Compressing Doubles |
2020 |
SIGMOD |
4.1905499e-05 |
| 3,741 |
DeepSqueeze: Deep Semantic Compression for Tabular Data |
2020 |
SIGMOD |
6.7952067e-05 |
| 7,431 |
CompressDB: Enabling Efficient Compressed Data Direct Processing for Various Databases |
2022 |
SIGMOD |
4.7274757e-05 |
| 10,748 |
Beyond Compression: A Comprehensive Evaluation of Lossless Floating-Point Compression |
2025 |
VLDB |
4.1905499e-05 |
| 9,408 |
CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models |
2024 |
SIGMOD |
4.3399748e-05 |
| 3,541 |
Similarity search in the blink of an eye with compressed indices |
2023 |
VLDB |
6.9910982e-05 |
| 1,970 |
Compressed Linear Algebra for Large-Scale Machine Learning |
2016 |
VLDB |
9.9024431e-05 |