BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector Databases
Summary: BigVectorBench introduces an integrated benchmark for heterogeneous data embedding and compound (multimodal/constrained) queries—capabilities missing from existing vector DB evaluations. Experiments expose bottlenecks in mainstream systems and validate these workload choices. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Guoxin Kang (Chinese Academy of Sciences)
- 2. Zhongxin Ge (Chinese Academy of Sciences)
- 3. Jingpei Hu (Chinese Academy of Sciences)
- 4. Xueya Zhang (University of Chinese Academy of Sciences)
- 5. Lei Wang (Chinese Academy of Sciences)
- 6. Jianfeng Zhan (Chinese Academy of Sciences)
BibTeX Citation
@article{kang_vldb25,
title = {{BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector Databases}},
author = {Kang, Guoxin and Ge, Zhongxin and Hu, Jingpei and Zhang, Xueya and Wang, Lei and Zhan, Jianfeng},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {5},
pages = {1536--1550},
doi = {10.14778/3718057.3718078},
url = {https://doi.org/10.14778/3718057.3718078},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,193 | An In-Depth Study of Filter-Agnostic Vector Search on a PostgreSQL Database System: [Experiments & Analysis] | 2026 | SIGMOD | 5.093636e-05 |
| 10,304 | VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis] | 2026 | SIGMOD | 5.093636e-05 |
| 10,493 | Reveal Hidden Pitfalls and Navigate Next Generation of Vector Similarity Search from Task-Centric Views: [Experiments & Analysis] | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 93 | Fast Approximate Nearest Neighbor Search With The Navigating Spreading-out Graph | 2019 | VLDB | 0.00034701237 |
| 286 | Milvus: A Purpose-Built Vector Data Management System | 2021 | SIGMOD | 0.00022357911 |
| 705 | HD-Index: Pushing the Scalability-Accuracy Boundary for Approximate kNN Search in High-Dimensional Spaces | 2018 | VLDB | 0.00014829964 |
| 1,402 | Creating Embeddings of Heterogeneous Relational Datasets for Data Integration Tasks | 2020 | SIGMOD | 0.00010888094 |
| 1,589 | Manu: A Cloud Native Vector Database Management System | 2022 | VLDB | 0.00010264469 |
| 4,886 | Chat2Data: An Interactive Data Analysis System with RAG, Vector Databases and LLMs | 2024 | VLDB | 6.4614955e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 4,025 | Beyond Macrobenchmarks: Microbenchmark-based Graph Database Evaluation | 2019 | VLDB |
| 2 | 1,631 | High-Throughput Vector Similarity Search in Knowledge Graphs | 2023 | SIGMOD |
| 3 | 1,693 | BigBench: Towards an Industry Standard Benchmark for Big Data Analytics | 2013 | SIGMOD |
| 4 | 10,449 | Efficient Vector Index Merging in Vector Databases | 2026 | SIGMOD |
| 5 | 10,381 | Integrating Vector Databases across Embedding Models | 2026 | SIGMOD |
| 6 | 10,493 | Reveal Hidden Pitfalls and Navigate Next Generation of Vector Similarity Search from Task-Centric Views: [Experiments & Analysis] | 2026 | SIGMOD |
| 7 | 3,240 | Graph-Based Vector Search: An Experimental Evaluation of the State-of-the-Art | 2025 | SIGMOD |
| 8 | 2,920 | SingleStore-V: An Integrated Vector Database System in SingleStore | 2024 | VLDB |
| 9 | 10,571 | An Experimental Evaluation of Hybrid Querying on Vectors | 2026 | VLDB |
| 10 | 10,304 | VecBench: A Controllable Benchmark for Filtered Vector Search: [Experiments & Analysis] | 2026 | SIGMOD |