DBScholar

Back to papers

DiskJoin: Large-scale Vector Similarity Join with SSD

Summary: DiskJoin: first disk-based similarity join for billion-scale vectors on one machine, leveraging NVMe SSDs to avoid costly cluster communication. It minimizes read amplification via SSD-aware access, uses dynamic cache+eviction policies, and probabilistic pruning to achieve 50×–1000× speedups. (summarized by gpt-5-mini on Feb 11 2026)

Paper ID
he1ba73d5e3ce5a67
Venue
SIGMOD
Year
2026
Pagerank
5.1103839e-05
Overall Rank
9,905 | 33.41%
DOI
10.1145/3769780

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{chen_sigmod26,
        title = {{DiskJoin: Large-scale Vector Similarity Join with SSD}},
        author = {Chen, Yanqi and Yan, Xiao and Meliou, Alexandra and Lo, Eric},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3769780},
        url = {https://dl.acm.org/doi/10.1145/3769780},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
10,910 TEngineDB-V: An OLAP-Native Vector Search System for Large-k Workloads at Tencent 2026 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
56 M-tree: An Efficient Access Method for Similarity Search in Metric Spaces 1997 VLDB 0.00040370171
146 Efficient Processing of Spatial Joins Using R-trees 1993 SIGMOD 0.00029061754
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027163517
298 Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search 2016 VLDB 0.00021833987
338 Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting 2012 SIGMOD 0.00020585187
667 SharedDB: Killing One Thousand Queries With One Stone 2012 VLDB 0.00014978213
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013020115
1,440 Speedup Graph Processing by Graph Ordering 2016 SIGMOD 0.00010641888
1,613 Starling: An I/O-Efficient Disk-Resident Graph Index Framework for High-Dimensional Vector Similarity Search on Data Segment 2024 SIGMOD 0.00010072237
1,957 SeRF: Segment Graph for Range-Filtering Approximate Nearest Neighbor Search 2024 SIGMOD 9.3159589e-05
1,996 Epsilon Grid Order: An Algorithm for the Similarity Join on Massive High-Dimensional Data 2001 SIGMOD 9.216723e-05
2,060 Similarity search in the blink of an eye with compressed indices 2023 VLDB 9.0983169e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2738285e-05
2,821 Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search 2025 SIGMOD 7.9711961e-05
2,944 Spatio-Textual Similarity Joins 2013 VLDB 7.828107e-05
3,255 Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme 2011 SIGMOD 7.4890984e-05
5,619 A Topology-Aware Localized Update Strategy for Graph-Based ANN Index 2026 VLDB 6.062554e-05
Previous Page 1 / 1 Next

Semantically Similar Papers