DBScholar

Back to papers

DiskJoin: Large-scale Vector Similarity Join with SSD

Summary: DiskJoin: first disk-based similarity join for billion-scale vectors on one machine, leveraging NVMe SSDs to avoid costly cluster communication. It minimizes read amplification via SSD-aware access, uses dynamic cache+eviction policies, and probabilistic pruning to achieve 50×–1000× speedups. (summarized by gpt-5-mini on Feb 11 2026)

Paper ID
he1ba73d5e3ce5a67
Venue
SIGMOD
Year
2026
Pagerank
5.1079647e-05
Overall Rank
9,912 | 33.38%
DOI
10.1145/3769780

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{chen_sigmod26,
        title = {{DiskJoin: Large-scale Vector Similarity Join with SSD}},
        author = {Chen, Yanqi and Yan, Xiao and Meliou, Alexandra and Lo, Eric},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3769780},
        url = {https://dl.acm.org/doi/10.1145/3769780},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
10,919 TEngineDB-V: An OLAP-Native Vector Search System for Large-k Workloads at Tencent 2026 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 17 of 17 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
56 M-tree: An Efficient Access Method for Similarity Search in Metric Spaces 1997 VLDB 0.00040363819
146 Efficient Processing of Spatial Joins Using R-trees 1993 SIGMOD 0.00029048509
168 Efficient Exact Set-Similarity Joins 2006 VLDB 0.00027151132
297 Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search 2016 VLDB 0.00021849337
338 Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting 2012 SIGMOD 0.00020600264
669 SharedDB: Killing One Thousand Queries With One Stone 2012 VLDB 0.00014975391
929 Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints 2008 VLDB 0.00013014029
1,438 Speedup Graph Processing by Graph Ordering 2016 SIGMOD 0.00010647473
1,603 Starling: An I/O-Efficient Disk-Resident Graph Index Framework for High-Dimensional Vector Similarity Search on Data Segment 2024 SIGMOD 0.00010100279
1,958 SeRF: Segment Graph for Range-Filtering Approximate Nearest Neighbor Search 2024 SIGMOD 9.3118048e-05
1,998 Epsilon Grid Order: An Algorithm for the Similarity Join on Massive High-Dimensional Data 2001 SIGMOD 9.2123795e-05
2,061 Similarity search in the blink of an eye with compressed indices 2023 VLDB 9.1002904e-05
2,577 ClusterJoin: A Similarity Joins Framework using Map-Reduce 2014 VLDB 8.2699584e-05
2,821 Practical and Asymptotically Optimal Quantization of High-Dimensional Vectors in Euclidean Space for Approximate Nearest Neighbor Search 2025 SIGMOD 7.9674226e-05
2,945 Spatio-Textual Similarity Joins 2013 VLDB 7.8244366e-05
3,256 Efficient Exact Edit Similarity Query Processing with the Asymmetric Signature Scheme 2011 SIGMOD 7.4855571e-05
5,620 A Topology-Aware Localized Update Strategy for Graph-Based ANN Index 2026 VLDB 6.0596841e-05
Previous Page 1 / 1 Next

Semantically Similar Papers