LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH
Summary: LSHAlign solves all-pair near-duplicate alignment by grouping subsequences sharing LSH signatures and representing each group in O(1) space. It achieves expected O((|T|+|S|)mL) time/space, excluding output, with order-of-magnitude speedups. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yuheng Zhang (Rutgers University)
- 2. Zhencan Peng (Rutgers University)
- 3. Miao Qiao (University of Auckland)
- 4. Wei Zhang (Alibaba)
- 5. Feifei Li (Alibaba)
- 6. Dong Deng (Rutgers University)
BibTeX Citation
@inproceedings{zhang_sigmod26,
title = {{LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH}},
author = {Zhang, Yuheng and Peng, Zhencan and Qiao, Miao and Zhang, Wei and Li, Feifei and Deng, Dong},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802083},
url = {https://dl.acm.org/doi/10.1145/3802083},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 287 | Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search | 2007 | VLDB |
| 2 | 7,692 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD |
| 3 | 4,052 | Local Similarity Search for Unstructured Text | 2016 | SIGMOD |
| 4 | 8,613 | Bidirectionally Densifying LSH Sketches with Empty Bins | 2021 | SIGMOD |
| 5 | 4,857 | Similarity Join Size Estimation using Locality Sensitive Hashing | 2011 | VLDB |
| 6 | 2,390 | Streaming Similarity Search over one Billion Tweets using Parallel Locality-Sensitive Hashing | 2013 | VLDB |
| 7 | 10,554 | Near-Duplicate Text Alignment under Weighted Jaccard Similarity | 2026 | VLDB |
| 8 | 8,522 | TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection | 2022 | SIGMOD |
| 9 | 7,745 | Near-Duplicate Text Alignment with One Permutation Hashing | 2024 | SIGMOD |
| 10 | 7,676 | Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts | 2021 | SIGMOD |