DBScholar

Back to papers

LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH

Summary: LSHAlign solves all-pair near-duplicate alignment by grouping subsequences sharing LSH signatures and representing each group in O(1) space. It achieves expected O((|T|+|S|)mL) time/space, excluding output, with order-of-magnitude speedups. (summarized by gpt-5.6-luna on Jul 26 2026)

Paper ID
h04ce17cb6ac5c5c8
Venue
SIGMOD
Year
2026
Pagerank
4.9793485e-05
Overall Rank
10,478 | 29.56%
DOI
10.1145/3802083

Incoming Non-self Citations Over Time

No non-self incoming citations found for this paper in this database.

Authors

BibTeX Citation

@inproceedings{zhang_sigmod26,
        title = {{LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH}},
        author = {Zhang, Yuheng and Peng, Zhencan and Qiao, Miao and Zhang, Wei and Li, Feifei and Deng, Dong},
        series = {{SIGMOD} '26},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3802083},
        url = {https://dl.acm.org/doi/10.1145/3802083},
        year = {2026}
}

Incoming Citations (Sorted by Pagerank)

Showing 0 of 0 citing papers.

Rank Citing Paper Year Venue Pagerank
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
20 Similarity Search in High Dimensions via Hashing 1999 VLDB 0.00057568153
280 Multi-Probe LSH: Efficient Indexing for High-Dimensional Similarity Search 2007 VLDB 0.0002230467
298 Query-Aware Locality-Sensitive Hashing for Approximate Nearest Neighbor Search 2016 VLDB 0.00021833987
338 Locality-Sensitive Hashing Scheme Based on Dynamic Collision Counting 2012 SIGMOD 0.00020585187
562 SRS: Solving c-Approximate Nearest Neighbor Queries in High Dimensional Euclidean Space with a Tiny Index 2015 VLDB 0.00016335405
576 Quality and Efficiency in High Dimensional Nearest Neighbor Search 2009 SIGMOD 0.00016121388
694 JOSIE: Overlap Set Similarity Search for Finding Joinable Tables in Data Lakes 2019 SIGMOD 0.00014727089
1,001 Winnowing: Local Algorithms for Document Fingerprinting 2003 SIGMOD 0.00012599856
1,306 Copy Detection Mechanisms for Digital Documents 1995 SIGMOD 0.00011084152
1,318 LSH Ensemble: Internet-Scale Domain Search 2016 VLDB 0.00011047393
7,833 Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts 2021 SIGMOD 5.443666e-05
7,906 Near-Duplicate Text Alignment with One Permutation Hashing 2024 SIGMOD 5.428873e-05
8,689 TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection 2022 SIGMOD 5.2905577e-05
10,214 Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation 2023 SIGMOD 5.0596605e-05
10,736 Near-Duplicate Text Alignment under Weighted Jaccard Similarity 2026 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Semantically Similar Papers