Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts
Summary: Allign uses a min-hash based method to align all-pair near-duplicate passages in two texts, avoiding O(n^2 m^2) enumeration via compact windows. It matches windows by shared min-hash, reports the longest and sentence-level near-duplicates, and outperforms prior alignment methods on real data. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Weiqi Feng (Shanghai Jiao Tong University)
- 2. Dong Deng (Rutgers University)
BibTeX Citation
@inproceedings{feng_sigmod21,
title = {{Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts}},
author = {Feng, Weiqi and Deng, Dong},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457548},
url = {https://dl.acm.org/doi/10.1145/3448016.3457548},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 7 of 7 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 7,745 | Near-Duplicate Text Alignment with One Permutation Hashing | 2024 | SIGMOD | 5.5534781e-05 |
| 8,522 | TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection | 2022 | SIGMOD | 5.4119882e-05 |
| 9,065 | R2D2: Reducing Redundancy and Duplication in Data Lakes | 2023 | SIGMOD | 5.3251649e-05 |
| 10,024 | Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation | 2023 | SIGMOD | 5.1757914e-05 |
| 10,265 | LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH | 2026 | SIGMOD | 5.093636e-05 |
| 10,533 | SeDA: Bridging the Gap between Efficient Syntactic and Precise Semantic Search of Similar Passages in Large Text Corpora | 2026 | VLDB | 5.093636e-05 |
| 10,554 | Near-Duplicate Text Alignment under Weighted Jaccard Similarity | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 911 | Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints | 2008 | VLDB |
| 2 | 10,024 | Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation | 2023 | SIGMOD |
| 3 | 10,086 | Efficient and Effective KNN Sequence Search with Approximate n-grams | 2014 | VLDB |
| 4 | 1,886 | Pass-Join: A Partition-based Method for Similarity Joins | 2012 | VLDB |
| 5 | 7,692 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD |
| 6 | 4,052 | Local Similarity Search for Unstructured Text | 2016 | SIGMOD |
| 7 | 10,554 | Near-Duplicate Text Alignment under Weighted Jaccard Similarity | 2026 | VLDB |
| 8 | 7,745 | Near-Duplicate Text Alignment with One Permutation Hashing | 2024 | SIGMOD |
| 9 | 8,522 | TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection | 2022 | SIGMOD |
| 10 | 10,265 | LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH | 2026 | SIGMOD |