Back to papers
Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts
Summary: Allign uses a min-hash based method to align all-pair near-duplicate passages in two texts, avoiding O(n^2 m^2) enumeration via compact windows. It matches windows by shared min-hash, reports the longest and sentence-level near-duplicates, and outperforms prior alignment methods on real data.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6238
- Venue
- SIGMOD
- Year
- 2021
- Pagerank
- 4.6863871e-05
- Overall Rank
- 7,636 | 46.93%
- DOI
-
10.1145/3448016.3457548
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 34 |
Similarity Search in High Dimensions via Hashing |
1999 |
VLDB |
0.00076824554 |
| 614 |
Copy Detection Mechanisms for Digital Documents |
1995 |
SIGMOD |
0.00019088537 |
| 703 |
Winnowing: Local Algorithms for Document Fingerprinting |
2003 |
SIGMOD |
0.0001784726 |
| 1,299 |
Bayesian Locality Sensitive Hashing for Fast Similarity Search |
2012 |
VLDB |
0.00012712766 |
| 1,396 |
Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search |
2012 |
SIGMOD |
0.00012215253 |
| 2,588 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
8.4872437e-05 |
| 3,583 |
Efficient Approximate Entity Extraction with Edit Distance Constraints |
2009 |
SIGMOD |
6.944299e-05 |
| 4,041 |
An Efficient Partition Based Method for Exact Set Similarity Joins |
2016 |
VLDB |
6.5048916e-05 |
| 4,247 |
Local Similarity Search for Unstructured Text |
2016 |
SIGMOD |
6.3180334e-05 |
| 4,350 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
6.2576191e-05 |
| 4,808 |
On the Complexity of Inner Product Similarity Join |
2016 |
PODS |
5.9040739e-05 |
| 5,071 |
Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction |
2011 |
SIGMOD |
5.7122974e-05 |
| 6,080 |
Pigeonring: A Principle for Faster Thresholded Similarity Search |
2019 |
VLDB |
5.219249e-05 |
| 6,730 |
A Pivotal Prefix Based Filtering Algorithm for String Similarity Search |
2014 |
SIGMOD |
4.9436522e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 8,037 |
A New Approach for Processing Ranked Subsequence Matching Based on Ranked Union |
2011 |
SIGMOD |
4.5965282e-05 |
| 1,232 |
Ed-Join: An Efficient Algorithm for Similarity Joins With Edit Distance Constraints |
2008 |
VLDB |
0.00013133604 |
| 9,875 |
Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation |
2023 |
SIGMOD |
4.2626861e-05 |
| 9,934 |
Efficient and Effective KNN Sequence Search with Approximate n-grams |
2014 |
VLDB |
4.245954e-05 |
| 2,588 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
8.4872437e-05 |
| 7,707 |
Efficient Top-k Algorithms for Approximate Substring Matching |
2013 |
SIGMOD |
4.6676985e-05 |
| 4,247 |
Local Similarity Search for Unstructured Text |
2016 |
SIGMOD |
6.3180334e-05 |
| 10,266 |
Near-Duplicate Text Alignment under Weighted Jaccard Similarity |
2026 |
VLDB |
4.1905499e-05 |
| 7,698 |
Near-Duplicate Text Alignment with One Permutation Hashing |
2024 |
SIGMOD |
4.6699546e-05 |
| 8,284 |
TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection |
2022 |
SIGMOD |
4.5392079e-05 |