Near-Duplicate Text Alignment under Weighted Jaccard Similarity
Summary: MonoActive enables weighted-Jaccard near-duplicate substring alignment via consistent weighted sampling, beyond heuristic or unweighted min-hash indexing. For raw-count TF, it is provably optimal in group complexity, while improving speed up to 4.7× and reducing index size 30%. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yuheng Zhang (Rutgers University)
- 2. Miao Qiao (University of Auckland)
- 3. Zhencan Peng (Rutgers University)
- 4. Dong Deng (Rutgers University)
BibTeX Citation
@article{zhang_vldb26,
title = {{Near-Duplicate Text Alignment under Weighted Jaccard Similarity}},
author = {Zhang, Yuheng and Qiao, Miao and Peng, Zhencan and Deng, Dong},
journal = {PVLDB},
series = {{VLDB} '26},
volume = {19},
number = {7},
pages = {1600--1613},
doi = {10.14778/3801059.3801072},
url = {https://doi.org/10.14778/3801059.3801072},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,478 | LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH | 2026 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,001 | Winnowing: Local Algorithms for Document Fingerprinting | 2003 | SIGMOD | 0.00012599856 |
| 1,306 | Copy Detection Mechanisms for Digital Documents | 1995 | SIGMOD | 0.00011084152 |
| 4,147 | Local Similarity Search for Unstructured Text | 2016 | SIGMOD | 6.7803631e-05 |
| 7,833 | Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts | 2021 | SIGMOD | 5.443666e-05 |
| 7,906 | Near-Duplicate Text Alignment with One Permutation Hashing | 2024 | SIGMOD | 5.428873e-05 |
| 8,689 | TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection | 2022 | SIGMOD | 5.2905577e-05 |
| 10,214 | Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation | 2023 | SIGMOD | 5.0596605e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,973 | A Scalable Index for Top-k Subtree Similarity Queries | 2019 | SIGMOD |
| 2 | 2,354 | Hashed Samples: Selectivity Estimators For Set Similarity Selection Queries | 2008 | VLDB |
| 3 | 10,635 | Efficient and Robust Out-Of-Distribution Vector Similarity Search with Cross-Distribution Monotonic Graph | 2026 | SIGMOD |
| 4 | 11,813 | TokenJoin: Efficient Filtering for Set Similarity Join with Maximum Weighted Bipartite Matching | 2023 | VLDB |
| 5 | 7,842 | Efficient Top-k Algorithms for Approximate Substring Matching | 2013 | SIGMOD |
| 6 | 4,147 | Local Similarity Search for Unstructured Text | 2016 | SIGMOD |
| 7 | 8,689 | TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection | 2022 | SIGMOD |
| 8 | 7,833 | Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts | 2021 | SIGMOD |
| 9 | 10,478 | LSHAlign: All-Pair Near-Duplicate Text Alignment via LSH | 2026 | SIGMOD |
| 10 | 7,906 | Near-Duplicate Text Alignment with One Permutation Hashing | 2024 | SIGMOD |