DBScholar

Back to papers

Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts

Summary: Allign uses a min-hash based method to align all-pair near-duplicate passages in two texts, avoiding O(n^2 m^2) enumeration via compact windows. It matches windows by shared min-hash, reports the longest and sentence-level near-duplicates, and outperforms prior alignment methods on real data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6299
Venue
SIGMOD
Year
2021
Pagerank
5.5686107e-05
Overall Rank
7,676 | 47.34%
DOI
10.1145/3448016.3457548

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{feng_sigmod21,
        title = {{Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts}},
        author = {Feng, Weiqi and Deng, Dong},
        series = {{SIGMOD} '21},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3448016.3457548},
        url = {https://dl.acm.org/doi/10.1145/3448016.3457548},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 7 of 7 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 14 of 14 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers