Back to papers
Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation
Summary: Proposes scalable near-duplicate sequence search to measure LLM memorization in trillion-token corpora. The approach groups min-hash values for all sequences with at least t tokens in linear time, uses inverted indexes and prefix filtering, and proves a bound 2^{(n+1)/(t+1)}−1, with real-world validation.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6683
- Venue
- SIGMOD
- Year
- 2023
- Pagerank
- 4.2626861e-05
- Overall Rank
- 9,875 | 31.37%
- DOI
-
10.1145/3589324
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 614 |
Copy Detection Mechanisms for Digital Documents |
1995 |
SIGMOD |
0.00019088537 |
| 703 |
Winnowing: Local Algorithms for Document Fingerprinting |
2003 |
SIGMOD |
0.0001784726 |
| 1,396 |
Can We Beat the Prefix Filtering? An Adaptive Framework for Similarity Join and Search |
2012 |
SIGMOD |
0.00012215253 |
| 1,675 |
Adaptive Parallel Aggregation Algorithms |
1995 |
SIGMOD |
0.00010937908 |
| 2,588 |
Pass-Join: A Partition-based Method for Similarity Joins |
2012 |
VLDB |
8.4872437e-05 |
| 4,041 |
An Efficient Partition Based Method for Exact Set Similarity Joins |
2016 |
VLDB |
6.5048916e-05 |
| 4,247 |
Local Similarity Search for Unstructured Text |
2016 |
SIGMOD |
6.3180334e-05 |
| 4,350 |
Overlap Set Similarity Joins with Theoretical Guarantees |
2018 |
SIGMOD |
6.2576191e-05 |
| 5,071 |
Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction |
2011 |
SIGMOD |
5.7122974e-05 |
| 6,730 |
A Pivotal Prefix Based Filtering Algorithm for String Similarity Search |
2014 |
SIGMOD |
4.9436522e-05 |
| 7,636 |
Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts |
2021 |
SIGMOD |
4.6863871e-05 |
| 8,284 |
TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection |
2022 |
SIGMOD |
4.5392079e-05 |
| 9,566 |
META: An Efficient Matching-Based Method for Error-Tolerant Autocompletion |
2016 |
VLDB |
4.3212967e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 9,786 |
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training |
2025 |
SIGMOD |
4.2799988e-05 |
| 7,636 |
Allign: Aligning All-Pair Near-Duplicate Passages in Long Texts |
2021 |
SIGMOD |
4.6863871e-05 |
| 10,064 |
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,462 |
ScaleLLM: A Technique for Scalable LLM-augmented Data Systems |
2025 |
SIGMOD |
4.1905499e-05 |
| 7,698 |
Near-Duplicate Text Alignment with One Permutation Hashing |
2024 |
SIGMOD |
4.6699546e-05 |
| 13,152 |
Database Perspective on LLM Inference Systems |
2025 |
VLDB |
- |
| 8,284 |
TxtAlign: Efficient Near-Duplicate Text Alignment Search via Bottom-k Sketches for Plagiarism Detection |
2022 |
SIGMOD |
4.5392079e-05 |
| 10,022 |
In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration |
2026 |
SIGMOD |
4.1905499e-05 |
| 11,061 |
LLM-PBE: Assessing Data Privacy in Large Language Models |
2024 |
VLDB |
4.1905499e-05 |
| 10,508 |
Privacy and Accuracy-Aware AI/ML Model Deduplication |
2025 |
SIGMOD |
4.1905499e-05 |