LLM-PBE: Assessing Data Privacy in Large Language Models
Summary: LLM-PBE: toolkit for systematic evaluation of training-data privacy leakage in LLMs across the model lifecycle, unifying diverse attacks, defenses, data modalities, and privacy metrics. Experiments show model scale, data properties, and temporal drift shape leakage; artifacts and benchmarks released for reproducible privacy research. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Qinbin Li (University of California Berkeley)
- 2. Junyuan Hong (University of Texas)
- 3. Chulin Xie (University of Illinois Urbana-Champaign)
- 4. Jeffrey Tan (University of California Berkeley)
- 5. Rachel Xin (University of California Berkeley)
- 6. Junyi Hou (National University of Singapore)
- 7. Xavier Yin (University of California Berkeley)
- 8. Zhun Wang (University of California Berkeley)
- 9. Dan Hendrycks (Center for AI Safety)
- 10. Zhangyang Wang (University of Texas)
- 11. Bo Li (University of Chicago)
- 12. Bingsheng He (National University of Singapore)
- 13. Dawn Song (University of California Berkeley)
BibTeX Citation
@article{li_vldb24,
title = {{LLM-PBE: Assessing Data Privacy in Large Language Models}},
author = {Li, Qinbin and Hong, Junyuan and Xie, Chulin and Tan, Jeffrey and Xin, Rachel and Hou, Junyi and Yin, Xavier and Wang, Zhun and Hendrycks, Dan and Wang, Zhangyang and Li, Bo and He, Bingsheng and Song, Dawn},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {11},
pages = {3201--3214},
doi = {10.14778/3681954.3681994},
url = {https://doi.org/10.14778/3681954.3681994},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,290 | Skyline Retrieval meets Set-Cover Chunk Merging: A Cost-Effective RAG-Sketch for Long-Context LLM QA | 2026 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 70 | Privacy-Preserving Data Mining | 2000 | SIGMOD | 0.0003804755 |
| 572 | Anatomy: Simple and Effective Privacy Preservation | 2006 | VLDB | 0.00016316092 |
| 865 | Natural language to SQL: Where are we today? | 2020 | VLDB | 0.00013521464 |
| 1,444 | CatSQL: Towards Real World Natural Language to SQL Applications | 2023 | VLDB | 0.00010769944 |
| 3,536 | How Large Language Models Will Disrupt Data Management | 2023 | VLDB | 7.3297343e-05 |
| 4,449 | From BERT to GPT-3 Codex: Harnessing the Potential of Very Large Language Models for Data Management | 2022 | VLDB | 6.697553e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,228 | Mind the Data Gap: Bridging LLMs to Enterprise Data Integration | 2025 | CIDR |
| 2 | 10,024 | Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation | 2023 | SIGMOD |
| 3 | 10,344 | Are Your LLM-based Text-to-SQL Models Secure? Exploring SQL Injection via Backdoor Attacks | 2026 | SIGMOD |
| 4 | 3,536 | How Large Language Models Will Disrupt Data Management | 2023 | VLDB |
| 5 | 8,906 | Unveiling Challenges for LLMs in Enterprise Data Engineering | 2026 | VLDB |
| 6 | 6,101 | LLM for Data Management | 2024 | VLDB |
| 7 | 8,908 | Privacy and Accuracy-Aware AI/ML Model Deduplication | 2025 | SIGMOD |
| 8 | 10,614 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB |
| 9 | 10,282 | Prism: Private Relational Data Synthesis with Language Models | 2026 | SIGMOD |
| 10 | 13,369 | EncChain: Enhancing Large Language Model Applications with Advanced Privacy Preservation Techniques | 2024 | VLDB |