Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
Summary: Bootleg: open-source, self-supervised NED using a simple transformer + hierarchical regularization to dramatically boost tail-entity disambiguation (up to +41.2 F1) and match/exceed SOTA on benchmarks. Calls out serving of entity embeddings and related data-management challenges. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Laurel Orr (Stanford University)
- 2. Megan Leszczynski (Stanford University)
- 3. Neel Guha (Stanford University)
- 4. Sen Wu (Stanford University)
- 5. Simran Arora (Stanford University)
- 6. Xiao Ling (Apple)
- 7. Christopher Ré (Stanford University)
BibTeX Citation
@inproceedings{orr_cidr21,
address = {Amsterdam, Netherlands},
series = {{CIDR} '21},
title = {{Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Orr, Laurel and Leszczynski, Megan and Guha, Neel and Wu, Sen and Arora, Simran and Ling, Xiao and Ré, Christopher},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,013 | Saga: A Platform for Continuous Construction and Serving of Knowledge At Scale | 2022 | SIGMOD | 7.7541678e-05 |
| 3,521 | Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins | 2022 | VLDB | 7.2351481e-05 |
| 5,506 | Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems | 2021 | VLDB | 6.1016716e-05 |
| 11,825 | Data Management Opportunities for Foundation Models | 2022 | CIDR | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 104 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00033690989 |
| 158 | Deep Learning for Entity Matching: A Design Space Exploration | 2018 | SIGMOD | 0.00028046388 |
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025181304 |
| 1,038 | ARDA: Automatic Relational Data Augmentation for Machine Learning | 2020 | VLDB | 0.00012370691 |
| 3,997 | Overton: A Data System for Monitoring and Improving Machine-Learned Products | 2020 | CIDR | 6.8655222e-05 |
| 4,952 | KBPearl: A Knowledge Base Population System Supported by Joint Entity and Relation Linking | 2020 | VLDB | 6.3417814e-05 |
| 8,012 | ItemSuggest: A Data Management Platform for Machine Learned Ranking Services | 2019 | CIDR | 5.4072178e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,708 | Ground Truth Inference for Weakly Supervised Entity Matching | 2023 | SIGMOD |
| 2 | 2,514 | Deep Learning for Blocking in Entity Matching: A Design Space Exploration | 2021 | VLDB |
| 3 | 5,540 | Pre-trained Embeddings for Entity Resolution: An Experimental Analysis | 2023 | VLDB |
| 4 | 5,064 | Supervised Meta-blocking | 2014 | VLDB |
| 5 | 4,033 | Dual-Objective Fine-Tuning of BERT for Entity Matching | 2021 | VLDB |
| 6 | 4,995 | Medical Entity Disambiguation Using Graph Neural Networks | 2021 | SIGMOD |
| 7 | 10,981 | NiceT: Named Entity Cleaning and Enhancement with Human-in-the-loop | 2026 | VLDB |
| 8 | 2,475 | A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching | 2020 | SIGMOD |
| 9 | 134 | Deep Entity Matching with Pre-Trained Language Models | 2021 | VLDB |
| 10 | 3,512 | Efficient Approximate Entity Extraction with Edit Distance Constraints | 2009 | SIGMOD |