Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation
Summary: Bootleg: open-source, self-supervised NED using a simple transformer + hierarchical regularization to dramatically boost tail-entity disambiguation (up to +41.2 F1) and match/exceed SOTA on benchmarks. Calls out serving of entity embeddings and related data-management challenges. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Laurel Orr (Stanford University)
- 2. Megan Leszczynski (Stanford University)
- 3. Neel Guha (Stanford University)
- 4. Sen Wu (Stanford University)
- 5. Simran Arora (Stanford University)
- 6. Xiao Ling (Apple)
- 7. Christopher Ré (Stanford University)
BibTeX Citation
@inproceedings{orr_cidr21,
address = {Amsterdam, Netherlands},
series = {{CIDR} '21},
title = {{Bootleg: Chasing the Tail with Self-Supervised Named Entity Disambiguation}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Orr, Laurel and Leszczynski, Megan and Guha, Neel and Wu, Sen and Arora, Simran and Ling, Xiao and Ré, Christopher},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 3,304 | Saga: A Platform for Continuous Construction and Serving of Knowledge At Scale | 2022 | SIGMOD | 7.5404346e-05 |
| 3,589 | Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins | 2022 | VLDB | 7.2812353e-05 |
| 5,925 | Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems | 2021 | VLDB | 6.0397888e-05 |
| 11,516 | Data Management Opportunities for Foundation Models | 2022 | CIDR | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 112 | HoloClean: Holistic Data Repairs with Probabilistic Inference | 2017 | VLDB | 0.00032801121 |
| 176 | Deep Learning for Entity Matching: A Design Space Exploration | 2018 | SIGMOD | 0.00027191081 |
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 1,121 | ARDA: Automatic Relational Data Augmentation for Machine Learning | 2020 | VLDB | 0.00012093059 |
| 3,947 | Overton: A Data System for Monitoring and Improving Machine-Learned Products | 2020 | CIDR | 7.0040437e-05 |
| 4,834 | KBPearl: A Knowledge Base Population System Supported by Joint Entity and Relation Linking | 2020 | VLDB | 6.4866871e-05 |
| 7,855 | ItemSuggest: A Data Management Platform for Machine Learned Ranking Services | 2019 | CIDR | 5.530673e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 9,560 | Ground Truth Inference for Weakly Supervised Entity Matching | 2023 | SIGMOD |
| 2 | 11,455 | Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages | 2023 | VLDB |
| 3 | 2,475 | Deep Learning for Blocking in Entity Matching: A Design Space Exploration | 2021 | VLDB |
| 4 | 6,192 | Pre-trained Embeddings for Entity Resolution: An Experimental Analysis | 2023 | VLDB |
| 5 | 4,938 | Supervised Meta-blocking | 2014 | VLDB |
| 6 | 4,447 | Dual-Objective Fine-Tuning of BERT for Entity Matching | 2021 | VLDB |
| 7 | 4,905 | Medical Entity Disambiguation Using Graph Neural Networks | 2021 | SIGMOD |
| 8 | 2,463 | A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching | 2020 | SIGMOD |
| 9 | 141 | Deep Entity Matching with Pre-Trained Language Models | 2021 | VLDB |
| 10 | 3,446 | Efficient Approximate Entity Extraction with Edit Distance Constraints | 2009 | SIGMOD |