Scalable Ad-hoc Entity Extraction from Text Collections
Summary: Introduces ad-hoc entity extraction where target entities come from a task-specific list, avoiding full-document processing. Proposes an inverted-index-driven pruning approach that identifies and processes only task-relevant documents, with empirical gains on real data. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Sanjay Agrawal (Microsoft)
- 2. Kaushik Chakrabarti (Microsoft)
- 3. Surajit Chaudhuri (Microsoft)
- 4. Venkatesh Ganti (Microsoft)
BibTeX Citation
@article{agrawal_vldb08,
title = {{Scalable Ad-hoc Entity Extraction from Text Collections}},
author = {Agrawal, Sanjay and Chakrabarti, Kaushik and Chaudhuri, Surajit and Ganti, Venkatesh},
journal = {PVLDB},
series = {{VLDB} '08},
volume = {1},
number = {1},
pages = {945--956},
doi = {10.14778/1454159.1454164},
url = {https://doi.org/10.14778/1454159.1454164},
year = {2008}
}
Incoming Citations (Sorted by Pagerank)
Showing 8 of 8 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,886 | Pass-Join: A Partition-based Method for Similarity Joins | 2012 | VLDB | 9.5358137e-05 |
| 3,446 | Efficient Approximate Entity Extraction with Edit Distance Constraints | 2009 | SIGMOD | 7.4087786e-05 |
| 4,052 | Local Similarity Search for Unstructured Text | 2016 | SIGMOD | 6.935988e-05 |
| 5,130 | Natural Language Data Management and Interfaces: Recent Development and Open Challenges | 2017 | SIGMOD | 6.354484e-05 |
| 5,279 | Mining Document Collections to Facilitate Accurate Approximate Entity Matching | 2009 | VLDB | 6.2849137e-05 |
| 5,412 | Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction | 2011 | SIGMOD | 6.2272563e-05 |
| 7,252 | Query Portals: Dynamically Generating Portals for Entity-Oriented Web Queries | 2010 | SIGMOD | 5.6630343e-05 |
| 9,200 | JENNER: Just-in-time Enrichment in Query Processing | 2022 | VLDB | 5.3058708e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 169 | Efficient Exact Set-Similarity Joins | 2006 | VLDB | 0.0002743469 |
| 200 | Efficient set joins on similarity predicates | 2004 | SIGMOD | 0.00025597287 |
| 974 | To Search or to Crawl? Towards a Query Optimizer for Text-Centric Tasks | 2006 | SIGMOD | 0.00012870746 |
| 3,442 | An Efficient Filter for Approximate Membership Checking | 2008 | SIGMOD | 7.4122197e-05 |
| 5,861 | Factorizing Complex Predicates in Queries to Exploit Indexes | 2003 | SIGMOD | 6.0636778e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,412 | Faerie: Efficient Filtering Algorithms for Approximate Dictionary-based Entity Extraction | 2011 | SIGMOD |
| 2 | 2,120 | Comparative Analysis of Approximate Blocking Techniques for Entity Resolution | 2016 | VLDB |
| 3 | 10,286 | ScaleDoc: Scaling LLM-based Predicates over Large Document Collections | 2026 | SIGMOD |
| 4 | 11,961 | Scalable Semantic Querying of Text | 2018 | VLDB |
| 5 | 7,005 | Effective and Efficient Retrieval of Structured Entities | 2020 | VLDB |
| 6 | 11,980 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD |
| 7 | 2,976 | Entity Search Engine: Towards Agile Best-Effort Information Integration over the Web | 2007 | CIDR |
| 8 | 12,048 | Automatic Entity Recognition and Typing in Massive Text Data | 2016 | SIGMOD |
| 9 | 3,446 | Efficient Approximate Entity Extraction with Edit Distance Constraints | 2009 | SIGMOD |
| 10 | 5,279 | Mining Document Collections to Facilitate Accurate Approximate Entity Matching | 2009 | VLDB |