Extracting and Querying a Comprehensive Web Database
Summary: Omnivore builds a comprehensive web-scale entity-relationship DB by running multiple domain-independent extractors (tables, text, relations) in parallel over a crawl and merging heterogeneous outputs to overcome model-specific blind spots. Provides SQL-like and search interfaces, supports user corrections, and automatically selects output model/schema to render results without prior metadata. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael J. Cafarella (University of Washington)
BibTeX Citation
@inproceedings{cafarella_cidr09,
address = {Amsterdam, Netherlands},
series = {{CIDR} '09},
title = {{Extracting and Querying a Comprehensive Web Database}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Cafarella, Michael J.},
year = {2009}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,011 | From Information to Knowledge: Harvesting Entities and Relationships from Web Sources | 2010 | PODS | 6.0105397e-05 |
| 9,588 | IQ: The Case for Iterative Querying for Knowledge | 2011 | CIDR | 5.2521703e-05 |
| 12,242 | Knowledge Harvesting in the Big-Data Era | 2013 | SIGMOD | 5.093636e-05 |
| 12,439 | DoCQS: A Prototype System for Supporting Data-oriented Content Query | 2010 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 65 | Freebase: A Collaboratively Created Graph Database For Structuring Human Knowledge | 2008 | SIGMOD | 0.00038697603 |
| 91 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB | 0.00034838835 |
| 183 | Applying Model Management to Classical Meta Data Problems | 2003 | CIDR | 0.00026283359 |
| 228 | Reference Reconciliation in Complex Information Spaces | 2005 | SIGMOD | 0.00023941271 |
| 1,728 | Structured Querying of Web Text: A Technical Challenge | 2007 | CIDR | 9.9098087e-05 |
| 1,848 | Building Structured Web Community Portals: A Top-Down, Compositional, and Incremental Approach | 2007 | VLDB | 9.6208418e-05 |
| 2,743 | Snowball: A Prototype System for Extracting Relations from Large Text Collections | 2001 | SIGMOD | 8.1801198e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 493 | Data Integration for the Relational Web | 2009 | VLDB |
| 2 | 91 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB |
| 3 | 2,322 | Extraction and Integration of Partially Overlapping Web Sources | 2013 | VLDB |
| 4 | 5,745 | DIADEM: Thousands of Websites to a Single Database | 2014 | VLDB |
| 5 | 12,045 | Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale | 2016 | SIGMOD |
| 6 | 3,690 | Toward Large Scale Integration: Building a MetaQuerier over Databases on the Web | 2005 | CIDR |
| 7 | 2,832 | Expressive and Flexible Access to Web-Extracted Data: A Keyword-based Structured Query Language | 2010 | SIGMOD |
| 8 | 12,453 | ObjectRunner: Lightweight, Targeted Extraction and Querying of Structured Web Data | 2010 | VLDB |
| 9 | 5,545 | A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration | 2009 | VLDB |
| 10 | 1,728 | Structured Querying of Web Text: A Technical Challenge | 2007 | CIDR |