Extracting and Querying a Comprehensive Web Database
Summary: Omnivore builds a comprehensive web-scale entity-relationship DB by running multiple domain-independent extractors (tables, text, relations) in parallel over a crawl and merging heterogeneous outputs to overcome model-specific blind spots. Provides SQL-like and search interfaces, supports user corrections, and automatically selects output model/schema to render results without prior metadata. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael J. Cafarella (University of Washington)
BibTeX Citation
@inproceedings{cafarella_cidr09,
address = {Amsterdam, Netherlands},
series = {{CIDR} '09},
title = {{Extracting and Querying a Comprehensive Web Database}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Cafarella, Michael J.},
year = {2009}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6,104 | From Information to Knowledge: Harvesting Entities and Relationships from Web Sources | 2010 | PODS | 5.886265e-05 |
| 9,767 | IQ: The Case for Iterative Querying for Knowledge | 2011 | CIDR | 5.1343712e-05 |
| 12,533 | Knowledge Harvesting in the Big-Data Era | 2013 | SIGMOD | 4.9793485e-05 |
| 12,730 | DoCQS: A Prototype System for Supporting Data-oriented Content Query | 2010 | SIGMOD | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 65 | Freebase: A Collaboratively Created Graph Database For Structuring Human Knowledge | 2008 | SIGMOD | 0.00038222149 |
| 89 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB | 0.00035129658 |
| 190 | Applying Model Management to Classical Meta Data Problems | 2003 | CIDR | 0.00025768763 |
| 230 | Reference Reconciliation in Complex Information Spaces | 2005 | SIGMOD | 0.00023871933 |
| 1,760 | Structured Querying of Web Text: A Technical Challenge | 2007 | CIDR | 9.7106555e-05 |
| 1,794 | Building Structured Web Community Portals: A Top-Down, Compositional, and Incremental Approach | 2007 | VLDB | 9.6220467e-05 |
| 2,737 | Snowball: A Prototype System for Extracting Relations from Large Text Collections | 2001 | SIGMOD | 8.0734152e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 484 | Data Integration for the Relational Web | 2009 | VLDB |
| 2 | 89 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB |
| 3 | 2,255 | Extraction and Integration of Partially Overlapping Web Sources | 2013 | VLDB |
| 4 | 5,866 | DIADEM: Thousands of Websites to a Single Database | 2014 | VLDB |
| 5 | 12,340 | Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale | 2016 | SIGMOD |
| 6 | 3,759 | Toward Large Scale Integration: Building a MetaQuerier over Databases on the Web | 2005 | CIDR |
| 7 | 2,880 | Expressive and Flexible Access to Web-Extracted Data: A Keyword-based Structured Query Language | 2010 | SIGMOD |
| 8 | 12,744 | ObjectRunner: Lightweight, Targeted Extraction and Querying of Structured Web Data | 2010 | VLDB |
| 9 | 5,678 | A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration | 2009 | VLDB |
| 10 | 1,760 | Structured Querying of Web Text: A Technical Challenge | 2007 | CIDR |