DBScholar

Back to papers

Extracting and Querying a Comprehensive Web Database

Summary: Omnivore builds a comprehensive web-scale entity-relationship DB by running multiple domain-independent extractors (tables, text, relations) in parallel over a crawl and merging heterogeneous outputs to overcome model-specific blind spots. Provides SQL-like and search interfaces, supports user corrections, and automatically selects output model/schema to render results without prior metadata. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
hbe8598472416b30f
Venue
CIDR
Year
2009
Pagerank
6.5391226e-05
Overall Rank
4,554 | 69.39%
DOI
-

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{cafarella_cidr09,
        address = {Amsterdam, Netherlands},
        series = {{CIDR} '09},
        title = {{Extracting and Querying a Comprehensive Web Database}},
        booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
        author = {Cafarella, Michael J.},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Rank Citing Paper Year Venue Pagerank
6,104 From Information to Knowledge: Harvesting Entities and Relationships from Web Sources 2010 PODS 5.886265e-05
9,767 IQ: The Case for Iterative Querying for Knowledge 2011 CIDR 5.1343712e-05
12,533 Knowledge Harvesting in the Big-Data Era 2013 SIGMOD 4.9793485e-05
12,730 DoCQS: A Prototype System for Supporting Data-oriented Content Query 2010 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers