DBScholar

Back to papers

From Information to Knowledge: Harvesting Entities and Relationships from Web Sources

Summary: Tutorial surveying methods to automatically harvest entities, classes, relations and temporal contexts from semi-structured and natural-language Web sources (e.g., Wikipedia, DBpedia, YAGO) into high-precision, high-recall knowledge bases. Discusses extraction/integration pipelines, maintenance, evaluation, and open research challenges. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
1507
Venue
PODS
Year
2010
Pagerank
6.0105397e-05
Overall Rank
6,011 | 58.77%
DOI
10.1145/1807085.1807097

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{weikum_pods10,
        address = {New York, NY, USA},
        series = {{PODS} '10},
        title = {{From Information to Knowledge: Harvesting Entities and Relationships from Web Sources}},
        url = {https://dl.acm.org/doi/10.1145/1807085.1807097},
        doi = {10.1145/1807085.1807097},
        booktitle = {Proceedings of the {ACM} {SIGMOD} Symposium on {Principles} of {Database} {Systems}},
        publisher = {Association for Computing Machinery},
        author = {Weikum, Gerhard and Theobald, Martin},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Rank Citing Paper Year Venue Pagerank
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016217563
6,861 An Efficient Publish/Subscribe Index for E-Commerce Databases 2014 VLDB 5.7517379e-05
12,242 Knowledge Harvesting in the Big-Data Era 2013 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 24 of 24 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
91 WebTables: Exploring the Power of Tables on the Web 2008 VLDB 0.00034838835
254 Record Linkage: Similarity Measures and Algorithms 2006 SIGMOD 0.00023199211
319 Declarative Information Extraction Using Datalog with Embedded Extraction Predicates 2007 VLDB 0.00021377065
493 Data Integration for the Relational Web 2009 VLDB 0.00017558709
503 RoadRunner: Towards Automatic Data Extraction from Large Web Sites 2001 VLDB 0.00017314037
588 Extracting Structured Data from Web Pages 2003 SIGMOD 0.00016092668
699 Data Integration with Uncertainty 2007 VLDB 0.0001487423
944 RDF-3X: a RISC-style Engine for RDF 2008 VLDB 0.00013067088
974 To Search or to Crawl? Towards a Query Optimizer for Text-Centric Tasks 2006 SIGMOD 0.00012870746
1,131 The Lixto Data Extraction Project - Back and Forth between Theory and Practice 2004 PODS 0.00012044716
1,162 Principles of Dataspace Systems 2006 PODS 0.00011866011
1,316 Harvesting Relational Tables from Lists on the Web 2009 VLDB 0.00011181216
1,689 EntityRank: Searching Entities Directly and Holistically 2007 VLDB 0.00010003102
1,848 Building Structured Web Community Portals: A Top-Down, Compositional, and Incremental Approach 2007 VLDB 9.6208418e-05
2,252 Leveraging Data and Structure in Ontology Integration 2007 SIGMOD 8.8675372e-05
2,405 DBLife: A Community Information Management Platform for the Database Research Community 2007 CIDR 8.6217909e-05
2,564 Visual Web Information Extraction with Lixto 2001 VLDB 8.4118854e-05
2,724 A Relational Approach to Incrementally Extracting and Querying Structure in Unstructured Data 2007 VLDB 8.2048095e-05
2,743 Snowball: A Prototype System for Extracting Relations from Large Text Collections 2001 SIGMOD 8.1801198e-05
3,973 Uncertainty Management in Rule-Based Information Extraction Systems 2009 SIGMOD 6.9810029e-05
4,461 Extracting and Querying a Comprehensive Web Database 2009 CIDR 6.6889881e-05
5,279 Mining Document Collections to Facilitate Accurate Approximate Entity Matching 2009 VLDB 6.2849137e-05
5,389 A First Tutorial on Dataspaces 2008 VLDB 6.2346998e-05
9,828 Optimizing Complex Extraction Programs over Evolving Text Data 2009 SIGMOD 5.2138468e-05
Previous Page 1 / 1 Next

Semantically Similar Papers