DBScholar

Back to papers

From Information to Knowledge: Harvesting Entities and Relationships from Web Sources

Summary: Tutorial surveying methods to automatically harvest entities, classes, relations and temporal contexts from semi-structured and natural-language Web sources (e.g., Wikipedia, DBpedia, YAGO) into high-precision, high-recall knowledge bases. Discusses extraction/integration pipelines, maintenance, evaluation, and open research challenges. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
he0be59b6aea3fcf0
Venue
PODS
Year
2010
Pagerank
5.886265e-05
Overall Rank
6,104 | 58.97%
DOI
10.1145/1807085.1807097

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{weikum_pods10,
        address = {New York, NY, USA},
        series = {{PODS} '10},
        title = {{From Information to Knowledge: Harvesting Entities and Relationships from Web Sources}},
        url = {https://dl.acm.org/doi/10.1145/1807085.1807097},
        doi = {10.1145/1807085.1807097},
        booktitle = {Proceedings of the {ACM} {SIGMOD} Symposium on {Principles} of {Database} {Systems}},
        publisher = {Association for Computing Machinery},
        author = {Weikum, Gerhard and Theobald, Martin},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 3 of 3 citing papers.

Rank Citing Paper Year Venue Pagerank
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016086569
7,007 An Efficient Publish/Subscribe Index for E-Commerce Databases 2014 VLDB 5.6226844e-05
12,533 Knowledge Harvesting in the Big-Data Era 2013 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 24 of 24 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
89 WebTables: Exploring the Power of Tables on the Web 2008 VLDB 0.00035129658
262 Record Linkage: Similarity Measures and Algorithms 2006 SIGMOD 0.00022821392
328 Declarative Information Extraction Using Datalog with Embedded Extraction Predicates 2007 VLDB 0.00020928928
484 Data Integration for the Relational Web 2009 VLDB 0.00017548541
515 RoadRunner: Towards Automatic Data Extraction from Large Web Sites 2001 VLDB 0.00016970049
599 Extracting Structured Data from Web Pages 2003 SIGMOD 0.0001575536
715 Data Integration with Uncertainty 2007 VLDB 0.00014559092
954 RDF-3X: a RISC-style Engine for RDF 2008 VLDB 0.00012867202
997 To Search or to Crawl? Towards a Query Optimizer for Text-Centric Tasks 2006 SIGMOD 0.00012633809
1,114 Principles of Dataspace Systems 2006 PODS 0.00011964697
1,154 The Lixto Data Extraction Project - Back and Forth between Theory and Practice 2004 PODS 0.00011793227
1,337 Harvesting Relational Tables from Lists on the Web 2009 VLDB 0.00010987014
1,711 EntityRank: Searching Entities Directly and Holistically 2007 VLDB 9.8182129e-05
1,794 Building Structured Web Community Portals: A Top-Down, Compositional, and Incremental Approach 2007 VLDB 9.6220467e-05
2,300 Leveraging Data and Structure in Ontology Integration 2007 SIGMOD 8.6724454e-05
2,462 DBLife: A Community Information Management Platform for the Database Research Community 2007 CIDR 8.4317833e-05
2,477 A Relational Approach to Incrementally Extracting and Querying Structure in Unstructured Data 2007 VLDB 8.4079432e-05
2,615 Visual Web Information Extraction with Lixto 2001 VLDB 8.2253533e-05
2,737 Snowball: A Prototype System for Extracting Relations from Large Text Collections 2001 SIGMOD 8.0734152e-05
4,057 Uncertainty Management in Rule-Based Information Extraction Systems 2009 SIGMOD 6.8257248e-05
4,554 Extracting and Querying a Comprehensive Web Database 2009 CIDR 6.5391226e-05
5,403 Mining Document Collections to Facilitate Accurate Approximate Entity Matching 2009 VLDB 6.1459021e-05
5,523 A First Tutorial on Dataspaces 2008 VLDB 6.0951795e-05
10,014 Optimizing Complex Extraction Programs over Evolving Text Data 2009 SIGMOD 5.0970738e-05
Previous Page 1 / 1 Next

Semantically Similar Papers