Snowball: A Prototype System for Extracting Relations from Large Text Collections
Summary: Snowball is a prototype system that extracts organization–location relations from large text collections, producing a structured relation (a materialized view) over unstructured documents. An interactive, low-human-participation demo that builds on Brin's DIPRE approach to iterative extraction. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Eugene Agichtein (Columbia University)
- 2. Luis Gravano (Columbia University)
- 3. Jeff Pavel (Columbia University)
- 4. Viktoriya Sokolova (Columbia University)
- 5. Aleksandr Voskoboynik (Columbia University)
BibTeX Citation
@inproceedings{agichtein_sigmod01,
title = {{Snowball: A Prototype System for Extracting Relations from Large Text Collections}},
author = {Agichtein, Eugene and Gravano, Luis and Pavel, Jeff and Sokolova, Viktoriya and Voskoboynik, Aleksandr},
series = {{SIGMOD} '01},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/375663.375774},
url = {https://dl.acm.org/doi/10.1145/375663.375774},
year = {2001}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 91 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB | 0.00034838835 |
| 4,461 | Extracting and Querying a Comprehensive Web Database | 2009 | CIDR | 6.6889881e-05 |
| 6,011 | From Information to Knowledge: Harvesting Entities and Relationships from Web Sources | 2010 | PODS | 6.0105397e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 0 of 0 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,728 | Structured Querying of Web Text: A Technical Challenge | 2007 | CIDR |
| 2 | 13,921 | QXtract: A Building Block for Efficient Information Extraction from Text Databases | 2003 | SIGMOD |
| 3 | 11,880 | PivotE: Revealing and Visualizing the Underlying Entity Structures for Exploration | 2019 | VLDB |
| 4 | 11,440 | Autonomously Computable Information Extraction | 2023 | VLDB |
| 5 | 13,602 | NewsNetExplorer: Automatic Construction and Exploration of News Information Networks | 2014 | SIGMOD |
| 6 | 4,834 | KBPearl: A Knowledge Base Population System Supported by Joint Entity and Relation Linking | 2020 | VLDB |
| 7 | 2,609 | The SphereSearch Engine for Unified Ranked Retrieval of Heterogeneous XML and Web Documents | 2005 | VLDB |
| 8 | 11,980 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD |
| 9 | 12,529 | iNextCube: Information Network-Enhanced Text Cube | 2009 | VLDB |
| 10 | 2,724 | A Relational Approach to Incrementally Extracting and Querying Structure in Unstructured Data | 2007 | VLDB |