Automatic segmentation of text into structured records
Summary: Automatic segmentation of unformatted text into structured records; datamold learns structure from a small seed set. Extends HMMs with multi-source cues (sequence, length, vocabulary, external dictionary) for robust address extraction; 90% Asian, 99% US accuracy, beating rule-based IE. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Vinayak Borkar (Indian Institute of Technology Mumbai)
- 2. Kaustubh Deshmukh (Indian Institute of Technology Mumbai; University of Washington)
- 3. Sunita Sarawagi (Indian Institute of Technology Mumbai)
BibTeX Citation
@inproceedings{borkar_sigmod01,
title = {{Automatic segmentation of text into structured records}},
author = {Borkar, Vinayak and Deshmukh, Kaustubh and Sarawagi, Sunita},
series = {{SIGMOD} '01},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/375663.375682},
url = {https://dl.acm.org/doi/10.1145/375663.375682},
year = {2001}
}
Incoming Citations (Sorted by Pagerank)
Showing 16 of 16 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 58 | The Merge/Purge Problem for Large Databases | 1995 | SIGMOD | 0.00040116748 |
| 812 | NoDoSE - A Tool for Semi-Automatically Extracting Structured and Semistructured Data from Text Documents. | 1998 | SIGMOD | 0.00013851743 |
| 1,178 | Building light-weight wrappers for legacy Web data-sources using W4F | 1999 | VLDB | 0.00011806558 |
| 2,329 | Record-Boundary Discovery in Web Documents | 1999 | SIGMOD | 8.7453177e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,811 | The Smallest Extraction Problem | 2021 | VLDB |
| 2 | 5,554 | Joint Unsupervised Structure Discovery and Information Extraction | 2011 | SIGMOD |
| 3 | 8,067 | Mining Quality Phrases from Massive Text Corpora | 2015 | SIGMOD |
| 4 | 588 | Extracting Structured Data from Web Pages | 2003 | SIGMOD |
| 5 | 4,752 | Automatic Rule Refinement for Information Extraction | 2010 | VLDB |
| 6 | 12,425 | ONDUX: On-Demand Unsupervised Learning for Information Extraction | 2010 | SIGMOD |
| 7 | 4,494 | Structured Annotations of Web Queries | 2010 | SIGMOD |
| 8 | 2,329 | Record-Boundary Discovery in Web Documents | 1999 | SIGMOD |
| 9 | 11,980 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD |
| 10 | 3,523 | Using the Structure of Web Sites for Automatic Segmentation of Tables | 2004 | SIGMOD |