DBScholar

Back to papers

Automatic segmentation of text into structured records

Summary: Automatic segmentation of unformatted text into structured records; datamold learns structure from a small seed set. Extends HMMs with multi-source cues (sequence, length, vocabulary, external dictionary) for robust address extraction; 90% Asian, 99% US accuracy, beating rule-based IE. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h5e254cf307563b25
Venue
SIGMOD
Year
2001
Pagerank
0.00014373044
Overall Rank
735 | 95.06%
DOI
10.1145/375663.375682

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{borkar_sigmod01,
        title = {{Automatic segmentation of text into structured records}},
        author = {Borkar, Vinayak and Deshmukh, Kaustubh and Sarawagi, Sunita},
        series = {{SIGMOD} '01},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/375663.375682},
        url = {https://dl.acm.org/doi/10.1145/375663.375682},
        year = {2001}
}

Incoming Citations (Sorted by Pagerank)

Showing 16 of 16 citing papers.

Rank Citing Paper Year Venue Pagerank
95 Potter's Wheel: An Interactive Data Cleaning System 2001 VLDB 0.00034382643
307 Eliminating Fuzzy Duplicates in Data Warehouses 2002 VLDB 0.00021499031
690 Creating Probabilistic Databases from Information Extraction Models 2006 VLDB 0.00014745444
1,337 Harvesting Relational Tables from Lists on the Web 2009 VLDB 0.00010987014
1,581 Example-driven Design of Efficient Record Matching Queries 2007 VLDB 0.00010180038
2,580 Tuning Schema Matching Software using Synthetic Scenarios 2005 VLDB 8.2716518e-05
3,588 Using the Structure of Web Sites for Automatic Segmentation of Tables 2004 SIGMOD 7.1870903e-05
3,593 Merging the Results of Approximate Match Operations 2004 VLDB 7.183446e-05
3,615 TEGRA: Table Extraction by Global Record Alignment 2015 SIGMOD 7.1591776e-05
4,010 Exploiting Content Redundancy for Web Information Extraction 2010 VLDB 6.8582009e-05
4,531 Database Principles in Information Extraction 2014 PODS 6.558436e-05
5,367 Entity Extraction, Linking, Classification, and Tagging for Social Media: A Wikipedia-Based Approach 2013 VLDB 6.1598727e-05
5,670 Joint Unsupervised Structure Discovery and Information Extraction 2011 SIGMOD 6.0445956e-05
7,565 A Probabilistic Approach for Automatically Filling Form-Based Web Interfaces 2011 VLDB 5.4953097e-05
8,059 A Grammar-based Entity Representation Framework for Data Cleaning 2009 SIGMOD 5.3965759e-05
12,716 ONDUX: On-Demand Unsupervised Learning for Information Extraction 2010 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 4 of 4 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers