Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages
Summary: LEAST self-trains IE models for semi-structured webpages, using a generative model to expand fewer than ten labeled pages per site. Uncertainty-aware training suppresses synthetic-label noise, achieving cross-domain/model gains and up to 11× lower annotation costs. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Ritesh Sarkhel (Amazon)
- 2. Binxuan Huang (Amazon)
- 3. Colin Lockard (Amazon)
- 4. Prashant Shiralkar (Amazon)
BibTeX Citation
@article{sarkhel_vldb23,
title = {{Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages}},
author = {Sarkhel, Ritesh and Huang, Binxuan and Lockard, Colin and Shiralkar, Prashant},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {11},
pages = {3098--3110},
doi = {10.14778/3611479.3611511},
url = {https://doi.org/10.14778/3611479.3611511},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 2,322 | Extraction and Integration of Partially Overlapping Web Sources | 2013 | VLDB | 8.7526172e-05 |
| 3,192 | Fonduer: Knowledge Base Construction from Richly Formatted Data | 2018 | SIGMOD | 7.65035e-05 |
| 3,699 | KBQA: Learning Question Answering over QA Corpora and Knowledge Bases | 2017 | VLDB | 7.1882315e-05 |
| 6,320 | CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web | 2018 | VLDB | 5.9150411e-05 |
| 8,677 | Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents | 2019 | SIGMOD | 5.3867222e-05 |
| 9,398 | Improving Information Extraction from Visually Rich Documents using Visual Span Representations | 2021 | VLDB | 5.2755515e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,186 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 2 | 10,614 | LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning | 2026 | VLDB |
| 3 | 12,884 | Toward Learning Based Web Query Processing | 2000 | VLDB |
| 4 | 7,239 | I4E: Interactive Investigation of Iterative Information Extraction | 2010 | SIGMOD |
| 5 | 11,980 | Building Structured Databases of Factual Knowledge from Massive Text Corpora | 2017 | SIGMOD |
| 6 | 588 | Extracting Structured Data from Web Pages | 2003 | SIGMOD |
| 7 | 4,494 | Structured Annotations of Web Queries | 2010 | SIGMOD |
| 8 | 3,523 | Using the Structure of Web Sites for Automatic Segmentation of Tables | 2004 | SIGMOD |
| 9 | 12,045 | Potential and Pitfalls of Domain-Specific Information Extraction at Web Scale | 2016 | SIGMOD |
| 10 | 7,811 | The Smallest Extraction Problem | 2021 | VLDB |