Extracting Logical Hierarchical Structure of HTML Documents Based on Headings
Summary: Heading-driven extraction of the logical HTML hierarchy, addressing mismatch between markup and semantics. Uses heading position, visual prominence, and level-based styling to derive hierarchical blocks with their associated headings, outperforming prior methods. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Tomohiro Manabe (Kyoto University)
- 2. Keishi Tajima (Kyoto University)
BibTeX Citation
@article{manabe_vldb15,
title = {{Extracting Logical Hierarchical Structure of HTML Documents Based on Headings}},
author = {Manabe, Tomohiro and Tajima, Keishi},
journal = {PVLDB},
series = {{VLDB} '15},
volume = {8},
number = {12},
pages = {1606},
doi = {10.14778/2824032.2824058},
url = {https://doi.org/10.14778/2824032.2824058},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,677 | Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents | 2019 | SIGMOD | 5.3867222e-05 |
| 9,398 | Improving Information Extraction from Visually Rich Documents using Visual Span Representations | 2021 | VLDB | 5.2755515e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 588 | Extracting Structured Data from Web Pages | 2003 | SIGMOD | 0.00016092668 |
| 2,415 | Schema Extraction for Tabular Data on the Web | 2013 | VLDB | 8.6069635e-05 |
| 2,609 | The SphereSearch Engine for Unified Ranked Retrieval of Heterogeneous XML and Web Documents | 2005 | VLDB | 8.3476234e-05 |
| 5,554 | Joint Unsupervised Structure Discovery and Information Extraction | 2011 | SIGMOD | 6.1775485e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,316 | Harvesting Relational Tables from Lists on the Web | 2009 | VLDB |
| 2 | 12,738 | A Framework for Processing Complex Document-centric XML with Overlapping Structures | 2005 | SIGMOD |
| 3 | 14,009 | A Method of Re-ranking Web Search Results Using their Hidden Hyperlink Structure | 2002 | VLDB |
| 4 | 3,944 | Robust Web Extraction: An Approach Based on a Probabilistic Tree-Edit Model | 2009 | SIGMOD |
| 5 | 7,811 | The Smallest Extraction Problem | 2021 | VLDB |
| 6 | 2,096 | An Analysis of Structured Data on the Web | 2012 | VLDB |
| 7 | 588 | Extracting Structured Data from Web Pages | 2003 | SIGMOD |
| 8 | 3,523 | Using the Structure of Web Sites for Automatic Segmentation of Tables | 2004 | SIGMOD |
| 9 | 2,329 | Record-Boundary Discovery in Web Documents | 1999 | SIGMOD |
| 10 | 5,545 | A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration | 2009 | VLDB |