Extracting Logical Hierarchical Structure of HTML Documents Based on Headings
Summary: Heading-driven extraction of the logical HTML hierarchy, addressing mismatch between markup and semantics. Uses heading position, visual prominence, and level-based styling to derive hierarchical blocks with their associated headings, outperforming prior methods. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Tomohiro Manabe
- 2. Keishi Tajima
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,456 | Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents | 2019 | SIGMOD | 4.5018004e-05 |
| 9,259 | Improving Information Extraction from Visually Rich Documents using Visual Span Representations | 2021 | VLDB | 4.3648789e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 586 | Extracting Structured Data from Web Pages | 2003 | SIGMOD | 0.0001963091 |
| 2,229 | The SphereSearch Engine for Unified Ranked Retrieval of Heterogeneous XML and Web Documents | 2005 | VLDB | 9.2422485e-05 |
| 2,638 | Schema Extraction for Tabular Data on the Web | 2013 | VLDB | 8.3995765e-05 |
| 5,408 | Joint Unsupervised Structure Discovery and Information Extraction | 2011 | SIGMOD | 5.5238635e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| Overall Rank | Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,316 | Harvesting Relational Tables from Lists on the Web | 2009 | VLDB | 0.00012616422 |
| 12,554 | A Framework for Processing Complex Document-centric XML with Overlapping Structures | 2005 | SIGMOD | 4.1905499e-05 |
| 13,822 | A Method of Re-ranking Web Search Results Using their Hidden Hyperlink Structure | 2002 | VLDB | - |
| 4,438 | Robust Web Extraction: An Approach Based on a Probabilistic Tree-Edit Model | 2009 | SIGMOD | 6.1819088e-05 |
| 7,831 | The Smallest Extraction Problem | 2021 | VLDB | 4.6372231e-05 |
| 1,855 | An Analysis of Structured Data on the Web | 2012 | VLDB | 0.00010319508 |
| 586 | Extracting Structured Data from Web Pages | 2003 | SIGMOD | 0.0001963091 |
| 3,286 | Using the Structure of Web Sites for Automatic Segmentation of Tables | 2004 | SIGMOD | 7.2692403e-05 |
| 2,011 | Record-Boundary Discovery in Web Documents | 1999 | SIGMOD | 9.8032193e-05 |
| 5,783 | A Hierarchical Approach to Model Web Query Interfaces for Web Source Integration | 2009 | VLDB | 5.3262443e-05 |