Improving Information Extraction from Visually Rich Documents using Visual Span Representations
Summary: Artemis, a visually aware IE method for heterogeneous visually rich documents, encodes visual+textual+layout context into fixed-length span representations. Minimal supervision for visual-span boundaries; multimodal embeddings boost IE, up to 17 F1 points on four datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Ritesh Sarkhel (Ohio State University)
- 2. Arnab Nandi (Ohio State University)
BibTeX Citation
@article{sarkhel_vldb21,
title = {{Improving Information Extraction from Visually Rich Documents using Visual Span Representations}},
author = {Sarkhel, Ritesh and Nandi, Arnab},
journal = {PVLDB},
series = {{VLDB} '21},
volume = {14},
number = {5},
pages = {822--834},
doi = {10.14778/3446095.3446104},
url = {https://doi.org/10.14778/3446095.3446104},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,414 | Visual Template Inference for Data Extraction from Documents | 2026 | SIGMOD | 5.093636e-05 |
| 11,455 | Self-Training for Label-Efficient Information Extraction from Semi-Structured Web-Pages | 2023 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025235185 |
| 3,192 | Fonduer: Knowledge Base Construction from Richly Formatted Data | 2018 | SIGMOD | 7.65035e-05 |
| 3,701 | Snorkel: Fast Training Set Generation for Information Extraction | 2017 | SIGMOD | 7.185321e-05 |
| 6,308 | Extracting Logical Hierarchical Structure of HTML Documents Based on Headings | 2015 | VLDB | 5.9193546e-05 |
| 6,320 | CERES: Distantly Supervised Relation Extraction from the Semi-Structured Web | 2018 | VLDB | 5.9150411e-05 |
| 8,677 | Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents | 2019 | SIGMOD | 5.3867222e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 2,384 | DeepLens: Towards a Visual Data Management System | 2019 | CIDR |
| 2 | 2,475 | Deep Learning for Blocking in Entity Matching: A Design Space Exploration | 2021 | VLDB |
| 3 | 4,494 | Structured Annotations of Web Queries | 2010 | SIGMOD |
| 4 | 10,248 | Generalized Entity Matching with Adaptivity via Large Language Models | 2026 | SIGMOD |
| 5 | 12,048 | Automatic Entity Recognition and Typing in Massive Text Data | 2016 | SIGMOD |
| 6 | 12,214 | Effective Multi-Modal Retrieval based on Stacked Auto-Encoders | 2014 | VLDB |
| 7 | 176 | Deep Learning for Entity Matching: A Design Space Exploration | 2018 | SIGMOD |
| 8 | 10,414 | Visual Template Inference for Data Extraction from Documents | 2026 | SIGMOD |
| 9 | 10,872 | OpenMEL: Unsupervised Multimodal Entity Linking Using Noise-Free Expanded Queries and Global Coherence | 2025 | VLDB |
| 10 | 8,677 | Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents | 2019 | SIGMOD |