Visual Template Inference for Data Extraction from Documents
Summary: TWIX infers latent visual templates for programmatically generated documents by clustering consistently co-located fields and enforcing alignment constraints (e.g., column/header and key/value alignment) to assemble templates for extraction. Template-driven extraction achieves >25% higher precision/recall than Evaporate/Textract/Azure/GPT-4-Vision on 34 datasets and is massively more scalable (≈520× faster, ≈3,786× cheaper) on large corpora. (summarized by gpt-5-mini on Feb 11 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Yiming Lin (University of California Berkeley)
- 2. Mawil Hasan (University of California Berkeley)
- 3. Rohan Kosalge (University of California Berkeley)
- 4. Alvin Cheung (University of California Berkeley)
- 5. Aditya G. Parameswaran (University of California Berkeley)
BibTeX Citation
@inproceedings{lin_sigmod26,
title = {{Visual Template Inference for Data Extraction from Documents}},
author = {Lin, Yiming and Hasan, Mawil and Kosalge, Rohan and Cheung, Alvin and Parameswaran, Aditya G.},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3769840},
url = {https://dl.acm.org/doi/10.1145/3769840},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,440 | Autonomously Computable Information Extraction | 2023 | VLDB |
| 2 | 13,339 | DocDB: A Database for Unstructured Document Analysis | 2025 | VLDB |
| 3 | 7,482 | Table Extraction and Understanding for Scientific and Enterprise Applications | 2020 | VLDB |
| 4 | 11,186 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 5 | 12,312 | Just-in-Time Information Extraction using Extraction Views | 2012 | SIGMOD |
| 6 | 10,719 | Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents | 2025 | SIGMOD |
| 7 | 9,398 | Improving Information Extraction from Visually Rich Documents using Visual Span Representations | 2021 | VLDB |
| 8 | 8,677 | Visual Segmentation for Information Extraction from Heterogeneous Visually Rich Documents | 2019 | SIGMOD |
| 9 | 11,589 | Blueprint: A Constraint-solving Approach For Document Extraction | 2022 | VLDB |
| 10 | 9,399 | Glean: Structured Extractions from Templatic Documents | 2021 | VLDB |