Blueprint: A Constraint-solving Approach For Document Extraction
Summary: Blueprint declaratively extracts document fields via fuzzy spatial, textual, semantic, and numerical constraints, offering interpretable alternatives to deep learning. Studio’s no-code labeling and program synthesis address usability, achieving comparable accuracy and development time. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Andrey Mishchenko (University of Michigan)
- 2. Dominique Danco (University of Amsterdam)
- 3. Abhilash Jindal (Indian Institute of Technology Delhi)
- 4. Adrian Blue (Instabase)
BibTeX Citation
@article{mishchenko_vldb22,
title = {{Blueprint: A Constraint-solving Approach For Document Extraction}},
author = {Mishchenko, Andrey and Danco, Dominique and Jindal, Abhilash and Blue, Adrian},
journal = {PVLDB},
series = {{VLDB} '22},
volume = {15},
number = {12},
pages = {3459--3471},
doi = {10.14778/3554821.3554836},
url = {https://doi.org/10.14778/3554821.3554836},
year = {2022}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 91 | WebTables: Exploring the Power of Tables on the Web | 2008 | VLDB | 0.00034838835 |
| 319 | Declarative Information Extraction Using Datalog with Embedded Extraction Predicates | 2007 | VLDB | 0.00021377065 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,195 | When Speed Has a Price: Fast Information Extraction Using Approximate Algorithms | 2013 | VLDB |
| 2 | 13,690 | The SystemT IDE: An Integrated Development Environment for Information Extraction Rules | 2011 | SIGMOD |
| 3 | 2,348 | Document Spanners for Extracting Incomplete Information: Expressiveness and Complexity | 2018 | PODS |
| 4 | 1,343 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB |
| 5 | 11,186 | Unstructured Data Fusion for Schema and Data Extraction | 2024 | SIGMOD |
| 6 | 13,339 | DocDB: A Database for Unstructured Document Analysis | 2025 | VLDB |
| 7 | 9,399 | Glean: Structured Extractions from Templatic Documents | 2021 | VLDB |
| 8 | 8,340 | Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents | 2025 | VLDB |
| 9 | 10,414 | Visual Template Inference for Data Extraction from Documents | 2026 | SIGMOD |
| 10 | 10,719 | Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents | 2025 | SIGMOD |