Back to papers
Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
Summary: Evaporate: an LLM-based system that converts heterogeneous documents into queryable tables using in‑context learning rather than domain-specific training. Evaporate‑Code+ ensembles many synthesized extractors with weak supervision to approach/exceed direct extraction quality while using a sublinear LLM pass (≈110× fewer document calls).
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13766
- Venue
- VLDB
- Year
- 2024
- Pagerank
- 0.00014158762
- Overall Rank
- 1,088 | 92.45%
- DOI
-
10.14778/3626292.3626294
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 28 of 28 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 1,839 |
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing |
2025 |
VLDB |
0.00010351287 |
| 2,013 |
Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing |
2025 |
CIDR |
9.7986166e-05 |
| 3,003 |
Chorus: Foundation Models for Unified Data Discovery and Exploration |
2024 |
VLDB |
7.7358219e-05 |
| 3,639 |
The Design of an LLM-powered Unstructured Analytics System |
2025 |
CIDR |
6.8886648e-05 |
| 5,429 |
Logical and Physical Optimizations for SQL Query Execution over Large Language Models |
2025 |
SIGMOD |
5.511638e-05 |
| 5,506 |
Can Large Language Models Predict Data Correlations from Column Names? |
2023 |
VLDB |
5.4711611e-05 |
| 5,669 |
Databases Unbound: Querying All of the World's Bytes with AI |
2024 |
VLDB |
5.3805024e-05 |
| 7,027 |
Mind the Data Gap: Bridging LLMs to Enterprise Data Integration |
2025 |
CIDR |
4.8524216e-05 |
| 7,369 |
ELEET: Efficient Learned Query Execution over Text and Tables |
2024 |
VLDB |
4.7452331e-05 |
| 7,676 |
E2ETune: End-to-End Knob Tuning via Fine-tuned Generative Language Model |
2025 |
VLDB |
4.6770108e-05 |
| 7,703 |
AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries |
2025 |
CIDR |
4.668568e-05 |
| 8,464 |
Semantic Operators and Their Optimization: Enabling LLM-Based Data Processing with Accuracy Guarantees in LOTUS |
2025 |
VLDB |
4.5003888e-05 |
| 8,479 |
Can Large Language Models Be Query Optimizer for Relational Databases? |
2026 |
SIGMOD |
4.4967983e-05 |
| 8,518 |
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs |
2025 |
VLDB |
4.4893996e-05 |
| 9,152 |
Doctopus: Budget-aware Structural Table Extraction from Unstructured Documents |
2025 |
VLDB |
4.380727e-05 |
| 9,971 |
KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration |
2026 |
CIDR |
4.1905499e-05 |
| 10,064 |
Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,115 |
ST-Raptor: LLM-Powered Semi-Structured Table Question Answering |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,126 |
Visual Template Inference for Data Extraction from Documents |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,215 |
Task Cascades for Efficient Unstructured Data Processing |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,285 |
Relational Deep Dive: Error-Aware Queries Over Unstructured Data |
2026 |
VLDB |
4.1905499e-05 |
| 10,448 |
Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,465 |
Sentence to Model: Cost-Effective Data Collection LLM Agent |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,466 |
SwellDB: Dynamic Query-Driven Table Generation with Large Language Models |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,603 |
Optimized Batch Prompting for Cost-effective LLMs |
2025 |
VLDB |
4.1905499e-05 |
| 10,720 |
CoLA: Model Collaboration for Log-based Anomaly Detection |
2025 |
VLDB |
4.1905499e-05 |
| 10,758 |
QUEST: Query Optimization in Unstructured Document Analysis |
2025 |
VLDB |
4.1905499e-05 |
| 11,071 |
Chameleon: Foundation Models for Fairness-aware Multi-modal Data Augmentation to Enhance Coverage of Minorities |
2024 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 10,976 |
Unstructured Data Fusion for Schema and Data Extraction |
2024 |
SIGMOD |
4.1905499e-05 |
| 10,603 |
Optimized Batch Prompting for Cost-effective LLMs |
2025 |
VLDB |
4.1905499e-05 |
| 8,157 |
Automated Data Visualization from Natural Language via Large Language Models: An Exploratory Study |
2024 |
SIGMOD |
4.5701714e-05 |
| 13,152 |
Database Perspective on LLM Inference Systems |
2025 |
VLDB |
- |
| 8,732 |
Unveiling Challenges for LLMs in Enterprise Data Engineering |
2026 |
VLDB |
4.4520434e-05 |
| 7,703 |
AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries |
2025 |
CIDR |
4.668568e-05 |
| 3,982 |
How Large Language Models Will Disrupt Data Management |
2023 |
VLDB |
6.5595332e-05 |
| 10,803 |
A Demonstration of QueryArtisan: Real-Time Data Lake Analysis via Dynamically Generated Data Manipulation Code |
2025 |
VLDB |
4.1905499e-05 |
| 9,960 |
QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data Lakes |
2025 |
VLDB |
4.2254157e-05 |
| 7,016 |
LLM for Data Management |
2024 |
VLDB |
4.8561622e-05 |