The Design of an LLM-powered Unstructured Analytics System
Summary: Aryn compiles NL queries into semantic plans executed by Sycamore, a distributed declarative engine exposing DocSets to analyze, enrich, and transform large unstructured document collections. Luna (NL→Sycamore) and DocParse (PDF→DocSet) improve accuracy over RAG on NTSB reports and surface explainable execution traces to build trust. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Eric Anderson (Aryn, Inc.)
- 2. Jonathan Fritz (Aryn, Inc.)
- 3. Austin Lee (Aryn, Inc.)
- 4. Bohou Li (Aryn, Inc.)
- 5. Mark Lindblad (Aryn, Inc.)
- 6. Henry Lindeman (Aryn, Inc.)
- 7. Alex Meyer (Aryn, Inc.)
- 8. Parthkumar Parmar (Aryn, Inc.)
- 9. Tanvi Ranade (Aryn, Inc.)
- 10. Mehul A. Shah (Aryn, Inc.)
- 11. Benjamin Sowell (Aryn, Inc.)
- 12. Dan Tecuci (Aryn, Inc.)
- 13. Vinayak Thapliyal (Aryn, Inc.)
- 14. Matt Welsh (Aryn, Inc.)
BibTeX Citation
@inproceedings{anderson_cidr25,
address = {Amsterdam, Netherlands},
series = {{CIDR} '25},
title = {{The Design of an LLM-powered Unstructured Analytics System}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Anderson, Eric and Fritz, Jonathan and Lee, Austin and Li, Bohou and Lindblad, Mark and Lindeman, Henry and Meyer, Alex and Parmar, Parthkumar and Ranade, Tanvi and Shah, Mehul A. and Sowell, Benjamin and Tecuci, Dan and Thapliyal, Vinayak and Welsh, Matt},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 14 of 14 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 141 | Deep Entity Matching with Pre-Trained Language Models | 2021 | VLDB | 0.0002964847 |
| 420 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00018789852 |
| 713 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB | 0.00014672521 |
| 865 | Natural language to SQL: Where are we today? | 2020 | VLDB | 0.00013521464 |
| 950 | CAESURA: Language Models as Multi-Modal Query Planners | 2024 | CIDR | 0.0001302491 |
| 1,343 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB | 0.00011095866 |
| 1,923 | Annotating Columns with Pre-trained Language Models | 2022 | SIGMOD | 9.4789109e-05 |
| 2,242 | CHORUS: Foundation Models for Unified Data Discovery and Exploration | 2024 | VLDB | 8.8823802e-05 |
| 2,809 | Text2SQL is Not Enough: Unifying AI and Databases with TAG | 2025 | CIDR | 8.0994951e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 13,375 | Reimagining Deep Learning Systems Through the Lens of Data Systems | 2024 | VLDB |
| 2 | 5,220 | Databases Unbound: Querying All of the World’s Bytes with AI | 2024 | VLDB |
| 3 | 10,286 | ScaleDoc: Scaling LLM-based Predicates over Large Document Collections | 2026 | SIGMOD |
| 4 | 10,733 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD |
| 5 | 1,343 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB |
| 6 | 9,841 | Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems | 2025 | VLDB |
| 7 | 7,439 | AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries | 2025 | CIDR |
| 8 | 13,339 | DocDB: A Database for Unstructured Document Analysis | 2025 | VLDB |
| 9 | 10,137 | Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics | 2026 | CIDR |
| 10 | 9,307 | Unify: A System For Unstructured Data Analytics | 2025 | VLDB |