The Design of an LLM-powered Unstructured Analytics System
Summary: Aryn compiles NL queries into semantic plans executed by Sycamore, a distributed declarative engine exposing DocSets to analyze, enrich, and transform large unstructured document collections. Luna (NL→Sycamore) and DocParse (PDF→DocSet) improve accuracy over RAG on NTSB reports and surface explainable execution traces to build trust. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Eric Anderson (Aryn, Inc.)
- 2. Jonathan Fritz (Aryn, Inc.)
- 3. Austin Lee (Aryn, Inc.)
- 4. Bohou Li (Aryn, Inc.)
- 5. Mark Lindblad (Aryn, Inc.)
- 6. Henry Lindeman (Aryn, Inc.)
- 7. Alex Meyer (Aryn, Inc.)
- 8. Parthkumar Parmar (Aryn, Inc.)
- 9. Tanvi Ranade (Aryn, Inc.)
- 10. Mehul A. Shah (Aryn, Inc.)
- 11. Benjamin Sowell (Aryn, Inc.)
- 12. Dan Tecuci (Aryn, Inc.)
- 13. Vinayak Thapliyal (Aryn, Inc.)
- 14. Matt Welsh (Aryn, Inc.)
BibTeX Citation
@inproceedings{anderson_cidr25,
address = {Amsterdam, Netherlands},
series = {{CIDR} '25},
title = {{The Design of an LLM-powered Unstructured Analytics System}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Anderson, Eric and Fritz, Jonathan and Lee, Austin and Li, Bohou and Lindblad, Mark and Lindeman, Henry and Meyer, Alex and Parmar, Parthkumar and Ranade, Tanvi and Shah, Mehul A. and Sowell, Benjamin and Tecuci, Dan and Thapliyal, Vinayak and Welsh, Matt},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 18 of 18 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 134 | Deep Entity Matching with Pre-Trained Language Models | 2021 | VLDB | 0.00030043481 |
| 329 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00020858443 |
| 501 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB | 0.00017267905 |
| 669 | CAESURA: Language Models as Multi-Modal Query Planners | 2024 | CIDR | 0.0001495987 |
| 683 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB | 0.00014817539 |
| 776 | Natural language to SQL: Where are we today? | 2020 | VLDB | 0.00014063545 |
| 1,780 | Annotating Columns with Pre-trained Language Models | 2022 | SIGMOD | 9.6560923e-05 |
| 1,932 | CHORUS: Foundation Models for Unified Data Discovery and Exploration | 2024 | VLDB | 9.34643e-05 |
| 2,606 | Text2SQL is Not Enough: Unifying AI and Databases with TAG | 2025 | CIDR | 8.2295646e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,161 | ScaleLLM: A Technique for Scalable LLM-augmented Data Systems | 2025 | SIGMOD |
| 2 | 11,004 | LLM-CER: An Interactive System for In-context Clustering-based Entity Resolution with Large Language Models | 2026 | VLDB |
| 3 | 683 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB |
| 4 | 7,885 | Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems | 2025 | VLDB |
| 5 | 4,142 | AOP: Automated and Interactive LLM Pipeline Orchestration for Answering Complex Queries | 2025 | CIDR |
| 6 | 11,038 | Semantic Data Systems: From Data Management to Data Understanding | 2026 | VLDB |
| 7 | 8,632 | DocDB: A Database for Unstructured Document Analysis | 2025 | VLDB |
| 8 | 10,144 | Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics | 2026 | CIDR |
| 9 | 8,680 | Unstructured Data Analysis Using LLMs: A Comprehensive Benchmark | 2026 | VLDB |
| 10 | 6,379 | Unify: A System For Unstructured Data Analytics | 2025 | VLDB |