The Design of an LLM-powered Unstructured Analytics System
Summary: Aryn compiles NL queries into semantic plans executed by Sycamore, a distributed declarative engine exposing DocSets to analyze, enrich, and transform large unstructured document collections. Luna (NL→Sycamore) and DocParse (PDF→DocSet) improve accuracy over RAG on NTSB reports and surface explainable execution traces to build trust. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Eric Anderson
- 2. Jonathan Fritz
- 3. Austin Lee
- 4. Bohou Li
- 5. Mark Lindblad
- 6. Henry Lindeman
- 7. Alex Meyer
- 8. Parthkumar Parmar
- 9. Tanvi Ranade
- 10. Mehul A. Shah
- 11. Benjamin Sowell
- 12. Dan Tecuci
- 13. Vinayak Thapliyal
- 14. Matt Welsh
Incoming Citations (Sorted by Pagerank)
Showing 13 of 13 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 219 | Deep Entity Matching with Pre-Trained Language Models | 2021 | VLDB | 0.00033354456 |
| 516 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00021194444 |
| 973 | Natural language to SQL: Where are we today? | 2020 | VLDB | 0.0001488435 |
| 997 | CAESURA: Language Models as Multi-Modal Query Planners | 2024 | CIDR | 0.00014726927 |
| 1,088 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB | 0.00014158762 |
| 1,839 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB | 0.00010351287 |
| 2,513 | Annotating Columns with Pre-trained Language Models | 2022 | SIGMOD | 8.6155767e-05 |
| 3,003 | Chorus: Foundation Models for Unified Data Discovery and Exploration | 2024 | VLDB | 7.7358219e-05 |
| 3,189 | Text2SQL is Not Enough: Unifying AI and Databases with TAG | 2025 | CIDR | 7.4140094e-05 |
Previous
Page 1 / 1
Next