Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing
Summary: PALIMPZEST: declarative language + system to express AI-powered analytics over unstructured corpora, automating orchestration of models, prompts, and data operations. A cost-based optimizer searches model/prompt/implementation choices to trade latency, cost, and accuracy, yielding up to 3.3x speedup and 2.9x cost reduction with improved F1. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Chunwei Liu (Massachusetts Institute of Technology)
- 2. Matthew Russo (Massachusetts Institute of Technology)
- 3. Michael Cafarella (Massachusetts Institute of Technology)
- 4. Lei Cao (University of Arizona)
- 5. Peter Baile Chen (Massachusetts Institute of Technology)
- 6. Zui Chen (Massachusetts Institute of Technology)
- 7. Michael Franklin (University of Chicago)
- 8. Tim Kraska (Massachusetts Institute of Technology)
- 9. Samuel Madden (Massachusetts Institute of Technology)
- 10. Rana Shahout (Harvard University)
- 11. Gerardo Vitagliano (Massachusetts Institute of Technology)
BibTeX Citation
@inproceedings{liu_cidr25,
address = {Amsterdam, Netherlands},
series = {{CIDR} '25},
title = {{Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Liu, Chunwei and Russo, Matthew and Cafarella, Michael and Cao, Lei and Chen, Peter Baile and Chen, Zui and Franklin, Michael and Kraska, Tim and Madden, Samuel and Shahout, Rana and Vitagliano, Gerardo},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 31 of 31 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 713 | Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes | 2024 | VLDB | 0.00014672521 |
| 950 | CAESURA: Language Models as Multi-Modal Query Planners | 2024 | CIDR | 0.0001302491 |
| 4,050 | Revisiting Prompt Engineering via Declarative Crowdsourcing | 2024 | CIDR | 6.9368666e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,279 | AI Meets AI: Leveraging Query Executions to Improve Index Recommendations | 2019 | SIGMOD |
| 2 | 11,166 | PLAQUE: Automated Predicate Learning at Query Time | 2024 | SIGMOD |
| 3 | 13,040 | Investigation of Algebraic Query Optimisation for Database Programming Languages | 1994 | VLDB |
| 4 | 11,845 | Query-Driven Learning for Next Generation Predictive Modeling & Analytics | 2019 | SIGMOD |
| 5 | 5,220 | Databases Unbound: Querying All of the World’s Bytes with AI | 2024 | VLDB |
| 6 | 4,045 | Logical and Physical Optimizations for SQL Query Execution over Large Language Models | 2025 | SIGMOD |
| 7 | 250 | Answering Queries using Humans, Algorithms and Databases | 2011 | CIDR |
| 8 | 10,137 | Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics | 2026 | CIDR |
| 9 | 6,371 | QUEST: Query Optimization in Unstructured Document Analysis | 2025 | VLDB |
| 10 | 7,694 | PalimpChat: Declarative and Interactive AI analytics | 2025 | SIGMOD |