spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines
Summary: spade synthesizes data-quality assertions for LLM pipelines by mining prompt-version histories to generate candidate assertion functions and selecting a minimal set meeting coverage and accuracy constraints. Yields fewer assertions and ~21% fewer false failures; deployed in LangSmith. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Shreya Shankar (University of California Berkeley)
- 2. Haotian Li (Hong Kong University of Science and Technology)
- 3. Parth Asawa (University of California Berkeley)
- 4. Madelon Hulsebos (University of California Berkeley)
- 5. Yiming Lin (University of California Berkeley)
- 6. J.D. Zamfirescu-Pereira (University of California Berkeley)
- 7. Harrison Chase (LangChain)
- 8. Will Fu-Hinthorn (LangChain)
- 9. Aditya G. Parameswaran (University of California Berkeley)
- 10. Eugene Wu (Columbia University)
BibTeX Citation
@article{shankar_vldb24,
title = {{spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines}},
author = {Shankar, Shreya and Li, Haotian and Asawa, Parth and Hulsebos, Madelon and Lin, Yiming and Zamfirescu-Pereira, J.D. and Chase, Harrison and Fu-Hinthorn, Will and Parameswaran, Aditya G. and Wu, Eugene},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {12},
pages = {4173--4186},
doi = {10.14778/3685800.3685835},
url = {https://doi.org/10.14778/3685800.3685835},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 7 of 7 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 683 | DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing | 2025 | VLDB | 0.00014817539 |
| 6,886 | Multi-Objective Agentic Rewrites for Unstructured Data Processing | 2026 | VLDB | 5.6551172e-05 |
| 7,885 | Semantic Integrity Constraints: Declarative Guardrails for AI-Augmented Data Processing Systems | 2025 | VLDB | 5.433531e-05 |
| 10,293 | D-Bot: An LLM-Powered DBA Copilot | 2025 | SIGMOD | 5.0431863e-05 |
| 10,353 | Making Prompts First-Class Citizens for Adaptive LLM Pipelines | 2026 | CIDR | 4.9793485e-05 |
| 11,303 | LLMLog: Advanced Log Template Generation via LLM-driven Multi-Round Annotation | 2025 | VLDB | 4.9793485e-05 |
| 13,633 | Prompt Editor: A Taxonomy-driven System for Guided LLM Prompt Development in Enterprise Settings | 2025 | SIGMOD | - |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 205 | Snorkel: Rapid Training Data Creation with Weak Supervision | 2018 | VLDB | 0.00025181304 |
| 329 | Can Foundation Models Wrangle Your Data? | 2023 | VLDB | 0.00020858443 |
| 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB | 0.0001107886 |
| 3,250 | How Large Language Models Will Disrupt Data Management | 2023 | VLDB | 7.4938546e-05 |
| 3,695 | Revisiting Prompt Engineering via Declarative Crowdsourcing | 2024 | CIDR | 7.092445e-05 |
| 4,672 | Data Platform for Machine Learning | 2019 | SIGMOD | 6.4759004e-05 |
| 6,552 | Finding Label and Model Errors in Perception Data With Learned Observation Assertions | 2022 | SIGMOD | 5.750057e-05 |
| 7,745 | Ease.ml/ci and Ease.ml/meter in Action: Towards Data Management for Statistical Generalization | 2019 | VLDB | 5.4606673e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,418 | BAT: Target-Instance-Free Data Preparation Synthesis via LLM-Driven Tree Search | 2026 | SIGMOD |
| 2 | 3,695 | Revisiting Prompt Engineering via Declarative Crowdsourcing | 2024 | CIDR |
| 3 | 11,468 | Welding Natural Language Queries to Analytics IRs with LLMs | 2024 | CIDR |
| 4 | 10,841 | Bolt-on, Verifiable Provenance for LLM-Powered Data Processing | 2026 | VLDB |
| 5 | 6,867 | AquaPipe: A Quality-Aware Pipeline for Knowledge Retrieval and Large Language Models | 2025 | SIGMOD |
| 6 | 1,308 | Automating Large-Scale Data Quality Verification | 2018 | VLDB |
| 7 | 11,284 | CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines | 2025 | VLDB |
| 8 | 8,997 | Cut Costs, Not Accuracy: LLM-Powered Data Processing with Guarantees | 2026 | SIGMOD |
| 9 | 10,893 | Developing and Benchmarking Verification Algorithms to Improve Text-to-SQL Generation | 2026 | VLDB |
| 10 | 10,353 | Making Prompts First-Class Citizens for Adaptive LLM Pipelines | 2026 | CIDR |