DBScholar

Back to papers

spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines

Summary: spade synthesizes data-quality assertions for LLM pipelines by mining prompt-version histories to generate candidate assertion functions and selecting a minimal set meeting coverage and accuracy constraints. Yields fewer assertions and ~21% fewer false failures; deployed in LangSmith. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
13804
Venue
VLDB
Year
2024
Pagerank
6.6024412e-05
Overall Rank
4,621 | 68.30%
DOI
10.14778/3685800.3685835

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shankar_vldb24,
        title = {{spade: Synthesizing Data Quality Assertions for Large Language Model Pipelines}},
        author = {Shankar, Shreya and Li, Haotian and Asawa, Parth and Hulsebos, Madelon and Lin, Yiming and Zamfirescu-Pereira, J.D. and Chase, Harrison and Fu-Hinthorn, Will and Parameswaran, Aditya G. and Wu, Eugene},
        journal = {PVLDB},
        series = {{VLDB} '24},
        volume = {17},
        number = {12},
        pages = {4173--4186},
        doi = {10.14778/3685800.3685835},
        url = {https://doi.org/10.14778/3685800.3685835},
        year = {2024}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 8 of 8 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers