DBScholar

Back to papers

Towards Observability for Production Machine Learning Pipelines

Summary: End-to-end observability for production ML pipelines to address post-deployment issues like data shift and silent failures. Proposes a bolt-on data-management architecture enabling detection, diagnosis, and reaction, wrapping existing tools to deliver ML observability. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
13099
Venue
VLDB
Year
2022
Pagerank
5.2992628e-05
Overall Rank
9,245 | 36.58%
DOI
10.14778/3565838.3565853

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shankar_vldb22,
        title = {{Towards Observability for Production Machine Learning Pipelines}},
        author = {Shankar, Shreya and Parameswaran, Aditya G.},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {13},
        pages = {4015--4022},
        doi = {10.14778/3565838.3565853},
        url = {https://doi.org/10.14778/3565838.3565853},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
8,949 Modyn: Data-Centric Machine Learning Pipeline Orchestration 2025 SIGMOD 5.3462965e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 27 of 27 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
54 On Random Sampling over Joins 1999 SIGMOD 0.00040810225
112 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00032801121
327 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00021091539
401 Deep Unsupervised Cardinality Estimation 2020 VLDB 0.00019092557
582 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00016148948
819 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013815639
1,147 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011974846
1,350 Automating Large-Scale Data Quality Verification 2018 VLDB 0.00011065626
1,351 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00011064851
1,670 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 0.00010045615
1,902 Ground: A Data Context Service 2017 CIDR 9.506714e-05
1,937 Elastic Machine Learning Algorithms in Amazon SageMaker 2020 SIGMOD 9.4524758e-05
2,041 Combining Quantitative and Logical Data Cleaning 2016 VLDB 9.2692538e-05
2,273 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.8230899e-05
3,947 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 7.0040437e-05
4,229 On Biased Reservoir Sampling in the Presence of Stream Evolution 2006 VLDB 6.8181027e-05
4,518 MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines 2021 SIGMOD 6.6474737e-05
4,592 Data Platform for Machine Learning 2019 SIGMOD 6.6139071e-05
5,404 Dagger: A Data (not code) Debugger 2020 CIDR 6.2294192e-05
5,743 Joins on Samples: A Theoretical Guide for Practitioners 2020 VLDB 6.1025457e-05
5,914 ReproZip: Computational Reproducibility With Ease 2016 SIGMOD 6.0439255e-05
6,206 Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing 2021 SIGMOD 5.9443409e-05
6,525 Hindsight Logging for Model Training 2021 VLDB 5.8514923e-05
8,050 Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science 2021 VLDB 5.5000099e-05
9,298 VisClean: Interactive Cleaning for Progressive Visualization 2020 VLDB 5.2901384e-05
11,512 Towards Observability for Machine Learning Pipelines 2022 CIDR 5.093636e-05
Previous Page 1 / 1 Next

Semantically Similar Papers