DBScholar

Back to papers

Towards Observability for Production Machine Learning Pipelines

Summary: End-to-end observability for production ML pipelines to address post-deployment issues like data shift and silent failures. Proposes a bolt-on data-management architecture enabling detection, diagnosis, and reaction, wrapping existing tools to deliver ML observability. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hf4a515ecd64ba641
Venue
VLDB
Year
2022
Pagerank
5.1803615e-05
Overall Rank
9,417 | 36.69%
DOI
10.14778/3565838.3565853

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{shankar_vldb22,
        title = {{Towards Observability for Production Machine Learning Pipelines}},
        author = {Shankar, Shreya and Parameswaran, Aditya G.},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {13},
        pages = {4015--4022},
        doi = {10.14778/3565838.3565853},
        url = {https://doi.org/10.14778/3565838.3565853},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 1 of 1 citing papers.

Rank Citing Paper Year Venue Pagerank
9,117 Modyn: Data-Centric Machine Learning Pipeline Orchestration 2025 SIGMOD 5.2263399e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 27 of 27 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
57 On Random Sampling over Joins 1999 SIGMOD 0.00040108301
104 HoloClean: Holistic Data Repairs with Probabilistic Inference 2017 VLDB 0.00033690989
336 The Aqua Approximate Query Answering System 1999 SIGMOD 0.00020657819
406 Deep Unsupervised Cardinality Estimation 2020 VLDB 0.00019045544
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017590977
840 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.0001354605
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011798912
1,308 Automating Large-Scale Data Quality Verification 2018 VLDB 0.0001107886
1,344 Detecting Data Errors: Where are we and what needs to be done? 2016 VLDB 0.00010956518
1,691 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 9.8570722e-05
1,880 Ground: A Data Context Service 2017 CIDR 9.4415148e-05
1,927 Elastic Machine Learning Algorithms in Amazon SageMaker 2020 SIGMOD 9.3607648e-05
2,055 Combining Quantitative and Logical Data Cleaning 2016 VLDB 9.1128336e-05
2,326 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6309237e-05
3,997 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 6.8655222e-05
4,324 On Biased Reservoir Sampling in the Presence of Stream Evolution 2006 VLDB 6.6656407e-05
4,606 MLINSPECT: A Data Distribution Debugger for Machine Learning Pipelines 2021 SIGMOD 6.5041225e-05
4,672 Data Platform for Machine Learning 2019 SIGMOD 6.4759004e-05
5,252 Dagger: A Data (not code) Debugger 2020 CIDR 6.2087063e-05
5,831 Joins on Samples: A Theoretical Guide for Practitioners 2020 VLDB 5.9782109e-05
6,037 ReproZip: Computational Reproducibility With Ease 2016 SIGMOD 5.908316e-05
6,221 Combining Aggregation and Sampling (Nearly) Optimally for Approximate Query Processing 2021 SIGMOD 5.8463347e-05
6,652 Hindsight Logging for Model Training 2021 VLDB 5.7202005e-05
7,212 Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science 2021 VLDB 5.5840773e-05
9,470 VisClean: Interactive Cleaning for Progressive Visualization 2020 VLDB 5.1714419e-05
11,821 Towards Observability for Machine Learning Pipelines 2022 CIDR 4.9793485e-05
Previous Page 1 / 1 Next

Semantically Similar Papers