Back to papers
Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities
Summary: Analyzes 3,000 production ML pipelines at Google via provenance graphs and 450k trainings to characterize lifespan, topology, and complexity. Introduces model graphlets, a data model for repeated components, and shows optimization opportunities—pruning wasted computation can cut costs by ~50% without delaying deployment.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h2094ded95d203b3f
Venue
SIGMOD
Year
2021
Pagerank
8.8854572e-05
Overall Rank
2,189 | 85.29%
DOI
10.1145/3448016.3457566
PDF
Download
(CC BY 4.0)
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{xin_sigmod21,
title = {{Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities}},
author = {Xin, Doris and Miao, Hui and Parameswaran, Aditya and Polyzotis, Neoklis},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457566},
url = {https://dl.acm.org/doi/10.1145/3448016.3457566},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 16 of 16 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
4,926
TPCx-AI - An Industry Standard Benchmark for Artificial Intelligence and Machine Learning Systems
2023
VLDB
6.3481677e-05
6,666
UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads
2022
VLDB
5.7144587e-05
6,924
CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning
2024
SIGMOD
5.6405901e-05
7,544
Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines
2023
SIGMOD
5.496984e-05
8,389
AWARE: Workload-aware, Redundancy-exploiting Linear Algebra
2023
SIGMOD
5.3396465e-05
8,978
Pipemizer: An Optimizer for Analytics Data Pipelines
2022
VLDB
5.2446911e-05
9,079
Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models
2025
SIGMOD
5.2258409e-05
9,126
Modyn: Data-Centric Machine Learning Pipeline Orchestration
2025
SIGMOD
5.2238659e-05
10,864
PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
2026
VLDB
4.9769913e-05
10,907
stratum: A System Infrastructure for Massive Agent-Centric ML Workloads
2026
VLDB
4.9769913e-05
10,941
IMLane: Composable Framework for Efficient AI Function Execution in Database Engine
2026
VLDB
4.9769913e-05
10,970
SemPiper: Interactive Code Synthesis for Semantic Operators in Machine Learning Pipelines
2026
VLDB
4.9769913e-05
11,416
APEX-DAG: Library and Language independent Pipeline EXtraction
2025
VLDB
4.9769913e-05
11,595
Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines
2024
VLDB
4.9769913e-05
11,737
Demystifying the QoS and QoE of Edge-hosted Video Streaming Applications in the Wild with SNESet
2023
SIGMOD
4.9769913e-05
11,760
Enabling Secure and Efficient Data Analytics Pipeline Evolution with Trusted Execution Environment
2023
VLDB
4.9769913e-05
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
105
The MADlib Analytics Library or MAD Skills, the SQL
2012
VLDB
0.00033633007
509
Goods: Organizing Google's Datasets
2016
SIGMOD
0.00017063491
1,153
Data Management Challenges in Production Machine Learning
2017
SIGMOD
0.00011793347
1,564
Titian: Data Provenance Support in Spark
2016
VLDB
0.00010219222
1,568
HELIX: Holistic Optimization for Accelerating Iterative Machine Learning
2019
VLDB
0.00010208225
1,767
Putting Lipstick on Pig: Enabling Database-style Workflow Provenance
2012
VLDB
9.6910812e-05
1,881
Ground: A Data Context Service
2017
CIDR
9.4372419e-05
1,928
Elastic Machine Learning Algorithms in Amazon SageMaker
2020
SIGMOD
9.3563363e-05
2,184
Extending Relational Query Processing with ML Inference
2020
CIDR
8.8953085e-05
2,447
noWorkflow: a Tool for Collecting, Analyzing, and Managing Provenance from Python Scripts
2017
VLDB
8.4530494e-05
4,161
The Relational Data Borg is Learning
2020
VLDB
6.7669004e-05
4,804
Improving Reproducibility of Data Science Pipelines through Transparent Provenance Capture
2020
VLDB
6.4074452e-05
5,749
An Optimal Labeling Scheme for Workflow Provenance Using Skeleton Labels
2010
SIGMOD
6.0051184e-05
6,223
Incremental View Maintenance For Collection Programming
2016
PODS
5.8443814e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
2,055
tf.data: A Machine Learning Data Processing Framework
2021
VLDB
2
3,601
Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines
2022
SIGMOD
3
12,131
Leveraging Organizational Resources to Adapt Models to New Data Modalities
2020
VLDB
4
5,794
Optimizing Machine Learning Workloads in Collaborative Environments
2020
SIGMOD
5
11,831
Data Management Opportunities for Foundation Models
2022
CIDR
6
11,827
Towards Observability for Machine Learning Pipelines
2022
CIDR
7
7,214
Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science
2021
VLDB
8
1,153
Data Management Challenges in Production Machine Learning
2017
SIGMOD
9
9,426
Towards Observability for Production Machine Learning Pipelines
2022
VLDB
10
6,222
Materialization and Reuse Optimizations for Production Data Science Pipelines
2022
SIGMOD