Back to papers
Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities
Summary: Analyzes 3,000 production ML pipelines at Google via provenance graphs and 450k trainings to characterize lifespan, topology, and complexity. Introduces model graphlets, a data model for repeated components, and shows optimization opportunities—pruning wasted computation can cut costs by ~50% without delaying deployment.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h2094ded95d203b3f
Venue
SIGMOD
Year
2021
Pagerank
8.8896655e-05
Overall Rank
2,187 | 85.30%
DOI
10.1145/3448016.3457566
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{xin_sigmod21,
title = {{Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities}},
author = {Xin, Doris and Miao, Hui and Parameswaran, Aditya and Polyzotis, Neoklis},
series = {{SIGMOD} '21},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3448016.3457566},
url = {https://dl.acm.org/doi/10.1145/3448016.3457566},
year = {2021}
}
Incoming Citations (Sorted by Pagerank)
Showing 16 of 16 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
4,925
TPCx-AI - An Industry Standard Benchmark for Artificial Intelligence and Machine Learning Systems
2023
VLDB
6.3511742e-05
6,662
UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads
2022
VLDB
5.7171651e-05
6,921
CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning
2024
SIGMOD
5.6432616e-05
7,538
Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines
2023
SIGMOD
5.4995874e-05
8,384
AWARE: Workload-aware, Redundancy-exploiting Linear Algebra
2023
SIGMOD
5.3421754e-05
8,968
Pipemizer: An Optimizer for Analytics Data Pipelines
2022
VLDB
5.2471751e-05
9,070
Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models
2025
SIGMOD
5.2283159e-05
9,117
Modyn: Data-Centric Machine Learning Pipeline Orchestration
2025
SIGMOD
5.2263399e-05
10,855
PipeLens: Identifying Interventions for Resolving Malfunctioning Data Science Pipelines
2026
VLDB
4.9793485e-05
10,898
stratum: A System Infrastructure for Massive Agent-Centric ML Workloads
2026
VLDB
4.9793485e-05
10,932
IMLane: Composable Framework for Efficient AI Function Execution in Database Engine
2026
VLDB
4.9793485e-05
10,961
SemPiper: Interactive Code Synthesis for Semantic Operators in Machine Learning Pipelines
2026
VLDB
4.9793485e-05
11,410
APEX-DAG: Library and Language independent Pipeline EXtraction
2025
VLDB
4.9793485e-05
11,589
Efficiently Mitigating the Impact of Data Drift on Machine Learning Pipelines
2024
VLDB
4.9793485e-05
11,731
Demystifying the QoS and QoE of Edge-hosted Video Streaming Applications in the Wild with SNESet
2023
SIGMOD
4.9793485e-05
11,754
Enabling Secure and Efficient Data Analytics Pipeline Evolution with Trusted Execution Environment
2023
VLDB
4.9793485e-05
Outgoing Citations (Sorted by Pagerank)
Showing 14 of 14 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
105
The MADlib Analytics Library or MAD Skills, the SQL
2012
VLDB
0.00033638251
509
Goods: Organizing Google's Datasets
2016
SIGMOD
0.00017071087
1,153
Data Management Challenges in Production Machine Learning
2017
SIGMOD
0.00011798912
1,564
Titian: Data Provenance Support in Spark
2016
VLDB
0.00010222394
1,568
HELIX: Holistic Optimization for Accelerating Iterative Machine Learning
2019
VLDB
0.0001021302
1,765
Putting Lipstick on Pig: Enabling Database-style Workflow Provenance
2012
VLDB
9.6955585e-05
1,880
Ground: A Data Context Service
2017
CIDR
9.4415148e-05
1,927
Elastic Machine Learning Algorithms in Amazon SageMaker
2020
SIGMOD
9.3607648e-05
2,182
Extending Relational Query Processing with ML Inference
2020
CIDR
8.8982998e-05
2,445
noWorkflow: a Tool for Collecting, Analyzing, and Managing Provenance from Python Scripts
2017
VLDB
8.4570447e-05
4,161
The Relational Data Borg is Learning
2020
VLDB
6.7700593e-05
4,801
Improving Reproducibility of Data Science Pipelines through Transparent Provenance Capture
2020
VLDB
6.4104797e-05
5,748
An Optimal Labeling Scheme for Workflow Provenance Using Skeleton Labels
2010
SIGMOD
6.0079624e-05
6,219
Incremental View Maintenance For Collection Programming
2016
PODS
5.8471493e-05
Semantically Similar Papers
#
Overall Rank
Paper
Year
Venue
1
2,053
tf.data: A Machine Learning Data Processing Framework
2021
VLDB
2
3,601
Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines
2022
SIGMOD
3
12,125
Leveraging Organizational Resources to Adapt Models to New Data Modalities
2020
VLDB
4
5,792
Optimizing Machine Learning Workloads in Collaborative Environments
2020
SIGMOD
5
11,825
Data Management Opportunities for Foundation Models
2022
CIDR
6
7,212
Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science
2021
VLDB
7
11,821
Towards Observability for Machine Learning Pipelines
2022
CIDR
8
1,153
Data Management Challenges in Production Machine Learning
2017
SIGMOD
9
9,417
Towards Observability for Production Machine Learning Pipelines
2022
VLDB
10
6,217
Materialization and Reuse Optimizations for Production Data Science Pipelines
2022
SIGMOD