Back to papers
Modyn: Data-Centric Machine Learning Pipeline Orchestration
Summary: Modyn is a data-centric ML platform for growing datasets, declaratively configuring training with data-selection and triggering policies. Composite models for fair evaluation; open benchmarks; high-throughput, sample-level data selection.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 7051
- Venue
- SIGMOD
- Year
- 2025
- Pagerank
- 4.3648789e-05
- Overall Rank
- 9,234 | 35.83%
- DOI
-
10.1145/3709705
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 1 of 1 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 12 of 12 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 44 |
The Design Of Postgres |
1986 |
SIGMOD |
0.00071946446 |
| 1,481 |
Automating Large-Scale Data Quality Verification |
2018 |
VLDB |
0.00011715754 |
| 2,175 |
tf.data: A Machine Learning Data Processing Framework |
2021 |
VLDB |
9.3745231e-05 |
| 2,456 |
Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities |
2021 |
SIGMOD |
8.7649259e-05 |
| 2,690 |
Accelerating Recommendation System Training by Leveraging Popular Choices |
2022 |
VLDB |
8.2911466e-05 |
| 4,113 |
Learning to Validate the Predictions of Black Box Classifiers on Unseen Data |
2020 |
SIGMOD |
6.4326771e-05 |
| 4,423 |
PrIU: A Provenance-Based Approach for Incrementally Updating Regression Models |
2020 |
SIGMOD |
6.1925724e-05 |
| 7,257 |
Incremental Tabular Learning on Heterogeneous Feature Space |
2023 |
SIGMOD |
4.7819761e-05 |
| 8,166 |
Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science |
2021 |
VLDB |
4.567959e-05 |
| 8,178 |
DORIAN in action: Assisted Design of Data Science Pipelines |
2022 |
VLDB |
4.5629474e-05 |
| 8,253 |
Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines |
2023 |
SIGMOD |
4.5444167e-05 |
| 9,116 |
Towards Observability for Production Machine Learning Pipelines |
2022 |
VLDB |
4.3886184e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 2,175 |
tf.data: A Machine Learning Data Processing Framework |
2021 |
VLDB |
9.3745231e-05 |
| 9,116 |
Towards Observability for Production Machine Learning Pipelines |
2022 |
VLDB |
4.3886184e-05 |
| 6,464 |
Materialization and Reuse Optimizations for Production Data Science Pipelines |
2022 |
SIGMOD |
5.0471003e-05 |
| 10,776 |
cedar: Optimized and Unified Machine Learning Input Data Pipelines |
2025 |
VLDB |
4.1905499e-05 |
| 2,122 |
SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle |
2020 |
CIDR |
9.4905306e-05 |
| 8,253 |
Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines |
2023 |
SIGMOD |
4.5444167e-05 |
| 11,315 |
Towards Observability for Machine Learning Pipelines |
2022 |
CIDR |
4.1905499e-05 |
| 2,456 |
Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities |
2021 |
SIGMOD |
8.7649259e-05 |
| 4,006 |
Data Platform for Machine Learning |
2019 |
SIGMOD |
6.5371762e-05 |
| 7,309 |
The Machine Learning Bazaar: Harnessing the ML Ecosystem for Effective System Development |
2020 |
SIGMOD |
4.7611148e-05 |