Back to papers
tf.data: A Machine Learning Data Processing Framework
Summary: tf.data offers composable, parameterized operators for efficient ML input pipelines; the runtime overlaps I/O and compute with minimal tuning. Google workloads show data processing dominates resources and motivates cross-job sharing and storage projection.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 12504
- Venue
- VLDB
- Year
- 2021
- Pagerank
- 9.3745231e-05
- Overall Rank
- 2,175 | 84.89%
- DOI
-
10.14778/3476311.3476374
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 16 of 16 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 3,409 |
End-to-end Optimization of Machine Learning Prediction Queries |
2022 |
SIGMOD |
7.1240791e-05 |
| 3,694 |
Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines |
2022 |
SIGMOD |
6.8316905e-05 |
| 4,175 |
FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline |
2023 |
VLDB |
6.3772575e-05 |
| 5,560 |
GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning |
2023 |
SIGMOD |
5.4350242e-05 |
| 6,059 |
Progressive Compressed Records: Taking a Byte out of Deep Learning Data |
2021 |
VLDB |
5.2272544e-05 |
| 7,470 |
Bullion: A Column Store for Machine Learning |
2025 |
CIDR |
4.7159125e-05 |
| 8,343 |
FusionFlow: Accelerating Data Preprocessing for Machine Learning with CPU-GPU Cooperation |
2024 |
VLDB |
4.5366487e-05 |
| 8,515 |
UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads |
2022 |
VLDB |
4.4901466e-05 |
| 8,731 |
TensorSocket: Shared Data Loading for Deep Learning Training |
2026 |
SIGMOD |
4.4520434e-05 |
| 8,733 |
Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models |
2025 |
SIGMOD |
4.4520434e-05 |
| 9,234 |
Modyn: Data-Centric Machine Learning Pipeline Orchestration |
2025 |
SIGMOD |
4.3648789e-05 |
| 9,786 |
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training |
2025 |
SIGMOD |
4.2799988e-05 |
| 10,183 |
Mixtera: A Data Plane for Foundation Model Training |
2026 |
SIGMOD |
4.1905499e-05 |
| 10,220 |
FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads |
2026 |
VLDB |
4.1905499e-05 |
| 10,589 |
GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research |
2025 |
VLDB |
4.1905499e-05 |
| 10,776 |
cedar: Optimized and Unified Machine Learning Input Data Pipelines |
2025 |
VLDB |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 6 of 6 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 8,343 |
FusionFlow: Accelerating Data Preprocessing for Machine Learning with CPU-GPU Cooperation |
2024 |
VLDB |
4.5366487e-05 |
| 4,006 |
Data Platform for Machine Learning |
2019 |
SIGMOD |
6.5371762e-05 |
| 5,560 |
GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning |
2023 |
SIGMOD |
5.4350242e-05 |
| 8,515 |
UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads |
2022 |
VLDB |
4.4901466e-05 |
| 2,179 |
Spinning Fast Iterative Data Flows |
2012 |
VLDB |
9.3632007e-05 |
| 2,456 |
Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities |
2021 |
SIGMOD |
8.7649259e-05 |
| 3,694 |
Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines |
2022 |
SIGMOD |
6.8316905e-05 |
| 10,776 |
cedar: Optimized and Unified Machine Learning Input Data Pipelines |
2025 |
VLDB |
4.1905499e-05 |
| 3,495 |
TensorFlow Data Validation: Data Analysis and Validation in Continuous ML Pipelines |
2020 |
SIGMOD |
7.0383507e-05 |
| 4,175 |
FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline |
2023 |
VLDB |
6.3772575e-05 |