DBScholar

Back to papers

tf.data: A Machine Learning Data Processing Framework

Summary: tf.data is a compositional framework and runtime for high-throughput ML input pipelines, hiding parallelism, caching, overlap, and tuning. Fleet-scale analysis shows diverse preprocessing consumes substantial training resources, motivating cross-job sharing and storage-level projection. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h8aab8372df9b0f76
Venue
VLDB
Year
2021
Pagerank
9.1140407e-05
Overall Rank
2,053 | 86.20%
DOI
10.14778/3476311.3476374

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{murray_vldb21,
        title = {{tf.data: A Machine Learning Data Processing Framework}},
        author = {Murray, Derek G. and Šimša, Jiří and Klimovic, Ana and Indyk, Ihor},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {12},
        pages = {2945--2958},
        doi = {10.14778/3476311.3476374},
        url = {https://doi.org/10.14778/3476311.3476374},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 16 of 16 citing papers.

Rank Citing Paper Year Venue Pagerank
2,662 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.1596229e-05
3,601 Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines 2022 SIGMOD 7.1757877e-05
3,811 FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline 2023 VLDB 7.0087478e-05
5,381 GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning 2023 SIGMOD 6.1550396e-05
6,311 Progressive Compressed Records: Taking a Byte out of Deep Learning Data 2021 VLDB 5.8175614e-05
6,662 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7171651e-05
7,350 Bullion: A Column Store for Machine Learning 2025 CIDR 5.5451228e-05
7,827 FusionFlow: Accelerating Data Preprocessing for Machine Learning with CPU-GPU Cooperation 2024 VLDB 5.4442509e-05
9,067 TensorSocket: Shared Data Loading for Deep Learning Training 2026 SIGMOD 5.2283159e-05
9,070 Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models 2025 SIGMOD 5.2283159e-05
9,117 Modyn: Data-Centric Machine Learning Pipeline Orchestration 2025 SIGMOD 5.2263399e-05
10,000 cedar: Optimized and Unified Machine Learning Input Data Pipelines 2025 VLDB 5.0979044e-05
10,042 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.0921006e-05
10,659 Mixtera: A Data Plane for Foundation Model Training 2026 SIGMOD 4.9793485e-05
10,694 FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads 2026 VLDB 4.9793485e-05
11,250 GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research 2025 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 6 of 6 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers