FusionFlow: Accelerating Data Preprocessing for Machine Learning with CPU-GPU Cooperation
Summary: FusionFlow accelerates dynamic DL augmentation via CPU-GPU co-scheduling, managing shared GPU memory to avoid overflow and interference with training. Adaptive scheduling reallocates resources across heterogeneous tasks, cutting CPU requirements 50–60% while improving throughput. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Taeyoon Kim (Ulsan National Institute of Science and Technology)
- 2. ChanHo Park (Ulsan National Institute of Science and Technology)
- 3. Mansur Mukimbekov (Ulsan National Institute of Science and Technology)
- 4. Heelim Hong (Ulsan National Institute of Science and Technology)
- 5. Minseok Kim (Ulsan National Institute of Science and Technology)
- 6. Ze Jin (ByteDance)
- 7. Changdae Kim (Electronics and Telecommunications Research Institute)
- 8. Ji-Yong Shin (Northeastern University)
- 9. Myeongjae Jeon (Ulsan National Institute of Science and Technology)
BibTeX Citation
@article{kim_vldb24,
title = {{FusionFlow: Accelerating Data Preprocessing for Machine Learning with CPU-GPU Cooperation}},
author = {Kim, Taeyoon and Park, ChanHo and Mukimbekov, Mansur and Hong, Heelim and Kim, Minseok and Jin, Ze and Kim, Changdae and Shin, Ji-Yong and Jeon, Myeongjae},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {4},
pages = {863--876},
doi = {10.14778/3636218.3636238},
url = {https://doi.org/10.14778/3636218.3636238},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,905 | TensorSocket: Shared Data Loading for Deep Learning Training | 2026 | SIGMOD | 5.3483178e-05 |
| 10,999 | cedar: Optimized and Unified Machine Learning Input Data Pipelines | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,446 | Analyzing and Mitigating Data Stalls in DNN Training | 2021 | VLDB | 0.0001076818 |
| 2,018 | tf.data: A Machine Learning Data Processing Framework | 2021 | VLDB | 9.3001937e-05 |
| 3,764 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB | 7.1446931e-05 |
| 4,956 | Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce | 2021 | SIGMOD | 6.4290135e-05 |
| 5,282 | GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning | 2023 | SIGMOD | 6.2840712e-05 |
Previous
Page 1 / 1
Next