DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization
Summary: DFLOP makes multimodal LLM pipeline parallelism data-aware by profiling input-induced cost variance and using predictive stage/microbatch scheduling. It mitigates heterogeneous-modality skew, improving utilization and throughput by up to 3.6× over existing frameworks. (summarized by gpt-5.6-luna on Jul 26 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Hyeonjun An (Yonsei University)
- 2. Sihyun Kim (Yonsei University)
- 3. Chaerim Lim (Yonsei University)
- 4. Hyunjoon Kim (Yonsei University)
- 5. Rathijit Sen (Microsoft)
- 6. Sangmin Jung (SK Telecom)
- 7. Hyeonsoo Lee (SK Telecom)
- 8. Dongwook Kim (SK Telecom)
- 9. Takki Yu (SK Telecom)
- 10. Jinkyu Jeong (Yonsei University)
- 11. Youngsok Kim (Yonsei University)
- 12. Kwanghyun Park (Yonsei University)
BibTeX Citation
@inproceedings{an_sigmod26,
title = {{DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization}},
author = {An, Hyeonjun and Kim, Sihyun and Lim, Chaerim and Kim, Hyunjoon and Sen, Rathijit and Jung, Sangmin and Lee, Hyeonsoo and Kim, Dongwook and Yu, Takki and Jeong, Jinkyu and Kim, Youngsok and Park, Kwanghyun},
series = {{SIGMOD} '26},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3802037},
url = {https://dl.acm.org/doi/10.1145/3802037},
year = {2026}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 521 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB | 0.0001713368 |
| 1,079 | Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML | 2014 | VLDB | 0.00012258469 |
| 2,473 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel | 2023 | VLDB | 8.5326287e-05 |
| 2,823 | Query Processing on Tensor Computation Runtimes | 2022 | VLDB | 8.0893814e-05 |
| 2,865 | End-to-end Optimization of Machine Learning Prediction Queries | 2022 | SIGMOD | 8.0180243e-05 |
| 5,003 | Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism | 2023 | VLDB | 6.4065691e-05 |
| 5,704 | Optimizing Data Pipelines for Machine Learning in Feature Stores | 2023 | VLDB | 6.1146371e-05 |
| 6,538 | UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads | 2022 | VLDB | 5.8477764e-05 |
Previous
Page 1 / 1
Next