GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research
Summary: GPEmu: a GPU emulator enabling DL systems prototyping/evaluation without real GPUs via time emulation, memory emulation, distributed execution and sharing support. Scales to 30+ models and 6 GPU configs, reproduces 9 papers' results and speeds micro-optimization prototyping. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Meng Wang (University of Chicago)
- 2. Gus Waldspurger (University of Chicago)
- 3. Naufal Ananda (Telkom University)
- 4. Yuyang Huang (University of Chicago)
- 5. Kemas Wiharja (Telkom University)
- 6. John Bent (Los Alamos National Laboratory)
- 7. Swaminathan Sundararaman (IBM)
- 8. Vijay Chidambaram (University of Texas)
- 9. Haryadi S. Gunawi (University of Chicago)
BibTeX Citation
@article{wang_vldb25,
title = {{GPEmu: A GPU Emulator for Faster and Cheaper Prototyping and Evaluation of Deep Learning System Research}},
author = {Wang, Meng and Waldspurger, Gus and Ananda, Naufal and Huang, Yuyang and Wiharja, Kemas and Bent, John and Sundararaman, Swaminathan and Chidambaram, Vijay and Gunawi, Haryadi S.},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {6},
pages = {1919--1932},
doi = {10.14778/3725688.3725716},
url = {https://doi.org/10.14778/3725688.3725716},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 8 of 8 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 521 | PyTorch Distributed: Experiences on Accelerating Data Parallel Training | 2020 | VLDB | 0.0001713368 |
| 1,446 | Analyzing and Mitigating Data Stalls in DNN Training | 2021 | VLDB | 0.0001076818 |
| 2,018 | tf.data: A Machine Learning Data Processing Framework | 2021 | VLDB | 9.3001937e-05 |
| 2,473 | PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel | 2023 | VLDB | 8.5326287e-05 |
| 2,688 | Accelerating Recommendation System Training by Leveraging Popular Choices | 2022 | VLDB | 8.2564305e-05 |
| 3,764 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB | 7.1446931e-05 |
| 5,282 | GoldMiner: Elastic Scaling of Training Data Pre-Processing Pipelines for Deep Learning | 2023 | SIGMOD | 6.2840712e-05 |
| 6,339 | In-Database Machine Learning with CorgiPile: Stochastic Gradient Descent without Full Data Shuffle | 2022 | SIGMOD | 5.907165e-05 |
Previous
Page 1 / 1
Next