NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task Parallelism
Summary: Introduce GNN task parallelism that partitions per-layer training tasks across GPUs (instead of graph partitioning), reducing neighbor replication, enabling intra-GPU shared neighbor-embedding reuse and overlapping subgraph computation. Combine with a task-decoupled training framework that releases intermediate data early to cut memory, enabling billion-scale full-graph multi-GPU training and 1.27×–5.47× speedups over NeutronStar/Sancus on 4×A5000. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Zhenbo Fu (Northeastern University)
- 2. Xin Ai (Northeastern University)
- 3. Qiange Wang (National University of Singapore)
- 4. Yanfeng Zhang (Northeastern University)
- 5. Shizhan Lu (Northeastern University)
- 6. Chaoyi Chen (Northeastern University)
- 7. Chunyu Cao (Northeastern University)
- 8. Hao Yuan (Northeastern University)
- 9. Zhewei Wei (Renmin University of China)
- 10. Yu Gu (Northeastern University)
- 11. Yingyou Wen (Neusoft AI Research)
- 12. Ge Yu (Northeastern University)
BibTeX Citation
@article{fu_vldb25,
title = {{NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task Parallelism}},
author = {Fu, Zhenbo and Ai, Xin and Wang, Qiange and Zhang, Yanfeng and Lu, Shizhan and Chen, Chaoyi and Cao, Chunyu and Yuan, Hao and Wei, Zhewei and Gu, Yu and Wen, Yingyou and Yu, Ge},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {6},
pages = {1705--1719},
doi = {10.14778/3725688.3725700},
url = {https://doi.org/10.14778/3725688.3725700},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.