DBScholar

Back to papers

NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task Parallelism

Summary: Introduce GNN task parallelism that partitions per-layer training tasks across GPUs (instead of graph partitioning), reducing neighbor replication, enabling intra-GPU shared neighbor-embedding reuse and overlapping subgraph computation. Combine with a task-decoupled training framework that releases intermediate data early to cut memory, enabling billion-scale full-graph multi-GPU training and 1.27×–5.47× speedups over NeutronStar/Sancus on 4×A5000. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
h71a59d88b6811c91
Venue
VLDB
Year
2025
Pagerank
5.499936e-05
Overall Rank
7,527 | 49.42%
DOI
10.14778/3725688.3725700
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{fu_vldb25,
        title = {{NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task Parallelism}},
        author = {Fu, Zhenbo and Ai, Xin and Wang, Qiange and Zhang, Yanfeng and Lu, Shizhan and Chen, Chaoyi and Cao, Chunyu and Yuan, Hao and Wei, Zhewei and Gu, Yu and Wen, Yingyou and Yu, Ge},
        journal = {PVLDB},
        series = {{VLDB} '25},
        volume = {18},
        number = {6},
        pages = {1705--1719},
        doi = {10.14778/3725688.3725700},
        url = {https://doi.org/10.14778/3725688.3725700},
        year = {2025}
}

Incoming Citations (Sorted by Pagerank)

Showing 2 of 2 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
211 AliGraph: A Comprehensive Graph Neural Network Platform 2019 VLDB 0.00024805216
1,134 SANCUS: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural Networks 2022 VLDB 0.0001188789
1,772 ByteGNN: Efficient Graph Neural Network Training at Large Scale 2022 VLDB 9.6746467e-05
2,281 NeutronStar: Distributed GNN Training with Hybrid Dependency Management 2022 SIGMOD 8.7021423e-05
2,506 DUCATI: A Dual-Cache Training System for Graph Neural Networks on Giant Graphs with the GPU 2023 SIGMOD 8.3735089e-05
2,598 Scalable and Efficient Full-Graph GNN Training for Large Graphs 2023 SIGMOD 8.238044e-05
4,457 LargeEA: Aligning Entities for Large-scale Knowledge Graphs 2022 VLDB 6.5896326e-05
4,827 DAHA: Accelerating GNN Training with Data and Hardware Aware Execution Planning 2024 VLDB 6.3893293e-05
4,855 NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous Environments 2024 VLDB 6.377545e-05
5,055 Decoupled Graph Neural Networks for Large Dynamic Graphs 2023 VLDB 6.293158e-05
5,901 Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics Engines 2023 VLDB 5.9517243e-05
6,673 Efficient Training of Graph Neural Networks on Large Graphs 2024 VLDB 5.7120259e-05
6,798 XGNN: Boosting Multi-GPU GNN Training via Global GNN Memory Store 2024 VLDB 5.6783443e-05
6,858 OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single Machine 2024 VLDB 5.6620796e-05
6,927 HongTu: Scalable Full-Graph GNN Training on Multiple GPUs 2023 SIGMOD 5.6399587e-05
8,158 NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism 2025 VLDB 5.3870998e-05
Previous Page 1 / 1 Next

Semantically Similar Papers