Back to papers
NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task Parallelism
Summary: Introduce GNN task parallelism that partitions per-layer training tasks across GPUs (instead of graph partitioning), reducing neighbor replication, enabling intra-GPU shared neighbor-embedding reuse and overlapping subgraph computation. Combine with a task-decoupled training framework that releases intermediate data early to cut memory, enabling billion-scale full-graph multi-GPU training and 1.27×–5.47× speedups over NeutronStar/Sancus on 4×A5000.
(summarized by gpt-5-mini on Feb 09 2026)
- Paper ID
- 13831
- Venue
- VLDB
- Year
- 2025
- Pagerank
- 4.1905499e-05
- Overall Rank
- 10,579 | 26.48%
- DOI
-
10.14778/3725688.3725700
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 271 |
AliGraph: A Comprehensive Graph Neural Network Platform |
2019 |
VLDB |
0.00029565193 |
| 1,162 |
Sancus: Staleness-Aware Communication-Avoiding Full-Graph Decentralized Training in Large-Scale Graph Neural Networks |
2022 |
VLDB |
0.00013573136 |
| 2,399 |
ByteGNN: Efficient Graph Neural Network Training at Large Scale |
2022 |
VLDB |
8.8869693e-05 |
| 2,425 |
DUCATI: A Dual-Cache Training System for Graph Neural Networks on Giant Graphs with the GPU |
2023 |
SIGMOD |
8.8414587e-05 |
| 3,028 |
NeutronStar: Distributed GNN Training with Hybrid Dependency Management |
2022 |
SIGMOD |
7.6833093e-05 |
| 3,092 |
Scalable and Efficient Full-Graph GNN Training for Large Graphs |
2023 |
SIGMOD |
7.5869574e-05 |
| 4,354 |
LargeEA: Aligning Entities for Large-scale Knowledge Graphs |
2022 |
VLDB |
6.2534656e-05 |
| 5,135 |
NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous Environments |
2024 |
VLDB |
5.6669017e-05 |
| 5,452 |
Decoupled Graph Neural Networks for Large Dynamic Graphs |
2023 |
VLDB |
5.4972958e-05 |
| 6,885 |
Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics Engines |
2023 |
VLDB |
4.8908367e-05 |
| 6,944 |
Efficient Training of Graph Neural Networks on Large Graphs |
2024 |
VLDB |
4.8875946e-05 |
| 6,979 |
OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single Machine |
2024 |
VLDB |
4.8697538e-05 |
| 7,087 |
HongTu: Scalable Full-Graph GNN Training on Multiple GPUs |
2023 |
SIGMOD |
4.8324242e-05 |
| 7,286 |
DAHA: Accelerating GNN Training with Data and Hardware Aware Execution Planning |
2024 |
VLDB |
4.7701372e-05 |
| 7,543 |
XGNN: Boosting Multi-GPU GNN Training via Global GNN Memory Store |
2024 |
VLDB |
4.7103673e-05 |
| 9,400 |
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism |
2025 |
VLDB |
4.3399748e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 7,087 |
HongTu: Scalable Full-Graph GNN Training on Multiple GPUs |
2023 |
SIGMOD |
4.8324242e-05 |
| 5,746 |
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective |
2024 |
VLDB |
5.3429324e-05 |
| 2,399 |
ByteGNN: Efficient Graph Neural Network Training at Large Scale |
2022 |
VLDB |
8.8869693e-05 |
| 5,356 |
NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph Streams |
2024 |
VLDB |
5.5514335e-05 |
| 3,092 |
Scalable and Efficient Full-Graph GNN Training for Large Graphs |
2023 |
SIGMOD |
7.5869574e-05 |
| 10,027 |
NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous Clusters |
2026 |
SIGMOD |
4.1905499e-05 |
| 3,028 |
NeutronStar: Distributed GNN Training with Hybrid Dependency Management |
2022 |
SIGMOD |
7.6833093e-05 |
| 10,310 |
NeutronCloud: Resource-Aware Distributed GNN Training in Fluctuating Cloud Environments |
2026 |
VLDB |
4.1905499e-05 |
| 5,135 |
NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous Environments |
2024 |
VLDB |
5.6669017e-05 |
| 9,400 |
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism |
2025 |
VLDB |
4.3399748e-05 |