OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single Machine
Summary: Identifies neighborhood and temporal redundancies in out-of-core sampling-based GNN training and reframes the bottleneck as excessive overall data-request volume rather than cache-hit optimization. OUTRE uses partition-based batch construction, a historical-embedding cache, and automatic cache-space management to de-redundancy I/O on a single machine, yielding 1.52–3.51× speedups vs. state-of-the-art. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Zeang Sheng (Peking University)
- 2. Wentao Zhang (Peking University)
- 3. Yangyu Tao (Tencent)
- 4. Bin Cui (Peking University)
BibTeX Citation
@article{sheng_vldb24,
title = {{OUTRE: An OUT-of-core De-REdundancy GNN Training Framework for Massive Graphs within A Single Machine}},
author = {Sheng, Zeang and Zhang, Wentao and Tao, Yangyu and Cui, Bin},
journal = {PVLDB},
series = {{VLDB} '24},
volume = {17},
number = {11},
pages = {2960--2973},
doi = {10.14778/3681954.3681976},
url = {https://doi.org/10.14778/3681954.3681976},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 3 of 3 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 10,323 | NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous Clusters | 2026 | SIGMOD | 5.093636e-05 |
| 10,835 | NeutronTask: Scalable and Efficient Multi-GPU GNN Training with Task Parallelism | 2025 | VLDB | 5.093636e-05 |
| 10,890 | Heta: Distributed Training of Heterogeneous Graph Neural Networks | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 223 | AliGraph: A Comprehensive Graph Neural Network Platform | 2019 | VLDB | 0.00024182473 |
| 1,234 | Large Graph Convolutional Network Training with GPU-Oriented Data Communication Architecture | 2021 | VLDB | 0.00011549432 |
| 2,953 | Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory Caching | 2022 | VLDB | 7.9237794e-05 |
| 5,243 | FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network Training | 2024 | VLDB | 6.3018825e-05 |
| 5,597 | Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses | 2024 | VLDB | 6.1547003e-05 |
Previous
Page 1 / 1
Next