Back to papers
MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training
Summary: Memo enables ultra-long context LLM training via fine-grained activation memory management: offloads activations to CPU after each layer and fetches them in backprop with token-wise recomputation. Bi-level MIP optimizes cross-layer memory reuse to curb fragmentation and communication, delivering MFU gains over Megatron-LM and DeepSpeed.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 7049
- Venue
- SIGMOD
- Year
- 2025
- Pagerank
- 4.2799988e-05
- Overall Rank
- 9,786 | 31.99%
- DOI
-
10.1145/3709703
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 20 of 20 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 1,500 |
Analyzing and Mitigating Data Stalls in DNN Training |
2021 |
VLDB |
0.00011636174 |
| 2,175 |
tf.data: A Machine Learning Data Processing Framework |
2021 |
VLDB |
9.3745231e-05 |
| 2,336 |
Concurrent Analytical Query Processing with GPUs |
2014 |
VLDB |
9.0106308e-05 |
| 2,357 |
MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud |
2023 |
VLDB |
8.9688205e-05 |
| 2,678 |
HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework |
2022 |
VLDB |
8.3224016e-05 |
| 2,907 |
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel |
2023 |
VLDB |
7.9322286e-05 |
| 3,694 |
Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines |
2022 |
SIGMOD |
6.8316905e-05 |
| 3,899 |
Efficient Join Algorithms For Large Database Tables in a Multi-GPU Environment |
2021 |
VLDB |
6.6513982e-05 |
| 4,054 |
Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees |
2023 |
SIGMOD |
6.4910804e-05 |
| 4,175 |
FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline |
2023 |
VLDB |
6.3772575e-05 |
| 4,700 |
Tensors: An abstraction for general data processing |
2021 |
VLDB |
5.9810592e-05 |
| 5,130 |
Memory Management Techniques for Large-Scale Persistent-Main-Memory Systems |
2017 |
VLDB |
5.6721688e-05 |
| 5,836 |
Tensor Relational Algebra for Distributed Machine Learning System Design |
2021 |
VLDB |
5.3079723e-05 |
| 6,157 |
Optimizing Tensor Programs on Flexible Storage |
2023 |
SIGMOD |
5.1755682e-05 |
| 6,361 |
Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism |
2023 |
VLDB |
5.0903244e-05 |
| 7,014 |
SIMPLE: Efficient Temporal Graph Neural Network Training at Scale with Dynamic Data Placement |
2024 |
SIGMOD |
4.8570865e-05 |
| 8,117 |
SDPipe: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training |
2023 |
VLDB |
4.5788485e-05 |
| 8,161 |
TOD: GPU-accelerated Outlier Detection via Tensor Operations |
2023 |
VLDB |
4.5688249e-05 |
| 9,408 |
CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models |
2024 |
SIGMOD |
4.3399748e-05 |
| 9,414 |
Experimental Analysis of Large-scale Learnable Vector Storage Compression |
2024 |
VLDB |
4.3399748e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 8,807 |
FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement |
2023 |
SIGMOD |
4.4413307e-05 |
| 13,101 |
OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration |
2026 |
VLDB |
- |
| 3,569 |
Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation |
2025 |
SIGMOD |
6.9588368e-05 |
| 6,350 |
PQCache: Product Quantization-based KVCache for Long Context LLM Inference |
2025 |
SIGMOD |
5.0957601e-05 |
| 10,143 |
Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache Management |
2026 |
SIGMOD |
4.1905499e-05 |
| 8,518 |
mLoRA: Fine-Tuning LoRA Adapters via Highly-Efficient Pipeline Parallelism in Multiple GPUs |
2025 |
VLDB |
4.4893996e-05 |
| 9,875 |
Near-Duplicate Sequence Search at Scale for Large Language Model Memorization Evaluation |
2023 |
SIGMOD |
4.2626861e-05 |
| 10,222 |
RetroInfer: A Vector Storage Engine for Scalable Long-Context LLM Inference |
2026 |
VLDB |
4.1905499e-05 |
| 9,174 |
MemFlow: Memory-Aware Distributed Deep Learning |
2020 |
SIGMOD |
4.3807157e-05 |
| 7,143 |
Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity |
2024 |
VLDB |
4.8143774e-05 |