DBScholar

Back to papers

HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework

Summary: HET scales huge embedding training with a cache-enabled distributed framework that exploits skewed popularity. Embedding-level consistency with write-time staleness enables cache coherence, yielding up to 88% comms reduction and 20.68x speedup. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12981
Venue
VLDB
Year
2022
Pagerank
8.5145736e-05
Overall Rank
2,485 | 82.96%
DOI
10.14778/3489496.3489511

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{miao_vldb22,
        title = {{HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework}},
        author = {Miao, Xupeng and Zhang, Hailin and Shi, Yining and Nie, Xiaonan and Yang, Zhi and Tao, Yangyu and Cui, Bin},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {2},
        pages = {312--320},
        doi = {10.14778/3489496.3489511},
        url = {https://doi.org/10.14778/3489496.3489511},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 22 of 22 citing papers.

Rank Citing Paper Year Venue Pagerank
3,237 BlindFL: Vertical Federated Machine Learning without Peeking into Your Data 2022 SIGMOD 7.6089416e-05
3,634 Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees 2023 SIGMOD 7.2358691e-05
4,878 NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph Streams 2024 VLDB 6.4684388e-05
4,891 DGC: Training Dynamic Graphs with Spatio-Temporal Non-Uniformity using Graph Partitioning by Chunks 2023 SIGMOD 6.4581865e-05
4,912 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 6.4481656e-05
5,003 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.4065691e-05
6,900 Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving 2025 SIGMOD 5.7430032e-05
7,166 Accelerating Graph Indexing for ANNS on Modern CPUs 2025 SIGMOD 5.6846045e-05
7,190 PetPS: Supporting Huge Embedding Models with Persistent Memory 2023 VLDB 5.6772818e-05
8,034 Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent 2023 VLDB 5.502946e-05
8,101 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.4870581e-05
8,883 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement 2023 SIGMOD 5.351513e-05
8,907 Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models 2025 SIGMOD 5.3483178e-05
9,164 FEC: Efficient Deep Recommendation Model Training with Flexible Embedding Communication 2023 SIGMOD 5.3101785e-05
9,520 Experimental Analysis of Large-scale Learnable Vector Storage Compression 2024 VLDB 5.2561283e-05
9,556 CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models 2024 SIGMOD 5.2528121e-05
9,729 Scalable Graph Convolutional Network Training on Distributed-Memory Systems 2023 VLDB 5.2289669e-05
9,880 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.2040783e-05
10,114 Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local Updates 2022 VLDB 5.1319012e-05
10,309 A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and Effectiveness 2026 SIGMOD 5.093636e-05
11,187 GE2: A General and Efficient Knowledge Graph Embedding Learning System 2024 SIGMOD 5.093636e-05
11,464 EmbedX: A Versatile, Efficient and Scalable Platform to Embed Both Graphs and High-Dimensional Sparse Data 2023 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 6 of 6 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers