DBScholar

Back to papers

HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework

Summary: HET scales huge embedding training with a cache-enabled distributed framework that exploits skewed popularity. Embedding-level consistency with write-time staleness enables cache coherence, yielding up to 88% comms reduction and 20.68x speedup. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hc3c6da0c9ca9eb11
Venue
VLDB
Year
2022
Pagerank
8.3521095e-05
Overall Rank
2,519 | 83.07%
DOI
10.14778/3489496.3489511

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{miao_vldb22,
        title = {{HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework}},
        author = {Miao, Xupeng and Zhang, Hailin and Shi, Yining and Nie, Xiaonan and Yang, Zhi and Tao, Yangyu and Cui, Bin},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {2},
        pages = {312--320},
        doi = {10.14778/3489496.3489511},
        url = {https://doi.org/10.14778/3489496.3489511},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 22 of 22 citing papers.

Rank Citing Paper Year Venue Pagerank
3,311 BlindFL: Vertical Federated Machine Learning without Peeking into Your Data 2022 SIGMOD 7.4392547e-05
3,399 Orca: Scalable Temporal Graph Neural Network Training with Theoretical Guarantees 2023 SIGMOD 7.3349076e-05
4,502 NeutronStream: A Dynamic GNN Training Framework with Sliding Window for Graph Streams 2024 VLDB 6.5740722e-05
4,945 HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training 2022 SIGMOD 6.3435891e-05
4,993 DGC: Training Dynamic Graphs with Spatio-Temporal Non-Uniformity using Graph Partitioning by Chunks 2023 SIGMOD 6.32357e-05
5,009 Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism 2023 VLDB 6.3158866e-05
6,166 Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving 2025 SIGMOD 5.8631131e-05
7,004 Accelerating Graph Indexing for ANNS on Modern CPUs 2025 SIGMOD 5.6238624e-05
7,333 PetPS: Supporting Huge Embedding Models with Persistent Memory 2023 VLDB 5.5498988e-05
8,197 Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent 2023 VLDB 5.3794747e-05
8,251 SDP_PIPE: A Semi-Decentralized Framework for Heterogeneity-aware Pipeline-parallel Training 2023 VLDB 5.3687311e-05
8,812 FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement 2023 SIGMOD 5.2722828e-05
9,070 Scheduling Data Processing Pipelines for Incremental Training on MLP-based Recommendation Models 2025 SIGMOD 5.2283159e-05
9,339 FEC: Efficient Deep Recommendation Model Training with Flexible Embedding Communication 2023 SIGMOD 5.1910676e-05
9,700 Experimental Analysis of Large-scale Learnable Vector Storage Compression 2024 VLDB 5.1381949e-05
9,734 CAFE: Towards Compact, Adaptive, and Fast Embedding for Large-scale Recommendation Models 2024 SIGMOD 5.1349531e-05
9,900 Scalable Graph Convolutional Network Training on Distributed-Memory Systems 2023 VLDB 5.1116429e-05
10,042 MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training 2025 SIGMOD 5.0921006e-05
10,342 Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local Updates 2022 VLDB 5.0167904e-05
10,520 A Comprehensive Benchmark on Spectral GNNs: The Impact on Efficiency, Memory, and Effectiveness 2026 SIGMOD 4.9793485e-05
11,530 GE2: A General and Efficient Knowledge Graph Embedding Learning System 2024 SIGMOD 4.9793485e-05
11,776 EmbedX: A Versatile, Efficient and Scalable Platform to Embed Both Graphs and High-Dimensional Sparse Data 2023 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 6 of 6 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers