A Distributed System for Large-scale n-gram Language Models at Tencent
Summary: Tencent's distributed system for large n-gram LMs; indexing, batching, caching reduce network traffic and speed nodes. Cascade fault-tolerance switches to smaller n-grams on failure; shown on 9 ASR datasets and deployed to WeChat at 100M messages/min. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Qiang Long (Tencent)
- 2. Wei Wang (National University of Singapore)
- 3. Jinfu Deng (Tencent)
- 4. Song Liu (Tencent)
- 5. Wenhao Huang (Tencent)
- 6. Fangying Chen (Tencent)
- 7. Sifan Liu (Tencent)
BibTeX Citation
@article{long_vldb19,
title = {{A Distributed System for Large-scale n-gram Language Models at Tencent}},
author = {Long, Qiang and Wang, Wei and Deng, Jinfu and Liu, Song and Huang, Wenhao and Chen, Fangying and Liu, Sifan},
journal = {PVLDB},
series = {{VLDB} '19},
volume = {12},
number = {12},
pages = {2206--2217},
doi = {10.14778/3352063.3352136},
url = {https://doi.org/10.14778/3352063.3352136},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 2 of 2 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 8,855 | ANN Softmax: Acceleration of Extreme Classification Training | 2022 | VLDB | 5.3573227e-05 |
| 11,196 | A Universal Sketch for Estimating Heavy Hitters and Per-Element Frequency Moments in Data Streams with Bounded Deletions | 2024 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 284 | NoScope: Optimizing Neural Network Queries over Video at Scale | 2017 | VLDB | 0.00022370521 |
| 295 | Accelerating Machine Learning Inference with Probabilistic Predicates | 2018 | SIGMOD | 0.00022238183 |
| 1,792 | MacroBase: Prioritizing Attention in Fast Data | 2017 | SIGMOD | 9.7436856e-05 |
| 4,083 | SketchML: Accelerating Distributed Machine Learning with Data Sketches | 2018 | SIGMOD | 6.9160949e-05 |
| 7,218 | Fast Failure Recovery in Distributed Graph Processing Systems | 2015 | VLDB | 5.6683391e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 20 | Distributed GraphLab: A Framework for Machine Learning and Data Mining in the Cloud | 2012 | VLDB |
| 2 | 2,640 | Scalable and Efficient Full-Graph GNN Training for Large Graphs | 2023 | SIGMOD |
| 3 | 8,454 | D3-GNN: Dynamic Distributed Dataflow for Streaming Graph Neural Networks | 2024 | VLDB |
| 4 | 14,018 | Communication-Efficient Distributed Mining of Association Rules | 2001 | SIGMOD |
| 5 | 4,912 | HET-GMP: A Graph-based System Approach to Scaling Large Embedding Model Training | 2022 | SIGMOD |
| 6 | 2,162 | Heterogeneity-aware Distributed Parameter Servers | 2017 | SIGMOD |
| 7 | 8,034 | Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent | 2023 | VLDB |
| 8 | 3,109 | Managing Large Dynamic Graphs Efficiently | 2012 | SIGMOD |
| 9 | 7,900 | Distributed Graph Embedding with Information-Oriented Random Walks | 2023 | VLDB |
| 10 | 1,875 | Large-Scale Distributed Graph Computing Systems: An Experimental Evaluation | 2015 | VLDB |