DBScholar

Back to papers

SketchML: Accelerating Distributed Machine Learning with Data Sketches

Summary: SketchML uses data sketches to compress distributed SGD gradients, targeting sparse gradients. Key ideas: quantile-sketch bucketization, MinMaxSketch collision-resolving hash tables, and delta-binary encoding; shows error bounds and 10x speedups on Tencent. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
5602
Venue
SIGMOD
Year
2018
Pagerank
6.9160949e-05
Overall Rank
4,083 | 71.99%
DOI
10.1145/3183713.3196894

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{jiang_sigmod18,
        title = {{SketchML: Accelerating Distributed Machine Learning with Data Sketches}},
        author = {Jiang, Jiawei and Fu, Fangcheng and Yang, Tong and Cui, Bin},
        series = {{SIGMOD} '18},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3183713.3196894},
        url = {https://dl.acm.org/doi/10.1145/3183713.3196894},
        year = {2018}
}

Incoming Citations (Sorted by Pagerank)

Showing 14 of 14 citing papers.

Rank Citing Paper Year Venue Pagerank
3,169 Towards Demystifying Serverless Machine Learning Training 2021 SIGMOD 7.6715222e-05
3,221 Camel: Managing Data for Efficient Stream Learning 2022 SIGMOD 7.6271601e-05
3,237 BlindFL: Vertical Federated Machine Learning without Peeking into Your Data 2022 SIGMOD 7.6089416e-05
3,898 BurstSketch: Finding Bursts in Data Streams 2021 SIGMOD 7.0367583e-05
4,956 Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce 2021 SIGMOD 6.4290135e-05
6,633 Agile and Accurate CTR Prediction Model Training for Massive-Scale Online Advertising Systems 2021 SIGMOD 5.8185371e-05
8,034 Angel-PTM: A Scalable and Economical Large-scale Pre-training System in Tencent 2023 VLDB 5.502946e-05
8,310 TreeSensing: Linearly Compressing Sketches with Flexibility 2023 SIGMOD 5.4556836e-05
8,794 AWARE: Workload-aware, Redundancy-exploiting Linear Algebra 2023 SIGMOD 5.370464e-05
8,869 A Distributed System for Large-scale n-gram Language Models at Tencent 2019 VLDB 5.3543688e-05
10,114 Towards Communication-efficient Vertical Federated Learning Training via Cache-enabled Local Updates 2022 VLDB 5.1319012e-05
10,613 CounterSnake: A lossless and generalized compression framework for diverse sketches 2026 VLDB 5.093636e-05
10,769 Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization 2025 SIGMOD 5.093636e-05
11,562 MinMax Sampling: A Near-optimal Global Summary for Aggregation in the Wide Area 2022 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 5 of 5 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
82 Space-Efficient Online Computation of Quantile Summaries 2001 SIGMOD 0.00036378991
1,150 DimmWitted: A Study of Main-Memory Statistical Analytics 2014 VLDB 0.00011943462
2,162 Heterogeneity-aware Distributed Parameter Servers 2017 SIGMOD 9.0581831e-05
4,529 TencentRec: Real-time Stream Recommendation in Practice 2015 SIGMOD 6.6446748e-05
11,999 LDA*: A Robust and Large-scale Topic Modeling System 2017 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Semantically Similar Papers