DistME: A Fast and Elastic Distributed Matrix Computation Engine using GPUs
Summary: Introduces DistME, a fast elastic distributed matrix computation engine built atop Spark, combining CuboidMM with GPU acceleration. CuboidMM partitions matrices into cuboids to minimize network traffic under memory constraints, while subcuboid GPU partitioning reduces PCIe costs; experiments show superior performance and scalability over existing distributed matrix mult methods. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Donghyoung Han (Daegu Gyeongbuk Institute of Science and Technology)
- 2. Yoon-Min Nam (Daegu Gyeongbuk Institute of Science and Technology)
- 3. Jihye Lee (Daegu Gyeongbuk Institute of Science and Technology)
- 4. Kyongseok Park (Korea Advanced Institute of Science and Technology)
- 5. Hyunwoo Kim (Korea Advanced Institute of Science and Technology)
- 6. Min-Soo Kim (Daegu Gyeongbuk Institute of Science and Technology)
BibTeX Citation
@inproceedings{han_sigmod19,
title = {{DistME: A Fast and Elastic Distributed Matrix Computation Engine using GPUs}},
author = {Han, Donghyoung and Nam, Yoon-Min and Lee, Jihye and Park, Kyongseok and Kim, Hyunwoo and Kim, Min-Soo},
series = {{SIGMOD} '19},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3299869.3319865},
url = {https://dl.acm.org/doi/10.1145/3299869.3319865},
year = {2019}
}
Incoming Citations (Sorted by Pagerank)
Showing 5 of 5 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,195 | Tensor Relational Algebra for Distributed Machine Learning System Design | 2021 | VLDB | 6.3232105e-05 |
| 8,350 | FuseME: Distributed Matrix Computation Engine based on Cuboid-based Fused Operator and Plan Generation | 2022 | SIGMOD | 5.4460082e-05 |
| 9,840 | Distributed Numerical and Machine Learning Computations via Two-Phase Execution of Aggregated Join Trees | 2021 | VLDB | 5.2103367e-05 |
| 11,537 | Redundancy Elimination in Distributed Matrix Computation | 2022 | SIGMOD | 5.093636e-05 |
| 11,669 | Hybrid Evaluation for Distributed Iterative Matrix Computation | 2021 | SIGMOD | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 9 of 9 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 239 | Overview of SciDB: Large Scale Array Storage, Processing and Analysis | 2010 | SIGMOD | 0.00023674329 |
| 415 | SystemML: Declarative Machine Learning on Spark | 2016 | VLDB | 0.0001888524 |
| 769 | A Comparison of Join Algorithms for Log Processing in MapReduce | 2010 | SIGMOD | 0.00014166872 |
| 1,235 | Towards Linear Algebra over Normalized Data | 2017 | VLDB | 0.00011548457 |
| 2,681 | Exploiting Matrix Dependency for Efficient Distributed Matrix Computation | 2015 | SIGMOD | 8.2632778e-05 |
| 3,459 | A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics | 2018 | VLDB | 7.3953716e-05 |
| 3,740 | GTS: A Fast and Scalable Graph Processing Method based on Streaming Topology to GPUs | 2016 | SIGMOD | 7.1593902e-05 |
| 6,030 | The BUDS Language for Distributed Bayesian Machine Learning | 2017 | SIGMOD | 6.0009079e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,537 | Redundancy Elimination in Distributed Matrix Computation | 2022 | SIGMOD |
| 2 | 5,587 | Distributed GPU Joins on Fast RDMA-capable Networks | 2023 | SIGMOD |
| 3 | 11,978 | Template Skycube Algorithms for Heterogeneous Parallelism on Multicore and GPU Architectures | 2017 | SIGMOD |
| 4 | 4,088 | Fast Sparse Matrix-Vector Multiplication on GPUs: Implications for Graph Mining | 2011 | VLDB |
| 5 | 3,303 | A Distributed Multi-GPU System for Fast Graph Processing | 2018 | VLDB |
| 6 | 11,669 | Hybrid Evaluation for Distributed Iterative Matrix Computation | 2021 | SIGMOD |
| 7 | 12,036 | An Efficient MapReduce Cube Algorithm for Varied Data Distributions | 2016 | SIGMOD |
| 8 | 6,046 | Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra | 2021 | SIGMOD |
| 9 | 8,350 | FuseME: Distributed Matrix Computation Engine based on Cuboid-based Fused Operator and Plan Generation | 2022 | SIGMOD |
| 10 | 2,681 | Exploiting Matrix Dependency for Efficient Distributed Matrix Computation | 2015 | SIGMOD |