Measuring and Optimizing Distributed Array Programs
Summary: Proposes KASEN, an array-based distributed model with a strict compute/communication contract for performance analysis. An optimizer applies transformations to programs, cutting memory I/O, buffers, and network traffic, delivering up to 5.82x speedup. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Mingxing Zhang (Research Institute of Tsinghua University in Shenzhen; Tsinghua University; Yangtze Delta Region Institute of Tsinghua University)
- 2. Yongwei Wu (Research Institute of Tsinghua University in Shenzhen; Tsinghua University; Yangtze Delta Region Institute of Tsinghua University)
- 3. Kang Chen (Research Institute of Tsinghua University in Shenzhen; Tsinghua University; Yangtze Delta Region Institute of Tsinghua University)
- 4. Teng Ma (Research Institute of Tsinghua University in Shenzhen; Tsinghua University; Yangtze Delta Region Institute of Tsinghua University)
- 5. Weimin Zheng (Research Institute of Tsinghua University in Shenzhen; Tsinghua University; Yangtze Delta Region Institute of Tsinghua University)
BibTeX Citation
@article{zhang_vldb16,
title = {{Measuring and Optimizing Distributed Array Programs}},
author = {Zhang, Mingxing and Wu, Yongwei and Chen, Kang and Ma, Teng and Zheng, Weimin},
journal = {PVLDB},
series = {{VLDB} '16},
volume = {9},
number = {12},
pages = {912--923},
doi = {10.14778/2994509.2994511},
url = {https://doi.org/10.14778/2994509.2994511},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,250 | Data Management in Machine Learning: Challenges, Techniques, and Systems | 2017 | SIGMOD | 0.00011485301 |
| 3,205 | On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML | 2018 | VLDB | 7.6386536e-05 |
| 3,284 | SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning | 2017 | CIDR | 7.5663058e-05 |
| 9,475 | BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach | 2023 | SIGMOD | 5.2634238e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 3 of 3 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 239 | Overview of SciDB: Large Scale Array Storage, Processing and Analysis | 2010 | SIGMOD | 0.00023674329 |
| 1,654 | Navigating the Maze of Graph Analytics Frameworks using Massive Graph Datasets | 2014 | SIGMOD | 0.00010104703 |
| 4,137 | Optimizing I/O for Big Array Analytics | 2012 | VLDB | 6.8783832e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,079 | Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML | 2014 | VLDB |
| 2 | 11,669 | Hybrid Evaluation for Distributed Iterative Matrix Computation | 2021 | SIGMOD |
| 3 | 8,394 | Topology-aware Parallel Data Processing: Models, Algorithms and Systems at Scale | 2020 | CIDR |
| 4 | 5,195 | Tensor Relational Algebra for Distributed Machine Learning System Design | 2021 | VLDB |
| 5 | 4,049 | Resource Elasticity for Large-Scale Machine Learning | 2015 | SIGMOD |
| 6 | 1,911 | Fast Iterative Graph Computation with Block Updates | 2013 | VLDB |
| 7 | 4,137 | Optimizing I/O for Big Array Analytics | 2012 | VLDB |
| 8 | 2,681 | Exploiting Matrix Dependency for Efficient Distributed Matrix Computation | 2015 | SIGMOD |
| 9 | 8,738 | Translation of Array-Based Loops to Distributed Data-Parallel Programs | 2020 | VLDB |
| 10 | 6,046 | Automatic Optimization of Matrix Implementations for Distributed Machine Learning and Linear Algebra | 2021 | SIGMOD |