DBScholar

Back to papers

Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML

Summary: SystemML combines task and data parallelism for declarative, large-scale ML via a generic ParFOR construct over MapReduce. A cost-based optimizer automatically selects multi-core and cluster execution plans, adapting to workloads and unknown data characteristics. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
ha5d406fbdd20049f
Venue
VLDB
Year
2014
Pagerank
0.00012118261
Overall Rank
1,082 | 92.73%
DOI
10.14778/2732296.2732302
PDF
Download (CC BY-NC-ND 3.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{boehm_vldb14,
        title = {{Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML}},
        author = {Boehm, Matthias and Tatikonda, Shirish and Reinwald, Berthold and Sen, Prithviraj and Tian, Yuanyuan and Burdick, Douglas R. and Vaithyanathan, Shivakumar},
        journal = {PVLDB},
        series = {{VLDB} '14},
        volume = {7},
        number = {7},
        pages = {553--564},
        doi = {10.14778/2732296.2732302},
        url = {https://doi.org/10.14778/2732296.2732302},
        year = {2014}
}

Incoming Citations (Sorted by Pagerank)

Showing 36 of 36 citing papers.

Rank Citing Paper Year Venue Pagerank
416 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00018650998
521 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.00016923519
1,152 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011796404
1,223 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011468426
1,257 Towards Linear Algebra over Normalized Data 2017 VLDB 0.0001130959
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
2,266 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.7248802e-05
2,329 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6268411e-05
2,724 Exploiting Matrix Dependency for Efficient Distributed Matrix Computation 2015 SIGMOD 8.090674e-05
2,978 In-Database Learning with Sparse Tensors 2018 PODS 7.7872011e-05
3,103 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6456038e-05
3,242 Towards Demystifying Serverless Machine Learning Training 2021 SIGMOD 7.4967268e-05
3,518 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.236638e-05
4,097 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.8064316e-05
4,117 Resource Elasticity for Large-Scale Machine Learning 2015 SIGMOD 6.7929814e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6537801e-05
5,308 Probabilistic Demand Forecasting at Scale 2017 VLDB 6.1844285e-05
5,574 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0773771e-05
5,999 BAGUA: Scaling up Distributed Learning with System Relaxations 2022 VLDB 5.9170896e-05
6,561 DeepBase: Deep Inspection of Neural Networks 2019 SIGMOD 5.7461874e-05
6,594 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7394502e-05
6,666 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7144587e-05
6,883 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines 2022 CIDR 5.6552024e-05
7,144 A Cost-based Optimizer for Gradient Descent Optimization 2017 SIGMOD 5.5979118e-05
8,078 Robust Recursive Query Parallelism in Graph Database Management Systems 2025 VLDB 5.3917406e-05
8,600 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage 2021 VLDB 5.3043615e-05
9,042 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.2310219e-05
9,564 Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning 2021 VLDB 5.1547835e-05
9,663 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.142891e-05
9,753 PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development 2018 SIGMOD 5.1325223e-05
10,447 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 4.9769913e-05
10,907 stratum: A System Infrastructure for Massive Agent-Centric ML Workloads 2026 VLDB 4.9769913e-05
11,251 Quantum Data Management in the NISQ Era 2025 VLDB 4.9769913e-05
11,555 Database Native Model Selection: Harnessing Deep Neural Networks in Database Systems 2024 VLDB 4.9769913e-05
11,982 Hybrid Evaluation for Distributed Iterative Matrix Computation 2021 SIGMOD 4.9769913e-05
12,359 dmapply: A functional primitive to express distributed machine learning algorithms in R 2016 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers