DBScholar

Back to papers

Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML

Summary: SystemML combines task and data parallelism for declarative, large-scale ML via a generic ParFOR construct over MapReduce. A cost-based optimizer automatically selects multi-core and cluster execution plans, adapting to workloads and unknown data characteristics. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
ha5d406fbdd20049f
Venue
VLDB
Year
2014
Pagerank
0.00012123917
Overall Rank
1,081 | 92.74%
DOI
10.14778/2732296.2732302

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{boehm_vldb14,
        title = {{Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML}},
        author = {Boehm, Matthias and Tatikonda, Shirish and Reinwald, Berthold and Sen, Prithviraj and Tian, Yuanyuan and Burdick, Douglas R. and Vaithyanathan, Shivakumar},
        journal = {PVLDB},
        series = {{VLDB} '14},
        volume = {7},
        number = {7},
        pages = {553--564},
        doi = {10.14778/2732296.2732302},
        url = {https://doi.org/10.14778/2732296.2732302},
        year = {2014}
}

Incoming Citations (Sorted by Pagerank)

Showing 36 of 36 citing papers.

Rank Citing Paper Year Venue Pagerank
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001865959
521 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.00016929744
1,152 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011801961
1,255 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011325762
1,256 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011314687
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
2,264 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.7289107e-05
2,326 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6309237e-05
2,723 Exploiting Matrix Dependency for Efficient Distributed Matrix Computation 2015 SIGMOD 8.0944934e-05
2,975 In-Database Learning with Sparse Tensors 2018 PODS 7.7907759e-05
3,101 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.649219e-05
3,240 Towards Demystifying Serverless Machine Learning Training 2021 SIGMOD 7.5002772e-05
3,518 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.2400627e-05
4,095 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.8095767e-05
4,116 Resource Elasticity for Large-Scale Machine Learning 2015 SIGMOD 6.7961306e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6569314e-05
5,304 Probabilistic Demand Forecasting at Scale 2017 VLDB 6.1873575e-05
5,572 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0802555e-05
5,997 BAGUA: Scaling up Distributed Learning with System Relaxations 2022 VLDB 5.9198921e-05
6,559 DeepBase: Deep Inspection of Neural Networks 2019 SIGMOD 5.7489089e-05
6,592 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7421684e-05
6,662 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7171651e-05
6,878 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines 2022 CIDR 5.657878e-05
7,142 A Cost-based Optimizer for Gradient Descent Optimization 2017 SIGMOD 5.600563e-05
8,071 Robust Recursive Query Parallelism in Graph Database Management Systems 2025 VLDB 5.3942942e-05
8,593 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage 2021 VLDB 5.3068556e-05
9,034 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.2334993e-05
9,556 Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning 2021 VLDB 5.1572248e-05
9,656 BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler Approach 2023 SIGMOD 5.1453267e-05
9,748 PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development 2018 SIGMOD 5.1349531e-05
10,435 DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization 2026 SIGMOD 4.9793485e-05
10,898 stratum: A System Infrastructure for Massive Agent-Centric ML Workloads 2026 VLDB 4.9793485e-05
11,243 Quantum Data Management in the NISQ Era 2025 VLDB 4.9793485e-05
11,549 Database Native Model Selection: Harnessing Deep Neural Networks in Database Systems 2024 VLDB 4.9793485e-05
11,976 Hybrid Evaluation for Distributed Iterative Matrix Computation 2021 SIGMOD 4.9793485e-05
12,353 dmapply: A functional primitive to express distributed machine learning algorithms in R 2016 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers