Back to papers
Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning
Summary: Proposes grouped learning, a GROUP BY-like abstraction for ML over subgroups. Presents Gradient Accumulation Parallelism (GAP) and a hybrid task/data-parallel approach in Kingpin on Ray, delivering up to 4x–14x speedups vs. state-of-the-art.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 12411
- Venue
- VLDB
- Year
- 2021
- Pagerank
- 4.3656789e-05
- Overall Rank
- 9,225 | 35.89%
- DOI
-
10.14778/3476249.3476284
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 19 of 19 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 139 |
The MADlib Analytics Library or MAD Skills, the SQL |
2012 |
VLDB |
0.00042320525 |
| 411 |
PyTorch Distributed: Experiences on Accelerating Data Parallel Training |
2020 |
VLDB |
0.00023881138 |
| 557 |
SystemML: Declarative Machine Learning on Spark |
2016 |
VLDB |
0.00020186115 |
| 638 |
Towards a Unified Architecture for in-RDBMS Analytics |
2012 |
SIGMOD |
0.00018810785 |
| 684 |
Cerebro: A Data System for Optimized Deep Learning Model Selection |
2020 |
VLDB |
0.00018152321 |
| 832 |
Learning Linear Regression Models over Factorized Joins |
2016 |
SIGMOD |
0.00016089705 |
| 851 |
Scaling Factorization Machines to Relational Data |
2013 |
VLDB |
0.00015909639 |
| 1,172 |
Learning Generalized Linear Models Over Normalized Data |
2015 |
SIGMOD |
0.00013504249 |
| 1,283 |
Towards Linear Algebra over Normalized Data |
2017 |
VLDB |
0.00012826013 |
| 1,393 |
Ease.ml: Towards Multi-tenant Resource Sharing for Machine Learning Workloads |
2018 |
VLDB |
0.00012223372 |
| 1,407 |
Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML |
2014 |
VLDB |
0.00012163413 |
| 2,122 |
SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle |
2020 |
CIDR |
9.4905306e-05 |
| 2,197 |
Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra |
2019 |
SIGMOD |
9.3117431e-05 |
| 3,920 |
On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML |
2018 |
VLDB |
6.6246708e-05 |
| 4,195 |
F: Regression Models over Factorized Views |
2016 |
VLDB |
6.3635322e-05 |
| 4,794 |
Demonstration of Santoku: Optimizing Machine Learning over Normalized Data |
2015 |
VLDB |
5.910645e-05 |
| 4,924 |
An Experimental Evaluation of Large Scale GBDT Systems |
2019 |
VLDB |
5.8211961e-05 |
| 8,864 |
Cerebro: A Layered Data Platform for Scalable Deep Learning |
2021 |
CIDR |
4.4283952e-05 |
| 9,115 |
Ease.ml in Action: Towards Multi-tenant Declarative Learning Services |
2018 |
VLDB |
4.3886513e-05 |
Semantically Similar Papers