DBScholar

Back to papers

Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning

Summary: Proposes grouped learning, a GROUP BY-like abstraction for ML over subgroups. Presents Gradient Accumulation Parallelism (GAP) and a hybrid task/data-parallel approach in Kingpin on Ray, delivering up to 4x–14x speedups vs. state-of-the-art. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12598
Venue
VLDB
Year
2021
Pagerank
5.275595e-05
Overall Rank
9,371 | 35.71%
DOI
10.14778/3476249.3476284

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{li_vldb21,
        title = {{Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning}},
        author = {Li, Side and Kumar, Arun},
        journal = {PVLDB},
        series = {{VLDB} '21},
        volume = {14},
        number = {11},
        pages = {2327--2340},
        doi = {10.14778/3476249.3476284},
        url = {https://doi.org/10.14778/3476249.3476284},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 4 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 19 of 19 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
106 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033539462
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001888524
518 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017167492
521 PyTorch Distributed: Experiences on Accelerating Data Parallel Training 2020 VLDB 0.0001713368
536 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.0001693369
715 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014655327
835 Scaling Factorization Machines to Relational Data 2013 VLDB 0.00013721583
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,157 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011924049
1,235 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011548457
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
2,029 Ease.ml: Towards Multi-tenant Resource Sharing for Machine Learning Workloads 2018 VLDB 9.2843642e-05
2,179 Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra 2019 SIGMOD 9.0146333e-05
3,205 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6386536e-05
3,334 F: Regression Models over Factorized Views 2016 VLDB 7.5110164e-05
5,557 Demonstration of Santoku: Optimizing Machine Learning over Normalized Data 2015 VLDB 6.17499e-05
5,589 An Experimental Evaluation of Large Scale GBDT Systems 2019 VLDB 6.1559057e-05
8,979 Cerebro: A Layered Data Platform for Scalable Deep Learning 2021 CIDR 5.3399615e-05
9,265 Ease.ml in Action: Towards Multi-tenant Declarative Learning Services 2018 VLDB 5.2965435e-05
Previous Page 1 / 1 Next

Semantically Similar Papers