DBScholar

Back to papers

The MADlib Analytics Library or MAD Skills, the SQL

Summary: MADlib is an open-source, extensible library of parallel SQL-native machine-learning, data-mining, and statistical methods, bringing CRAN-like community analytics into database engines. It emphasizes in-database execution at scale, avoiding data movement and supporting research contributions across platforms. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
h9358f138642ac4e4
Venue
VLDB
Year
2012
Pagerank
0.00033638251
Overall Rank
105 | 99.30%
DOI
10.14778/2367502.2367519

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{hellerstein_vldb12,
        title = {{The MADlib Analytics Library or MAD Skills, the SQL}},
        author = {Hellerstein, Joseph M. and Ré, Christoper and Schoppmann, Florian and Wang, Daisy Zhe and Fratkin, Eugene and Gorajek, Aleksander and Ng, Kee Siong and Welton, Caleb and Feng, Xixuan and Li, Kun and Kumar, Arun},
        journal = {PVLDB},
        series = {{VLDB} '12},
        volume = {5},
        number = {12},
        pages = {1700--1711},
        doi = {10.14778/2367502.2367519},
        url = {https://doi.org/10.14778/2367502.2367519},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 108 citing papers.

Rank Citing Paper Year Venue Pagerank
4,095 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.8095767e-05
4,161 The Relational Data Borg is Learning 2020 VLDB 6.7700593e-05
4,275 Combining Databases and Signal Processing in Plato 2015 CIDR 6.6917477e-05
4,278 QuickFOIL: Scalable Inductive Logic Programming 2015 VLDB 6.6905317e-05
4,300 Efficient and Portable Einstein Summation in SQL 2023 SIGMOD 6.6785834e-05
4,781 Learned Approximate Query Processing: Make it Light, Accurate and Fast 2021 CIDR 6.4162085e-05
4,879 Accelerating Generalized Linear Models with MLWeaving: A One-Size-Fits-All System for Any-Precision Learning 2019 VLDB 6.3715334e-05
4,901 Scalable Asynchronous Gradient Descent Optimization for Out-of-Core Models 2017 VLDB 6.3647386e-05
5,685 SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained Environments 2024 VLDB 6.037958e-05
5,702 Large-scale Predictive Analytics in Vertica: Fast Data Transfer, Distributed Model Creation, and In-database Prediction 2015 SIGMOD 6.0312071e-05
5,776 The Tensor Data Platform: Towards an AI-centric Database System 2023 CIDR 5.9981216e-05
5,834 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 5.9759267e-05
5,877 BlinkML: Efficient Maximum Likelihood Estimation with Probabilistic Guarantees 2019 SIGMOD 5.9627218e-05
6,086 Demonstration of Nimbus: Model-based Pricing for Machine Learning in a Data Marketplace 2019 SIGMOD 5.8912658e-05
6,143 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.8721471e-05
6,144 The BUDS Language for Distributed Bayesian Machine Learning 2017 SIGMOD 5.8717563e-05
6,246 ColumnML: Column-Store Machine Learning with On-The-Fly Data Transformation 2019 VLDB 5.8370793e-05
6,458 In-Database Machine Learning with CorgiPile: Stochastic Gradient Descent without Full Data Shuffle 2022 SIGMOD 5.7773842e-05
6,550 JoinBoost: Grow Trees Over Normalized Data Using Only SQL 2023 VLDB 5.7503043e-05
6,559 DeepBase: Deep Inspection of Neural Networks 2019 SIGMOD 5.7489089e-05
6,592 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7421684e-05
6,631 A Relational Matrix Algebra and its Implementation in a Column Store 2020 SIGMOD 5.7270666e-05
6,742 Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine 2025 SIGMOD 5.6910432e-05
6,912 EAGr: Supporting Continuous Ego-centric Aggregate Queries over Large Dynamic Graphs 2014 SIGMOD 5.6484964e-05
7,025 GeoDeepDive: Statistical Inference using Familiar Data-Processing Languages 2013 SIGMOD 5.6184467e-05
7,142 A Cost-based Optimizer for Gradient Descent Optimization 2017 SIGMOD 5.600563e-05
7,201 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 5.5871656e-05
7,819 You Say ‘What’, I Hear ‘Where’ and ‘Why’ — (Mis-)Interpreting SQL to Derive Fine-Grained Provenance 2018 VLDB 5.4466474e-05
8,151 Schema Independent Relational Learning 2017 SIGMOD 5.3899758e-05
8,230 Adda: Towards Efficient in-Database Feature Generation via LLM-based Agents 2025 SIGMOD 5.372368e-05
8,544 Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? 2021 SIGMOD 5.3186267e-05
8,593 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage 2021 VLDB 5.3068556e-05
8,602 UDA-GIST: An In-database Framework to Unify Data-Parallel and State-Parallel Analytics 2015 VLDB 5.3050579e-05
8,847 Machine Learning Meets Big Spatial Data 2019 VLDB 5.2646682e-05
9,019 Complaint-Driven Training Data Debugging at Interactive Speeds 2022 SIGMOD 5.2364886e-05
9,071 Privacy and Accuracy-Aware AI/ML Model Deduplication 2025 SIGMOD 5.2283159e-05
9,144 Cerebro: A Layered Data Platform for Scalable Deep Learning 2021 CIDR 5.2201471e-05
9,212 Powering In-Database Dynamic Model Slicing for Structured Data Analytics 2024 VLDB 5.2060505e-05
9,278 Ontological Pathfinding: Mining First-Order Knowledge from Large Knowledge Bases 2016 SIGMOD 5.203976e-05
9,556 Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning 2021 VLDB 5.1572248e-05
9,570 Determining Exact Quantiles with Randomized Summaries 2024 SIGMOD 5.1571823e-05
9,724 Database as Runtime: Compiling LLMs to SQL for In-database Model Serving 2025 SIGMOD 5.1349531e-05
9,828 Transforming ML Predictive Pipelines into SQL with MASQ 2021 SIGMOD 5.1249969e-05
10,181 In-Database Data Imputation 2024 SIGMOD 5.0653015e-05
10,305 Fast and Scalable Data Transfer Across Data Systems 2025 SIGMOD 5.0400722e-05
10,446 Eliminating Redundant Feature Tests in Decision Tree and Random Forest Inference on SQL Predicates 2026 SIGMOD 4.9793485e-05
10,582 NeurStore: Efficient In-database Deep Learning Model Management System 2026 SIGMOD 4.9793485e-05
10,674 Quantile Estimation with Duplicates 2026 SIGMOD 4.9793485e-05
10,798 NeurIDA: Dynamic Modeling for Effective In-Database Analytics 2026 VLDB 4.9793485e-05
10,926 Benchmarking Native In-Database TPCx-AI at 100 Terabytes in Ocient Hyperscale Data Warehouse: An Eight-Use-Case Study of Classical and Statistical ML on Relational Primitives 2026 VLDB 4.9793485e-05
Previous Page 2 / 3 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers