DBScholar

Back to papers

MAD Skills: New Analysis Practices for Big Data

Summary: MAD: Magnetic, Agile, Deep data analysis marks a radical shift from EDW/BI for big data. It presents data-parallel density methods and SQL/MapReduce-enabled workflows on Greenplum to enable agile analytics for advertising networks. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hb488806ad54f3530
Venue
VLDB
Year
2009
Pagerank
0.00028568843
Overall Rank
154 | 98.97%
DOI
10.14778/1687553.1687576

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{cohen_vldb09,
        title = {{MAD Skills: New Analysis Practices for Big Data}},
        author = {Cohen, Jeffrey and Dolan, Brian and Dunlap, Mark and Hellerstein, Joseph M. and Welton, Caleb},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687553.1687576},
        url = {https://doi.org/10.14778/1687553.1687576},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 54 citing papers.

Rank Citing Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055384955
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.0004552807
105 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033633007
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018331051
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017195428
548 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016572863
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00015096817
675 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00014880686
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013642066
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012959992
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012371105
1,070 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.00012179575
1,082 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012118261
1,223 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011468426
1,371 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.00010894415
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010067153
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
1,774 ArrayStore: A Storage Manager for Complex Parallel Array Processing 2011 SIGMOD 9.6661993e-05
2,176 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.9150466e-05
2,249 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7544468e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8816694e-05
3,040 Knowledge Expansion over Probabilistic Knowledge Bases 2014 SIGMOD 7.7229256e-05
3,103 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6456038e-05
3,739 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.0595382e-05
3,790 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.0162259e-05
3,801 All-in-One: Graph Processing in RDBMSs Revisited 2017 SIGMOD 7.0128676e-05
4,117 Resource Elasticity for Large-Scale Machine Learning 2015 SIGMOD 6.7929814e-05
4,302 Efficient and Portable Einstein Summation in SQL 2023 SIGMOD 6.6754218e-05
5,002 GLADE: Big Data Analytics Made Easy 2012 SIGMOD 6.3183539e-05
5,392 PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics 2013 VLDB 6.1494698e-05
5,653 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.048773e-05
5,799 MCDB-R: Risk Analysis in the Database 2010 VLDB 5.9879586e-05
5,944 Bridging Two Worlds with RICE: Integrating R into the SAP In-Memory Computing Engine 2011 VLDB 5.9360073e-05
6,402 Functional-Style SQL UDFs With a Capital 'F' 2020 SIGMOD 5.7953646e-05
6,711 Machine Learning, Linear Algebra, and More: Is SQL All You Need? 2022 CIDR 5.7011955e-05
7,203 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 5.5845207e-05
7,263 Entity Resolution On-Demand 2022 VLDB 5.5693872e-05
7,302 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 5.5584215e-05
7,383 Online Expansion of Large-scale Data Warehouses 2011 VLDB 5.5366489e-05
7,843 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4406331e-05
8,600 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage 2021 VLDB 5.3043615e-05
8,609 UDA-GIST: An In-database Framework to Unify Data-Parallel and State-Parallel Analytics 2015 VLDB 5.3025486e-05
9,545 GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example 2023 SIGMOD 5.1600923e-05
9,761 Storing Matrices on Disk: Theory and Practice Revisited 2011 VLDB 5.1325223e-05
9,829 BlockJoin: Efficient Matrix Partitioning Through Joins 2017 VLDB 5.1230568e-05
9,952 On Efficient Large Sparse Matrix Chain Multiplication 2024 SIGMOD 5.101836e-05
10,312 Fast and Scalable Data Transfer Across Data Systems 2025 SIGMOD 5.0376863e-05
10,935 Benchmarking Native In-Database TPCx-AI at 100 Terabytes in Ocient Hyperscale Data Warehouse: An Eight-Use-Case Study of Classical and Statistical ML on Relational Primitives 2026 VLDB 4.9769913e-05
11,292 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9769913e-05
11,555 Database Native Model Selection: Harnessing Deep Neural Networks in Database Systems 2024 VLDB 4.9769913e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers