DBScholar

Back to papers

MAD Skills: New Analysis Practices for Big Data

Summary: MAD: Magnetic, Agile, Deep data analysis marks a radical shift from EDW/BI for big data. It presents data-parallel density methods and SQL/MapReduce-enabled workflows on Greenplum to enable agile analytics for advertising networks. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
10048
Venue
VLDB
Year
2009
Pagerank
0.00028713176
Overall Rank
155 | 98.94%
DOI
10.14778/1687553.1687576

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{cohen_vldb09,
        title = {{MAD Skills: New Analysis Practices for Big Data}},
        author = {Cohen, Jeffrey and Dolan, Brian and Dunlap, Mark and Hellerstein, Joseph M. and Welton, Caleb},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687553.1687576},
        url = {https://doi.org/10.14778/1687553.1687576},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 53 citing papers.

Rank Citing Paper Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
44 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00046055057
106 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033539462
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
518 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017167492
549 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016692839
640 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00015409494
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
803 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013899943
923 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00013189886
1,021 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012606673
1,070 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.0001232307
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,250 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011485301
1,333 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.0001112858
1,644 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010132912
1,742 ArrayStore: A Storage Manager for Complex Parallel Array Processing 2011 SIGMOD 9.8748669e-05
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
2,159 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 9.061086e-05
2,217 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.9332438e-05
2,850 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 8.0499212e-05
2,990 Knowledge Expansion over Probabilistic Knowledge Bases 2014 SIGMOD 7.8883704e-05
3,205 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6386536e-05
3,682 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.2035518e-05
3,714 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.1764857e-05
3,872 All-in-One: Graph Processing in RDBMSs Revisited 2017 SIGMOD 7.0580243e-05
4,049 Resource Elasticity for Large-Scale Machine Learning 2015 SIGMOD 6.9369379e-05
4,253 Efficient and Portable Einstein Summation in SQL 2023 SIGMOD 6.8029311e-05
4,897 GLADE: Big Data Analytics Made Easy 2012 SIGMOD 6.4546134e-05
5,287 PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics 2013 VLDB 6.2815429e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
5,672 MCDB-R: Risk Analysis in the Database 2010 VLDB 6.1261535e-05
5,855 Bridging Two Worlds with RICE: Integrating R into the SAP In-Memory Computing Engine 2011 VLDB 6.064726e-05
6,328 Functional-Style SQL UDFs With a Capital 'F' 2020 SIGMOD 5.9113895e-05
6,604 Machine Learning, Linear Algebra, and More: Is SQL All You Need? 2022 CIDR 5.8240599e-05
7,112 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 5.6990782e-05
7,179 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 5.6802479e-05
7,243 Online Expansion of Large-scale Data Warehouses 2011 VLDB 5.6640533e-05
7,687 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.5671645e-05
7,743 Entity Resolution On-Demand 2022 VLDB 5.5545622e-05
8,432 UDA-GIST: An In-database Framework to Unify Data-Parallel and State-Parallel Analytics 2015 VLDB 5.4259582e-05
8,450 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage 2021 VLDB 5.4231788e-05
9,436 GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example 2023 SIGMOD 5.2692207e-05
9,580 Storing Matrices on Disk: Theory and Practice Revisited 2011 VLDB 5.2528121e-05
9,675 BlockJoin: Efficient Matrix Partitioning Through Joins 2017 VLDB 5.2380072e-05
9,766 On Efficient Large Sparse Matrix Chain Multiplication 2024 SIGMOD 5.2214067e-05
10,761 Fast and Scalable Data Transfer Across Data Systems 2025 SIGMOD 5.093636e-05
10,882 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 5.093636e-05
11,209 Database Native Model Selection: Harnessing Deep Neural Networks in Database Systems 2024 VLDB 5.093636e-05
11,955 An Authorization Model for Multi Provider Queries 2018 VLDB 5.093636e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers