DBScholar

Back to papers

MAD Skills: New Analysis Practices for Big Data

Summary: MAD: Magnetic, Agile, Deep data analysis marks a radical shift from EDW/BI for big data. It presents data-parallel density methods and SQL/MapReduce-enabled workflows on Greenplum to enable agile analytics for advertising networks. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hb488806ad54f3530
Venue
VLDB
Year
2009
Pagerank
0.00028579704
Overall Rank
154 | 98.97%
DOI
10.14778/1687553.1687576

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{cohen_vldb09,
        title = {{MAD Skills: New Analysis Practices for Big Data}},
        author = {Cohen, Jeffrey and Dolan, Brian and Dunlap, Mark and Hellerstein, Joseph M. and Welton, Caleb},
        journal = {PVLDB},
        series = {{VLDB} '09},
        doi = {10.14778/1687553.1687576},
        url = {https://doi.org/10.14778/1687553.1687576},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 54 citing papers.

Rank Citing Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
105 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033638251
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018339357
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017202276
547 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016578131
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.0001510357
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012964445
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012376819
1,069 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.00012185253
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,255 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011325762
1,371 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.00010899041
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010071891
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
1,774 ArrayStore: A Storage Manager for Complex Parallel Array Processing 2011 SIGMOD 9.6688719e-05
2,173 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.91924e-05
2,247 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7585767e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8853204e-05
3,039 Knowledge Expansion over Probabilistic Knowledge Bases 2014 SIGMOD 7.7262528e-05
3,101 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.649219e-05
3,737 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.0628666e-05
3,788 Large-Scale Machine Learning at Twitter 2012 SIGMOD 7.0195464e-05
3,798 All-in-One: Graph Processing in RDBMSs Revisited 2017 SIGMOD 7.0161889e-05
4,116 Resource Elasticity for Large-Scale Machine Learning 2015 SIGMOD 6.7961306e-05
4,300 Efficient and Portable Einstein Summation in SQL 2023 SIGMOD 6.6785834e-05
5,001 GLADE: Big Data Analytics Made Easy 2012 SIGMOD 6.3197648e-05
5,386 PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics 2013 VLDB 6.1522468e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
5,797 MCDB-R: Risk Analysis in the Database 2010 VLDB 5.9907755e-05
5,944 Bridging Two Worlds with RICE: Integrating R into the SAP In-Memory Computing Engine 2011 VLDB 5.9387238e-05
6,400 Functional-Style SQL UDFs With a Capital 'F' 2020 SIGMOD 5.7981078e-05
6,707 Machine Learning, Linear Algebra, and More: Is SQL All You Need? 2022 CIDR 5.7038956e-05
7,201 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 5.5871656e-05
7,260 Entity Resolution On-Demand 2022 VLDB 5.571581e-05
7,298 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 5.561054e-05
7,381 Online Expansion of Large-scale Data Warehouses 2011 VLDB 5.5392533e-05
7,839 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4432099e-05
8,593 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage 2021 VLDB 5.3068556e-05
8,602 UDA-GIST: An In-database Framework to Unify Data-Parallel and State-Parallel Analytics 2015 VLDB 5.3050579e-05
9,549 GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example 2023 SIGMOD 5.1599622e-05
9,756 Storing Matrices on Disk: Theory and Practice Revisited 2011 VLDB 5.1349531e-05
9,822 BlockJoin: Efficient Matrix Partitioning Through Joins 2017 VLDB 5.1254832e-05
9,944 On Efficient Large Sparse Matrix Chain Multiplication 2024 SIGMOD 5.1042523e-05
10,305 Fast and Scalable Data Transfer Across Data Systems 2025 SIGMOD 5.0400722e-05
10,926 Benchmarking Native In-Database TPCx-AI at 100 Terabytes in Ocient Hyperscale Data Warehouse: An Eight-Use-Case Study of Classical and Statistical ML on Relational Primitives 2026 VLDB 4.9793485e-05
11,284 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9793485e-05
11,549 Database Native Model Selection: Harnessing Deep Neural Networks in Database Systems 2024 VLDB 4.9793485e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers