Database Paper Browser

Back to papers

MAD Skills: New Analysis Practices for Big Data

Summary: MAD: Magnetic, Agile, Deep data analysis marks a radical shift from EDW/BI for big data. It presents data-parallel density methods and SQL/MapReduce-enabled workflows on Greenplum to enable agile analytics for advertising networks. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
9858
Venue
VLDB
Year
2009
Pagerank
0.00038967371
Overall Rank
169 | 98.83%
DOI
-

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 53 citing papers.

Rank Citing Paper Year Venue Pagerank
42 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00073570328
66 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00061707583
139 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00042320525
539 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00020615453
638 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00018810785
656 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00018590675
758 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00017053915
789 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00016602215
947 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00015112344
1,048 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00014442178
1,265 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012933537
1,346 NoDB: Efficient Query Execution on Raw Data Files 2012 SIGMOD 0.00012472598
1,407 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012163413
1,494 Ricardo: Integrating R and Hadoop 2010 SIGMOD 0.00011677793
1,534 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011462072
1,878 ArrayStore: A Storage Manager for Complex Parallel Array Processing 2011 SIGMOD 0.00010229691
1,970 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 9.9024431e-05
2,122 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.4905306e-05
2,340 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 9.001663e-05
2,672 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.3334428e-05
3,070 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.6157391e-05
3,084 Knowledge Expansion over Probabilistic Knowledge Bases 2014 SIGMOD 7.5967738e-05
3,606 Large-Scale Machine Learning at Twitter 2012 SIGMOD 6.9246371e-05
3,920 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 6.6246708e-05
3,984 All-in-One: Graph Processing in RDBMSs Revisited 2017 SIGMOD 6.5587512e-05
4,040 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 6.5052227e-05
4,548 Efficient and Portable Einstein Summation in SQL 2023 SIGMOD 6.089487e-05
4,807 Resource Elasticity for Large-Scale Machine Learning 2015 SIGMOD 5.9045148e-05
5,300 GLADE: Big Data Analytics Made Easy 2012 SIGMOD 5.5751386e-05
5,699 PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics 2013 VLDB 5.3655096e-05
5,908 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 5.2731311e-05
5,969 Bridging Two Worlds with RICE: Integrating R into the SAP In-Memory Computing Engine 2011 VLDB 5.2469372e-05
5,976 MCDB-R: Risk Analysis in the Database 2010 VLDB 5.2438744e-05
6,646 Functional-Style SQL UDFs With a Capital 'F' 2020 SIGMOD 4.9734285e-05
6,989 Machine Learning, Linear Algebra, and More: Is SQL All You Need? 2022 CIDR 4.8658293e-05
7,180 Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning 2023 VLDB 4.8032775e-05
7,233 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 4.788267e-05
7,260 Online Expansion of Large-scale Data Warehouses 2011 VLDB 4.781278e-05
7,702 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 4.6689015e-05
7,930 Entity Resolution On-Demand 2022 VLDB 4.6089604e-05
8,397 UDA-GIST: An In-database Framework to Unify Data-Parallel and State-Parallel Analytics 2015 VLDB 4.5208365e-05
8,435 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage 2021 VLDB 4.5075971e-05
9,355 GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example 2023 SIGMOD 4.3484264e-05
9,432 Storing Matrices on Disk: Theory and Practice Revisited 2011 VLDB 4.3399748e-05
9,442 BlockJoin: Efficient Matrix Partitioning Through Joins 2017 VLDB 4.3384032e-05
9,670 On Efficient Large Sparse Matrix Chain Multiplication 2024 SIGMOD 4.3024878e-05
10,492 Fast and Scalable Data Transfer Across Data Systems 2025 SIGMOD 4.1905499e-05
10,636 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.1905499e-05
11,001 Database Native Model Selection: Harnessing Deep Neural Networks in Database Systems 2024 VLDB 4.1905499e-05
11,757 An Authorization Model for Multi-Provider Queries 2018 VLDB 4.1905499e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers