DBScholar

Back to papers

A Comparison of Approaches to Large-Scale Data Analysis

Summary: Compare MapReduce with parallel DBMSs for large-scale data analysis, tying MR to decades of parallel-SQL work. A 100-node benchmark finds DBMSs load/tune longer but run faster than MR; discusses causes and future-system implications. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
ha21de992bbd6e4e8
Venue
SIGMOD
Year
2009
Pagerank
0.00045546775
Overall Rank
43 | 99.72%
DOI
10.1145/1559845.1559865

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{pavlo_sigmod09,
        title = {{A Comparison of Approaches to Large-Scale Data Analysis}},
        author = {Pavlo, Andrew and Paulson, Erik and Rasin, Alexander and Abadi, Daniel J. and DeWitt, David J. and Madden, Samuel and Stonebraker, Michael},
        series = {{SIGMOD} '09},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/1559845.1559865},
        url = {https://dl.acm.org/doi/10.1145/1559845.1559865},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 71 citing papers.

Rank Citing Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055406774
31 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00049839909
52 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00041219077
105 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033638251
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.000311132
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020009936
384 HaLoop: Efficient Iterative Data Processing on Large Clusters 2010 VLDB 0.0001948031
432 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018339357
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
685 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014782777
799 A Comparison of Join Algorithms for Log Processing in MapReduce 2010 SIGMOD 0.00013889081
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
865 Processing Theta-Joins using MapReduce* 2011 SIGMOD 0.0001338765
892 Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance 2010 VLDB 0.00013227162
1,072 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012168947
1,080 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012131978
1,281 Automatic Optimization for MapReduce Programs 2011 VLDB 0.00011213384
1,322 An Evaluation of Alternative Architectures for Transaction Processing in the Cloud 2010 SIGMOD 0.00011034098
1,466 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010569837
1,541 HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics 2016 VLDB 0.00010313459
1,702 BigBench: Towards an Industry Standard Benchmark for Big Data Analytics 2013 SIGMOD 9.8361594e-05
1,931 Instant Loading for Main Memory Databases 2013 VLDB 9.3474242e-05
1,995 Track Join: Distributed Joins with Minimal Network Traffic 2014 SIGMOD 9.2169073e-05
2,173 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 8.91924e-05
2,232 Shark: Fast Data Analysis Using Coarse-grained Distributed Memory 2012 SIGMOD 8.7960223e-05
2,264 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.7289107e-05
2,282 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.6964781e-05
2,410 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.5114032e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3425785e-05
2,672 A Latency and Fault-Tolerance Optimizer for Online Parallel Query Plans 2011 SIGMOD 8.1491683e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8853204e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.8810111e-05
3,157 Energy Management for MapReduce Clusters 2010 VLDB 7.5834595e-05
3,393 CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster 2012 SIGMOD 7.3471344e-05
3,402 RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! - 2018 VLDB 7.3304477e-05
3,747 TIRAMOLA: Elastic NoSQL Provisioning Through a Cloud Management Platform 2012 SIGMOD 7.0552485e-05
3,823 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.0029055e-05
3,840 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 6.9920793e-05
4,069 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 6.821366e-05
4,289 Can the Elephants Handle the NoSQL Onslaught? 2012 VLDB 6.6838669e-05
4,308 Comparative Evaluation of Big-Data Systems on Scientific Image Analytics Workloads 2017 VLDB 6.6735191e-05
4,984 Fast In-Memory SQL Analytics on Typed Graphs 2017 VLDB 6.3268363e-05
5,001 GLADE: Big Data Analytics Made Easy 2012 SIGMOD 6.3197648e-05
5,274 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.2005706e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
5,958 Cheetah: Accelerating Database Queries with Switch Pruning 2020 SIGMOD 5.9338314e-05
5,967 Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems 2019 VLDB 5.9324372e-05
6,057 Towards Energy-Efficient Database Cluster Design 2012 VLDB 5.900701e-05
6,276 HadoopDB in Action: Building Real World Applications 2010 SIGMOD 5.8268854e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers