DBScholar

Back to papers

A Comparison of Approaches to Large-Scale Data Analysis

Summary: Compare MapReduce with parallel DBMSs for large-scale data analysis, tying MR to decades of parallel-SQL work. A 100-node benchmark finds DBMSs load/tune longer but run faster than MR; discusses causes and future-system implications. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
4179
Venue
SIGMOD
Year
2009
Pagerank
0.00046055057
Overall Rank
44 | 99.70%
DOI
10.1145/1559845.1559865

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{pavlo_sigmod09,
        title = {{A Comparison of Approaches to Large-Scale Data Analysis}},
        author = {Pavlo, Andrew and Paulson, Erik and Rasin, Alexander and Abadi, Daniel J. and DeWitt, David J. and Madden, Samuel and Stonebraker, Michael},
        series = {{SIGMOD} '09},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/1559845.1559865},
        url = {https://dl.acm.org/doi/10.1145/1559845.1559865},
        year = {2009}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 71 citing papers.

Rank Citing Paper Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
32 Hive - A Warehousing Solution Over a Map-Reduce Framework 2009 VLDB 0.00050111008
66 The Snowflake Elastic Data Warehouse 2016 SIGMOD 0.00038561587
106 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033539462
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.00031680027
356 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020303289
372 HaLoop: Efficient Iterative Data Processing on Large Clusters 2010 VLDB 0.0001981521
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
660 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.00015198804
710 Trill: A High-Performance Incremental Query Processor for Diverse Analytics 2015 VLDB 0.00014715033
769 A Comparison of Join Algorithms for Log Processing in MapReduce 2010 SIGMOD 0.00014166872
803 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013899943
843 Processing Theta-Joins using MapReduce* 2011 SIGMOD 0.00013666161
895 Runtime Measurements in the Cloud: Observing, Analyzing, and Reducing Variance 2010 VLDB 0.00013357681
1,054 Interactive Analytical Processing in Big Data Systems: A Cross-Industry Study of MapReduce Workloads 2012 VLDB 0.00012390673
1,067 Hadoop-GIS: A High Performance Spatial Data Warehousing System over MapReduce 2013 VLDB 0.00012327784
1,257 Automatic Optimization for MapReduce Programs 2011 VLDB 0.000114432
1,348 An Evaluation of Alternative Architectures for Transaction Processing in the Cloud 2010 SIGMOD 0.00011071078
1,436 The Performance of MapReduce: An In-depth Study 2010 VLDB 0.00010797443
1,666 HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics 2016 VLDB 0.00010068964
1,693 BigBench: Towards an Industry Standard Benchmark for Big Data Analytics 2013 SIGMOD 9.9965799e-05
1,903 Instant Loading for Main Memory Databases 2013 VLDB 9.5049156e-05
1,974 Track Join: Distributed Joins with Minimal Network Traffic 2014 SIGMOD 9.3658402e-05
2,159 Efficient Processing of Data Warehousing Queries in a Split Execution Environment 2011 SIGMOD 9.061086e-05
2,207 Shark: Fast Data Analysis Using Coarse-grained Distributed Memory 2012 SIGMOD 8.9565862e-05
2,239 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.8875753e-05
2,265 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.8398946e-05
2,354 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.7060612e-05
2,486 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.5143189e-05
2,622 A Latency and Fault-Tolerance Optimizer for Online Parallel Query Plans 2011 SIGMOD 8.3330136e-05
2,642 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.3059948e-05
2,849 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 8.053191e-05
2,850 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 8.0499212e-05
3,104 Energy Management for MapReduce Clusters 2010 VLDB 7.7552953e-05
3,351 RHEEM: Enabling Cross-Platform Data Processing - May The Big Data Be With You! - 2018 VLDB 7.4937347e-05
3,368 CloudRAMSort: Fast and Efficient Large-Scale Distributed RAM Sort on Shared-Nothing Cluster 2012 SIGMOD 7.4713287e-05
3,674 TIRAMOLA: Elastic NoSQL Provisioning Through a Cloud Management Platform 2012 SIGMOD 7.211109e-05
3,749 Integrating Hadoop and Parallel DBMS 2010 SIGMOD 7.1539955e-05
3,783 MISO: Souping Up Big Data Query Processing with a Multistore System 2014 SIGMOD 7.1298683e-05
4,224 Comparative Evaluation of Big-Data Systems on Scientific Image Analytics Workloads 2017 VLDB 6.8214304e-05
4,230 Can the Elephants Handle the NoSQL Onslaught? 2012 VLDB 6.8178353e-05
4,481 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 6.6754521e-05
4,892 Fast In-Memory SQL Analytics on Typed Graphs 2017 VLDB 6.4574091e-05
4,897 GLADE: Big Data Analytics Made Easy 2012 SIGMOD 6.4546134e-05
5,156 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.3410921e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
5,865 Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems 2019 VLDB 6.0619534e-05
5,871 Cheetah: Accelerating Database Queries with Switch Pruning 2020 SIGMOD 6.0604381e-05
5,935 Towards Energy-Efficient Database Cluster Design 2012 VLDB 6.0361286e-05
6,147 HadoopDB in Action: Building Real World Applications 2010 SIGMOD 5.9601328e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers