DBScholar

Back to papers

The Performance of MapReduce: An In-depth Study

Summary: A 100-node EC2 study isolates five Hadoop/MapReduce performance factors across parallelism levels. Careful tuning yields 2.5–3.5× speedups, narrowing the gap with parallel databases while retaining elastic, cost-conscious scalability. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
ha78d49573e09220a
Venue
VLDB
Year
2010
Pagerank
0.00010565007
Overall Rank
1,466 | 90.15%
DOI
10.14778/1920841.1920902

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jiang_vldb10,
        title = {{The Performance of MapReduce: An In-depth Study}},
        author = {Jiang, Dawei and Ooi, Beng Chin and Shi, Lei and Wu, Sai},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {472--483},
        doi = {10.14778/1920841.1920902},
        url = {https://doi.org/10.14778/1920841.1920902},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
754 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014231311
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012959992
1,708 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.8184303e-05
2,175 Efficient Processing of k Nearest Neighbor Joins using MapReduce 2012 VLDB 8.918268e-05
2,285 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.6924793e-05
2,411 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.5073756e-05
2,688 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1262948e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8816694e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.877293e-05
2,932 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.8368679e-05
3,009 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7570485e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4045097e-05
5,278 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.1976381e-05
5,653 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.048773e-05
5,969 Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems 2019 VLDB 5.9296488e-05
6,076 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8918218e-05
8,164 ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] 2014 VLDB 5.3850352e-05
8,631 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.298422e-05
9,699 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.1375205e-05
12,195 An Experimental Evaluation of Garbage Collectors on Big Data Applications 2019 VLDB 4.9769913e-05
12,482 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index 2014 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers