DBScholar

Back to papers

The Performance of MapReduce: An In-depth Study

Summary: A 100-node EC2 study isolates five Hadoop/MapReduce performance factors across parallelism levels. Careful tuning yields 2.5–3.5× speedups, narrowing the gap with parallel databases while retaining elastic, cost-conscious scalability. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
10211
Venue
VLDB
Year
2010
Pagerank
0.00010797443
Overall Rank
1,436 | 90.15%
DOI
10.14778/1920841.1920902

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jiang_vldb10,
        title = {{The Performance of MapReduce: An In-depth Study}},
        author = {Jiang, Dawei and Ooi, Beng Chin and Shi, Lei and Wu, Sai},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {472--483},
        doi = {10.14778/1920841.1920902},
        url = {https://doi.org/10.14778/1920841.1920902},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
735 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014522606
923 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00013189886
1,694 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.9953156e-05
2,137 Efficient Processing of k Nearest Neighbor Joins using MapReduce 2012 VLDB 9.110238e-05
2,265 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.8398946e-05
2,354 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.7060612e-05
2,642 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.3059948e-05
2,849 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 8.053191e-05
2,850 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 8.0499212e-05
2,887 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.9952432e-05
2,942 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.9358593e-05
3,275 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.5747814e-05
5,156 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.3410921e-05
5,522 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.1871697e-05
5,865 Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems 2019 VLDB 6.0619534e-05
6,044 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.995949e-05
7,990 ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] 2014 VLDB 5.5107935e-05
8,455 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.4226e-05
9,510 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.2576928e-05
11,889 An Experimental Evaluation of Garbage Collectors on Big Data Applications 2019 VLDB 5.093636e-05
12,185 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index 2014 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers