DBScholar

Back to papers

The Performance of MapReduce: An In-depth Study

Summary: A 100-node EC2 study isolates five Hadoop/MapReduce performance factors across parallelism levels. Careful tuning yields 2.5–3.5× speedups, narrowing the gap with parallel databases while retaining elastic, cost-conscious scalability. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
ha78d49573e09220a
Venue
VLDB
Year
2010
Pagerank
0.00010569837
Overall Rank
1,466 | 90.15%
DOI
10.14778/1920841.1920902

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{jiang_vldb10,
        title = {{The Performance of MapReduce: An In-depth Study}},
        author = {Jiang, Dawei and Ooi, Beng Chin and Shi, Lei and Wu, Sai},
        journal = {PVLDB},
        series = {{VLDB} '10},
        volume = {3},
        number = {1},
        pages = {472--483},
        doi = {10.14778/1920841.1920902},
        url = {https://doi.org/10.14778/1920841.1920902},
        year = {2010}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
753 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014237583
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012964445
1,708 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.8224587e-05
2,176 Efficient Processing of k Nearest Neighbor Joins using MapReduce 2012 VLDB 8.9159001e-05
2,282 A Platform for Scalable One-Pass Analytics using MapReduce 2011 SIGMOD 8.6964781e-05
2,410 epiC: an Extensible and Scalable System for Processing Big Data 2014 VLDB 8.5114032e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8853204e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.8810111e-05
2,931 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.8405483e-05
3,007 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7607173e-05
3,339 Optimizing Analytic Data Flows for Multiple Execution Engines 2012 SIGMOD 7.4080114e-05
5,274 Only Aggressive Elephants are Fast Elephants 2012 VLDB 6.2005706e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
5,967 Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems 2019 VLDB 5.9324372e-05
6,074 Exploiting Soft and Hard Correlations in Big Data Query Optimization 2016 VLDB 5.8946121e-05
8,158 ScalaGiST: Scalable Generalized Search Trees for MapReduce Systems [Innovative Systems Paper] 2014 VLDB 5.3875848e-05
8,624 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.3009314e-05
9,693 Efficient Big Data Processing in Hadoop MapReduce 2012 VLDB 5.1399537e-05
12,189 An Experimental Evaluation of Garbage Collectors on Big Data Applications 2019 VLDB 4.9793485e-05
12,476 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index 2014 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 9 of 9 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers