DBScholar

Back to papers

Minimal MapReduce Algorithms

Summary: Introduces the 'minimal algorithm' notion for MapReduce, optimizing load balancing, space, CPU, I/O, and network cost within a small constant factor. Shows existence of elegant minimal algorithms for fundamental database problems, validated by extensive experiments. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h35532cf700e0ccfd
Venue
SIGMOD
Year
2013
Pagerank
8.2821647e-05
Overall Rank
2,573 | 82.71%
DOI
10.1145/2463676.2463719

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{tao_sigmod13,
        title = {{Minimal MapReduce Algorithms}},
        author = {Tao, Yufei and Lin, Wenqing and Xiao, Xiaokui},
        series = {{SIGMOD} '13},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2463676.2463719},
        url = {https://dl.acm.org/doi/10.1145/2463676.2463719},
        year = {2013}
}

Incoming Citations (Sorted by Pagerank)

Showing 11 of 11 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 36 of 36 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00043160717
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.000311132
360 Efficient Parallel Set-Similarity Joins Using MapReduce 2010 SIGMOD 0.00020009936
561 Densest Subgraph in Streaming and MapReduce 2012 VLDB 0.00016396211
673 Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) 2010 VLDB 0.0001488755
753 Profiling, What-if Analysis, and Cost-based Optimization of MapReduce Programs 2011 VLDB 0.00014237583
799 A Comparison of Join Algorithms for Log Processing in MapReduce 2010 SIGMOD 0.00013889081
823 MRShare: Sharing Across Multiple Queries in MapReduce 2010 VLDB 0.00013648332
865 Processing Theta-Joins using MapReduce* 2011 SIGMOD 0.0001338765
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013309176
962 Fast Personalized PageRank on MapReduce 2011 SIGMOD 0.00012820478
977 Parallel Evaluation of Conjunctive Queries 2011 PODS 0.00012731794
1,022 Online Aggregation for Large MapReduce Jobs 2011 VLDB 0.00012438826
1,037 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis 2011 VLDB 0.00012376819
1,281 Automatic Optimization for MapReduce Programs 2011 VLDB 0.00011213384
1,351 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00010934347
1,421 V-SMART-Join: A Scalable MapReduce Framework for All-Pair Similarity Joins of Multisets and Vectors 2012 VLDB 0.00010726757
1,688 ParaTimer: A Progress Indicator for MapReduce DAGs 2010 SIGMOD 9.8645784e-05
1,708 PerfXplain: Debugging MapReduce Job Performance 2012 VLDB 9.8224587e-05
1,837 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5350058e-05
1,924 ReStore: Reusing Results of MapReduce Jobs 2012 VLDB 9.3687009e-05
2,176 Efficient Processing of k Nearest Neighbor Joins using MapReduce 2012 VLDB 8.9159001e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3425785e-05
2,603 PLANET: Massively Parallel Learning of Tree Ensembles with MapReduce 2009 VLDB 8.2312028e-05
2,687 CoHadoop: Flexible Data Placement and Its Exploitation in Hadoop 2011 VLDB 8.1301312e-05
2,899 Social Content Matching in MapReduce 2011 VLDB 7.8847529e-05
2,903 Column-Oriented Storage Techniques for MapReduce 2011 VLDB 7.8810111e-05
2,931 Efficient Multi-way Theta-Join Processing Using MapReduce 2012 VLDB 7.8405483e-05
3,007 Llama: Leveraging Columnar Storage for Scalable Join Processing in the MapReduce Framework 2011 SIGMOD 7.7607173e-05
3,145 Early Accurate Results for Advanced Analytics on MapReduce 2012 VLDB 7.5965257e-05
3,157 Energy Management for MapReduce Clusters 2010 VLDB 7.5834595e-05
3,779 M3R: Increased Performance for In-Memory Hadoop Jobs 2012 VLDB 7.0256928e-05
4,141 Behavioral Simulations in MapReduce 2010 VLDB 6.7841596e-05
5,652 Building Wavelet Histograms on Large Data in MapReduce 2012 VLDB 6.0515502e-05
Previous Page 1 / 1 Next

Semantically Similar Papers