DBScholar

Back to papers

SkewTune: Mitigating Skew in MapReduce Applications

Summary: SkewTune automatically mitigates MapReduce skew without extra user input, as a drop-in Hadoop extension. It uses idle-node detection to repartition a straggler's unprocessed data, preserves input order for concatenation-based output reconstruction, and incurs minimal overhead when skew is absent. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
4571
Venue
SIGMOD
Year
2012
Pagerank
0.00011175005
Overall Rank
1,319 | 90.96%
DOI
10.1145/2213836.2213840

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{kwon_sigmod12,
        title = {{SkewTune: Mitigating Skew in MapReduce Applications}},
        author = {Kwon, YongChul and Balazinska, Magdalena and Howe, Bill and Rolia, Jerome},
        series = {{SIGMOD} '12},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2213836.2213840},
        url = {https://dl.acm.org/doi/10.1145/2213836.2213840},
        year = {2012}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
1,514 Upper and Lower Bounds on the Cost of a Map-Reduce Computation 2013 VLDB 0.00010527649
2,539 Minimal MapReduce Algorithms 2013 SIGMOD 8.4526595e-05
3,645 A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories 2017 VLDB 7.2293286e-05
4,480 LocationSpark: A Distributed In-Memory Data Management System for Big Spatial Data 2016 VLDB 6.67565e-05
5,122 A Padded Encoding Scheme to Accelerate Scans by Leveraging Skew 2015 SIGMOD 6.3573169e-05
5,676 Scalable Progressive Analytics on Big Data in the Cloud 2013 VLDB 6.1251441e-05
5,748 Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture 2013 SIGMOD 6.1014813e-05
6,736 Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads 2013 VLDB 5.787547e-05
6,762 SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning 2017 VLDB 5.7814194e-05
7,111 Submodularity of Distributed Join Computation 2018 SIGMOD 5.69924e-05
7,546 MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs 2014 VLDB 5.6024593e-05
8,668 Toward Progress Indicators on Steroids for Big Data Systems 2013 CIDR 5.3882971e-05
8,958 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.3449654e-05
9,128 Dalton: Learned Partitioning for Distributed Data Streams 2023 VLDB 5.3189314e-05
9,137 SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory 2014 SIGMOD 5.3162511e-05
11,728 Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared Clusters 2021 VLDB 5.093636e-05
11,889 An Experimental Evaluation of Garbage Collectors on Big Data Applications 2019 VLDB 5.093636e-05
12,131 FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed Data 2015 VLDB 5.093636e-05
12,147 Big Data Research: Will Industry Solve all the Problems? 2015 VLDB 5.093636e-05
12,336 SkewTune in Action: Mitigating Skew in MapReduce Applications 2012 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 16 of 16 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers