Back to papers
SkewTune: Mitigating Skew in MapReduce Applications
Summary: SkewTune automatically mitigates MapReduce skew without extra user input, as a drop-in Hadoop extension. It uses idle-node detection to repartition a straggler's unprocessed data, preserves input order for concatenation-based output reconstruction, and incurs minimal overhead when skew is absent.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h6bd1b7236e7dc74f
Venue
SIGMOD
Year
2012
Pagerank
0.00010934347
Overall Rank
1,351 | 90.92%
DOI
10.1145/2213836.2213840
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{kwon_sigmod12,
title = {{SkewTune: Mitigating Skew in MapReduce Applications}},
author = {Kwon, YongChul and Balazinska, Magdalena and Howe, Bill and Rolia, Jerome},
series = {{SIGMOD} '12},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2213836.2213840},
url = {https://dl.acm.org/doi/10.1145/2213836.2213840},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 21 of 21 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
432
Shark: SQL and Rich Analytics at Scale
2013
SIGMOD
0.00018339357
1,545
Upper and Lower Bounds on the Cost of a Map-Reduce Computation
2013
VLDB
0.00010302384
2,573
Minimal MapReduce Algorithms
2013
SIGMOD
8.2821647e-05
3,696
A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories
2017
VLDB
7.0918941e-05
4,513
LocationSpark: A Distributed In-Memory Data Management System for Big Spatial Data
2016
VLDB
6.565997e-05
5,213
A Padded Encoding Scheme to Accelerate Scans by Leveraging Skew
2015
SIGMOD
6.2250048e-05
5,745
Scalable Progressive Analytics on Big Data in the Cloud
2013
VLDB
6.0090513e-05
5,828
Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture
2013
SIGMOD
5.9786171e-05
6,786
Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads
2013
VLDB
5.6820666e-05
6,858
SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning
2017
VLDB
5.6631651e-05
7,259
Submodularity of Distributed Join Computation
2018
SIGMOD
5.5716049e-05
7,632
MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs
2014
VLDB
5.4793438e-05
8,825
Toward Progress Indicators on Steroids for Big Data Systems
2013
CIDR
5.268646e-05
9,034
The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap –
2021
SIGMOD
5.2334993e-05
9,200
Dalton: Learned Partitioning for Distributed Data Streams
2023
VLDB
5.2084806e-05
9,307
SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory
2014
SIGMOD
5.197526e-05
12,031
Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared Clusters
2021
VLDB
4.9793485e-05
12,189
An Experimental Evaluation of Garbage Collectors on Big Data Applications
2019
VLDB
4.9793485e-05
12,424
FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed Data
2015
VLDB
4.9793485e-05
12,439
Big Data Research: Will Industry Solve all the Problems?
2015
VLDB
4.9793485e-05
12,627
SkewTune in Action: Mitigating Skew in MapReduce Applications
2012
VLDB
4.9793485e-05
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
9
Online Aggregation
1997
SIGMOD
0.00076195956
384
HaLoop: Efficient Iterative Data Processing on Large Clusters
2010
VLDB
0.0001948031
489
Practical Skew Handling in Parallel Joins
1992
VLDB
0.00017441895
651
Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience
2009
VLDB
0.00015130782
673
Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)
2010
VLDB
0.0001488755
833
A Taxonomy and Performance Model of Data Skew Effects in Parallel Joins
1991
VLDB
0.00013583955
940
Starfish: A Self-tuning System for Big Data Analytics
2011
CIDR
0.00012964445
977
Parallel Evaluation of Conjunctive Queries
2011
PODS
0.00012731794
1,228
Highly Available, Fault-Tolerant, Parallel Dataflows
2004
SIGMOD
0.00011425028
1,628
Adaptive Parallel Aggregation Algorithms
1995
SIGMOD
0.00010037937
2,282
A Platform for Scalable One-Pass Analytics using MapReduce
2011
SIGMOD
8.6964781e-05
2,672
A Latency and Fault-Tolerance Optimizer for Online Parallel Query Plans
2011
SIGMOD
8.1491683e-05
2,978
Clustera: An Integrated Computation And Data Management System
2008
VLDB
7.7894665e-05
3,967
Estimation of Query-Result Distribution and its Application in Parallel-Join Load Balancing
1996
VLDB
6.8905715e-05
5,542
Efficient outer join data skew handling in parallel DBMS
2009
VLDB
6.0891064e-05
12,627
SkewTune in Action: Mitigating Skew in MapReduce Applications
2012
VLDB
4.9793485e-05
Semantically Similar Papers