Back to papers
SkewTune: Mitigating Skew in MapReduce Applications
Summary: SkewTune automatically mitigates MapReduce skew without extra user input, as a drop-in Hadoop extension. It uses idle-node detection to repartition a straggler's unprocessed data, preserves input order for concatenation-based output reconstruction, and incurs minimal overhead when skew is absent.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
h6bd1b7236e7dc74f
Venue
SIGMOD
Year
2012
Pagerank
0.00010929229
Overall Rank
1,351 | 90.93%
DOI
10.1145/2213836.2213840
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{kwon_sigmod12,
title = {{SkewTune: Mitigating Skew in MapReduce Applications}},
author = {Kwon, YongChul and Balazinska, Magdalena and Howe, Bill and Rolia, Jerome},
series = {{SIGMOD} '12},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2213836.2213840},
url = {https://dl.acm.org/doi/10.1145/2213836.2213840},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 21 of 21 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
432
Shark: SQL and Rich Analytics at Scale
2013
SIGMOD
0.00018331051
1,546
Upper and Lower Bounds on the Cost of a Map-Reduce Computation
2013
VLDB
0.0001029751
2,573
Minimal MapReduce Algorithms
2013
SIGMOD
8.2782871e-05
3,697
A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories
2017
VLDB
7.0885369e-05
4,514
LocationSpark: A Distributed In-Memory Data Management System for Big Spatial Data
2016
VLDB
6.5629213e-05
5,213
A Padded Encoding Scheme to Accelerate Scans by Leveraging Skew
2015
SIGMOD
6.222614e-05
5,744
Scalable Progressive Analytics on Big Data in the Cloud
2013
VLDB
6.0072829e-05
5,830
Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture
2013
SIGMOD
5.9757869e-05
6,792
Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads
2013
VLDB
5.6794384e-05
6,863
SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning
2017
VLDB
5.6604842e-05
7,266
Submodularity of Distributed Join Computation
2018
SIGMOD
5.5689674e-05
7,638
MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs
2014
VLDB
5.4767699e-05
8,834
Toward Progress Indicators on Steroids for Big Data Systems
2013
CIDR
5.2661519e-05
9,042
The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap –
2021
SIGMOD
5.2310219e-05
9,209
Dalton: Learned Partitioning for Distributed Data Streams
2023
VLDB
5.2060149e-05
9,316
SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory
2014
SIGMOD
5.1951189e-05
12,037
Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared Clusters
2021
VLDB
4.9769913e-05
12,195
An Experimental Evaluation of Garbage Collectors on Big Data Applications
2019
VLDB
4.9769913e-05
12,430
FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed Data
2015
VLDB
4.9769913e-05
12,445
Big Data Research: Will Industry Solve all the Problems?
2015
VLDB
4.9769913e-05
12,633
SkewTune in Action: Mitigating Skew in MapReduce Applications
2012
VLDB
4.9769913e-05
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
9
Online Aggregation
1997
SIGMOD
0.00076265429
384
HaLoop: Efficient Iterative Data Processing on Large Clusters
2010
VLDB
0.00019471648
490
Practical Skew Handling in Parallel Joins
1992
VLDB
0.00017433989
651
Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience
2009
VLDB
0.00015123701
675
Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)
2010
VLDB
0.00014880686
833
A Taxonomy and Performance Model of Data Skew Effects in Parallel Joins
1991
VLDB
0.00013577649
940
Starfish: A Self-tuning System for Big Data Analytics
2011
CIDR
0.00012959992
978
Parallel Evaluation of Conjunctive Queries
2011
PODS
0.00012725823
1,230
Highly Available, Fault-Tolerant, Parallel Dataflows
2004
SIGMOD
0.00011419665
1,628
Adaptive Parallel Aggregation Algorithms
1995
SIGMOD
0.00010033698
2,285
A Platform for Scalable One-Pass Analytics using MapReduce
2011
SIGMOD
8.6924793e-05
2,673
A Latency and Fault-Tolerance Optimizer for Online Parallel Query Plans
2011
SIGMOD
8.1453443e-05
2,980
Clustera: An Integrated Computation And Data Management System
2008
VLDB
7.7858466e-05
3,969
Estimation of Query-Result Distribution and its Application in Parallel-Join Load Balancing
1996
VLDB
6.8874168e-05
5,544
Efficient outer join data skew handling in parallel DBMS
2009
VLDB
6.0862382e-05
12,633
SkewTune in Action: Mitigating Skew in MapReduce Applications
2012
VLDB
4.9769913e-05
Semantically Similar Papers