Back to papers
SkewTune: Mitigating Skew in MapReduce Applications
Summary: SkewTune automatically mitigates MapReduce skew without extra user input, as a drop-in Hadoop extension. It uses idle-node detection to repartition a straggler's unprocessed data, preserves input order for concatenation-based output reconstruction, and incurs minimal overhead when skew is absent.
(summarized by gpt-5-nano on Feb 09 2026)
Paper ID
4571
Venue
SIGMOD
Year
2012
Pagerank
0.00011175005
Overall Rank
1,319 | 90.96%
DOI
10.1145/2213836.2213840
Incoming Non-self Citations Over Time
BibTeX Citation
Copy BibTeX
@inproceedings{kwon_sigmod12,
title = {{SkewTune: Mitigating Skew in MapReduce Applications}},
author = {Kwon, YongChul and Balazinska, Magdalena and Howe, Bill and Rolia, Jerome},
series = {{SIGMOD} '12},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2213836.2213840},
url = {https://dl.acm.org/doi/10.1145/2213836.2213840},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 21 of 21 citing papers.
Rank
Citing Paper
Year
Venue
Pagerank
425
Shark: SQL and Rich Analytics at Scale
2013
SIGMOD
0.00018704491
1,514
Upper and Lower Bounds on the Cost of a Map-Reduce Computation
2013
VLDB
0.00010527649
2,539
Minimal MapReduce Algorithms
2013
SIGMOD
8.4526595e-05
3,645
A General and Parallel Platform for Mining Co-Movement Patterns over Large-scale Trajectories
2017
VLDB
7.2293286e-05
4,480
LocationSpark: A Distributed In-Memory Data Management System for Big Spatial Data
2016
VLDB
6.67565e-05
5,122
A Padded Encoding Scheme to Accelerate Scans by Leveraging Skew
2015
SIGMOD
6.3573169e-05
5,676
Scalable Progressive Analytics on Big Data in the Cloud
2013
VLDB
6.1251441e-05
5,748
Fast Data in the Era of Big Data: Twitter's Real-Time Related Query Suggestion Architecture
2013
SIGMOD
6.1014813e-05
6,736
Hadoop's Adolescence: An analysis of Hadoop usage in scientific workloads
2013
VLDB
5.787547e-05
6,762
SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning
2017
VLDB
5.7814194e-05
7,111
Submodularity of Distributed Join Computation
2018
SIGMOD
5.69924e-05
7,546
MRTuner: A Toolkit to Enable Holistic Optimization for MapReduce Jobs
2014
VLDB
5.6024593e-05
8,668
Toward Progress Indicators on Steroids for Big Data Systems
2013
CIDR
5.3882971e-05
8,958
The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap –
2021
SIGMOD
5.3449654e-05
9,128
Dalton: Learned Partitioning for Distributed Data Streams
2023
VLDB
5.3189314e-05
9,137
SpongeFiles: Mitigating Data Skew in MapReduce Using Distributed Memory
2014
SIGMOD
5.3162511e-05
11,728
Fangorn: Adaptive Execution Framework for Heterogeneous Workloads on Shared Clusters
2021
VLDB
5.093636e-05
11,889
An Experimental Evaluation of Garbage Collectors on Big Data Applications
2019
VLDB
5.093636e-05
12,131
FP-Hadoop: Efficient Execution of Parallel Jobs Over Skewed Data
2015
VLDB
5.093636e-05
12,147
Big Data Research: Will Industry Solve all the Problems?
2015
VLDB
5.093636e-05
12,336
SkewTune in Action: Mitigating Skew in MapReduce Applications
2012
VLDB
5.093636e-05
Outgoing Citations (Sorted by Pagerank)
Showing 16 of 16 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Rank
Cited Paper
Year
Venue
Pagerank
9
Online Aggregation
1997
SIGMOD
0.00077458002
372
HaLoop: Efficient Iterative Data Processing on Large Clusters
2010
VLDB
0.0001981521
481
Practical Skew Handling in Parallel Joins
1992
VLDB
0.00017780716
642
Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience
2009
VLDB
0.00015395331
660
Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing)
2010
VLDB
0.00015198804
811
A Taxonomy and Performance Model of Data Skew Effects in Parallel Joins
1991
VLDB
0.00013859761
923
Starfish: A Self-tuning System for Big Data Analytics
2011
CIDR
0.00013189886
954
Parallel Evaluation of Conjunctive Queries
2011
PODS
0.00012997301
1,209
Highly Available, Fault-Tolerant, Parallel Dataflows
2004
SIGMOD
0.00011657178
1,596
Adaptive Parallel Aggregation Algorithms
1995
SIGMOD
0.00010247091
2,265
A Platform for Scalable One-Pass Analytics using MapReduce
2011
SIGMOD
8.8398946e-05
2,622
A Latency and Fault-Tolerance Optimizer for Online Parallel Query Plans
2011
SIGMOD
8.3330136e-05
2,921
Clustera: An Integrated Computation And Data Management System
2008
VLDB
7.9626308e-05
3,893
Estimation of Query-Result Distribution and its Application in Parallel-Join Load Balancing
1996
VLDB
7.04191e-05
5,411
Efficient outer join data skew handling in parallel DBMS
2009
VLDB
6.2273741e-05
12,336
SkewTune in Action: Mitigating Skew in MapReduce Applications
2012
VLDB
5.093636e-05
Semantically Similar Papers