Starfish: A Self-tuning System for Big Data Analytics
Summary: Starfish automatically tunes Hadoop MapReduce workflows to improve runtime, resource utilization, and cloud cost without manual knob fiddling. It adapts self-tuning DB techniques—cost models, profiling and reconfiguration—to MapReduce’s workload variability and pay‑as‑you‑go environments. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Herodotos Herodotou (Duke University)
- 2. Harold Lim (Duke University)
- 3. Gang Luo (Duke University)
- 4. Nedyalko Borisov (Duke University)
- 5. Liang Dong (Duke University)
- 6. Fatma Bilgen Cetin (Duke University)
- 7. Shivnath Babu (Duke University)
BibTeX Citation
@inproceedings{herodotou_cidr11,
address = {Amsterdam, Netherlands},
series = {{CIDR} '11},
title = {{Starfish: A Self-tuning System for Big Data Analytics}},
booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
author = {Herodotou, Herodotos and Lim, Harold and Luo, Gang and Borisov, Nedyalko and Dong, Liang and Cetin, Fatma Bilgen and Babu, Shivnath},
year = {2011}
}
Incoming Citations (Sorted by Pagerank)
Showing 32 of 32 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 5 of 5 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 120 | HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads | 2009 | VLDB | 0.00031680027 |
| 155 | MAD Skills: New Analysis Practices for Big Data | 2009 | VLDB | 0.00028713176 |
| 660 | Hadoop++: Making a Yellow Elephant Run Like a Cheetah (Without It Even Noticing) | 2010 | VLDB | 0.00015198804 |
| 803 | MRShare: Sharing Across Multiple Queries in MapReduce | 2010 | VLDB | 0.00013899943 |
| 1,436 | The Performance of MapReduce: An In-depth Study | 2010 | VLDB | 0.00010797443 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 8,276 | CARTILAGE: Adding Flexibility to the Hadoop Skeleton | 2013 | SIGMOD |
| 2 | 8,455 | Piranha: Optimizing Short Jobs in Hadoop | 2013 | VLDB |
| 3 | 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |
| 4 | 2,265 | A Platform for Scalable One-Pass Analytics using MapReduce | 2011 | SIGMOD |
| 5 | 9,640 | Supporting Scalable Analytics with Latency Constraints | 2015 | VLDB |
| 6 | 13,557 | Big Data Science Needs Big Data Middleware | 2015 | CIDR |
| 7 | 2,207 | Shark: Fast Data Analysis Using Coarse-grained Distributed Memory | 2012 | SIGMOD |
| 8 | 425 | Shark: SQL and Rich Analytics at Scale | 2013 | SIGMOD |
| 9 | 5,865 | Speedup Your Analytics: Automatic Parameter Tuning for Databases and Big Data Systems | 2019 | VLDB |
| 10 | 8,497 | MapReduce Programming and Cost-based Optimization? Crossing this Chasm with Starfish | 2011 | VLDB |