Shark: Fast Data Analysis Using Coarse-grained Distributed Memory
Summary: Shark is a data-analysis system built on a coarse-grained distributed shared-memory abstraction, unifying SQL querying with near-data analytics. Scales to thousands of fault-tolerant nodes; delivers 40x faster queries vs Hive and 25x faster ML vs MapReduce on large datasets. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Cliff Engle (University of California Berkeley)
- 2. Antonio Lupher (University of California Berkeley)
- 3. Reynold Xin (University of California Berkeley)
- 4. Matei Zaharia (University of California Berkeley)
- 5. Michael J. Franklin (University of California Berkeley)
- 6. Scott Shenker (University of California Berkeley)
- 7. Ion Stoica (University of California Berkeley)
BibTeX Citation
@inproceedings{engle_sigmod12,
title = {{Shark: Fast Data Analysis Using Coarse-grained Distributed Memory}},
author = {Engle, Cliff and Lupher, Antonio and Xin, Reynold and Zaharia, Matei and Franklin, Michael J. and Shenker, Scott and Stoica, Ion},
series = {{SIGMOD} '12},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2213836.2213934},
url = {https://dl.acm.org/doi/10.1145/2213836.2213934},
year = {2012}
}
Incoming Citations (Sorted by Pagerank)
Showing 12 of 12 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 1 of 1 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 44 | A Comparison of Approaches to Large-Scale Data Analysis | 2009 | SIGMOD | 0.00046055057 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,043 | SHARQL: Shape Analysis of Recursive SPARQL Queries | 2020 | SIGMOD |
| 2 | 9,953 | Amoeba: A Shape changing Storage System for Big Data | 2016 | VLDB |
| 3 | 8,455 | Piranha: Optimizing Short Jobs in Hadoop | 2013 | VLDB |
| 4 | 10,605 | SHARD: A Scalable and Resize-optimized Hash Index on Disaggregated Memory | 2026 | VLDB |
| 5 | 7,932 | Hone: “Scaling Down” Hadoop on Shared-Memory Systems | 2013 | VLDB |
| 6 | 3,789 | Husky: Towards a More Efficient and Expressive Distributed Computing Framework | 2016 | VLDB |
| 7 | 1,227 | Blink and It's Done: Interactive Queries on Very Large Data | 2012 | VLDB |
| 8 | 923 | Starfish: A Self-tuning System for Big Data Analytics | 2011 | CIDR |
| 9 | 4,806 | SharkDB: An In-Memory Storage System for Massive Trajectory Data | 2015 | SIGMOD |
| 10 | 425 | Shark: SQL and Rich Analytics at Scale | 2013 | SIGMOD |