Shark: SQL and Rich Analytics at Scale
Summary: Shark unifies SQL and analytics on clusters via a distributed memory abstraction into a single scalable engine. In-memory columnar storage, replanning, and fault tolerance enable SQL and ML, 100x faster than Hive/Hadoop, competitive with MPP. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Reynold S. Xin (University of California Berkeley)
- 2. Josh Rosen (University of California Berkeley)
- 3. Matei Zaharia (University of California Berkeley)
- 4. Michael J. Franklin (University of California Berkeley)
- 5. Scott Shenker (University of California Berkeley)
- 6. Ion Stoica (University of California Berkeley)
BibTeX Citation
@inproceedings{xin_sigmod13,
title = {{Shark: SQL and Rich Analytics at Scale}},
author = {Xin, Reynold S. and Rosen, Josh and Zaharia, Matei and Franklin, Michael J. and Shenker, Scott and Stoica, Ion},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2465288},
url = {https://dl.acm.org/doi/10.1145/2463676.2465288},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 50 of 54 citing papers.
Outgoing Citations (Sorted by Pagerank)
Showing 18 of 18 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,983 | Adaptive and Robust Query Execution for Lakehouses at Scale | 2024 | VLDB |
| 2 | 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 3 | 7,709 | Quill: Efficient, Transferable, and Rich Analytics at Scale | 2016 | VLDB |
| 4 | 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD |
| 5 | 2,850 | HAWQ: A Massively Parallel Processing SQL Engine in Hadoop | 2014 | SIGMOD |
| 6 | 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |
| 7 | 1,227 | Blink and It's Done: Interactive Queries on Very Large Data | 2012 | VLDB |
| 8 | 4,806 | SharkDB: An In-Memory Storage System for Massive Trajectory Data | 2015 | SIGMOD |
| 9 | 923 | Starfish: A Self-tuning System for Big Data Analytics | 2011 | CIDR |
| 10 | 2,207 | Shark: Fast Data Analysis Using Coarse-grained Distributed Memory | 2012 | SIGMOD |