Shark: SQL and Rich Analytics at Scale
Summary: Shark unifies SQL and analytics on clusters via a distributed memory abstraction into a single scalable engine. In-memory columnar storage, replanning, and fault tolerance enable SQL and ML, 100x faster than Hive/Hadoop, competitive with MPP. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Reynold S. Xin (University of California Berkeley)
- 2. Josh Rosen (University of California Berkeley)
- 3. Matei Zaharia (University of California Berkeley)
- 4. Michael J. Franklin (University of California Berkeley)
- 5. Scott Shenker (University of California Berkeley)
- 6. Ion Stoica (University of California Berkeley)
BibTeX Citation
@inproceedings{xin_sigmod13,
title = {{Shark: SQL and Rich Analytics at Scale}},
author = {Xin, Reynold S. and Rosen, Josh and Zaharia, Matei and Franklin, Michael J. and Shenker, Scott and Stoica, Ion},
series = {{SIGMOD} '13},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2463676.2465288},
url = {https://dl.acm.org/doi/10.1145/2463676.2465288},
year = {2013}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 54 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 12,146 | Tutorial: SQL-on-Hadoop Systems | 2015 | VLDB | 5.093636e-05 |
| 12,172 | DoomDB - Kill the Query | 2014 | SIGMOD | 5.093636e-05 |
| 12,191 | A Partitioning Framework for Aggressive Data Skipping | 2014 | VLDB | 5.093636e-05 |
| 12,197 | Getting Your Big Data Priorities Straight: A Demonstration of Priority-based QoS using Social-network-driven Stock Recommendation | 2014 | VLDB | 5.093636e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 18 of 18 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,983 | Adaptive and Robust Query Execution for Lakehouses at Scale | 2024 | VLDB |
| 2 | 1,889 | SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures | 2014 | VLDB |
| 3 | 7,709 | Quill: Efficient, Transferable, and Rich Analytics at Scale | 2016 | VLDB |
| 4 | 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD |
| 5 | 2,850 | HAWQ: A Massively Parallel Processing SQL Engine in Hadoop | 2014 | SIGMOD |
| 6 | 3,555 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing | 2019 | SIGMOD |
| 7 | 1,227 | Blink and It's Done: Interactive Queries on Very Large Data | 2012 | VLDB |
| 8 | 4,806 | SharkDB: An In-Memory Storage System for Massive Trajectory Data | 2015 | SIGMOD |
| 9 | 923 | Starfish: A Self-tuning System for Big Data Analytics | 2011 | CIDR |
| 10 | 2,207 | Shark: Fast Data Analysis Using Coarse-grained Distributed Memory | 2012 | SIGMOD |