DBScholar

Back to papers

Shark: SQL and Rich Analytics at Scale

Summary: Shark unifies SQL and analytics on clusters via a distributed memory abstraction into a single scalable engine. In-memory columnar storage, replanning, and fault tolerance enable SQL and ML, 100x faster than Hive/Hadoop, competitive with MPP. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hb5edf1e8c00cfaf4
Venue
SIGMOD
Year
2013
Pagerank
0.00018339357
Overall Rank
432 | 97.10%
DOI
10.1145/2463676.2465288

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{xin_sigmod13,
        title = {{Shark: SQL and Rich Analytics at Scale}},
        author = {Xin, Reynold S. and Rosen, Josh and Zaharia, Matei and Franklin, Michael J. and Shenker, Scott and Stoica, Ion},
        series = {{SIGMOD} '13},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2463676.2465288},
        url = {https://dl.acm.org/doi/10.1145/2463676.2465288},
        year = {2013}
}

Incoming Citations (Sorted by Pagerank)

Showing 4 of 54 citing papers.

Rank Citing Paper Year Venue Pagerank
12,438 Tutorial: SQL-on-Hadoop Systems 2015 VLDB 4.9793485e-05
12,463 DoomDB - Kill the Query 2014 SIGMOD 4.9793485e-05
12,482 A Partitioning Framework for Aggressive Data Skipping 2014 VLDB 4.9793485e-05
12,488 Getting Your Big Data Priorities Straight: A Demonstration of Priority-based QoS using Social-network-driven Stock Recommendation 2014 VLDB 4.9793485e-05
Previous Page 2 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pregel: A System for Large-Scale Graph Processing 2010 SIGMOD 0.0012092602
12 C-Store: A Column-oriented DBMS 2005 VLDB 0.00068998927
22 Distributed GraphLab: A Framework for Machine Learning and Data Mining in the Cloud 2012 VLDB 0.00055962491
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050495102
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.00043160717
53 Eddies: Continuously Adaptive Query Processing 2000 SIGMOD 0.00040860054
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.000311132
149 Efficient Mid-Query Re-Optimization of Sub-Optimal Query Execution Plans 1998 SIGMOD 0.00028981723
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028579704
384 HaLoop: Efficient Iterative Data Processing on Large Clusters 2010 VLDB 0.0001948031
424 Cost-based Query Scrambling for Initial Delays 1998 SIGMOD 0.0001848836
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017202276
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013309176
1,210 Processing a Trillion Cells per Mouse Click 2012 VLDB 0.00011527605
1,351 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00010934347
1,810 Distributed Data-Parallel Computing Using a High-Level Programming Language 2009 SIGMOD 9.5892079e-05
1,837 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5350058e-05
Previous Page 1 / 1 Next

Semantically Similar Papers