DBScholar

Back to papers

Shark: SQL and Rich Analytics at Scale

Summary: Shark unifies SQL and analytics on clusters via a distributed memory abstraction into a single scalable engine. In-memory columnar storage, replanning, and fault tolerance enable SQL and ML, 100x faster than Hive/Hadoop, competitive with MPP. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
hb5edf1e8c00cfaf4
Venue
SIGMOD
Year
2013
Pagerank
0.00018331051
Overall Rank
432 | 97.10%
DOI
10.1145/2463676.2465288

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{xin_sigmod13,
        title = {{Shark: SQL and Rich Analytics at Scale}},
        author = {Xin, Reynold S. and Rosen, Josh and Zaharia, Matei and Franklin, Michael J. and Shenker, Scott and Stoica, Ion},
        series = {{SIGMOD} '13},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2463676.2465288},
        url = {https://dl.acm.org/doi/10.1145/2463676.2465288},
        year = {2013}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 54 citing papers.

Rank Citing Paper Year Venue Pagerank
23 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00055384955
1,036 Fine-grained Partitioning for Aggressive Data Skipping 2014 SIGMOD 0.00012372946
1,178 Simba: Efficient In-Memory Spatial Analytics 2016 SIGMOD 0.00011627256
1,292 From Theory to Practice: Efficient Join Query Evaluation in a Parallel Database System 2015 SIGMOD 0.00011147959
1,381 Scuba: Diving into Data at Facebook 2013 VLDB 0.00010856731
1,428 Knowing When You’re Wrong: Building Fast and Reliable Approximate Query Processing Systems 2014 SIGMOD 0.0001069161
1,482 Skew in Parallel Query Processing 2014 PODS 0.00010534147
1,623 Mesa: Geo-Replicated, Near Real-Time, Scalable Data Warehousing 2014 VLDB 0.00010047769
1,885 SQL-on-Hadoop: Full Circle Back to Shared-Nothing Database Architectures 2014 VLDB 9.4345374e-05
2,139 Quickstep: A Data Platform Based on the Scaling-Up Approach 2018 VLDB 8.9735524e-05
2,424 BigDansing: A System for Big Data Cleansing 2015 SIGMOD 8.483813e-05
2,499 Towards Scalable Real-time Analytics: An Architecture for Scale-out of OLxP Workloads 2015 VLDB 8.3795899e-05
2,628 WideTable: An Accelerator for Analytical Data Processing 2014 VLDB 8.2027948e-05
2,898 HAWQ: A Massively Parallel Processing SQL Engine in Hadoop 2014 SIGMOD 7.8816694e-05
2,954 Column Sketches: A Scan Accelerator for Rapid and Robust Predicate Evaluation 2018 SIGMOD 7.8119682e-05
3,241 Locality-aware Partitioning in Parallel Database Systems 2015 SIGMOD 7.4972383e-05
3,558 WANalytics: Analytics for a Geo-Distributed Data-Intensive World 2015 CIDR 7.2050361e-05
3,597 Access Path Selection in Main-Memory Optimized Data Systems: Should I Scan or Should I Probe? 2017 SIGMOD 7.1759026e-05
3,739 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.0595382e-05
4,069 Dynamically Optimizing Queries over Large Scale Data Platforms 2014 SIGMOD 6.8184364e-05
4,396 WANalytics: Geo-Distributed Analytics for a Data Intensive World 2015 SIGMOD 6.6152509e-05
4,875 Design Tradeoffs of Data Access Methods 2016 SIGMOD 6.3699172e-05
5,062 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing 2022 VLDB 6.2896995e-05
5,354 Adaptive and Robust Query Execution for Lakehouses at Scale 2024 VLDB 6.1661231e-05
5,836 Opportunistic Physical Design for Big Data Analytics 2014 SIGMOD 5.9732317e-05
6,262 A Performance Study of Big Data on Small Nodes 2015 VLDB 5.8301898e-05
6,291 Elastic Pipelining in an In-Memory Database Cluster 2016 SIGMOD 5.8200348e-05
6,453 Towards General and Efficient Online Tuning for Spark 2023 VLDB 5.7790502e-05
6,552 JoinBoost: Grow Trees Over Normalized Data Using Only SQL 2023 VLDB 5.7475822e-05
6,769 Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters 2013 VLDB 5.6847727e-05
6,788 Adaptive Data Skipping in Main-Memory Systems 2016 SIGMOD 5.6808999e-05
6,819 Liquid: Unifying Nearline and Offline Big Data Integration 2015 CIDR 5.6713167e-05
6,960 JetScope: Reliable and Interactive Analytics at Cloud Scale 2015 VLDB 5.6313571e-05
6,978 SparkR: Scaling R Programs with Spark 2016 SIGMOD 5.6272603e-05
7,272 Decentralized Actor Scheduling and Reference-based Storage in Xorbits: a Native Scalable Data Science Engine 2025 VLDB 5.5668569e-05
7,315 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 5.5529695e-05
7,365 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 5.5401198e-05
7,840 Quill: Efficient, Transferable, and Rich Analytics at Scale 2016 VLDB 5.4409637e-05
7,901 Using VDMS to Index and Search 100M Images 2021 VLDB 5.4291648e-05
8,009 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 5.4063491e-05
8,349 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft 2021 VLDB 5.3485189e-05
8,623 Parallel-Correctness and Transferability for Conjunctive Queries 2015 PODS 5.2999107e-05
8,631 Piranha: Optimizing Short Jobs in Hadoop 2013 VLDB 5.298422e-05
9,264 QMapper for Smart Grid: Migrating SQL-based Application to Hive 2015 SIGMOD 5.2032182e-05
9,778 Cost-based Fault-tolerance for Parallel Data Processing 2015 SIGMOD 5.1294772e-05
9,985 Introduction to Spark 2.0 for Database Researchers 2016 SIGMOD 5.0989958e-05
11,135 Dynamic Pruning for Recursive Joins 2025 SIGMOD 4.9769913e-05
12,191 Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology 2019 VLDB 4.9769913e-05
12,333 Logical Aspects of Massively Parallel and Distributed Systems 2016 PODS 4.9769913e-05
12,381 Parallel Evaluation of Multi-Semi-Joins 2016 VLDB 4.9769913e-05
Previous Page 1 / 2 Next

Outgoing Citations (Sorted by Pagerank)

Showing 18 of 18 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pregel: A System for Large-Scale Graph Processing 2010 SIGMOD 0.0012087459
12 C-Store: A Column-oriented DBMS 2005 VLDB 0.0006897844
22 Distributed GraphLab: A Framework for Machine Learning and Data Mining in the Cloud 2012 VLDB 0.00055938421
30 SCOPE: Easy and Efficient Parallel Processing of Massive Data Sets 2008 VLDB 0.00050475202
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.0004552807
49 Dremel: Interactive Analysis of Web-Scale Datasets 2010 VLDB 0.0004314366
53 Eddies: Continuously Adaptive Query Processing 2000 SIGMOD 0.000408505
120 HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads 2009 VLDB 0.00031099083
149 Efficient Mid-Query Re-Optimization of Sub-Optimal Query Execution Plans 1998 SIGMOD 0.00028977821
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028568843
384 HaLoop: Efficient Iterative Data Processing on Large Clusters 2010 VLDB 0.00019471648
423 Cost-based Query Scrambling for Initial Delays 1998 SIGMOD 0.00018487497
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017195428
876 Tenzing: A SQL Implementation On The MapReduce Framework 2011 VLDB 0.00013304112
1,210 Processing a Trillion Cells per Mouse Click 2012 VLDB 0.0001152237
1,351 SkewTune: Mitigating Skew in MapReduce Applications 2012 SIGMOD 0.00010929229
1,810 Distributed Data-Parallel Computing Using a High-Level Programming Language 2009 SIGMOD 9.584952e-05
1,838 Cheetah: A High Performance, Custom Data Warehouse on Top of MapReduce 2010 VLDB 9.5305061e-05
Previous Page 1 / 1 Next

Semantically Similar Papers