DBScholar

Back to papers

Spark SQL: Relational Data Processing in Spark

Summary: Relational processing integrated into Spark via DataFrame API, unifying SQL queries with Spark's functional workflow. Catalyst, a Scala-based extensible optimizer, enables composable rules, code generation, JSON schema inference, and federation to databases. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
5084
Venue
SIGMOD
Year
2015
Pagerank
0.00054865648
Overall Rank
24 | 99.84%
DOI
10.1145/2723372.2742797

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{armbrust_sigmod15,
        title = {{Spark SQL: Relational Data Processing in Spark}},
        author = {Armbrust, Michael and Xin, Reynold S. and Lian, Cheng and Huai, Yin and Liu, Davies and Bradley, Joseph K. and Meng, Xiangrui and Kaftan, Tomer and Franklin, Michael J. and Ghodsi, Ali and Zaharia, Matei},
        series = {{SIGMOD} '15},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2723372.2742797},
        url = {https://dl.acm.org/doi/10.1145/2723372.2742797},
        year = {2015}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 207 citing papers.

Rank Citing Paper Year Venue Pagerank
6,344 Towards General and Efficient Online Tuning for Spark 2023 VLDB 5.9060457e-05
6,375 MIFO: A Query-Semantic Aware Resource Allocation Policy 2019 SIGMOD 5.8932222e-05
6,377 Scalable Querying of Nested Data 2021 VLDB 5.8931544e-05
6,407 Shared Foundations: Modernizing Meta's Data Lakehouse 2023 CIDR 5.8850639e-05
6,482 AStream: Ad-hoc Shared Stream Processing 2019 SIGMOD 5.8668582e-05
6,614 Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine 2025 SIGMOD 5.8216658e-05
6,628 ConnectorX: Accelerating Data Loading From Databases to Dataframes 2022 VLDB 5.8197537e-05
6,632 DistME: A Fast and Elastic Distributed Matrix Computation Engine using GPUs 2019 SIGMOD 5.8190054e-05
6,762 SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning 2017 VLDB 5.7814194e-05
6,801 Unit Testing Data with Deequ 2019 SIGMOD 5.7690726e-05
6,822 JetScope: Reliable and Interactive Analytics at Cloud Scale 2015 VLDB 5.7627651e-05
6,865 SparkR: Scaling R Programs with Spark 2016 SIGMOD 5.7508365e-05
6,939 Multi-Tenant Cloud Data Services: State-of-the-Art, Challenges and Opportunities 2022 SIGMOD 5.7338637e-05
7,091 Simple & Optimal Quantile Sketch: Combining Greenwald-Khanna with Khanna-Greenwald 2024 PODS 5.7062634e-05
7,139 Selection Pushdown in Column Stores using Bit Manipulation Instructions 2023 SIGMOD 5.6932481e-05
7,176 Interactive Demonstration of Probabilistic Predicates 2018 SIGMOD 5.6816132e-05
7,179 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 5.6802479e-05
7,182 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 5.6793679e-05
7,238 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 5.664642e-05
7,369 SmartBench: A Benchmark For Data Management In Smart Spaces 2020 VLDB 5.6314085e-05
7,401 Membrane - Safe and Performant Data Access Controls in Apache Spark in the Presence of Imperative Code 2024 VLDB 5.6255291e-05
7,551 Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams 2022 VLDB 5.6006414e-05
7,687 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.5671645e-05
7,709 Quill: Efficient, Transferable, and Rich Analytics at Scale 2016 VLDB 5.5628268e-05
7,733 Mind the Gap: Bridging Multi-Domain Query Workloads with EmptyHeaded 2017 VLDB 5.5564458e-05
7,777 S2RDF: RDF Querying with SPARQL on Spark 2016 VLDB 5.5461212e-05
7,780 Petabyte-Scale Row-Level Operations in Data Lakehouses 2024 VLDB 5.5457298e-05
7,798 A Survey and Experimental Comparison of Distributed SPARQL Engines for Very Large RDF Data 2017 VLDB 5.542286e-05
7,856 AJoin: Ad-hoc Stream Joins at Scale 2020 VLDB 5.5306709e-05
8,001 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.508791e-05
8,057 Architecting a Query Compiler for Spatial Workloads 2020 SIGMOD 5.4982831e-05
8,145 You Say ‘What’, I Hear ‘Where’ and ‘Why’ — (Mis-)Interpreting SQL to Derive Fine-Grained Provenance 2018 VLDB 5.4790624e-05
8,175 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft 2021 VLDB 5.4737932e-05
8,272 Shasta: Interactive Reporting At Scale 2016 SIGMOD 5.4574671e-05
8,318 Optimizing Declarative Graph Queries at Large Scale 2019 SIGMOD 5.4548766e-05
8,368 Excalibur: A Virtual Machine for Adaptive Fine-grained JIT-Compiled Query Execution based on VOILA 2023 VLDB 5.4419148e-05
8,438 Flare & Lantern: Efficiently Swapping Horses Midstream 2019 VLDB 5.4251194e-05
8,439 New Query Optimization Techniques in the Spark Engine of Azure Synapse 2022 VLDB 5.4248071e-05
8,465 Predicate Pushdown for Data Science Pipelines 2023 SIGMOD 5.4194578e-05
8,589 Optimizing Video Selection LIMIT Queries With Commonsense Knowledge 2024 VLDB 5.4065627e-05
8,615 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 5.4005602e-05
8,712 Handling Environments in a Nested Relational Algebra with Combinators and an Implementation in a Verified Query Compiler 2017 SIGMOD 5.3779634e-05
8,721 Accelerate Distributed Joins with Predicate Transfer 2025 SIGMOD 5.3772617e-05
8,738 Translation of Array-Based Loops to Distributed Data-Parallel Programs 2020 VLDB 5.3766157e-05
8,927 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.3483178e-05
8,958 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.3449654e-05
8,999 HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics Queries 2021 SIGMOD 5.3354529e-05
9,120 Making Data Engineering Declarative 2023 CIDR 5.3194157e-05
9,127 Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads 2025 SIGMOD 5.3189314e-05
9,167 Dynamic Speculative Optimizations for SQL Compilation in Apache Spark 2020 VLDB 5.3099969e-05
Previous Page 3 / 5 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers