Database Paper Browser

Back to papers

Spark SQL: Relational Data Processing in Spark

Summary: Relational processing integrated into Spark via DataFrame API, unifying SQL queries with Spark's functional workflow. Catalyst, a Scala-based extensible optimizer, enables composable rules, code generation, JSON schema inference, and federation to databases. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
5023
Venue
SIGMOD
Year
2015
Pagerank
0.00055280049
Overall Rank
25 | 99.83%
DOI
10.1145/2723372.2742797

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 206 citing papers.

Rank Citing Paper Year Venue Pagerank
6,248 Towards General and Efficient Online Tuning for Spark 2023 VLDB 6.0022621e-05
6,289 Scalable Querying of Nested Data 2021 VLDB 5.9881025e-05
6,296 MIFO: A Query-Semantic Aware Resource Allocation Policy 2019 SIGMOD 5.984487e-05
6,316 Shared Foundations: Modernizing Meta's Data Lakehouse 2023 CIDR 5.9759793e-05
6,403 AStream: Ad-hoc Shared Stream Processing 2019 SIGMOD 5.952729e-05
6,429 DistME: A Fast and Elastic Distributed Matrix Computation Engine using GPUs 2019 SIGMOD 5.9439698e-05
6,512 Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine 2025 SIGMOD 5.91183e-05
6,688 SquirrelJoin: Network-Aware Distributed Join Processing with Lazy Partitioning 2017 VLDB 5.8604776e-05
6,702 Unit Testing Data with Deequ 2019 SIGMOD 5.856543e-05
6,715 JetScope: Reliable and Interactive Analytics at Cloud Scale 2015 VLDB 5.8520075e-05
6,726 Membrane - Safe and Performant Data Access Controls in Apache Spark in the Presence of Imperative Code 2024 VLDB 5.8476884e-05
6,756 ConnectorX: Accelerating Data Loading From Databases to Dataframes 2022 VLDB 5.8399894e-05
6,785 SparkR: Scaling R Programs with Spark 2016 SIGMOD 5.8328772e-05
6,821 Multi-Tenant Cloud Data Services: State-of-the-Art, Challenges and Opportunities 2022 SIGMOD 5.8216165e-05
7,025 Selection Pushdown in Column Stores using Bit Manipulation Instructions 2023 SIGMOD 5.77988e-05
7,037 Bubble Execution: Resource-aware Reliable Analytics at Cloud Scale 2018 VLDB 5.7720294e-05
7,045 Interactive Demonstration of Probabilistic Predicates 2018 SIGMOD 5.7696025e-05
7,054 CleanM: An Optimizable Query Language for Unified Scale-Out Data Cleaning 2017 VLDB 5.7683138e-05
7,120 Kodiak: Leveraging Materialized Views For Very Low-Latency Analytics Over High-Dimensional Web-Scale Data 2016 VLDB 5.7519336e-05
7,241 SmartBench: A Benchmark For Data Management In Smart Spaces 2020 VLDB 5.7186261e-05
7,416 Enabling Efficient and General Subpopulation Analytics in Multidimensional Data Streams 2022 VLDB 5.6873825e-05
7,559 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.6533871e-05
7,583 Quill: Efficient, Transferable, and Rich Analytics at Scale 2016 VLDB 5.6487011e-05
7,600 Mind the Gap: Bridging Multi-Domain Query Workloads with EmptyHeaded 2017 VLDB 5.6441526e-05
7,646 S2RDF: RDF Querying with SPARQL on Spark 2016 VLDB 5.6320114e-05
7,650 Petabyte-Scale Row-Level Operations in Data Lakehouses 2024 VLDB 5.6316205e-05
7,669 A Survey and Experimental Comparison of Distributed SPARQL Engines for Very Large RDF Data 2017 VLDB 5.6279273e-05
7,729 AJoin: Ad-hoc Stream Joins at Scale 2020 VLDB 5.6161884e-05
7,763 Architecting a Query Compiler for Spatial Workloads 2020 SIGMOD 5.6079812e-05
7,879 Pangea: Monolithic Distributed Storage for Data Analytics 2019 VLDB 5.5941095e-05
7,959 Simple & Optimal Quantile Sketch: Combining Greenwald-Khanna with Khanna-Greenwald 2024 PODS 5.5727796e-05
8,019 You Say 'What', I Hear 'Where' and 'Why' - (Mis-)Interpreting SQL to Derive Fine-Grained Provenance 2018 VLDB 5.5637512e-05
8,046 SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft 2021 VLDB 5.5585663e-05
8,137 Shasta: Interactive Reporting At Scale 2016 SIGMOD 5.5419908e-05
8,179 Optimizing Declarative Graph Queries at Large Scale 2019 SIGMOD 5.5389901e-05
8,209 Flare & Lantern: Efficiently Swapping Horses Midstream 2019 VLDB 5.5345733e-05
8,245 Excalibur: A Virtual Machine for Adaptive Fine-grained JIT-Compiled Query Execution based on VOILA 2023 VLDB 5.5249762e-05
8,323 New Query Optimization Techniques in the Spark Engine of Azure Synapse 2022 VLDB 5.5075269e-05
8,347 Predicate Pushdown for Data Science Pipelines 2023 SIGMOD 5.5033928e-05
8,460 Optimizing Video Selection LIMIT Queries With Commonsense Knowledge 2024 VLDB 5.4902979e-05
8,472 A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning 2024 VLDB 5.4888877e-05
8,584 Accelerate Distributed Joins with Predicate Transfer 2025 SIGMOD 5.467762e-05
8,607 Handling Environments in a Nested Relational Algebra with Combinators and an Implementation in a Verified Query Compiler 2017 SIGMOD 5.4600294e-05
8,623 Translation of Array-Based Loops to Distributed Data-Parallel Programs 2020 VLDB 5.4598872e-05
8,797 Hyperspace: The Indexing Subsystem of Azure Synapse 2021 VLDB 5.4311509e-05
8,865 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.4200866e-05
8,899 HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics Queries 2021 SIGMOD 5.4123498e-05
8,987 Making Data Engineering Declarative 2023 CIDR 5.4018012e-05
8,994 Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads 2025 SIGMOD 5.4013095e-05
9,031 Dynamic Speculative Optimizations for SQL Compilation in Apache Spark 2020 VLDB 5.3924717e-05
Previous Page 3 / 5 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers