Database Paper Browser

Back to papers

Spark SQL: Relational Data Processing in Spark

Summary: Relational processing integrated into Spark via DataFrame API, unifying SQL queries with Spark's functional workflow. Catalyst, a Scala-based extensible optimizer, enables composable rules, code generation, JSON schema inference, and federation to databases. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
5023
Venue
SIGMOD
Year
2015
Pagerank
0.00055280049
Overall Rank
25 | 99.83%
DOI
10.1145/2723372.2742797

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 50 of 206 citing papers.

Rank Citing Paper Year Venue Pagerank
296 Accelerating Machine Learning Inference with Probabilistic Predicates 2018 SIGMOD 0.00022115632
452 Apache Calcite: A Foundational Framework for Optimized Query Processing Over Heterogeneous Data Sources 2018 SIGMOD 0.00018289051
472 AnalyticDB-V: A Hybrid Analytical Engine Towards Query Fusion for Structured and Unstructured Data 2020 VLDB 0.0001798859
513 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017314526
524 NeuroCard: One Cardinality Estimator for All Tables 2021 VLDB 0.00017177356
611 Wander Join: Online Aggregation via Random Walks 2016 SIGMOD 0.00015865071
833 Quickr: Lazily Approximating Complex AdHoc Queries in BigData Clusters 2016 SIGMOD 0.00013775566
839 Random Sampling over Joins Revisited 2018 SIGMOD 0.00013751264
1,098 Approximate Query Processing: No Silver Bullet 2017 SIGMOD 0.00012272786
1,139 Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark 2018 SIGMOD 0.0001209681
1,143 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012076462
1,162 Simba: Efficient In-Memory Spatial Analytics 2016 SIGMOD 0.00011985579
1,306 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.00011313192
1,337 Automating Large-Scale Data Quality Verification 2018 VLDB 0.00011211329
1,354 Hybrid Transactional/Analytical Processing: A Survey 2017 SIGMOD 0.00011165852
1,417 Towards Scalable Dataframe Systems 2020 VLDB 0.00010936782
1,503 Procella: Unifying serving and analytical data at YouTube 2019 VLDB 0.00010641071
1,536 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010490973
1,578 Titian: Data Provenance Support in Spark 2016 VLDB 0.00010366201
1,787 How to Architect a Query Compiler 2016 SIGMOD 9.8371029e-05
1,819 Axiomatic Foundations and Algorithms for Deciding Semantic Equivalences of SQL Queries 2018 VLDB 9.7751598e-05
1,856 Photon: A Fast Query Engine for Lakehouse Systems 2022 SIGMOD 9.699164e-05
1,901 DBEst: Revisiting Approximate Query Processing Engines with Machine Learning Models 2019 SIGMOD 9.6066443e-05
1,936 Tuplex: Data Science in Python at Native Code Speed 2021 SIGMOD 9.5283932e-05
1,973 POLARIS: The Distributed SQL Engine in Azure Synapse 2020 VLDB 9.4689067e-05
1,976 FLAT: Fast, Lightweight and Accurate Method for Cardinality Estimation 2021 VLDB 9.4645971e-05
1,978 Database Learning: Toward a Database that Becomes Smarter Every Time 2017 SIGMOD 9.4626567e-05
2,006 ModelarDB: Modular Model-Based Time Series Management with Spark and Cassandra 2018 VLDB 9.3994615e-05
2,118 DIFF: A Relational Interface for Large-Scale Data Explanation 2019 VLDB 9.2000813e-05
2,131 Quickstep: A Data Platform Based on the Scaling-Up Approach 2018 VLDB 9.1824992e-05
2,142 DUALSIM: Parallel Subgraph Enumeration in a Massive Graph on a Single Machine 2016 SIGMOD 9.1646667e-05
2,274 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 8.8988246e-05
2,334 DITA: Distributed In-Memory Trajectory Analytics 2018 SIGMOD 8.81355e-05
2,356 Mison: A Fast JSON Parser for Data Analytics 2017 VLDB 8.7744892e-05
2,384 Filter Before You Parse: Faster Analytics on Raw Data with Sparser 2018 VLDB 8.735059e-05
2,423 How to Architect a Query Compiler, Revisited 2018 SIGMOD 8.6700245e-05
2,530 AnalyticDB: Real-time OLAP Database System at Alibaba Cloud 2019 VLDB 8.5295823e-05
2,553 Big Data Analytics with Datalog Queries on Spark 2016 SIGMOD 8.4879606e-05
2,590 SQLShare: Results from a Multi-Year SQL-as-a-Service Experiment 2016 SIGMOD 8.448209e-05
2,794 AIDA - Abstraction for Advanced In-Database Analytics 2018 VLDB 8.1803827e-05
2,820 F1 Query: Declarative Querying at Scale 2018 VLDB 8.140112e-05
2,883 RaSQL: Greater Power and Performance for Big Data Analytics with Recursive-aggregate-SQL on Spark 2019 SIGMOD 8.0671021e-05
2,940 F1 Lightning: HTAP as a Service 2020 VLDB 8.002527e-05
2,960 OceanBase: A 707 Million tpmC Distributed Relational Database System 2022 VLDB 7.9703183e-05
2,961 Chi: A Scalable and Programmable Control Plane for Distributed Stream Processing Systems 2018 VLDB 7.9686889e-05
3,039 How to Win a Hot Dog Eating Contest: Distributed Incremental View Maintenance with Batch Updates 2016 SIGMOD 7.8829044e-05
3,046 Speculative Distributed CSV Data Parsing for Big Data Analytics 2019 SIGMOD 7.8803739e-05
3,079 Skipping-oriented Partitioning for Columnar Layouts 2017 VLDB 7.847438e-05
3,099 Helix: Accelerating Human-in-the-loop Machine Learning 2018 VLDB 7.8252921e-05
3,102 Pushing Data-Induced Predicates Through Joins in Big-Data Clusters 2020 VLDB 7.8221928e-05
Previous Page 1 / 5 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers