Spark SQL: Relational Data Processing in Spark
Summary: Relational processing integrated into Spark via DataFrame API, unifying SQL queries with Spark's functional workflow. Catalyst, a Scala-based extensible optimizer, enables composable rules, code generation, JSON schema inference, and federation to databases. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael Armbrust (Databricks)
- 2. Reynold S. Xin (Databricks)
- 3. Cheng Lian (Databricks)
- 4. Yin Huai (Databricks)
- 5. Davies Liu (Databricks)
- 6. Joseph K. Bradley (Databricks)
- 7. Xiangrui Meng (Databricks)
- 8. Tomer Kaftan (University of California Berkeley)
- 9. Michael J. Franklin (Databricks; University of California Berkeley)
- 10. Ali Ghodsi (Databricks)
- 11. Matei Zaharia (Databricks; Massachusetts Institute of Technology)
BibTeX Citation
@inproceedings{armbrust_sigmod15,
title = {{Spark SQL: Relational Data Processing in Spark}},
author = {Armbrust, Michael and Xin, Reynold S. and Lian, Cheng and Huai, Yin and Liu, Davies and Bradley, Joseph K. and Meng, Xiangrui and Kaftan, Tomer and Franklin, Michael J. and Ghodsi, Ali and Zaharia, Matei},
series = {{SIGMOD} '15},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2723372.2742797},
url = {https://dl.acm.org/doi/10.1145/2723372.2742797},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 7 of 207 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 11,885 | Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology | 2019 | VLDB | 5.093636e-05 |
| 11,889 | An Experimental Evaluation of Garbage Collectors on Big Data Applications | 2019 | VLDB | 5.093636e-05 |
| 11,955 | An Authorization Model for Multi Provider Queries | 2018 | VLDB | 5.093636e-05 |
| 11,959 | Effective Temporal Dependence Discovery in Time Series Data | 2018 | VLDB | 5.093636e-05 |
| 11,979 | Query Processing Techniques for Big Spatial-Keyword Data | 2017 | SIGMOD | 5.093636e-05 |
| 12,146 | Tutorial: SQL-on-Hadoop Systems | 2015 | VLDB | 5.093636e-05 |
| 13,301 | Blink Twice - Automatic Workload Pinning and Regression Detection for Versionless Apache Spark using Retries | 2025 | SIGMOD | - |
Outgoing Citations (Sorted by Pagerank)
Showing 15 of 15 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB |
| 2 | 425 | Shark: SQL and Rich Analytics at Scale | 2013 | SIGMOD |
| 3 | 415 | SystemML: Declarative Machine Learning on Spark | 2016 | VLDB |
| 4 | 2,594 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD |
| 5 | 11,772 | RASQL: A Powerful Language and its System for Big Data Applications | 2020 | SIGMOD |
| 6 | 1,190 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark | 2018 | SIGMOD |
| 7 | 6,865 | SparkR: Scaling R Programs with Spark | 2016 | SIGMOD |
| 8 | 9,803 | Introduction to Spark 2.0 for Database Researchers | 2016 | SIGMOD |
| 9 | 9,167 | Dynamic Speculative Optimizations for SQL Compilation in Apache Spark | 2020 | VLDB |
| 10 | 2,935 | RaSQL: Greater Power and Performance for Big Data Analytics with Recursive-aggregate-SQL on Spark | 2019 | SIGMOD |