DBScholar

Back to papers

Spark SQL: Relational Data Processing in Spark

Summary: Relational processing integrated into Spark via DataFrame API, unifying SQL queries with Spark's functional workflow. Catalyst, a Scala-based extensible optimizer, enables composable rules, code generation, JSON schema inference, and federation to databases. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h2541b5993b725084
Venue
SIGMOD
Year
2015
Pagerank
0.00055406774
Overall Rank
23 | 99.85%
DOI
10.1145/2723372.2742797

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{armbrust_sigmod15,
        title = {{Spark SQL: Relational Data Processing in Spark}},
        author = {Armbrust, Michael and Xin, Reynold S. and Lian, Cheng and Huai, Yin and Liu, Davies and Bradley, Joseph K. and Meng, Xiangrui and Kaftan, Tomer and Franklin, Michael J. and Ghodsi, Ali and Zaharia, Matei},
        series = {{SIGMOD} '15},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/2723372.2742797},
        url = {https://dl.acm.org/doi/10.1145/2723372.2742797},
        year = {2015}
}

Incoming Citations (Sorted by Pagerank)

Showing 50 of 219 citing papers.

Rank Citing Paper Year Venue Pagerank
9,157 HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics Queries 2021 SIGMOD 5.2169683e-05
9,282 Making Data Engineering Declarative 2023 CIDR 5.2034134e-05
9,321 Dynamic Speculative Optimizations for SQL Compilation in Apache Spark 2020 VLDB 5.1943472e-05
9,348 Evaluating Query Languages and Systems for High-Energy Physics Data 2022 VLDB 5.1899225e-05
9,472 Query Compilation Without Regrets 2024 SIGMOD 5.1711207e-05
9,549 GIO: Generating Efficient Matrix and Frame Readers for Custom Data Formats by Example 2023 SIGMOD 5.1599622e-05
9,600 Language-Agnostic Integrated Queries in a Managed Polyglot Runtime 2021 VLDB 5.1548991e-05
9,615 In-Browser Interactive SQL Analytics with Afterburner 2017 SIGMOD 5.1511784e-05
9,657 SODA: A Set of Fast Oblivious Algorithms in Distributed Secure Data Analytics 2023 VLDB 5.1453267e-05
9,659 Parallel Query Processing: To Separate Communication from Computation 2022 SIGMOD 5.1453267e-05
9,740 TreeToaster: Towards an IVM-Optimized Compiler 2021 SIGMOD 5.1349531e-05
9,748 PlinyCompute: A Platform for High-Performance, Distributed, Data-Intensive Tool Development 2018 SIGMOD 5.1349531e-05
9,832 [Demo] Low-latency Spark Queries on Updatable Data 2019 SIGMOD 5.1245795e-05
9,921 Polyglot Data Management: State of the Art & Open Challenges 2022 VLDB 5.1103839e-05
9,988 Introduction to Spark 2.0 for Database Researchers 2016 SIGMOD 5.0998353e-05
9,995 Enzyme Demo: Incremental View Maintenance for Data Engineering 2026 VLDB 5.0979044e-05
10,000 cedar: Optimized and Unified Machine Learning Input Data Pipelines 2025 VLDB 5.0979044e-05
10,047 Tuplex: Robust, Efficient Analytics When Python Rules 2019 VLDB 5.091453e-05
10,187 Saving Money for Analytical Workloads in the Cloud 2024 VLDB 5.0651993e-05
10,227 TreeCat: Standalone Catalog Engine for Large Data Systems 2025 VLDB 5.0571508e-05
10,280 Chukonu: A Fully-Featured High-Performance Big Data Framework that Integrates a Native Compute Engine into Spark 2022 VLDB 5.0455234e-05
10,309 Intra-Query Runtime Elasticity for Cloud-Native Data Analysis 2025 SIGMOD 5.0386264e-05
10,468 HotHash: Hotness-Aware Consistent Hashing for Cloud Databases 2026 SIGMOD 4.9793485e-05
10,622 Beyond Relational: Semantic-Aware Multi-Modal Analytics with LLM-Native Query Optimization 2026 SIGMOD 4.9793485e-05
10,727 SIDLE: Tree-structure Aware Indexes for CXL-based Heterogeneous Memory 2026 VLDB 4.9793485e-05
10,802 BaCon: Efficient Batch Processing of Counting Queries 2026 VLDB 4.9793485e-05
10,887 MFTune: An Efficient Multi-fidelity Framework for Spark SQL Configuration Tuning 2026 VLDB 4.9793485e-05
10,889 Rhyme Native: Efficient Code Generation for Structured and Semi-Structured Workloads 2026 VLDB 4.9793485e-05
10,901 Overlay Bitmap Encoding for Efficient Consumption of Apache Parquet Files 2026 VLDB 4.9793485e-05
10,911 Hermes at Scale: Powering Distributed Queries with a Unified Memory Fabric 2026 VLDB 4.9793485e-05
10,914 FastCompose: Eliminating Compilation Cold Starts in Query Execution with Composition 2026 VLDB 4.9793485e-05
10,915 Reaching the Pinnacle of TPC-DS: Co-design of Architecture, Executor, and Storage in TDSQL 2026 VLDB 4.9793485e-05
10,927 A Decade of Apache Spark Structured Streaming: How We Evolved The Architecture To Meet Real-World Needs 2026 VLDB 4.9793485e-05
10,935 OmniTable: A Unified Wide-Table System for Petabyte-Scale LLM Data Curation and Exploration 2026 VLDB 4.9793485e-05
10,944 From Presto to Prestissimo: A Velox-Powered Modernization Journey 2026 VLDB 4.9793485e-05
11,037 Future-Proof Data Systems 2026 VLDB 4.9793485e-05
11,113 Optimizing Block Skipping for High-Dimensional Data with Learned Adaptive Curve 2025 SIGMOD 4.9793485e-05
11,128 Managed Resource Scaling in Amazon EMR 2025 SIGMOD 4.9793485e-05
11,130 OpenMLDB: A Real-Time Relational Data Feature Computation System for Online ML 2025 SIGMOD 4.9793485e-05
11,242 QOVIS: Understanding and Diagnosing Query Optimizer via a Visualization-assisted Approach 2025 VLDB 4.9793485e-05
11,258 Accio: Bolt-on Query Federation 2025 VLDB 4.9793485e-05
11,306 ArrayMorph: Optimizing Hyperslab Queries on the Cloud for Machine Learning Pipelines 2025 VLDB 4.9793485e-05
11,344 Towards Designing Future-Proof Data Processing Systems 2025 VLDB 4.9793485e-05
11,375 GRewriter: Practical Query Rewriting with Automatic Rule Set Expansion in GaussDB 2025 VLDB 4.9793485e-05
11,433 CloudGlide: Deconstructing the Landscape of Cloud-Based Analytics 2025 VLDB 4.9793485e-05
11,443 LEAP: A Low-cost Spark SQL Query Optimizer using Pairwise Comparison 2025 VLDB 4.9793485e-05
11,456 IcedTea: Efficient and Responsive Time-Travel Debugging in Dataflow Systems 2025 VLDB 4.9793485e-05
11,592 Agile-Ant: Self-managing Distributed Cache Management for Cost Optimization of Big Data Applications 2024 VLDB 4.9793485e-05
11,611 Large-Scale Metric Computation in Online Controlled Experiment Platform 2024 VLDB 4.9793485e-05
11,670 Raising the Level of Abstraction for Time-State Analytics With the Timeline Framework 2023 CIDR 4.9793485e-05
Previous Page 4 / 5 Next

Outgoing Citations (Sorted by Pagerank)

Showing 15 of 15 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers