SparkR: Scaling R Programs with Spark
Summary: R's single-threaded nature and memory limits curb interactive analytics. SparkR adds an R frontend to Apache Spark, enabling scalable, distributed data analysis from the R shell via Spark's DataFrame API and distributed computation. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Shivaram Venkataraman (University of California Berkeley)
- 2. Zongheng Yang (University of California Berkeley)
- 3. Davies Liu (Databricks)
- 4. Eric Liang (Databricks)
- 5. Hossein Falaki (Databricks)
- 6. Xiangrui Meng (Databricks)
- 7. Reynold Xin (Databricks)
- 8. Ali Ghodsi (Databricks)
- 9. Michael Franklin (University of California Berkeley)
- 10. Ion Stoica (Databricks; University of California Berkeley)
- 11. Matei Zaharia (Databricks; Massachusetts Institute of Technology)
BibTeX Citation
@inproceedings{venkataraman_sigmod16,
title = {{SparkR: Scaling R Programs with Spark}},
author = {Venkataraman, Shivaram and Yang, Zongheng and Liu, Davies and Liang, Eric and Falaki, Hossein and Meng, Xiangrui and Xin, Reynold and Ghodsi, Ali and Franklin, Michael and Stoica, Ion and Zaharia, Matei},
series = {{SIGMOD} '16},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2882903.2903740},
url = {https://dl.acm.org/doi/10.1145/2882903.2903740},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 6 of 6 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,138 | Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics | 2021 | CIDR | 0.00012023643 |
| 1,250 | Data Management in Machine Learning: Challenges, Techniques, and Systems | 2017 | SIGMOD | 0.00011485301 |
| 2,179 | Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra | 2019 | SIGMOD | 9.0146333e-05 |
| 5,301 | Babelfish: Efficient Execution of Polyglot Queries | 2022 | VLDB | 6.2750553e-05 |
| 9,803 | Introduction to Spark 2.0 for Database Researchers | 2016 | SIGMOD | 5.21582e-05 |
| 10,072 | Query Compilation Without Regrets | 2024 | SIGMOD | 5.1624689e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 7 of 7 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 330 | Impala: A Modern, Open-Source SQL Engine for Hadoop | 2015 | CIDR | 0.0002104801 |
| 425 | Shark: SQL and Rich Analytics at Scale | 2013 | SIGMOD | 0.00018704491 |
| 1,333 | Ricardo: Integrating R and Hadoop | 2010 | SIGMOD | 0.0001112858 |
| 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB | 7.436229e-05 |
| 5,584 | Large-scale Predictive Analytics in Vertica: Fast Data Transfer, Distributed Model Creation, and In-database Prediction | 2015 | SIGMOD | 6.1604087e-05 |
| 7,197 | Stat! - An Interactive Analytics Environment for Big Data | 2013 | SIGMOD | 5.6763539e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,190 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark | 2018 | SIGMOD |
| 2 | 4,208 | Clash of the Titans: MapReduce vs. Spark for Large Scale Data Analytics | 2015 | VLDB |
| 3 | 2,594 | Big Data Analytics with Datalog Queries on Spark | 2016 | SIGMOD |
| 4 | 11,399 | QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark | 2023 | SIGMOD |
| 5 | 2,935 | RaSQL: Greater Power and Performance for Big Data Analytics with Recursive-aggregate-SQL on Spark | 2019 | SIGMOD |
| 6 | 6,850 | Bridging the Gap Between HPC and Big Data Frameworks | 2017 | VLDB |
| 7 | 415 | SystemML: Declarative Machine Learning on Spark | 2016 | VLDB |
| 8 | 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD |
| 9 | 12,060 | dmapply: A functional primitive to express distributed machine learning algorithms in R | 2016 | VLDB |
| 10 | 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB |