Big Data Analytics with Datalog Queries on Spark
Summary: BigDatalog enables concise declarative Datalog queries for large-scale analytics on Spark. It uses compilation and optimization to efficiently support recursion on Spark, with empirical comparisons against top Datalog systems showing Spark-based analytics viable. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Alexander Shkapsky (University of California Los Angeles)
- 2. Mohan Yang (University of California Los Angeles)
- 3. Matteo Interlandi (University of California Los Angeles)
- 4. Hsuan Chiu (University of California Los Angeles)
- 5. Tyson Condie (University of California Los Angeles)
- 6. Carlo Zaniolo (University of California Los Angeles)
BibTeX Citation
@inproceedings{shkapsky_sigmod16,
title = {{Big Data Analytics with Datalog Queries on Spark}},
author = {Shkapsky, Alexander and Yang, Mohan and Interlandi, Matteo and Chiu, Hsuan and Condie, Tyson and Zaniolo, Carlo},
series = {{SIGMOD} '16},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/2882903.2915229},
url = {https://dl.acm.org/doi/10.1145/2882903.2915229},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 23 of 23 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 21 of 21 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 1,350 | Automating Large-Scale Data Quality Verification | 2018 | VLDB |
| 2 | 11,399 | QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark | 2023 | SIGMOD |
| 3 | 5,570 | Debugging Big Data Analytics in Spark with BigDebug | 2017 | SIGMOD |
| 4 | 8,615 | A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning | 2024 | VLDB |
| 5 | 8,858 | Optimizing Parallel Recursive Datalog Evaluation on Multicore Machines | 2022 | SIGMOD |
| 6 | 2,935 | RaSQL: Greater Power and Performance for Big Data Analytics with Recursive-aggregate-SQL on Spark | 2019 | SIGMOD |
| 7 | 9,167 | Dynamic Speculative Optimizations for SQL Compilation in Apache Spark | 2020 | VLDB |
| 8 | 8,129 | Large-scale Complex Analytics on Semi-structured Datasets using AsterixDB and Spark | 2016 | VLDB |
| 9 | 11,772 | RASQL: A Powerful Language and its System for Big Data Applications | 2020 | SIGMOD |
| 10 | 5,783 | Scaling-Up In-Memory Datalog Processing: Observations and Techniques | 2019 | VLDB |