Back to papers
Big Data Analytics with Datalog Queries on Spark
Summary: BigDatalog enables concise declarative Datalog queries for large-scale analytics on Spark. It uses compilation and optimization to efficiently support recursion on Spark, with empirical comparisons against top Datalog systems showing Spark-based analytics viable.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 5257
- Venue
- SIGMOD
- Year
- 2016
- Pagerank
- 7.3847098e-05
- Overall Rank
- 3,207 | 77.72%
- DOI
-
10.1145/2882903.2915229
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 23 of 23 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 2,922 |
RaSQL: Greater Power and Performance for Big Data Analytics with Recursive-aggregate-SQL on Spark |
2019 |
SIGMOD |
7.897179e-05 |
| 3,984 |
All-in-One: Graph Processing in RDBMSs Revisited |
2017 |
SIGMOD |
6.5587512e-05 |
| 4,700 |
Tensors: An abstraction for general data processing |
2021 |
VLDB |
5.9810592e-05 |
| 4,925 |
Shared Arrangements: practical inter-query sharing for streaming dataflows |
2020 |
VLDB |
5.818597e-05 |
| 5,267 |
On the Optimization of Recursive Relational Queries: Application to Graph Queries |
2020 |
SIGMOD |
5.5930569e-05 |
| 5,716 |
Datalog Unchained |
2021 |
PODS |
5.3569788e-05 |
| 6,209 |
Automating Incremental and Asynchronous Evaluation for Recursive Aggregate Data Processing |
2020 |
SIGMOD |
5.1508692e-05 |
| 6,276 |
Scaling-Up In-Memory Datalog Processing: Observations and Techniques |
2019 |
VLDB |
5.1265189e-05 |
| 6,592 |
Complete Event Trend Detection in High-Rate Event Streams |
2017 |
SIGMOD |
4.9959015e-05 |
| 7,338 |
Optimizing Recursive Queries with Program Synthesis |
2022 |
SIGMOD |
4.7531793e-05 |
| 8,394 |
Optimizing Declarative Graph Queries at Large Scale |
2019 |
SIGMOD |
4.5233733e-05 |
| 8,884 |
Optimizing Parallel Recursive Datalog Evaluation on Multicore Machines |
2022 |
SIGMOD |
4.4243024e-05 |
| 9,000 |
Automatic Index Selection for Large-Scale Datalog Computation |
2019 |
VLDB |
4.4087102e-05 |
| 9,335 |
Parallel Query Processing: To Separate Communication from Computation |
2022 |
SIGMOD |
4.351469e-05 |
| 9,812 |
Datalog with First-Class Facts |
2025 |
VLDB |
4.2742278e-05 |
| 9,813 |
Optimizing Nested Recursive Queries |
2024 |
SIGMOD |
4.2742278e-05 |
| 10,296 |
FlowLog: Efficient and Extensible Datalog via Incrementality |
2026 |
VLDB |
4.1905499e-05 |
| 10,415 |
Dynamic Pruning for Recursive Joins |
2025 |
SIGMOD |
4.1905499e-05 |
| 11,056 |
Efficient Enumeration of Recursive Plans in Transformation-based Query Optimizers |
2024 |
VLDB |
4.1905499e-05 |
| 11,133 |
The Vadalog Parallel System: Distributed Reasoning with Datalog+/- |
2024 |
VLDB |
4.1905499e-05 |
| 11,157 |
Templating Shuffles |
2023 |
CIDR |
4.1905499e-05 |
| 11,343 |
Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data Applications |
2022 |
SIGMOD |
4.1905499e-05 |
| 11,652 |
Ariadne: Online Provenance for Big Graph Analytics |
2019 |
SIGMOD |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 21 of 21 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank |
Cited Paper |
Year |
Venue |
Pagerank |
| 4 |
Pregel: A System for Large-Scale Graph Processing |
2010 |
SIGMOD |
0.0019040811 |
| 39 |
Distributed GraphLab: A Framework for Machine Learning and Data Mining in the Cloud |
2012 |
VLDB |
0.00075263552 |
| 64 |
A Message Passing Framework for Logical Query Evaluation |
1986 |
SIGMOD |
0.00063295728 |
| 66 |
Spark SQL: Relational Data Processing in Spark |
2015 |
SIGMOD |
0.00061707583 |
| 70 |
Hive - A Warehousing Solution Over a Map-Reduce Framework |
2009 |
VLDB |
0.00059744625 |
| 774 |
Declarative Networking: Language, Execution and Optimization |
2006 |
SIGMOD |
0.00016775162 |
| 776 |
Building a High-Level Dataflow System on top of Map-Reduce: The Pig Experience |
2009 |
VLDB |
0.00016765694 |
| 1,323 |
Distributed SociaLite: A Datalog-Based Language for Large-Scale Graph Analysis |
2013 |
VLDB |
0.00012595443 |
| 1,377 |
Relational Transducers for Declarative Networking |
2011 |
PODS |
0.00012299784 |
| 2,087 |
A Framework for the Parallel Processing of Datalog Queries |
1990 |
SIGMOD |
9.5715836e-05 |
| 2,179 |
Spinning Fast Iterative Data Flows |
2012 |
VLDB |
9.3632007e-05 |
| 2,227 |
A New Paradigm For Parallel And Distributed Rule-Processing |
1990 |
SIGMOD |
9.2521507e-05 |
| 2,463 |
REX: Recursive, Delta-Based Data-Centric Computation |
2012 |
VLDB |
8.7438403e-05 |
| 3,075 |
Evita Raced: Metacompilation for Declarative Networks |
2008 |
VLDB |
7.6078165e-05 |
| 4,220 |
Monotonic Aggregation in Deductive Databases |
1992 |
PODS |
6.3414331e-05 |
| 4,225 |
Why A Single Parallelization Strategy Is Not Enough In Knowledge Bases |
1989 |
PODS |
6.3396067e-05 |
| 4,367 |
Distributed Processing Of Logic Programs |
1988 |
SIGMOD |
6.2423944e-05 |
| 4,694 |
Asynchronous and Fault-Tolerant Recursive Datalog Evaluation in Shared-Nothing Engines |
2015 |
VLDB |
5.985448e-05 |
| 5,003 |
Graph Queries in a Next-Generation Datalog System |
2013 |
VLDB |
5.7606385e-05 |
| 5,927 |
Parallelizing Datalog Programs by Generalized Pivoting |
1991 |
PODS |
5.2666195e-05 |
| 9,085 |
Collaborative Access Control in WebdamLog |
2015 |
SIGMOD |
4.395091e-05 |
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 1,481 |
Automating Large-Scale Data Quality Verification |
2018 |
VLDB |
0.00011715754 |
| 11,199 |
QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in Spark |
2023 |
SIGMOD |
4.1905499e-05 |
| 5,108 |
Debugging Big Data Analytics in Spark with BigDebug |
2017 |
SIGMOD |
5.6872497e-05 |
| 8,585 |
A Spark Optimizer for Adaptive, Fine-Grained Parameter Tuning |
2024 |
VLDB |
4.4856045e-05 |
| 8,884 |
Optimizing Parallel Recursive Datalog Evaluation on Multicore Machines |
2022 |
SIGMOD |
4.4243024e-05 |
| 2,922 |
RaSQL: Greater Power and Performance for Big Data Analytics with Recursive-aggregate-SQL on Spark |
2019 |
SIGMOD |
7.897179e-05 |
| 9,122 |
Dynamic Speculative Optimizations for SQL Compilation in Apache Spark |
2020 |
VLDB |
4.3877539e-05 |
| 7,796 |
Large-scale Complex Analytics on Semi-structured Datasets using AsterixDB and Spark |
2016 |
VLDB |
4.6438468e-05 |
| 11,580 |
RASQL: A Powerful Language and its System for Big Data Applications |
2020 |
SIGMOD |
4.1905499e-05 |
| 6,276 |
Scaling-Up In-Memory Datalog Processing: Observations and Techniques |
2019 |
VLDB |
5.1265189e-05 |