Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark
Summary: Declarative Spark Structured Streaming; incrementalizes SQL/DataFrame queries, not user-built DAG. End-to-end real-time apps unifying streaming with batch analytics; code generation yields high performance; rollbacks and mixed execution in production. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Michael Armbrust (Databricks)
- 2. Tathagata Das (Databricks)
- 3. Joseph Torres (Databricks)
- 4. Burak Yavuz (Databricks)
- 5. Shixiong Zhu (Databricks)
- 6. Reynold Xin (Databricks)
- 7. Ali Ghodsi (Databricks)
- 8. Ion Stoica (Databricks)
- 9. Matei Zaharia (Stanford University)
BibTeX Citation
@inproceedings{armbrust_sigmod18,
title = {{Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark}},
author = {Armbrust, Michael and Das, Tathagata and Torres, Joseph and Yavuz, Burak and Zhu, Shixiong and Xin, Reynold and Ghodsi, Ali and Stoica, Ion and Zaharia, Matei},
series = {{SIGMOD} '18},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3183713.3190664},
url = {https://dl.acm.org/doi/10.1145/3183713.3190664},
year = {2018}
}
Incoming Citations (Sorted by Pagerank)
Showing 35 of 35 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 39 | Efficiently Updating Materialized Views | 1986 | SIGMOD | 0.00047309646 |
| 127 | The Design of the Borealis Stream Processing Engine | 2005 | CIDR | 0.00030738755 |
| 361 | The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing | 2015 | VLDB | 0.00020138717 |
| 391 | View Maintenance in a Warehousing Environment | 1995 | SIGMOD | 0.00019370653 |
| 455 | Differential dataflow | 2013 | CIDR | 0.00018133241 |
| 505 | TelegraphCQ: Continuous Dataflow Processing | 2003 | SIGMOD | 0.00017285498 |
| 710 | Trill: A High-Performance Incremental Query Processor for Diverse Analytics | 2015 | VLDB | 0.00014715033 |
| 1,240 | Consistency Analysis in Bloom: a CALM and Collected Approach | 2011 | CIDR | 0.00011536847 |
| 1,945 | Continuous Analytics Over Discontinuous Streams | 2010 | SIGMOD | 9.4348569e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 6,482 | AStream: Ad-hoc Shared Stream Processing | 2019 | SIGMOD |
| 2 | 2,196 | Spinning Fast Iterative Data Flows | 2012 | VLDB |
| 3 | 4,534 | Shared Arrangements: practical inter-query sharing for streaming dataflows | 2020 | VLDB |
| 4 | 6,502 | Watermarks in Stream Processing Systems: Semantics and Comparative Analysis of Apache Flink and Google Cloud Dataflow | 2021 | VLDB |
| 5 | 9,167 | Dynamic Speculative Optimizations for SQL Compilation in Apache Spark | 2020 | VLDB |
| 6 | 415 | SystemML: Declarative Machine Learning on Spark | 2016 | VLDB |
| 7 | 8,258 | Correctness in Stream Processing: Challenges and Opportunities | 2022 | CIDR |
| 8 | 3,411 | Scaling Spark in the Real World: Performance and Usability | 2015 | VLDB |
| 9 | 11,933 | Challenges and Experiences in Building an Efficient Apache Beam Runner For IBM Streams | 2018 | VLDB |
| 10 | 4,907 | One SQL to Rule Them All – an Efficient and Syntactically Idiomatic Approach to Management of Streams and Tables | 2019 | SIGMOD |