Ursa: A Lakehouse-Native Data Streaming Engine for Kafka
Summary: Ursa: a leaderless, Kafka‑compatible, cloud‑native engine that writes directly to lakehouse tables on object storage, eliminating leader-based replication, broker disk storage, and external connectors. Matches Kafka throughput with exactly-once semantics and up to 10x lower infra cost. (summarized by gpt-5-mini on Feb 09 2026)
Incoming Non-self Citations Over Time
No non-self incoming citations found for this paper in this database.
Authors
- 1. Matteo Merli (StreamNative, Inc.)
- 2. Sijie Guo (StreamNative, Inc.)
- 3. Penghui Li (StreamNative, Inc.)
- 4. Hang Chen (StreamNative, Inc.)
- 5. Neng Lu (StreamNative, Inc.)
BibTeX Citation
@article{merli_vldb25,
title = {{Ursa: A Lakehouse-Native Data Streaming Engine for Kafka}},
author = {Merli, Matteo and Guo, Sijie and Li, Penghui and Chen, Hang and Lu, Neng},
journal = {PVLDB},
series = {{VLDB} '25},
volume = {18},
number = {12},
pages = {5184--5196},
doi = {10.14778/3750601.3750636},
url = {https://doi.org/10.14778/3750601.3750636},
year = {2025}
}
Incoming Citations (Sorted by Pagerank)
Showing 0 of 0 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 2 of 2 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 520 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores | 2020 | VLDB | 0.00017136828 |
| 1,138 | Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics | 2021 | CIDR | 0.00012023643 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 520 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores | 2020 | VLDB |
| 2 | 4,457 | Analyzing and Comparing Lakehouse Storage Systems | 2023 | CIDR |
| 3 | 12,255 | Resa: Realtime Elastic Streaming Analytics in the Cloud | 2013 | SIGMOD |
| 4 | 3,226 | Taurus Database: How to be Fast, Available, and Frugal in the Cloud | 2020 | SIGMOD |
| 5 | 11,683 | Real-time Data Infrastructure at Uber | 2021 | SIGMOD |
| 6 | 6,004 | Data Ingestion for the Connected World | 2017 | CIDR |
| 7 | 1,895 | Samza: Stateful Scalable Stream Processing at LinkedIn | 2017 | VLDB |
| 8 | 4,389 | Consistency and Completeness: Rethinking Distributed Stream Processing in Apache Kafka | 2021 | SIGMOD |
| 9 | 11,558 | KafkaDirect: Zero-copy Data Access for Apache Kafka over RDMA Networks | 2022 | SIGMOD |
| 10 | 9,562 | Kora: A Cloud-Native Event Streaming Platform For Kafka | 2023 | VLDB |