Titian: Data Provenance Support in Spark
Summary: Titian embeds fine-grained data provenance in Apache Spark, enabling interactive backward tracing from erroneous or outlier results to root-cause inputs. Its optimized lineage capture is orders of magnitude faster than alternatives, with typically ≤30% runtime overhead. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Matteo Interlandi (University of California Los Angeles)
- 2. Kshitij Shah (University of California Los Angeles)
- 3. Sai Deep Tetali (University of California Los Angeles)
- 4. Muhammad Ali Gulzar (University of California Los Angeles)
- 5. Seunghyun Yoo (University of California Los Angeles)
- 6. Miryung Kim (University of California Los Angeles)
- 7. Todd Millstein (University of California Los Angeles)
- 8. Tyson Condie (University of California Los Angeles)
BibTeX Citation
@article{interlandi_vldb16,
title = {{Titian: Data Provenance Support in Spark}},
author = {Interlandi, Matteo and Shah, Kshitij and Tetali, Sai Deep and Gulzar, Muhammad Ali and Yoo, Seunghyun and Kim, Miryung and Millstein, Todd and Condie, Tyson},
journal = {PVLDB},
series = {{VLDB} '16},
volume = {9},
number = {3},
pages = {216--227},
doi = {10.14778/2850583.2850593},
url = {https://doi.org/10.14778/2850583.2850593},
year = {2016}
}
Incoming Citations (Sorted by Pagerank)
Showing 21 of 21 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 11 of 11 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 6 | Pig Latin: A Not-So-Foreign Language for Data Processing | 2008 | SIGMOD | 0.0010686205 |
| 24 | Spark SQL: Relational Data Processing in Spark | 2015 | SIGMOD | 0.00054865648 |
| 32 | Hive - A Warehousing Solution Over a Map-Reduce Framework | 2009 | VLDB | 0.00050111008 |
| 585 | Lineage Tracing for General Data Warehouse Transformations | 2001 | VLDB | 0.00016139897 |
| 1,366 | Provenance for Generalized Map and Reduce Workflows | 2011 | CIDR | 0.00011016972 |
| 1,786 | Putting Lipstick on Pig: Enabling Database-style Workflow Provenance | 2012 | VLDB | 9.7690421e-05 |
| 1,871 | Efficient Lineage Tracking For Scientific Workflows | 2008 | SIGMOD | 9.5783449e-05 |
| 1,912 | Querying Data Provenance | 2010 | SIGMOD | 9.4946756e-05 |
| 3,323 | Efficient Querying and Maintenance of Network Provenance at Internet-Scale | 2010 | SIGMOD | 7.5211754e-05 |
| 4,693 | Inspector Gadget: A Framework for Custom Monitoring and Debugging of Distributed Dataflows | 2011 | VLDB | 6.560102e-05 |
| 5,574 | Lineage-driven Fault Injection | 2015 | SIGMOD | 6.1681803e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 11,603 | SparkCAD: Caching Anomalies Detector for Spark Applications | 2022 | VLDB |
| 2 | 11,949 | DfAnalyzer: Runtime Dataflow Analysis of Scientific Applications using Provenance | 2018 | VLDB |
| 3 | 9,843 | Debugging Missing Answers for Spark Queries over Nested Data with Breadcrumb | 2021 | VLDB |
| 4 | 11,594 | DPDS: Assisting Data Science with Data Provenance | 2022 | VLDB |
| 5 | 4,756 | Improving Reproducibility of Data Science Pipelines through Transparent Provenance Capture | 2020 | VLDB |
| 6 | 8,050 | Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science | 2021 | VLDB |
| 7 | 11,860 | Ursprung: Provenance for Large-Scale Analytics Environments | 2019 | SIGMOD |
| 8 | 11,842 | Ariadne: Online Provenance for Big Graph Analytics | 2019 | SIGMOD |
| 9 | 5,570 | Debugging Big Data Analytics in Spark with BigDebug | 2017 | SIGMOD |
| 10 | 11,857 | Capturing and Querying Structural Provenance in Spark with Pebble | 2019 | SIGMOD |