Back to papers
LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems
Summary: Fine-grained lineage tracing and reuse in ML systems (LIMA) to break coarse, black-box limits. Multi-level traces, loop/function dedup, and cross-hierarchy reuse enable low-overhead provenance with versioning, compatible with task parallelism and operator fusion, delivering up to 12.4x speedups.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 6070
- Venue
- SIGMOD
- Year
- 2021
- Pagerank
- 5.9259373e-05
- Overall Rank
- 4,779 | 66.79%
- DOI
-
10.1145/3448016.3452788
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 15 of 15 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 7,303 |
DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines |
2022 |
CIDR |
4.7632836e-05 |
| 7,481 |
Provenance-Enabled Explainable AI |
2024 |
SIGMOD |
4.7135369e-05 |
| 7,656 |
Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets |
2022 |
SIGMOD |
4.6826896e-05 |
| 7,702 |
ExDRa: Exploratory Data Science on Federated Raw Data |
2021 |
SIGMOD |
4.6689015e-05 |
| 8,096 |
Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications |
2023 |
SIGMOD |
4.583522e-05 |
| 8,515 |
UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads |
2022 |
VLDB |
4.4901466e-05 |
| 9,787 |
The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format |
2024 |
SIGMOD |
4.2799988e-05 |
| 9,911 |
ElasticNotebook: Enabling Live Migration for Computational Notebooks |
2024 |
VLDB |
4.2524493e-05 |
| 10,252 |
CAPS: Cost-Aware ML Pipeline Selection |
2026 |
VLDB |
4.1905499e-05 |
| 10,303 |
Morphing-based Compression for Data-centric ML Pipelines |
2026 |
VLDB |
4.1905499e-05 |
| 10,429 |
Unified Lineage System: Tracking Data Provenance at Scale |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,479 |
Alsatian: Optimizing Model Search for Deep Transfer Learning |
2025 |
SIGMOD |
4.1905499e-05 |
| 10,636 |
CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines |
2025 |
VLDB |
4.1905499e-05 |
| 10,846 |
ML-Asset Management: Curation, Discovery, and Utilization |
2025 |
VLDB |
4.1905499e-05 |
| 11,341 |
Redundancy Elimination in Distributed Matrix Computation |
2022 |
SIGMOD |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 54 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 9,225 |
Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning |
2021 |
VLDB |
4.3656789e-05 |
| 8,096 |
Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications |
2023 |
SIGMOD |
4.583522e-05 |
| 1,763 |
Efficient Lineage Tracking For Scientific Workflows |
2008 |
SIGMOD |
0.00010626896 |
| 2,355 |
An Intermediate Representation for Optimizing Machine Learning Pipelines |
2019 |
VLDB |
8.9727612e-05 |
| 6,290 |
Lightweight Inspection of Data Preprocessing in Native Machine Learning Pipelines |
2021 |
CIDR |
5.1220786e-05 |
| 8,166 |
Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science |
2021 |
VLDB |
4.567959e-05 |
| 3,920 |
On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML |
2018 |
VLDB |
6.6246708e-05 |
| 6,061 |
Optimizing Machine Learning Workloads in Collaborative Environments |
2020 |
SIGMOD |
5.2270653e-05 |
| 2,456 |
Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities |
2021 |
SIGMOD |
8.7649259e-05 |
| 6,464 |
Materialization and Reuse Optimizations for Production Data Science Pipelines |
2022 |
SIGMOD |
5.0471003e-05 |