DBScholar

Back to papers

LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems

Summary: Fine-grained lineage tracing and reuse in ML systems (LIMA) to break coarse, black-box limits. Multi-level traces, loop/function dedup, and cross-hierarchy reuse enable low-overhead provenance with versioning, compatible with task parallelism and operator fusion, delivering up to 12.4x speedups. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h0a8564e76e97bc86
Venue
SIGMOD
Year
2021
Pagerank
6.6537801e-05
Overall Rank
4,334 | 70.88%
DOI
10.1145/3448016.3452788

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{phani_sigmod21,
        title = {{LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems}},
        author = {Phani, Arnab and Rath, Benjamin and Boehm, Matthias},
        series = {{SIGMOD} '21},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3448016.3452788},
        url = {https://dl.acm.org/doi/10.1145/3448016.3452788},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 16 of 16 citing papers.

Rank Citing Paper Year Venue Pagerank
5,574 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0773771e-05
6,666 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7144587e-05
6,883 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines 2022 CIDR 5.6552024e-05
7,816 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5.446446e-05
7,843 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4406331e-05
8,083 Provenance-Enabled Explainable AI 2024 SIGMOD 5.3917406e-05
10,048 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0896901e-05
10,285 ElasticNotebook: Enabling Live Migration for Computational Notebooks 2024 VLDB 5.0431349e-05
10,732 CAPS: Cost-Aware ML Pipeline Selection 2026 VLDB 4.9769913e-05
10,907 stratum: A System Infrastructure for Massive Agent-Centric ML Workloads 2026 VLDB 4.9769913e-05
10,954 Morphing-based Compression for Data-centric ML Pipelines 2026 VLDB 4.9769913e-05
11,144 Unified Lineage System: Tracking Data Provenance at Scale 2025 SIGMOD 4.9769913e-05
11,185 Alsatian: Optimizing Model Search for Deep Transfer Learning 2025 SIGMOD 4.9769913e-05
11,292 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9769913e-05
11,432 ML-Asset Management: Curation, Discovery, and Utilization 2025 VLDB 4.9769913e-05
11,852 Redundancy Elimination in Distributed Matrix Computation 2022 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 54 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0010515896
17 Provenance Semirings 2007 PODS 0.00059813669
88 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035340164
129 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.00030395767
252 Database Cracking 2007 CIDR 0.00023101361
389 Why Not? 2009 SIGMOD 0.00019313101
416 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00018650998
417 MauveDB: Supporting Model-based User Views in Database Systems 2006 SIGMOD 0.00018625112
509 Goods: Organizing Google's Datasets 2016 SIGMOD 0.00017063491
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016083582
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00015096817
976 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.0001274453
1,040 The DataPath System: A Data-Centric Analytic Processing Engine for Large Data Warehouses 2010 SIGMOD 0.00012358804
1,082 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012118261
1,086 DataHub: Collaborative Data Science & Dataset Version Management at Scale 2015 CIDR 0.0001209772
1,132 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011893781
1,389 Provenance for Generalized Map and Reduce Workflows 2011 CIDR 0.00010818511
1,445 An Architecture for Compiling UDF-centric Workflows 2015 VLDB 0.00010628379
1,502 VisTrails: Visualization meets Data Management 2006 SIGMOD 0.00010454991
1,554 Update Exchange with Mappings and Provenance 2007 VLDB 0.00010265583
1,564 Titian: Data Provenance Support in Spark 2016 VLDB 0.00010219222
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010208225
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
1,692 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 9.8526476e-05
1,713 SMOKE: Fine-grained Lineage at Interactive Speed 2018 VLDB 9.8129981e-05
1,747 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7303647e-05
1,767 Putting Lipstick on Pig: Enabling Database-style Workflow Provenance 2012 VLDB 9.6910812e-05
1,847 Data Market Platforms: Trading Data Assets to Solve Data Problems 2020 VLDB 9.5091361e-05
1,918 Predictable Performance for Unpredictable Workloads 2009 VLDB 9.375192e-05
1,928 Elastic Machine Learning Algorithms in Amazon SageMaker 2020 SIGMOD 9.3563363e-05
2,086 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 9.0677901e-05
2,087 LINVIEW: Incremental View Maintenance for Complex Analytical Queries 2014 SIGMOD 9.0661817e-05
2,249 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7544468e-05
2,251 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 8.750953e-05
2,266 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.7248802e-05
2,835 Fine-Grained, Secure and Efficient Data Provenance on Blockchain Systems 2019 VLDB 7.9516187e-05
3,042 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 7.7179591e-05
3,103 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6456038e-05
3,114 Incremental View Maintenance with Triple Lock Factorization Benefits 2018 SIGMOD 7.6321464e-05
3,331 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.4138851e-05
3,685 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.0972826e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7576159e-05
4,300 The Missing Piece in Complex Analytics: Low Latency, Scalable Model Management and Serving with Velox 2015 CIDR 6.676283e-05
4,494 Juneau: Data Lake Management for Jupyter 2019 VLDB 6.5747499e-05
4,975 SPORES: Sum-Product Optimization via Relational Equality Saturation for Large Scale Linear Algebra 2020 VLDB 6.3289021e-05
5,794 Optimizing Machine Learning Workloads in Collaborative Environments 2020 SIGMOD 5.9906542e-05
5,809 Incrementally Maintaining Classification using an RDBMS 2011 VLDB 5.9848544e-05
6,048 Your notebook is not crumby enough, REPLace it 2020 CIDR 5.9024167e-05
6,193 Fine-Grained Lineage for Safer Notebook Interactions 2021 VLDB 5.8538103e-05
6,325 "Amnesia" - A Selection of Machine Learning Models That Can Forget User Data Very Fast 2020 CIDR 5.8121503e-05
Previous Page 1 / 2 Next

Semantically Similar Papers