DBScholar

Back to papers

LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems

Summary: Fine-grained lineage tracing and reuse in ML systems (LIMA) to break coarse, black-box limits. Multi-level traces, loop/function dedup, and cross-hierarchy reuse enable low-overhead provenance with versioning, compatible with task parallelism and operator fusion, delivering up to 12.4x speedups. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h0a8564e76e97bc86
Venue
SIGMOD
Year
2021
Pagerank
6.6569314e-05
Overall Rank
4,334 | 70.87%
DOI
10.1145/3448016.3452788

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{phani_sigmod21,
        title = {{LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems}},
        author = {Phani, Arnab and Rath, Benjamin and Boehm, Matthias},
        series = {{SIGMOD} '21},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3448016.3452788},
        url = {https://dl.acm.org/doi/10.1145/3448016.3452788},
        year = {2021}
}

Incoming Citations (Sorted by Pagerank)

Showing 16 of 16 citing papers.

Rank Citing Paper Year Venue Pagerank
5,572 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 6.0802555e-05
6,662 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 5.7171651e-05
6,878 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines 2022 CIDR 5.657878e-05
7,809 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5.4490255e-05
7,839 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4432099e-05
8,076 Provenance-Enabled Explainable AI 2024 SIGMOD 5.3942942e-05
10,043 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0921006e-05
10,279 ElasticNotebook: Enabling Live Migration for Computational Notebooks 2024 VLDB 5.0455234e-05
10,722 CAPS: Cost-Aware ML Pipeline Selection 2026 VLDB 4.9793485e-05
10,898 stratum: A System Infrastructure for Massive Agent-Centric ML Workloads 2026 VLDB 4.9793485e-05
10,945 Morphing-based Compression for Data-centric ML Pipelines 2026 VLDB 4.9793485e-05
11,135 Unified Lineage System: Tracking Data Provenance at Scale 2025 SIGMOD 4.9793485e-05
11,176 Alsatian: Optimizing Model Search for Deep Transfer Learning 2025 SIGMOD 4.9793485e-05
11,284 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.9793485e-05
11,426 ML-Asset Management: Curation, Discovery, and Utilization 2025 VLDB 4.9793485e-05
11,846 Redundancy Elimination in Distributed Matrix Computation 2022 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 54 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
6 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.001052036
17 Provenance Semirings 2007 PODS 0.00059752575
88 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00035351639
129 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.0003040756
253 Database Cracking 2007 CIDR 0.00023042111
391 Why Not? 2009 SIGMOD 0.00019238698
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001865959
417 MauveDB: Supporting Model-based User Views in Database Systems 2006 SIGMOD 0.00018631353
509 Goods: Organizing Google's Datasets 2016 SIGMOD 0.00017071087
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016086569
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.0001510357
975 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.00012750518
1,040 The DataPath System: A Data-Centric Analytic Processing Engine for Large Data Warehouses 2010 SIGMOD 0.00012364063
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,086 DataHub: Collaborative Data Science & Dataset Version Management at Scale 2015 CIDR 0.0001210322
1,131 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.0001189909
1,388 Provenance for Generalized Map and Reduce Workflows 2011 CIDR 0.00010821908
1,444 An Architecture for Compiling UDF-centric Workflows 2015 VLDB 0.0001063181
1,502 VisTrails: Visualization meets Data Management 2006 SIGMOD 0.00010456416
1,555 Update Exchange with Mappings and Provenance 2007 VLDB 0.00010269848
1,564 Titian: Data Provenance Support in Spark 2016 VLDB 0.00010222394
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.0001021302
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
1,691 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 9.8570722e-05
1,745 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 9.7343818e-05
1,746 SMOKE: Fine-grained Lineage at Interactive Speed 2018 VLDB 9.7341914e-05
1,765 Putting Lipstick on Pig: Enabling Database-style Workflow Provenance 2012 VLDB 9.6955585e-05
1,845 Data Market Platforms: Trading Data Assets to Solve Data Problems 2020 VLDB 9.5136397e-05
1,918 Predictable Performance for Unpredictable Workloads 2009 VLDB 9.3789552e-05
1,927 Elastic Machine Learning Algorithms in Amazon SageMaker 2020 SIGMOD 9.3607648e-05
2,084 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 9.0716512e-05
2,087 LINVIEW: Incremental View Maintenance for Complex Analytical Queries 2014 SIGMOD 9.068879e-05
2,247 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7585767e-05
2,249 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 8.7549752e-05
2,264 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.7289107e-05
2,835 Fine-Grained, Secure and Efficient Data Provenance on Blockchain Systems 2019 VLDB 7.9553846e-05
3,040 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 7.7216143e-05
3,101 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.649219e-05
3,112 Incremental View Maintenance with Triple Lock Factorization Benefits 2018 SIGMOD 7.6357579e-05
3,330 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.4173693e-05
3,683 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.1006425e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7608137e-05
4,302 The Missing Piece in Complex Analytics: Low Latency, Scalable Model Management and Serving with Velox 2015 CIDR 6.6778695e-05
4,491 Juneau: Data Lake Management for Jupyter 2019 VLDB 6.5778095e-05
4,974 SPORES: Sum-Product Optimization via Relational Equality Saturation for Large Scale Linear Algebra 2020 VLDB 6.3318399e-05
5,792 Optimizing Machine Learning Workloads in Collaborative Environments 2020 SIGMOD 5.99349e-05
5,813 Incrementally Maintaining Classification using an RDBMS 2011 VLDB 5.9861029e-05
6,046 Your notebook is not crumby enough, REPLace it 2020 CIDR 5.9052121e-05
6,190 Fine-Grained Lineage for Safer Notebook Interactions 2021 VLDB 5.8565828e-05
6,321 "Amnesia" - A Selection of Machine Learning Models That Can Forget User Data Very Fast 2020 CIDR 5.814903e-05
Previous Page 1 / 2 Next

Semantically Similar Papers