Database Paper Browser

Back to papers

LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems

Summary: Fine-grained lineage tracing and reuse in ML systems (LIMA) to break coarse, black-box limits. Multi-level traces, loop/function dedup, and cross-hierarchy reuse enable low-overhead provenance with versioning, compatible with task parallelism and operator fusion, delivering up to 12.4x speedups. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
6070
Venue
SIGMOD
Year
2021
Pagerank
5.9259373e-05
Overall Rank
4,779 | 66.79%
DOI
10.1145/3448016.3452788

Incoming Non-self Citations Over Time

Authors

Incoming Citations (Sorted by Pagerank)

Showing 15 of 15 citing papers.

Rank Citing Paper Year Venue Pagerank
7,303 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines 2022 CIDR 4.7632836e-05
7,481 Provenance-Enabled Explainable AI 2024 SIGMOD 4.7135369e-05
7,656 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 4.6826896e-05
7,702 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 4.6689015e-05
8,096 Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications 2023 SIGMOD 4.583522e-05
8,515 UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads 2022 VLDB 4.4901466e-05
9,787 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 4.2799988e-05
9,911 ElasticNotebook: Enabling Live Migration for Computational Notebooks 2024 VLDB 4.2524493e-05
10,252 CAPS: Cost-Aware ML Pipeline Selection 2026 VLDB 4.1905499e-05
10,303 Morphing-based Compression for Data-centric ML Pipelines 2026 VLDB 4.1905499e-05
10,429 Unified Lineage System: Tracking Data Provenance at Scale 2025 SIGMOD 4.1905499e-05
10,479 Alsatian: Optimizing Model Search for Deep Transfer Learning 2025 SIGMOD 4.1905499e-05
10,636 CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines 2025 VLDB 4.1905499e-05
10,846 ML-Asset Management: Curation, Discovery, and Utilization 2025 VLDB 4.1905499e-05
11,341 Redundancy Elimination in Distributed Matrix Computation 2022 SIGMOD 4.1905499e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 54 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
3 Pig Latin: A Not-So-Foreign Language for Data Processing 2008 SIGMOD 0.0024217964
31 Provenance Semirings 2007 PODS 0.00078516827
160 Automated Selection of Materialized Views and Indexes for SQL Databases 2000 VLDB 0.00040053897
179 Efficient and Extensible Algorithms for Multi Query Optimization 2000 SIGMOD 0.00037637319
407 Database Cracking 2007 CIDR 0.00023941779
468 MauveDB: Supporting Model-based User Views in Database Systems 2006 SIGMOD 0.00022407392
487 Why Not? 2009 SIGMOD 0.00022030123
557 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00020186115
609 Goods: Organizing Google's Datasets 2016 SIGMOD 0.00019223217
668 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00018428925
758 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00017053915
917 Democratizing Data Science through Interactive Curation of ML Pipelines 2019 SIGMOD 0.00015324193
1,278 DataHub: Collaborative Data Science & Dataset Version Management at Scale 2015 CIDR 0.00012851949
1,296 The DataPath System: A Data-Centric Analytic Processing Engine for Large Data Warehouses 2010 SIGMOD 0.00012742585
1,407 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012163413
1,412 VisTrails: Visualization meets Data Management 2006 SIGMOD 0.00012112718
1,441 Provenance for Generalized Map and Reduce Workflows 2011 CIDR 0.00011950114
1,475 Efficient Exploitation of Similar Subexpressions for Query Processing 2007 SIGMOD 0.00011765071
1,666 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010955907
1,868 Update Exchange with Mappings and Provenance 2007 VLDB 0.0001026365
1,875 An Architecture for Compiling UDF-centric Workflows 2015 VLDB 0.00010243959
1,921 Selecting Subexpressions to Materialize at Datacenter Scale 2018 VLDB 0.00010085899
2,030 Titian: Data Provenance Support in Spark 2016 VLDB 9.734332e-05
2,031 Putting Lipstick on Pig: Enabling Database-style Workflow Provenance 2012 VLDB 9.7341007e-05
2,122 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.4905306e-05
2,157 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 9.4153917e-05
2,164 Elastic Machine Learning Algorithms in Amazon SageMaker 2020 SIGMOD 9.3953268e-05
2,260 LINVIEW: Incremental View Maintenance for Complex Analytical Queries 2014 SIGMOD 9.1771829e-05
2,286 SMOKE: Fine-grained Lineage at Interactive Speed 2018 VLDB 9.102574e-05
2,355 An Intermediate Representation for Optimizing Machine Learning Pipelines 2019 VLDB 8.9727612e-05
2,366 Data Market Platforms: Trading Data Assets to Solve Data Problems 2020 VLDB 8.9521259e-05
2,372 Predictable Performance for Unpredictable Workloads 2009 VLDB 8.940791e-05
2,672 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.3334428e-05
2,695 An Architecture for Recycling Intermediates in a Column-store 2009 SIGMOD 8.2827669e-05
2,871 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 7.9800983e-05
2,904 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 7.9403097e-05
3,158 Fine-Grained, Secure and Efficient Data Provenance on Blockchain Systems 2019 VLDB 7.466958e-05
3,866 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 6.6795497e-05
3,920 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 6.6246708e-05
4,200 Incremental View Maintenance with Triple Lock Factorization Benefits 2018 SIGMOD 6.3618329e-05
4,508 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 6.1261819e-05
4,575 The Missing Piece in Complex Analytics: Low Latency, Scalable Model Management and Serving with Velox 2015 CIDR 6.0662145e-05
4,589 Juneau: Data Lake Management for Jupyter 2019 VLDB 6.0565861e-05
4,596 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.0540725e-05
5,444 "Amnesia" - A Selection of Machine Learning Models That Can Forget User Data Very Fast 2020 CIDR 5.4998709e-05
5,497 SPORES: Sum-Product Optimization via Relational Equality Saturation for Large Scale Linear Algebra 2020 VLDB 5.4741034e-05
5,884 Incrementally Maintaining Classification using an RDBMS 2011 VLDB 5.2864334e-05
6,061 Optimizing Machine Learning Workloads in Collaborative Environments 2020 SIGMOD 5.2270653e-05
6,290 Lightweight Inspection of Data Preprocessing in Native Machine Learning Pipelines 2021 CIDR 5.1220786e-05
6,295 Your notebook is not crumby enough, REPLace it 2020 CIDR 5.1200006e-05
Previous Page 1 / 2 Next

Semantically Similar Papers