DBScholar

Back to papers

An Intermediate Representation for Optimizing Machine Learning Pipelines

Summary: Lara, a declarative DSL, provides an IR for end-to-end ML pipelines, unifying preprocessing, UDFs, control flow, and training. Monads enable cross-boundary pushdown/fusion; combinators encode domain operators to optimize data access, with up to 10× speedups on dense and sparse data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h428c19b163755452
Venue
VLDB
Year
2019
Pagerank
8.7248802e-05
Overall Rank
2,266 | 84.78%
DOI
10.14778/3342263.3342633
PDF
Download (CC BY-NC-ND 4.0)

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{kunft_vldb19,
        title = {{An Intermediate Representation for Optimizing Machine Learning Pipelines}},
        author = {Kunft, Andreas and Katsifodimos, Asterios and Schelter, Sebastian and Breß, Sebastian and Rabl, Tilmann and Markl, Volker},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {11},
        pages = {1553--1567},
        doi = {10.14778/3342263.3342633},
        url = {https://doi.org/10.14778/3342263.3342633},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
1,669 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9324573e-05
2,184 Extending Relational Query Processing with ML Inference 2020 CIDR 8.8953085e-05
2,661 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.1568473e-05
3,685 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.0972826e-05
3,971 Optimizing Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 6.8868815e-05
4,097 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.8064316e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6537801e-05
5,234 Babelfish: Efficient Execution of Polyglot Queries 2022 VLDB 6.2159562e-05
5,271 InferDB: In-Database Machine Learning Inference Using Indexes 2024 VLDB 6.1991633e-05
5,349 The NebulaStream Platform: Data and Application Management for the Internet of Things 2020 CIDR 6.1679453e-05
5,933 Doing More with Less: Characterizing Dataset Downsampling for AutoML 2021 VLDB 5.9393751e-05
7,388 Lachesis: Automatic Partitioning for UDF-Centric Analytics 2021 VLDB 5.5355684e-05
8,806 Towards A Polyglot Framework for Factorized ML 2021 VLDB 5.2721353e-05
9,166 HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics Queries 2021 SIGMOD 5.2144986e-05
10,005 cedar: Optimized and Unified Machine Learning Input Data Pipelines 2025 VLDB 5.0954911e-05
10,048 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0896901e-05
10,459 EncoderForge: Generating Efficient SQL for Encoders in Machine Learning Inference Pipelines 2026 SIGMOD 4.9769913e-05
10,664 InferF: Declarative Factorization of AI/ML Inferences over Joins 2026 SIGMOD 4.9769913e-05
10,907 stratum: A System Infrastructure for Massive Agent-Centric ML Workloads 2026 VLDB 4.9769913e-05
11,852 Redundancy Elimination in Distributed Matrix Computation 2022 SIGMOD 4.9769913e-05
12,021 TraNCE: Transforming Nested Collections Efficiently 2021 VLDB 4.9769913e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
21 Efficiently Compiling Efficient Query Plans for Modern Hardware 2011 VLDB 0.00056835296
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.0004552807
73 Including Group-By in Query Optimization 1994 VLDB 0.0003750677
416 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.00018650998
537 MLbase: A Distributed Machine-learning System 2013 CIDR 0.00016762761
731 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014400356
1,082 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012118261
1,193 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.00011582619
1,257 Towards Linear Algebra over Normalized Data 2017 VLDB 0.0001130959
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010067153
2,197 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8740089e-05
2,228 Spinning Fast Iterative Data Flows 2012 VLDB 8.7996087e-05
2,251 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 8.750953e-05
2,420 How to Architect a Query Compiler, Revisited 2018 SIGMOD 8.4883426e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3386546e-05
2,753 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0513412e-05
3,103 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6456038e-05
3,331 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.4138851e-05
3,518 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.236638e-05
5,308 Probabilistic Demand Forecasting at Scale 2017 VLDB 6.1844285e-05
8,381 Meta-Dataflows: Efficient Exploratory Dataflow Jobs 2018 SIGMOD 5.3411788e-05
9,829 BlockJoin: Efficient Matrix Partitioning Through Joins 2017 VLDB 5.1230568e-05
12,342 Emma in Action: Declarative Dataflows for Scalable Data Analysis 2016 SIGMOD 4.9769913e-05
Previous Page 1 / 1 Next

Semantically Similar Papers