DBScholar

Back to papers

An Intermediate Representation for Optimizing Machine Learning Pipelines

Summary: Lara, a declarative DSL, provides an IR for end-to-end ML pipelines, unifying preprocessing, UDFs, control flow, and training. Monads enable cross-boundary pushdown/fusion; combinators encode domain operators to optimize data access, with up to 10× speedups on dense and sparse data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h428c19b163755452
Venue
VLDB
Year
2019
Pagerank
8.7289107e-05
Overall Rank
2,264 | 84.78%
DOI
10.14778/3342263.3342633

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{kunft_vldb19,
        title = {{An Intermediate Representation for Optimizing Machine Learning Pipelines}},
        author = {Kunft, Andreas and Katsifodimos, Asterios and Schelter, Sebastian and Breß, Sebastian and Rabl, Tilmann and Markl, Volker},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {11},
        pages = {1553--1567},
        doi = {10.14778/3342263.3342633},
        url = {https://doi.org/10.14778/3342263.3342633},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 21 of 21 citing papers.

Rank Citing Paper Year Venue Pagerank
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
2,182 Extending Relational Query Processing with ML Inference 2020 CIDR 8.8982998e-05
2,662 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.1596229e-05
3,683 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.1006425e-05
3,969 Optimizing Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 6.8899861e-05
4,095 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.8095767e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6569314e-05
5,230 Babelfish: Efficient Execution of Polyglot Queries 2022 VLDB 6.2189001e-05
5,272 InferDB: In-Database Machine Learning Inference Using Indexes 2024 VLDB 6.2010954e-05
5,343 The NebulaStream Platform: Data and Application Management for the Internet of Things 2020 CIDR 6.1708665e-05
5,933 Doing More with Less: Characterizing Dataset Downsampling for AutoML 2021 VLDB 5.942188e-05
7,386 Lachesis: Automatic Partitioning for UDF-Centric Analytics 2021 VLDB 5.538189e-05
8,798 Towards A Polyglot Framework for Factorized ML 2021 VLDB 5.2746322e-05
9,157 HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics Queries 2021 SIGMOD 5.2169683e-05
10,000 cedar: Optimized and Unified Machine Learning Input Data Pipelines 2025 VLDB 5.0979044e-05
10,043 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.0921006e-05
10,448 EncoderForge: Generating Efficient SQL for Encoders in Machine Learning Inference Pipelines 2026 SIGMOD 4.9793485e-05
10,653 InferF: Declarative Factorization of AI/ML Inferences over Joins 2026 SIGMOD 4.9793485e-05
10,898 stratum: A System Infrastructure for Massive Agent-Centric ML Workloads 2026 VLDB 4.9793485e-05
11,846 Redundancy Elimination in Distributed Matrix Computation 2022 SIGMOD 4.9793485e-05
12,015 TraNCE: Transforming Nested Collections Efficiently 2021 VLDB 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
21 Efficiently Compiling Efficient Query Plans for Modern Hardware 2011 VLDB 0.00056855599
43 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00045546775
73 Including Group-By in Query Optimization 1994 VLDB 0.00037522101
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001865959
537 MLbase: A Distributed Machine-learning System 2013 CIDR 0.00016768109
730 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014406936
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,193 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.0001158809
1,256 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011314687
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010071891
2,195 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 8.8781177e-05
2,227 Spinning Fast Iterative Data Flows 2012 VLDB 8.8021772e-05
2,249 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 8.7549752e-05
2,419 How to Architect a Query Compiler, Revisited 2018 SIGMOD 8.4923505e-05
2,527 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.3425785e-05
2,754 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0534972e-05
3,101 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.649219e-05
3,330 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.4173693e-05
3,518 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.2400627e-05
5,304 Probabilistic Demand Forecasting at Scale 2017 VLDB 6.1873575e-05
8,376 Meta-Dataflows: Efficient Exploratory Dataflow Jobs 2018 SIGMOD 5.3437059e-05
9,822 BlockJoin: Efficient Matrix Partitioning Through Joins 2017 VLDB 5.1254832e-05
12,336 Emma in Action: Declarative Dataflows for Scalable Data Analysis 2016 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Semantically Similar Papers