DBScholar

Back to papers

An Intermediate Representation for Optimizing Machine Learning Pipelines

Summary: Lara, a declarative DSL, provides an IR for end-to-end ML pipelines, unifying preprocessing, UDFs, control flow, and training. Monads enable cross-boundary pushdown/fusion; combinators encode domain operators to optimize data access, with up to 10× speedups on dense and sparse data. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
12035
Venue
VLDB
Year
2019
Pagerank
8.8875753e-05
Overall Rank
2,239 | 84.64%
DOI
10.14778/3342263.3342633

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{kunft_vldb19,
        title = {{An Intermediate Representation for Optimizing Machine Learning Pipelines}},
        author = {Kunft, Andreas and Katsifodimos, Asterios and Schelter, Sebastian and Breß, Sebastian and Rabl, Tilmann and Markl, Volker},
        journal = {PVLDB},
        series = {{VLDB} '19},
        volume = {12},
        number = {11},
        pages = {1553--1567},
        doi = {10.14778/3342263.3342633},
        url = {https://doi.org/10.14778/3342263.3342633},
        year = {2019}
}

Incoming Citations (Sorted by Pagerank)

Showing 20 of 20 citing papers.

Rank Citing Paper Year Venue Pagerank
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
2,293 Extending Relational Query Processing with ML Inference 2020 CIDR 8.7949378e-05
2,865 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.0180243e-05
3,614 Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML 2020 CIDR 7.2568185e-05
4,067 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.9293511e-05
4,114 Optimizing Machine Learning Inference Queries with Correlative Proxy Models 2022 VLDB 6.8941194e-05
4,240 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.809685e-05
5,215 The NebulaStream Platform: Data and Application Management for the Internet of Things 2020 CIDR 6.3124972e-05
5,301 Babelfish: Efficient Execution of Polyglot Queries 2022 VLDB 6.2750553e-05
5,456 InferDB: In-Database Machine Learning Inference Using Indexes 2024 VLDB 6.2131252e-05
5,839 Doing More with Less: Characterizing Dataset Downsampling for AutoML 2021 VLDB 6.0699037e-05
7,311 Lachesis: Automatic Partitioning for UDF-Centric Analytics 2021 VLDB 5.6491618e-05
8,638 Towards A Polyglot Framework for Factorized ML 2021 VLDB 5.395289e-05
8,999 HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics Queries 2021 SIGMOD 5.3354529e-05
9,881 The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format 2024 SIGMOD 5.2040783e-05
10,232 EncoderForge: Generating Efficient SQL for Encoders in Machine Learning Inference Pipelines 2026 SIGMOD 5.093636e-05
10,466 InferF: Declarative Factorization of AI/ML Inferences over Joins 2026 SIGMOD 5.093636e-05
10,999 cedar: Optimized and Unified Machine Learning Input Data Pipelines 2025 VLDB 5.093636e-05
11,537 Redundancy Elimination in Distributed Matrix Computation 2022 SIGMOD 5.093636e-05
11,711 TraNCE: Transforming Nested Collections Efficiently 2021 VLDB 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 23 of 23 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
23 Efficiently Compiling Efficient Query Plans for Modern Hardware 2011 VLDB 0.00054886415
44 A Comparison of Approaches to Large-Scale Data Analysis 2009 SIGMOD 0.00046055057
71 Including Group-By in Query Optimization 1994 VLDB 0.00038021159
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001888524
532 MLbase: A Distributed Machine-learning System 2013 CIDR 0.00017072641
715 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014655327
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,229 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.00011578425
1,235 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011548457
1,644 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010132912
2,164 Opening the Black Boxes in Data Flow Optimization 2012 VLDB 9.0521951e-05
2,196 Spinning Fast Iterative Data Flows 2012 VLDB 8.9704984e-05
2,316 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 8.7596739e-05
2,486 Stubby: A Transformation-based Optimizer for MapReduce Workflows 2012 VLDB 8.5143189e-05
2,488 How to Architect a Query Compiler, Revisited 2018 SIGMOD 8.5091578e-05
2,717 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.2102313e-05
3,205 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6386536e-05
3,284 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.5663058e-05
3,459 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.3953716e-05
5,182 Probabilistic Demand Forecasting at Scale 2017 VLDB 6.3287692e-05
8,237 Meta-Dataflows: Efficient Exploratory Dataflow Jobs 2018 SIGMOD 5.4601966e-05
9,675 BlockJoin: Efficient Matrix Partitioning Through Joins 2017 VLDB 5.2380072e-05
12,041 Emma in Action: Declarative Dataflows for Scalable Data Analysis 2016 SIGMOD 5.093636e-05
Previous Page 1 / 1 Next

Semantically Similar Papers