UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads
Summary: UPLIFT adaptively parallelizes costly feature transformations using data-aware fine-grained task graphs and cache-conscious execution, including for multi-pass workloads. On its FTBench benchmark, it achieves up to 31.6× speedup (9.27× average) over state-of-the-art ML systems. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Arnab Phani (Graz University of Technology)
- 2. Lukas Erlbacher (Graz University of Technology)
- 3. Matthias Boehm (Graz University of Technology)
BibTeX Citation
@article{phani_vldb22,
title = {{UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads}},
author = {Phani, Arnab and Erlbacher, Lukas and Boehm, Matthias},
journal = {PVLDB},
series = {{VLDB} '22},
volume = {15},
number = {11},
pages = {2929--2938},
doi = {10.14778/3551793.3551842},
url = {https://doi.org/10.14778/3551793.3551842},
year = {2022}
}
Incoming Citations (Sorted by Pagerank)
Showing 7 of 7 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 5,773 | Optimizing Data Pipelines for Machine Learning in Feature Stores | 2023 | VLDB | 5.9987664e-05 |
| 7,538 | Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines | 2023 | SIGMOD | 5.4995874e-05 |
| 10,043 | The Image Calculator: 10x Faster Image-AI Inference by Replacing JPEG with Self-designing Storage Format | 2024 | SIGMOD | 5.0921006e-05 |
| 10,435 | DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization | 2026 | SIGMOD | 4.9793485e-05 |
| 10,448 | EncoderForge: Generating Efficient SQL for Encoders in Machine Learning Inference Pipelines | 2026 | SIGMOD | 4.9793485e-05 |
| 10,898 | stratum: A System Infrastructure for Massive Agent-Centric ML Workloads | 2026 | VLDB | 4.9793485e-05 |
| 10,945 | Morphing-based Compression for Data-centric ML Pipelines | 2026 | VLDB | 4.9793485e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 45 of 45 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 3,811 | FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline | 2023 | VLDB |
| 2 | 9,556 | Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning | 2021 | VLDB |
| 3 | 5,773 | Optimizing Data Pipelines for Machine Learning in Feature Stores | 2023 | VLDB |
| 4 | 6,217 | Materialization and Reuse Optimizations for Production Data Science Pipelines | 2022 | SIGMOD |
| 5 | 2,227 | Spinning Fast Iterative Data Flows | 2012 | VLDB |
| 6 | 7,827 | FusionFlow: Accelerating Data Preprocessing for Machine Learning with CPU-GPU Cooperation | 2024 | VLDB |
| 7 | 7,538 | Automating and Optimizing Data-Centric What-If Analyses on Native Machine Learning Pipelines | 2023 | SIGMOD |
| 8 | 3,601 | Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines | 2022 | SIGMOD |
| 9 | 1,081 | Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML | 2014 | VLDB |
| 10 | 2,053 | tf.data: A Machine Learning Data Processing Framework | 2021 | VLDB |