DBScholar

Back to papers

UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads

Summary: UPLIFT adaptively parallelizes costly feature transformations using data-aware fine-grained task graphs and cache-conscious execution, including for multi-pass workloads. On its FTBench benchmark, it achieves up to 31.6× speedup (9.27× average) over state-of-the-art ML systems. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
12964
Venue
VLDB
Year
2022
Pagerank
5.8477764e-05
Overall Rank
6,538 | 55.15%
DOI
10.14778/3551793.3551842

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{phani_vldb22,
        title = {{UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads}},
        author = {Phani, Arnab and Erlbacher, Lukas and Boehm, Matthias},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {11},
        pages = {2929--2938},
        doi = {10.14778/3551793.3551842},
        url = {https://doi.org/10.14778/3551793.3551842},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 6 of 6 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 45 of 45 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
14 MonetDB/X100: Hyper-Pipelining Query Execution 2005 CIDR 0.0006312782
23 Efficiently Compiling Efficient Query Plans for Modern Hardware 2011 VLDB 0.00054886415
205 Snorkel: Rapid Training Data Creation with Weak Supervision 2018 VLDB 0.00025235185
209 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024932174
252 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023242719
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001888524
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018725853
426 The LDBC Social Network Benchmark: Interactive Workload 2015 SIGMOD 0.00018692185
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016217563
640 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.00015409494
689 On Synopses for Distinct-Value Estimation Under Multiset Operations 2007 SIGMOD 0.00014940023
1,079 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012258469
1,094 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.00012214617
1,147 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011974846
1,157 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011924049
1,265 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011415709
1,350 Automating Large-Scale Data Quality Verification 2018 VLDB 0.00011065626
1,431 Towards Scalable Dataframe Systems 2020 VLDB 0.00010807221
1,569 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.00010335423
1,621 Orca: A Modular Query Optimizer Architecture for Big Data 2014 SIGMOD 0.00010203114
1,644 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010132912
1,693 BigBench: Towards an Industry Standard Benchmark for Big Data Analytics 2013 SIGMOD 9.9965799e-05
1,756 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.8172465e-05
2,018 tf.data: A Machine Learning Data Processing Framework 2021 VLDB 9.3001937e-05
2,657 Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities 2021 SIGMOD 8.2887895e-05
2,681 Exploiting Matrix Dependency for Efficient Distributed Matrix Computation 2015 SIGMOD 8.2632778e-05
2,949 GenBase: A Complex Analytics Genomics Benchmark 2014 SIGMOD 7.9281519e-05
2,962 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9170451e-05
3,205 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.6386536e-05
3,284 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.5663058e-05
3,459 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.3953716e-05
3,726 Parallelizing Query Optimization 2008 VLDB 7.1697834e-05
3,905 Automated Feature Engineering for Algorithmic Fairness 2021 VLDB 7.029145e-05
3,947 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 7.0040437e-05
4,079 Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System 2022 VLDB 6.9188052e-05
4,223 Accelerating Queries with Group-By and Join by Groupjoin 2011 VLDB 6.8224393e-05
4,240 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.809685e-05
4,409 MNC: Structure-Exploiting Sparsity Estimation for Matrix Expressions 2019 SIGMOD 6.7178579e-05
5,051 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.385354e-05
5,699 Optimizing Machine Learning Workloads in Collaborative Environments 2020 SIGMOD 6.1170243e-05
5,925 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems 2021 VLDB 6.0397888e-05
7,156 The Case for Deep Query Optimisation 2020 CIDR 5.686096e-05
7,687 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.5671645e-05
7,733 Mind the Gap: Bridging Multi-Domain Query Workloads with EmptyHeaded 2017 VLDB 5.5564458e-05
8,958 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.3449654e-05
Previous Page 1 / 1 Next

Semantically Similar Papers