DBScholar

Back to papers

UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads

Summary: UPLIFT adaptively parallelizes costly feature transformations using data-aware fine-grained task graphs and cache-conscious execution, including for multi-pass workloads. On its FTBench benchmark, it achieves up to 31.6× speedup (9.27× average) over state-of-the-art ML systems. (summarized by gpt-5.6-luna on Jul 24 2026)

Paper ID
hff971f6079cdf7a3
Venue
VLDB
Year
2022
Pagerank
5.7171651e-05
Overall Rank
6,662 | 55.21%
DOI
10.14778/3551793.3551842

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@article{phani_vldb22,
        title = {{UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning Workloads}},
        author = {Phani, Arnab and Erlbacher, Lukas and Boehm, Matthias},
        journal = {PVLDB},
        series = {{VLDB} '22},
        volume = {15},
        number = {11},
        pages = {2929--2938},
        doi = {10.14778/3551793.3551842},
        url = {https://doi.org/10.14778/3551793.3551842},
        year = {2022}
}

Incoming Citations (Sorted by Pagerank)

Showing 7 of 7 citing papers.

Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 45 of 45 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
14 MonetDB/X100: Hyper-Pipelining Query Execution 2005 CIDR 0.00064031282
21 Efficiently Compiling Efficient Query Plans for Modern Hardware 2011 VLDB 0.00056855599
205 Snorkel: Rapid Training Data Creation with Weak Supervision 2018 VLDB 0.00025181304
210 Sort vs. Hash Revisited: Fast Join Implementation on Modern Multi-Core CPUs 2009 VLDB 0.00024851502
251 Multi-Core, Main-Memory Joins: Sort vs. Hash Revisited 2014 VLDB 0.00023143736
387 The LDBC Social Network Benchmark: Interactive Workload 2015 SIGMOD 0.00019426275
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001865959
423 Massively Parallel Sort-Merge Joins in Main Memory Multi-Core Database Systems 2012 VLDB 0.00018491327
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016086569
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.0001510357
707 On Synopses for Distinct-Value Estimation Under Multiset Operations 2007 SIGMOD 0.00014640173
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,120 Snuba: Automating Weak Supervision to Label Training Data 2019 VLDB 0.00011946047
1,152 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011801961
1,153 Data Management Challenges in Production Machine Learning 2017 SIGMOD 0.00011798912
1,266 An Experimental Comparison of Thirteen Relational Equi-Joins in Main Memory 2016 SIGMOD 0.00011269175
1,276 Orca: A Modular Query Optimizer Architecture for Big Data 2014 SIGMOD 0.00011239266
1,308 Automating Large-Scale Data Quality Verification 2018 VLDB 0.0001107886
1,369 Towards Scalable Dataframe Systems 2020 VLDB 0.00010899832
1,568 HELIX: Holistic Optimization for Accelerating Iterative Machine Learning 2019 VLDB 0.0001021302
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010071891
1,668 SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle 2020 CIDR 9.9371612e-05
1,702 BigBench: Towards an Industry Standard Benchmark for Big Data Analytics 2013 SIGMOD 9.8361594e-05
2,053 tf.data: A Machine Learning Data Processing Framework 2021 VLDB 9.1140407e-05
2,187 Production Machine Learning Pipelines: Empirical Analysis and Optimization Opportunities 2021 SIGMOD 8.8896655e-05
2,723 Exploiting Matrix Dependency for Efficient Distributed Matrix Computation 2015 SIGMOD 8.0944934e-05
2,818 To Partition, or Not to Partition, That is the Join Question in a Real System 2021 SIGMOD 7.9739791e-05
3,008 GenBase: A Complex Analytics Genomics Benchmark 2014 SIGMOD 7.7595064e-05
3,101 On Optimizing Operator Fusion Plans for Large-Scale Machine Learning in SystemML 2018 VLDB 7.649219e-05
3,330 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.4173693e-05
3,518 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.2400627e-05
3,745 Automated Feature Engineering for Algorithmic Fairness 2021 VLDB 7.055875e-05
3,775 Parallelizing Query Optimization 2008 VLDB 7.0272615e-05
3,930 Flexible Rule-Based Decomposition and Metadata Independence in Modin: A Parallel Dataframe System 2022 VLDB 6.9185978e-05
3,997 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 6.8655222e-05
4,209 Accelerating Queries with Group-By and Join by Groupjoin 2011 VLDB 6.7339503e-05
4,330 MNC: Structure-Exploiting Sparsity Estimation for Matrix Expressions 2019 SIGMOD 6.6595681e-05
4,334 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.6569314e-05
5,100 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.27349e-05
5,506 Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems 2021 VLDB 6.1016716e-05
5,792 Optimizing Machine Learning Workloads in Collaborative Environments 2020 SIGMOD 5.99349e-05
7,029 The Case for Deep Query Optimisation 2020 CIDR 5.6174811e-05
7,839 ExDRa: Exploratory Data Science on Federated Raw Data 2021 SIGMOD 5.4432099e-05
7,867 Mind the Gap: Bridging Multi-Domain Query Workloads with EmptyHeaded 2017 VLDB 5.4365458e-05
9,034 The Power of Nested Parallelism in Big Data Processing – Hitting Three Flies with One Slap – 2021 SIGMOD 5.2334993e-05
Previous Page 1 / 1 Next

Semantically Similar Papers