Optimizing Data Pipelines for Machine Learning in Feature Stores
Summary: Introduces database-style optimizations for point-in-time joins, a critical feature-store pipeline operation. Implemented in Feathr and validated on TPCx-AI and retail workloads, achieving up to 3× speedups over state-of-the-art baselines. (summarized by gpt-5.6-luna on Jul 24 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Rui Liu (University of Chicago)
- 2. Kwanghyun Park (Yonsei University)
- 3. Fotis Psallidas (Microsoft)
- 4. Xiaoyong Zhu (Microsoft)
- 5. Jinghui Mo (LinkedIn)
- 6. Rathijit Sen (Microsoft)
- 7. Matteo Interlandi (Microsoft)
- 8. Konstantinos Karanasos (Meta)
- 9. Yuanyuan Tian (Microsoft)
- 10. Jesús Camacho-Rodríguez (Microsoft)
BibTeX Citation
@article{liu_vldb23,
title = {{Optimizing Data Pipelines for Machine Learning in Feature Stores}},
author = {Liu, Rui and Park, Kwanghyun and Psallidas, Fotis and Zhu, Xiaoyong and Mo, Jinghui and Sen, Rathijit and Interlandi, Matteo and Karanasos, Konstantinos and Tian, Yuanyuan and Camacho-Rodríguez, Jesús},
journal = {PVLDB},
series = {{VLDB} '23},
volume = {16},
number = {13},
pages = {4230--4239},
doi = {10.14778/3625054.3625060},
url = {https://doi.org/10.14778/3625054.3625060},
year = {2023}
}
Incoming Citations (Sorted by Pagerank)
Showing 4 of 4 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,385 | The Hopsworks Feature Store for Machine Learning | 2024 | SIGMOD | 5.2755515e-05 |
| 10,219 | DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization | 2026 | SIGMOD | 5.093636e-05 |
| 10,531 | TPCx-AI under the Microscope: A Benchmarking Debt Analysis | 2026 | VLDB | 5.093636e-05 |
| 10,540 | CAPS: Cost-Aware ML Pipeline Selection | 2026 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 21 of 21 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 5,340 | Machine Learning for Databases | 2021 | VLDB |
| 2 | 640 | Materialization Optimizations for Feature Selection Workloads | 2014 | SIGMOD |
| 3 | 6,143 | ColumnML: Column-Store Machine Learning with On-The-Fly Data Transformation | 2019 | VLDB |
| 4 | 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |
| 5 | 5,297 | An Integrated Development Environment for Faster Feature Engineering | 2014 | VLDB |
| 6 | 9,962 | Structure-Aware Machine Learning over Multi-Relational Databases | 2021 | SIGMOD |
| 7 | 6,309 | Materialization and Reuse Optimizations for Production Data Science Pipelines | 2022 | SIGMOD |
| 8 | 5,925 | Managing ML Pipelines: Feature Stores and the Coming Wave of Embedding Ecosystems | 2021 | VLDB |
| 9 | 9,433 | FEBench: A Benchmark for Real-Time Relational Data Feature Extraction | 2023 | VLDB |
| 10 | 9,385 | The Hopsworks Feature Store for Machine Learning | 2024 | SIGMOD |