Demonstration of Santoku: Optimizing Machine Learning over Normalized Data
Summary: Demonstrates Santoku, toolkit optimizing ML on normalized data with factorized learning and auto-decisions to denormalize or push via joins. Leverages FDs to surface feature insights and ships as an R library for ML on normalized data. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Arun Kumar (University of Wisconsin)
- 2. Mona Jalal (University of Wisconsin)
- 3. Boqun Yan (University of Wisconsin)
- 4. Jeffrey Naughton (University of Wisconsin)
- 5. Jignesh M. Patel (University of Wisconsin)
BibTeX Citation
@article{kumar_vldb15,
title = {{Demonstration of Santoku: Optimizing Machine Learning over Normalized Data}},
author = {Kumar, Arun and Jalal, Mona and Yan, Boqun and Naughton, Jeffrey and Patel, Jignesh M.},
journal = {PVLDB},
series = {{VLDB} '15},
volume = {8},
number = {12},
pages = {1864},
doi = {10.14778/2824032.2824087},
url = {https://doi.org/10.14778/2824032.2824087},
year = {2015}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 1,235 | Towards Linear Algebra over Normalized Data | 2017 | VLDB | 0.00011548457 |
| 1,250 | Data Management in Machine Learning: Challenges, Techniques, and Systems | 2017 | SIGMOD | 0.00011485301 |
| 2,179 | Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra | 2019 | SIGMOD | 9.0146333e-05 |
| 3,334 | F: Regression Models over Factorized Views | 2016 | VLDB | 7.5110164e-05 |
| 3,681 | Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? | 2018 | VLDB | 7.2037388e-05 |
| 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB | 5.6990782e-05 |
| 8,638 | Towards A Polyglot Framework for Factorized ML | 2021 | VLDB | 5.395289e-05 |
| 8,979 | Cerebro: A Layered Data Platform for Scalable Deep Learning | 2021 | CIDR | 5.3399615e-05 |
| 9,371 | Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning | 2021 | VLDB | 5.275595e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 4 of 4 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
| Rank | Cited Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 640 | Materialization Optimizations for Feature Selection Workloads | 2014 | SIGMOD | 0.00015409494 |
| 715 | Learning Generalized Linear Models Over Normalized Data | 2015 | SIGMOD | 0.00014655327 |
| 2,604 | Brainwash: A Data System for Feature Engineering | 2013 | CIDR | 8.3524514e-05 |
| 7,403 | Feature Selection in Enterprise Analytics: A Demonstration using an R-based Data Analytics System | 2013 | VLDB | 5.6249895e-05 |
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,112 | Coresets over Multiple Tables for Feature-rich and Data-efficient Machine Learning | 2023 | VLDB |
| 2 | 11,481 | Demonstration of OpenDBML, a Framework for Democratizing In-Database Machine Learning | 2023 | VLDB |
| 3 | 2,179 | Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra | 2019 | SIGMOD |
| 4 | 2,315 | SANTOS: Relationship-based Semantic Table Union Search | 2023 | SIGMOD |
| 5 | 2,223 | Sato: Contextual Semantic Type Detection in Tables | 2020 | VLDB |
| 6 | 640 | Materialization Optimizations for Feature Selection Workloads | 2014 | SIGMOD |
| 7 | 9,962 | Structure-Aware Machine Learning over Multi-Relational Databases | 2021 | SIGMOD |
| 8 | 764 | To Join or Not to Join? Thinking Twice about Joins before Feature Selection | 2016 | SIGMOD |
| 9 | 715 | Learning Generalized Linear Models Over Normalized Data | 2015 | SIGMOD |
| 10 | 1,235 | Towards Linear Algebra over Normalized Data | 2017 | VLDB |