| 105 |
The MADlib Analytics Library or MAD Skills, the SQL |
2012 |
VLDB |
0.00033638251 |
| 503 |
Towards a Unified Architecture for in-RDBMS Analytics |
2012 |
SIGMOD |
0.00017202276 |
| 654 |
Materialization Optimizations for Feature Selection Workloads |
2014 |
SIGMOD |
0.0001510357 |
| 730 |
Learning Generalized Linear Models Over Normalized Data |
2015 |
SIGMOD |
0.00014406936 |
| 777 |
To Join or Not to Join? Thinking Twice about Joins before Feature Selection |
2016 |
SIGMOD |
0.00014054709 |
| 1,152 |
Cerebro: A Data System for Optimized Deep Learning Model Selection |
2020 |
VLDB |
0.00011801961 |
| 1,255 |
Data Management in Machine Learning: Challenges, Techniques, and Systems |
2017 |
SIGMOD |
0.00011325762 |
| 1,256 |
Towards Linear Algebra over Normalized Data |
2017 |
VLDB |
0.00011314687 |
| 1,362 |
Towards Model-based Pricing for Machine Learning in a Data Marketplace |
2019 |
SIGMOD |
0.0001091784 |
| 2,199 |
Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra |
2019 |
SIGMOD |
8.8750296e-05 |
| 2,586 |
Brainwash: A Data System for Feature Engineering |
2013 |
CIDR |
8.2560072e-05 |
| 2,710 |
Panorama: A Data System for Unbounded Vocabulary Querying over Video |
2020 |
VLDB |
8.1041775e-05 |
| 3,040 |
Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations |
2019 |
SIGMOD |
7.7216143e-05 |
| 3,518 |
A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics |
2018 |
VLDB |
7.2400627e-05 |
| 3,630 |
VISTA: Optimized System for Declarative Feature Transfer from Deep CNNs at Scale |
2020 |
SIGMOD |
7.1496182e-05 |
| 3,670 |
Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? |
2018 |
VLDB |
7.1108704e-05 |
| 3,711 |
Bolt-on Differential Privacy for Scalable Stochastic Gradient Descent-based Analytics |
2017 |
SIGMOD |
7.0783969e-05 |
| 3,737 |
In-RDBMS Hardware Acceleration of Advanced Analytics |
2018 |
VLDB |
7.0628666e-05 |
| 3,937 |
Understanding and Benchmarking the Impact of GDPR on Database Systems |
2020 |
VLDB |
6.9158079e-05 |
| 4,095 |
Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches |
2021 |
VLDB |
6.8095767e-05 |
| 4,694 |
Demonstration of SpeakQL: Speech-driven Multimodal Querying of Structured Data |
2019 |
SIGMOD |
6.4654925e-05 |
| 4,924 |
SNAILS: Schema Naming Assessments for Improved LLM-Based SQL Inference |
2025 |
SIGMOD |
6.3512934e-05 |
| 5,100 |
Towards Benchmarking Feature Type Inference for AutoML Platforms |
2021 |
SIGMOD |
6.27349e-05 |
| 5,682 |
Demonstration of Santoku: Optimizing Machine Learning over Normalized Data |
2015 |
VLDB |
6.0397904e-05 |
| 5,862 |
SpeakQL: Towards Speech-driven Multimodal Querying of Structured Data |
2020 |
SIGMOD |
5.9675576e-05 |
| 5,899 |
Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics Engines |
2023 |
VLDB |
5.9545431e-05 |
| 6,086 |
Demonstration of Nimbus: Model-based Pricing for Machine Learning in a Data Marketplace |
2019 |
SIGMOD |
5.8912658e-05 |
| 6,341 |
The future of data(base) education: Is the "cow book" dead? |
2021 |
VLDB |
5.8092399e-05 |
| 6,592 |
Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent |
2019 |
SIGMOD |
5.7421684e-05 |
| 7,068 |
How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses |
2024 |
VLDB |
5.6083188e-05 |
| 7,069 |
Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads |
2024 |
VLDB |
5.6083188e-05 |
| 7,536 |
Feature Selection in Enterprise Analytics: A Demonstration using an R-based Data Analytics System |
2013 |
VLDB |
5.5002435e-05 |
| 7,809 |
Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets |
2022 |
SIGMOD |
5.4490255e-05 |
| 8,544 |
Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? |
2021 |
SIGMOD |
5.3186267e-05 |
| 8,747 |
Probabilistic Management of OCR Data using an RDBMS |
2012 |
VLDB |
5.2865969e-05 |
| 8,798 |
Towards A Polyglot Framework for Factorized ML |
2021 |
VLDB |
5.2746322e-05 |
| 9,144 |
Cerebro: A Layered Data Platform for Scalable Deep Learning |
2021 |
CIDR |
5.2201471e-05 |
| 9,556 |
Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning |
2021 |
VLDB |
5.1572248e-05 |
| 9,557 |
Intermittent Human-in-the-Loop Model Selection using Cerebro: A Demonstration |
2021 |
VLDB |
5.1572248e-05 |
| 10,766 |
A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases |
2026 |
VLDB |
4.9793485e-05 |
| 13,689 |
Reimagining Deep Learning Systems Through the Lens of Data Systems |
2024 |
VLDB |
- |
| 13,787 |
Errata for "Cerebro: A Data System for Optimized Deep Learning Model Selection" |
2021 |
VLDB |
- |
| 13,827 |
Demonstration of Krypton: Optimized CNN Inference for Occlusion-based Deep CNN Explanations |
2019 |
VLDB |
- |