| 105 |
The MADlib Analytics Library or MAD Skills, the SQL |
2012 |
VLDB |
108 |
0.00033633007 |
| 503 |
Towards a Unified Architecture for in-RDBMS Analytics |
2012 |
SIGMOD |
49 |
0.00017195428 |
| 654 |
Materialization Optimizations for Feature Selection Workloads |
2014 |
SIGMOD |
46 |
0.00015096817 |
| 731 |
Learning Generalized Linear Models Over Normalized Data |
2015 |
SIGMOD |
56 |
0.00014400356 |
| 779 |
To Join or Not to Join? Thinking Twice about Joins before Feature Selection |
2016 |
SIGMOD |
41 |
0.00014048128 |
| 1,152 |
Cerebro: A Data System for Optimized Deep Learning Model Selection |
2020 |
VLDB |
27 |
0.00011796404 |
| 1,223 |
Data Management in Machine Learning: Challenges, Techniques, and Systems |
2017 |
SIGMOD |
33 |
0.00011468426 |
| 1,257 |
Towards Linear Algebra over Normalized Data |
2017 |
VLDB |
35 |
0.0001130959 |
| 1,363 |
Towards Model-based Pricing for Machine Learning in a Data Marketplace |
2019 |
SIGMOD |
22 |
0.00010912672 |
| 2,201 |
Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra |
2019 |
SIGMOD |
15 |
8.8708356e-05 |
| 2,587 |
Brainwash: A Data System for Feature Engineering |
2013 |
CIDR |
23 |
8.2523942e-05 |
| 2,705 |
Panorama: A Data System for Unbounded Vocabulary Querying over Video |
2020 |
VLDB |
15 |
8.1096249e-05 |
| 3,042 |
Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations |
2019 |
SIGMOD |
13 |
7.7179591e-05 |
| 3,518 |
A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics |
2018 |
VLDB |
15 |
7.236638e-05 |
| 3,632 |
VISTA: Optimized System for Declarative Feature Transfer from Deep CNNs at Scale |
2020 |
SIGMOD |
10 |
7.146238e-05 |
| 3,673 |
Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? |
2018 |
VLDB |
9 |
7.1075403e-05 |
| 3,713 |
Bolt-on Differential Privacy for Scalable Stochastic Gradient Descent-based Analytics |
2017 |
SIGMOD |
6 |
7.075046e-05 |
| 3,739 |
In-RDBMS Hardware Acceleration of Advanced Analytics |
2018 |
VLDB |
12 |
7.0595382e-05 |
| 3,938 |
Understanding and Benchmarking the Impact of GDPR on Database Systems |
2020 |
VLDB |
12 |
6.9125342e-05 |
| 4,097 |
Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches |
2021 |
VLDB |
14 |
6.8064316e-05 |
| 4,696 |
Demonstration of SpeakQL: Speech-driven Multimodal Querying of Structured Data |
2019 |
SIGMOD |
5 |
6.4624318e-05 |
| 4,925 |
SNAILS: Schema Naming Assessments for Improved LLM-Based SQL Inference |
2025 |
SIGMOD |
6 |
6.3482868e-05 |
| 5,102 |
Towards Benchmarking Feature Type Inference for AutoML Platforms |
2021 |
SIGMOD |
9 |
6.2705617e-05 |
| 5,683 |
Demonstration of Santoku: Optimizing Machine Learning over Normalized Data |
2015 |
VLDB |
9 |
6.0369434e-05 |
| 5,864 |
SpeakQL: Towards Speech-driven Multimodal Querying of Structured Data |
2020 |
SIGMOD |
3 |
5.9647326e-05 |
| 5,901 |
Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics Engines |
2023 |
VLDB |
7 |
5.9517243e-05 |
| 6,088 |
Demonstration of Nimbus: Model-based Pricing for Machine Learning in a Data Marketplace |
2019 |
SIGMOD |
4 |
5.8884769e-05 |
| 6,345 |
The future of data(base) education: Is the "cow book" dead? |
2021 |
VLDB |
1 |
5.8064898e-05 |
| 6,594 |
Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent |
2019 |
SIGMOD |
8 |
5.7394502e-05 |
| 7,070 |
How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses |
2024 |
VLDB |
2 |
5.6056639e-05 |
| 7,071 |
Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads |
2024 |
VLDB |
3 |
5.6056639e-05 |
| 7,542 |
Feature Selection in Enterprise Analytics: A Demonstration using an R-based Data Analytics System |
2013 |
VLDB |
9 |
5.4976824e-05 |
| 7,816 |
Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets |
2022 |
SIGMOD |
5 |
5.446446e-05 |
| 8,551 |
Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? |
2021 |
SIGMOD |
1 |
5.3161212e-05 |
| 8,755 |
Probabilistic Management of OCR Data using an RDBMS |
2012 |
VLDB |
2 |
5.2841044e-05 |
| 8,806 |
Towards A Polyglot Framework for Factorized ML |
2021 |
VLDB |
2 |
5.2721353e-05 |
| 9,153 |
Cerebro: A Layered Data Platform for Scalable Deep Learning |
2021 |
CIDR |
10 |
5.217676e-05 |
| 9,564 |
Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning |
2021 |
VLDB |
4 |
5.1547835e-05 |
| 9,565 |
Intermittent Human-in-the-Loop Model Selection using Cerebro: A Demonstration |
2021 |
VLDB |
2 |
5.1547835e-05 |
| 10,776 |
A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases |
2026 |
VLDB |
0 |
4.9769913e-05 |
| 13,694 |
Reimagining Deep Learning Systems Through the Lens of Data Systems |
2024 |
VLDB |
0 |
- |
| 13,792 |
Errata for "Cerebro: A Data System for Optimized Deep Learning Model Selection" |
2021 |
VLDB |
1 |
- |
| 13,832 |
Demonstration of Krypton: Optimized CNN Inference for Occlusion-based Deep CNN Explanations |
2019 |
VLDB |
3 |
- |