Back to papers
Democratizing Data Science through Interactive Curation of ML Pipelines
Summary: Interactive AutoML for scientists via curated ML pipelines. Uses query-optimization, cost-based bandits, and Bayesian optimization to achieve interactive latency and beat expert solutions on unseen data across 300+ datasets.
(summarized by gpt-5-nano on Feb 09 2026)
- Paper ID
- 5676
- Venue
- SIGMOD
- Year
- 2019
- Pagerank
- 0.00015324193
- Overall Rank
- 917 | 93.63%
- DOI
-
10.1145/3299869.3319863
Incoming Non-self Citations Over Time
Incoming Citations (Sorted by Pagerank)
Showing 27 of 27 citing papers.
| Rank |
Citing Paper |
Year |
Venue |
Pagerank |
| 1,462 |
ARDA: Automatic Relational Data Augmentation for Machine Learning |
2020 |
VLDB |
0.00011866333 |
| 1,742 |
Auctus: A Dataset Search Engine for Data Discovery and Augmentation |
2021 |
VLDB |
0.00010695388 |
| 2,122 |
SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle |
2020 |
CIDR |
9.4905306e-05 |
| 2,325 |
DBPal: A Fully Pluggable NL2SQL Training Pipeline |
2020 |
SIGMOD |
9.0277894e-05 |
| 3,938 |
SimpleTS: An Efficient and Universal Model Selection Framework for Time Series Forecasting |
2023 |
VLDB |
6.6111703e-05 |
| 4,455 |
AutoOD: Automatic Outlier Detection |
2023 |
SIGMOD |
6.1644904e-05 |
| 4,553 |
A Demonstration of AutoOD: A Self-Tuning Anomaly Detection System |
2022 |
VLDB |
6.0852762e-05 |
| 4,601 |
Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches |
2021 |
VLDB |
6.05274e-05 |
| 4,779 |
LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems |
2021 |
SIGMOD |
5.9259373e-05 |
| 4,962 |
Doing More with Less: Characterizing Dataset Downsampling for AutoML |
2021 |
VLDB |
5.7979872e-05 |
| 5,439 |
DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data |
2023 |
SIGMOD |
5.5034427e-05 |
| 6,061 |
Optimizing Machine Learning Workloads in Collaborative Environments |
2020 |
SIGMOD |
5.2270653e-05 |
| 7,309 |
The Machine Learning Bazaar: Harnessing the ML Ecosystem for Effective System Development |
2020 |
SIGMOD |
4.7611148e-05 |
| 7,702 |
ExDRa: Exploratory Data Science on Federated Raw Data |
2021 |
SIGMOD |
4.6689015e-05 |
| 8,096 |
Saga: A Scalable Framework for Optimizing Data Cleaning Pipelines for Machine Learning Applications |
2023 |
SIGMOD |
4.583522e-05 |
| 8,166 |
Capturing and Querying Fine-grained Provenance of Preprocessing Pipelines in Data Science |
2021 |
VLDB |
4.567959e-05 |
| 8,178 |
DORIAN in action: Assisted Design of Data Science Pipelines |
2022 |
VLDB |
4.5629474e-05 |
| 8,739 |
CtxPipe: Context-aware Data Preparation Pipeline Construction for Machine Learning |
2024 |
SIGMOD |
4.4520434e-05 |
| 8,828 |
HAIPipe: Combining Human-generated and Machine-generated Pipelines for Data Preparation |
2023 |
SIGMOD |
4.4364918e-05 |
| 9,196 |
Hyper-Tune: Towards Efficient Hyper-parameter Tuning at Scale |
2022 |
VLDB |
4.3723457e-05 |
| 10,252 |
CAPS: Cost-Aware ML Pipeline Selection |
2026 |
VLDB |
4.1905499e-05 |
| 10,569 |
A Systematic Study on Early Stopping Metrics in HPO and the Implications of Uncertainty |
2025 |
VLDB |
4.1905499e-05 |
| 10,636 |
CatDB: Data-catalog-guided, LLM-based Generation of Data-centric ML Pipelines |
2025 |
VLDB |
4.1905499e-05 |
| 10,690 |
AutoPrep: Natural Language Question-Aware Data Preparation with a Multi-Agent Framework |
2025 |
VLDB |
4.1905499e-05 |
| 11,218 |
Demystifying the QoS and QoE of Edge-hosted Video Streaming Applications in the Wild with SNESet |
2023 |
SIGMOD |
4.1905499e-05 |
| 11,480 |
Enforcing Constraints for Machine Learning Systems via Declarative Feature Selection: An Experimental Study |
2021 |
SIGMOD |
4.1905499e-05 |
| 11,553 |
Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop |
2020 |
CIDR |
4.1905499e-05 |
Outgoing Citations (Sorted by Pagerank)
Showing 10 of 10 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Semantically Similar Papers
| Overall Rank |
Paper |
Year |
Venue |
Pagerank |
| 2,122 |
SystemDS: A Declarative Machine Learning System for the End-to-End Data Science Lifecycle |
2020 |
CIDR |
9.4905306e-05 |
| 13,198 |
ML2DAC: Meta-learning to Democratize AutoML for Clustering Analyses |
2023 |
SIGMOD |
- |
| 13,112 |
Demonstrating CatDB: LLM-based Generation of Data-centric ML Pipelines |
2025 |
SIGMOD |
- |
| 10,821 |
mlidea: Interactively Improving ML Data Preparation Code via "Shadow Pipelines" |
2025 |
VLDB |
4.1905499e-05 |
| 3,076 |
Explore-by-Example: An Automatic Query Steering Framework for Interactive Data Exploration |
2014 |
SIGMOD |
7.6063803e-05 |
| 8,178 |
DORIAN in action: Assisted Design of Data Science Pipelines |
2022 |
VLDB |
4.5629474e-05 |
| 5,306 |
A Scalable AutoML Approach Based on Graph Neural Networks |
2022 |
VLDB |
5.5725759e-05 |
| 4,755 |
Optimization for Active Learning-based Interactive Database Exploration |
2019 |
VLDB |
5.9375171e-05 |
| 7,309 |
The Machine Learning Bazaar: Harnessing the ML Ecosystem for Effective System Development |
2020 |
SIGMOD |
4.7611148e-05 |
| 11,553 |
Active Reinforcement Learning for Data Preparation: Learn2Clean with Human-In-The-Loop |
2020 |
CIDR |
4.1905499e-05 |