DBScholar

Back to papers

Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML

Summary: Predicts enterprise adoption of DBMS-native ML that tightly integrates model lifecycle with rigorous data governance, privacy, and security at scale. Identifies unmet requirements and DB research challenges (provenance, auditing, access control, federated/secure training, deployment/auto-tuning) and sketches early system-building steps. (summarized by gpt-5-mini on Feb 09 2026)

Paper ID
386
Venue
CIDR
Year
2020
Pagerank
7.2568185e-05
Overall Rank
3,614 | 75.21%
DOI
-

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{agrawal_cidr20,
        address = {Amsterdam, Netherlands},
        series = {{CIDR} '20},
        title = {{Cloudy with High Chance of DBMS: A 10-year Prediction for Enterprise-Grade ML}},
        booktitle = {Proceedings of the {Conference} on {Innovative} {Data} {Systems} {Research}},
        author = {Agrawal, Ashvin and Chatterjee, Rony and Curino, Carlo and Floratou, Avrilia and Gowdal, Neha and Interlandi, Matteo and Jindal, Alekh and Karanasos, Konstantinos and Krishnan, Subru and Kroth, Brian and Leeka, Jyoti and Park, Kwanghyun and Patel, Hiren and Poppe, Olga and Psallidas, Fotis and Ramakrishnan, Raghu and Roy, Abhishek and Saur, Karla and Sen, Rathijit and Weimer, Markus and Wright, Travis and Zhu, Yiwen},
        year = {2020}
}

Incoming Citations (Sorted by Pagerank)

Showing 17 of 17 citing papers.

Rank Citing Paper Year Venue Pagerank
2,293 Extending Relational Query Processing with ML Inference 2020 CIDR 8.7949378e-05
2,347 Vertica-ML: Distributed Machine Learning in Vertica Database 2020 SIGMOD 8.7157552e-05
2,584 Complaint-driven Training Data Debugging for Query 2.0 2020 SIGMOD 8.3783546e-05
2,865 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.0180243e-05
4,067 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.9293511e-05
4,240 LIMA: Fine-grained Lineage Tracing and Reuse in Machine Learning Systems 2021 SIGMOD 6.809685e-05
4,756 Improving Reproducibility of Data Science Pipelines through Transparent Provenance Capture 2020 VLDB 6.5221401e-05
5,704 Optimizing Data Pipelines for Machine Learning in Feature Stores 2023 VLDB 6.1146371e-05
6,121 The Cosmos Big Data Platform at Microsoft: Over a Decade of Progress and a Decade to Look Forward 2021 VLDB 5.9688569e-05
6,614 Mitigating the Impedance Mismatch between Prediction Query Execution and Database Engine 2025 SIGMOD 5.8216658e-05
8,193 Towards Building Autonomous Data Services on Azure 2023 SIGMOD 5.4696038e-05
8,861 Optimizing the cloud? Don't train models. Build oracles! 2024 CIDR 5.355716e-05
9,649 Transforming ML Predictive Pipelines into SQL with MASQ 2021 SIGMOD 5.2430158e-05
9,835 Share the Tensor Tea: How Databases can Leverage the Machine Learning Ecosystem 2022 VLDB 5.2112879e-05
10,130 Does A Fish Need a Bicycle? The Case for On-Chip NPUs in DBMS 2026 CIDR 5.093636e-05
11,355 Git is for Data 2023 CIDR 5.093636e-05
11,512 Towards Observability for Machine Learning Pipelines 2022 CIDR 5.093636e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 7 of 7 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Previous Page 1 / 1 Next

Semantically Similar Papers