DBScholar

Back to papers

Data Management in Machine Learning: Challenges, Techniques, and Systems

Summary: Survey of data-management challenges and systems for ML workloads. Three lines of work: integrating ML with DBMS; adapting DB techniques to ML (queries, partitioning, compression); and combining data-management with ML lifecycles, plus open directions. (summarized by gpt-5-nano on Feb 09 2026)

Paper ID
h3600472ff61e2e81
Venue
SIGMOD
Year
2017
Pagerank
0.00011325762
Overall Rank
1,255 | 91.57%
DOI
10.1145/3035918.3054775

Incoming Non-self Citations Over Time

Authors

BibTeX Citation

@inproceedings{kumar_sigmod17,
        title = {{Data Management in Machine Learning: Challenges, Techniques, and Systems}},
        author = {Kumar, Arun and Boehm, Matthias and Yang, Jun},
        series = {{SIGMOD} '17},
        booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
        publisher = {Association for Computing Machinery},
        doi = {10.1145/3035918.3054775},
        url = {https://dl.acm.org/doi/10.1145/3035918.3054775},
        year = {2017}
}

Incoming Citations (Sorted by Pagerank)

Showing 32 of 32 citing papers.

Rank Citing Paper Year Venue Pagerank
1,152 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011801961
1,746 SMOKE: Fine-grained Lineage at Interactive Speed 2018 VLDB 9.7341914e-05
2,326 SliceLine: Fast, Linear-Algebra-based Slice Finding for ML Model Debugging 2021 SIGMOD 8.6309237e-05
2,460 Query Processing on Tensor Computation Runtimes 2022 VLDB 8.4348335e-05
2,662 End-to-end Optimization of Machine Learning Prediction Queries 2022 SIGMOD 8.1596229e-05
2,778 AIDA - Abstraction for Advanced In-Database Analytics 2018 VLDB 8.0299506e-05
2,908 AI Meets Database: AI4DB and DB4AI 2021 SIGMOD 7.8742664e-05
2,975 In-Database Learning with Sparse Tensors 2018 PODS 7.7907759e-05
3,112 Incremental View Maintenance with Triple Lock Factorization Benefits 2018 SIGMOD 7.6357579e-05
3,226 Opportunities for Quantum Acceleration of Databases: Optimization of Queries and Transaction Schedules 2023 VLDB 7.5095541e-05
3,737 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.0628666e-05
3,997 Overton: A Data System for Monitoring and Improving Machine-Learned Products 2020 CIDR 6.8655222e-05
4,161 The Relational Data Borg is Learning 2020 VLDB 6.7700593e-05
4,171 Data Integration and Machine Learning: A Natural Synergy 2018 SIGMOD 6.7608137e-05
4,330 MNC: Structure-Exploiting Sparsity Estimation for Matrix Expressions 2019 SIGMOD 6.6595681e-05
5,071 Data Collection and Quality Challenges for Deep Learning 2020 VLDB 6.2879908e-05
5,076 Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and Beyond 2021 SIGMOD 6.2860582e-05
6,143 Efficient Construction of Approximate Ad-Hoc ML models Through Materialization and Reuse 2018 VLDB 5.8721471e-05
6,246 ColumnML: Column-Store Machine Learning with On-The-Fly Data Transformation 2019 VLDB 5.8370793e-05
6,400 Functional-Style SQL UDFs With a Capital 'F' 2020 SIGMOD 5.7981078e-05
6,559 DeepBase: Deep Inspection of Neural Networks 2019 SIGMOD 5.7489089e-05
6,878 DAPHNE: An Open and Extensible System Infrastructure for Integrated Data Analysis Pipelines 2022 CIDR 5.657878e-05
7,897 Using VDMS to Index and Search 100M Images 2021 VLDB 5.4317361e-05
8,012 ItemSuggest: A Data Management Platform for Machine Learned Ranking Services 2019 CIDR 5.4072178e-05
8,411 SHiFT: An Efficient, Flexible Search Engine for Transfer Learning 2023 VLDB 5.336291e-05
8,847 Machine Learning Meets Big Spatial Data 2019 VLDB 5.2646682e-05
9,144 Cerebro: A Layered Data Platform for Scalable Deep Learning 2021 CIDR 5.2201471e-05
9,157 HADAD: A Lightweight Approach for Optimizing Hybrid Complex Analytics Queries 2021 SIGMOD 5.2169683e-05
9,286 ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUs 2021 VLDB 5.2025256e-05
10,181 In-Database Data Imputation 2024 SIGMOD 5.0653015e-05
11,846 Redundancy Elimination in Distributed Matrix Computation 2022 SIGMOD 4.9793485e-05
11,981 Enforcing Constraints for Machine Learning Systems via Declarative Feature Selection: An Experimental Study 2021 SIGMOD 4.9793485e-05
Previous Page 1 / 1 Next

Outgoing Citations (Sorted by Pagerank)

Showing 50 of 63 cited papers.

Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.

Rank Cited Paper Year Venue Pagerank
22 Distributed GraphLab: A Framework for Machine Learning and Data Mining in the Cloud 2012 VLDB 0.00055962491
105 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033638251
154 MAD Skills: New Analysis Practices for Big Data 2009 VLDB 0.00028579704
224 Self-Driving Database Management Systems 2017 CIDR 0.00024013745
357 FAQ: Questions Asked Frequently 2016 PODS 0.00020020639
415 SystemML: Declarative Machine Learning on Spark 2016 VLDB 0.0001865959
417 MauveDB: Supporting Model-based User Views in Database Systems 2006 SIGMOD 0.00018631353
483 ActiveClean: Interactive Data Cleaning For Statistical Modeling 2016 VLDB 0.00017590977
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017202276
521 Learning Linear Regression Models over Factorized Joins 2016 SIGMOD 0.00016929744
537 MLbase: A Distributed Machine-learning System 2013 CIDR 0.00016768109
547 ERACER: A Database Approach for Statistical Inference and Data Cleaning 2010 SIGMOD 0.00016578131
579 Incremental Knowledge Base Construction Using DeepDive 2015 VLDB 0.00016086569
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.0001510357
672 The TileDB Array Data Storage Manager 2017 VLDB 0.0001489638
730 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014406936
777 To Join or Not to Join? Thinking Twice about Joins before Feature Selection 2016 SIGMOD 0.00014054709
851 Scaling Factorization Machines to Relational Data 2013 VLDB 0.00013457975
940 Starfish: A Self-tuning System for Big Data Analytics 2011 CIDR 0.00012964445
1,009 Simulation of Database-Valued Markov Chains Using SimSQL 2013 SIGMOD 0.00012539827
1,044 RIOT: I/O-Efficient Numerical Computing without SQL 2009 CIDR 0.00012332595
1,062 Tuffy: Scaling up Statistical Inference in Markov Logic Networks using an RDBMS 2011 VLDB 0.0001221201
1,081 Hybrid Parallelization Strategies for Large-Scale Machine Learning in SystemML 2014 VLDB 0.00012123917
1,167 DimmWitted: A Study of Main-Memory Statistical Analytics 2014 VLDB 0.00011729888
1,256 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011314687
1,444 An Architecture for Compiling UDF-centric Workflows 2015 VLDB 0.0001063181
1,614 Compressed Linear Algebra for Large-Scale Machine Learning 2016 VLDB 0.00010071891
1,774 ArrayStore: A Storage Manager for Complex Parallel Array Processing 2011 SIGMOD 9.6688719e-05
1,833 MacroBase: Prioritizing Attention in Fast Data 2017 SIGMOD 9.5405247e-05
1,842 On Predictive Modeling for Optimizing Transaction Execution in Parallel OLTP Systems 2012 VLDB 9.526552e-05
2,087 LINVIEW: Incremental View Maintenance for Complex Analytical Queries 2014 SIGMOD 9.068879e-05
2,096 Vizdom: Interactive Analytics through Pen and Touch 2015 VLDB 9.0550231e-05
2,209 The Case for Predictive Database Systems: Opportunities and Challenges 2011 CIDR 8.8326479e-05
2,227 Spinning Fast Iterative Data Flows 2012 VLDB 8.8021772e-05
2,247 Cumulon: Optimizing Statistical Data Analysis in the Cloud 2013 SIGMOD 8.7585767e-05
2,586 Brainwash: A Data System for Feature Engineering 2013 CIDR 8.2560072e-05
2,754 Implicit Parallelism through Deep Language Embedding 2015 SIGMOD 8.0534972e-05
2,981 WiSeDB: A Learning-based Workload Management Advisor for Cloud Databases 2016 VLDB 7.7851845e-05
3,008 GenBase: A Complex Analytics Genomics Benchmark 2014 SIGMOD 7.7595064e-05
3,330 SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning 2017 CIDR 7.4173693e-05
3,421 Ava: From Data to Insights Through Conversation 2017 CIDR 7.3146932e-05
3,565 Processing Forecasting Queries 2007 VLDB 7.2039591e-05
3,586 A Comparison of Platforms for Implementing and Running Very Large Scale Machine Learning Algorithms 2014 SIGMOD 7.1889578e-05
4,029 Towards High-Throughput Gibbs Sampling at Scale: A Study across Storage Managers 2013 SIGMOD 6.8419186e-05
4,116 Resource Elasticity for Large-Scale Machine Learning 2015 SIGMOD 6.7961306e-05
4,212 Optimizing I/O for Big Array Analytics 2012 VLDB 6.7298655e-05
4,302 The Missing Piece in Complex Analytics: Low Latency, Scalable Model Management and Serving with Velox 2015 CIDR 6.6778695e-05
4,938 Machine Learning for Big Data 2013 SIGMOD 6.346081e-05
5,001 GLADE: Big Data Analytics Made Easy 2012 SIGMOD 6.3197648e-05
5,682 Demonstration of Santoku: Optimizing Machine Learning over Normalized Data 2015 VLDB 6.0397904e-05
Previous Page 1 / 2 Next

Semantically Similar Papers