DBScholar

Back to authors

Arun Kumar

Author ID
u871
ORCID
-
Links
(found by gpt-5.6-luna on jul 24 2026)
Most Frequent Institution
University of California San Diego
Pagerank
0.35554089
Overall Rank
125 | 99.43%
Paper Count
43

Affiliation Timeline

Incoming Non-self Citations Over Time

Total yearly non-self incoming citations across all papers by this author.

Publications by Paper Pagerank

Showing 43 of 43 publications.

Rank Title Year Venue Pagerank
105 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 0.00033638251
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 0.00017202276
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 0.0001510357
730 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 0.00014406936
777 To Join or Not to Join? Thinking Twice about Joins before Feature Selection 2016 SIGMOD 0.00014054709
1,152 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 0.00011801961
1,255 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 0.00011325762
1,256 Towards Linear Algebra over Normalized Data 2017 VLDB 0.00011314687
1,362 Towards Model-based Pricing for Machine Learning in a Data Marketplace 2019 SIGMOD 0.0001091784
2,199 Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra 2019 SIGMOD 8.8750296e-05
2,586 Brainwash: A Data System for Feature Engineering 2013 CIDR 8.2560072e-05
2,710 Panorama: A Data System for Unbounded Vocabulary Querying over Video 2020 VLDB 8.1041775e-05
3,040 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 7.7216143e-05
3,518 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 7.2400627e-05
3,630 VISTA: Optimized System for Declarative Feature Transfer from Deep CNNs at Scale 2020 SIGMOD 7.1496182e-05
3,670 Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? 2018 VLDB 7.1108704e-05
3,711 Bolt-on Differential Privacy for Scalable Stochastic Gradient Descent-based Analytics 2017 SIGMOD 7.0783969e-05
3,737 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 7.0628666e-05
3,937 Understanding and Benchmarking the Impact of GDPR on Database Systems 2020 VLDB 6.9158079e-05
4,095 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 6.8095767e-05
4,694 Demonstration of SpeakQL: Speech-driven Multimodal Querying of Structured Data 2019 SIGMOD 6.4654925e-05
4,924 SNAILS: Schema Naming Assessments for Improved LLM-Based SQL Inference 2025 SIGMOD 6.3512934e-05
5,100 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 6.27349e-05
5,682 Demonstration of Santoku: Optimizing Machine Learning over Normalized Data 2015 VLDB 6.0397904e-05
5,862 SpeakQL: Towards Speech-driven Multimodal Querying of Structured Data 2020 SIGMOD 5.9675576e-05
5,899 Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics Engines 2023 VLDB 5.9545431e-05
6,086 Demonstration of Nimbus: Model-based Pricing for Machine Learning in a Data Marketplace 2019 SIGMOD 5.8912658e-05
6,341 The future of data(base) education: Is the "cow book" dead? 2021 VLDB 5.8092399e-05
6,592 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 5.7421684e-05
7,068 How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses 2024 VLDB 5.6083188e-05
7,069 Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads 2024 VLDB 5.6083188e-05
7,536 Feature Selection in Enterprise Analytics: A Demonstration using an R-based Data Analytics System 2013 VLDB 5.5002435e-05
7,809 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5.4490255e-05
8,544 Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? 2021 SIGMOD 5.3186267e-05
8,747 Probabilistic Management of OCR Data using an RDBMS 2012 VLDB 5.2865969e-05
8,798 Towards A Polyglot Framework for Factorized ML 2021 VLDB 5.2746322e-05
9,144 Cerebro: A Layered Data Platform for Scalable Deep Learning 2021 CIDR 5.2201471e-05
9,556 Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning 2021 VLDB 5.1572248e-05
9,557 Intermittent Human-in-the-Loop Model Selection using Cerebro: A Demonstration 2021 VLDB 5.1572248e-05
10,766 A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases 2026 VLDB 4.9793485e-05
13,689 Reimagining Deep Learning Systems Through the Lens of Data Systems 2024 VLDB -
13,787 Errata for "Cerebro: A Data System for Optimized Deep Learning Model Selection" 2021 VLDB -
13,827 Demonstration of Krypton: Optimized CNN Inference for Occlusion-based Deep CNN Explanations 2019 VLDB -
Previous Page 1 / 1 Next

Frequent Co-authors

Co-authored at least 5 papers.

Co-author Shared Papers Rank Pagerank
Supun Nakandala 8 797 0.088713447
Jeffrey Naughton 6 9 1.0157552
Christopher RĂ© 6 63 0.50362508
Yuhao Zhang 6 1,566 0.050116096
Jignesh Patel 5 31 0.69363359
Vraj Shah 5 1,647 0.047404617
Side Li 5 1,676 0.046649453
Lingjiao Chen 5 1,716 0.045759804