DBScholar

Back to authors

Arun Kumar

Author ID
u871
ORCID
-
Links
(found by gpt-5.6-luna on jul 24 2026)
Most Frequent Institution
University of California San Diego
Pagerank
0.35553538
Overall Rank
125 | 99.43%
Paper Count
43

Affiliation Timeline

Incoming Non-self Citations Over Time

Total yearly non-self incoming citations across all papers by this author.

Publications by Paper Pagerank

Showing all 43 publications. Total citations include self and non-self citations.

Rank Title Year Venue Total Citations Pagerank
105 The MADlib Analytics Library or MAD Skills, the SQL 2012 VLDB 108 0.00033633007
503 Towards a Unified Architecture for in-RDBMS Analytics 2012 SIGMOD 49 0.00017195428
654 Materialization Optimizations for Feature Selection Workloads 2014 SIGMOD 46 0.00015096817
731 Learning Generalized Linear Models Over Normalized Data 2015 SIGMOD 56 0.00014400356
779 To Join or Not to Join? Thinking Twice about Joins before Feature Selection 2016 SIGMOD 41 0.00014048128
1,152 Cerebro: A Data System for Optimized Deep Learning Model Selection 2020 VLDB 27 0.00011796404
1,223 Data Management in Machine Learning: Challenges, Techniques, and Systems 2017 SIGMOD 33 0.00011468426
1,257 Towards Linear Algebra over Normalized Data 2017 VLDB 35 0.0001130959
1,363 Towards Model-based Pricing for Machine Learning in a Data Marketplace 2019 SIGMOD 22 0.00010912672
2,201 Enabling and Optimizing Non-linear Feature Interactions in Factorized Linear Algebra 2019 SIGMOD 15 8.8708356e-05
2,587 Brainwash: A Data System for Feature Engineering 2013 CIDR 23 8.2523942e-05
2,705 Panorama: A Data System for Unbounded Vocabulary Querying over Video 2020 VLDB 15 8.1096249e-05
3,042 Incremental and Approximate Inference for Faster Occlusion-based Deep CNN Explanations 2019 SIGMOD 13 7.7179591e-05
3,518 A Comparative Evaluation of Systems for Scalable Linear Algebra-based Analytics 2018 VLDB 15 7.236638e-05
3,632 VISTA: Optimized System for Declarative Feature Transfer from Deep CNNs at Scale 2020 SIGMOD 10 7.146238e-05
3,673 Are Key-Foreign Key Joins Safe to Avoid when Learning High-Capacity Classifiers? 2018 VLDB 9 7.1075403e-05
3,713 Bolt-on Differential Privacy for Scalable Stochastic Gradient Descent-based Analytics 2017 SIGMOD 6 7.075046e-05
3,739 In-RDBMS Hardware Acceleration of Advanced Analytics 2018 VLDB 12 7.0595382e-05
3,938 Understanding and Benchmarking the Impact of GDPR on Database Systems 2020 VLDB 12 6.9125342e-05
4,097 Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches 2021 VLDB 14 6.8064316e-05
4,696 Demonstration of SpeakQL: Speech-driven Multimodal Querying of Structured Data 2019 SIGMOD 5 6.4624318e-05
4,925 SNAILS: Schema Naming Assessments for Improved LLM-Based SQL Inference 2025 SIGMOD 6 6.3482868e-05
5,102 Towards Benchmarking Feature Type Inference for AutoML Platforms 2021 SIGMOD 9 6.2705617e-05
5,683 Demonstration of Santoku: Optimizing Machine Learning over Normalized Data 2015 VLDB 9 6.0369434e-05
5,864 SpeakQL: Towards Speech-driven Multimodal Querying of Structured Data 2020 SIGMOD 3 5.9647326e-05
5,901 Lotan: Bridging the Gap between GNNs and Scalable Graph Analytics Engines 2023 VLDB 7 5.9517243e-05
6,088 Demonstration of Nimbus: Model-based Pricing for Machine Learning in a Data Marketplace 2019 SIGMOD 4 5.8884769e-05
6,345 The future of data(base) education: Is the "cow book" dead? 2021 VLDB 1 5.8064898e-05
6,594 Tuple-oriented Compression for Large-scale Mini-batch Stochastic Gradient Descent 2019 SIGMOD 8 5.7394502e-05
7,070 How do Categorical Duplicates Affect ML? A New Benchmark and Empirical Analyses 2024 VLDB 2 5.6056639e-05
7,071 Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads 2024 VLDB 3 5.6056639e-05
7,542 Feature Selection in Enterprise Analytics: A Demonstration using an R-based Data Analytics System 2013 VLDB 9 5.4976824e-05
7,816 Nautilus: An Optimized System for Deep Transfer Learning over Evolving Training Datasets 2022 SIGMOD 5 5.446446e-05
8,551 Automation of Data Prep, ML, and Data Science: New Cure or Snake Oil? 2021 SIGMOD 1 5.3161212e-05
8,755 Probabilistic Management of OCR Data using an RDBMS 2012 VLDB 2 5.2841044e-05
8,806 Towards A Polyglot Framework for Factorized ML 2021 VLDB 2 5.2721353e-05
9,153 Cerebro: A Layered Data Platform for Scalable Deep Learning 2021 CIDR 10 5.217676e-05
9,564 Towards an Optimized GROUP BY Abstraction for Large-Scale Machine Learning 2021 VLDB 4 5.1547835e-05
9,565 Intermittent Human-in-the-Loop Model Selection using Cerebro: A Demonstration 2021 VLDB 2 5.1547835e-05
10,776 A Comparative Evaluation of Schema Subsetting for LLM-based NL-to-SQL over Large-Schema Databases 2026 VLDB 0 4.9769913e-05
13,694 Reimagining Deep Learning Systems Through the Lens of Data Systems 2024 VLDB 0 -
13,792 Errata for "Cerebro: A Data System for Optimized Deep Learning Model Selection" 2021 VLDB 1 -
13,832 Demonstration of Krypton: Optimized CNN Inference for Occlusion-based Deep CNN Explanations 2019 VLDB 3 -

Frequent Co-authors

Co-authored at least 5 papers.

Co-author Shared Papers Rank Pagerank
Supun Nakandala 8 798 0.08869248
Jeffrey Naughton 6 9 1.0155479
Christopher RĂ© 6 63 0.50355869
Yuhao Zhang 6 1,569 0.050109451
Jignesh Patel 5 31 0.69348996
Vraj Shah 5 1,650 0.047393443
Side Li 5 1,679 0.046638414
Lingjiao Chen 5 1,719 0.045749098