DBScholar

Back to authors

Matei Zaharia

Author ID
1157
ORCID
-
Links
(found by gpt-5.6-luna on jul 24 2026)
Most Frequent Institution
Stanford University
Pagerank
0.41165639
Overall Rank
91 | 99.57%
Paper Count
48

Affiliation Timeline

Incoming Non-self Citations Over Time

Total yearly non-self incoming citations across all papers by this author.

Publications by Paper Pagerank

Showing 48 of 48 publications.

Rank Title Year Venue Pagerank
24 Spark SQL: Relational Data Processing in Spark 2015 SIGMOD 0.00054865648
284 NoScope: Optimizing Neural Network Queries over Video at Scale 2017 VLDB 0.00022370521
425 Shark: SQL and Rich Analytics at Scale 2013 SIGMOD 0.00018704491
520 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores 2020 VLDB 0.00017136828
569 BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics 2020 VLDB 0.00016348191
1,138 Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics 2021 CIDR 0.00012023643
1,190 Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark 2018 SIGMOD 0.00011743246
1,229 Weld: A Common Runtime for High Performance Data Analytics 2017 CIDR 0.00011578425
1,469 Voodoo - A Vector Algebra for Portable Database Performance on Modern Hardware 2016 VLDB 0.00010678751
1,515 ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data 2024 SIGMOD 0.00010521317
1,607 Challenges and Opportunities in DNN-Based Video Analytics: A Demonstration of the BlazeIt Video Query Engine 2019 CIDR 0.0001022751
1,670 MISTIQUE: A System to Store and Query Model Intermediates for Model Diagnosis 2018 SIGMOD 0.00010045615
1,824 Photon: A Fast Query Engine for Lakehouse Systems 2022 SIGMOD 9.6734544e-05
2,161 DIFF: A Relational Interface for Large-Scale Data Explanation 2019 VLDB 9.0606664e-05
2,207 Shark: Fast Data Analysis Using Coarse-grained Distributed Memory 2012 SIGMOD 8.9565862e-05
2,316 Evaluating End-to-End Optimization for Data Analytics Applications in Weld 2018 VLDB 8.7596739e-05
2,414 Filter Before You Parse: Faster Analytics on Raw Data with Sparser 2018 VLDB 8.6078841e-05
2,809 Text2SQL is Not Enough: Unifying AI and Databases with TAG 2025 CIDR 8.0994951e-05
2,898 Approximate Selection with Guarantees using Proxies 2020 VLDB 7.978725e-05
3,253 Jointly Optimizing Preprocessing and Inference for DNN-based Visual Analytics 2021 VLDB 7.5936939e-05
3,411 Scaling Spark in the Real World: Performance and Usability 2015 VLDB 7.436229e-05
3,438 A Demonstration of Willump: A Statistically-Aware End-to-end Optimizer for Machine Learning Inference 2020 VLDB 7.415647e-05
3,549 DBOS: A DBMS-oriented Operating System 2022 VLDB 7.3212765e-05
3,766 TASTI: Semantic Indexes for Machine Learning-based Queries over Unstructured Data 2022 SIGMOD 7.1430942e-05
3,907 Accelerating Approximate Aggregation Queries with Expensive Predicates 2021 VLDB 7.0278233e-05
4,007 Optimizing Video Analytics with Declarative Model Relationships 2023 VLDB 6.9632395e-05
4,256 VIVA: An End-to-End System for Interactive Video Analytics 2022 CIDR 6.8018439e-05
4,457 Analyzing and Comparing Lakehouse Storage Systems 2023 CIDR 6.6913393e-05
5,983 Adaptive and Robust Query Execution for Lakehouses at Scale 2024 VLDB 6.0206841e-05
6,210 Supporting Our AI Overlords: Redesigning Data Systems to be Agent-First 2026 CIDR 5.9425753e-05
6,431 Finding Label and Model Errors in Perception Data With Learned Observation Assertions 2022 SIGMOD 5.8802622e-05
6,865 SparkR: Scaling R Programs with Spark 2016 SIGMOD 5.7508365e-05
7,568 Semantic Operators and Their Optimization: Enabling LLM-Based Data Processing with Accuracy Guarantees in LOTUS 2025 VLDB 5.5953379e-05
7,582 A Progress Report on DBOS: A Database-oriented Operating System 2022 CIDR 5.5918081e-05
7,648 Accelerating Aggregation Queries on Unstructured Streams of Data 2023 VLDB 5.575838e-05
7,767 Epoxy: ACID Transactions Across Diverse Data Stores 2023 VLDB 5.5485804e-05
8,091 Parallelism-Optimizing Data Placement for Faster Data-Parallel Computations 2023 VLDB 5.4889128e-05
8,524 Challenges and Opportunities for Autonomous Vehicle Query Systems 2021 CIDR 5.4119882e-05
8,605 Unity Catalog: Open and Universal Governance for the Lakehouse and Beyond 2025 SIGMOD 5.4028924e-05
8,696 Transactions Make Debugging Easy 2023 CIDR 5.3830471e-05
9,127 Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads 2025 SIGMOD 5.3189314e-05
9,694 Bringing the Operational and Analytical Worlds Together with Lakebase 2025 VLDB 5.2351259e-05
9,803 Introduction to Spark 2.0 for Database Researchers 2016 SIGMOD 5.21582e-05
11,454 R3: Record-Replay-Retroaction for Database-Backed Applications 2023 VLDB 5.093636e-05
13,301 Blink Twice - Automatic Workload Pinning and Regression Detection for Versionless Apache Spark using Retries 2025 SIGMOD -
13,329 Delta Sharing: An Open Protocol for Cross-Platform Data Sharing 2025 VLDB -
13,430 Cloud Data Systems: What are the Opportunities for the Database Research Community? 2022 VLDB -
13,471 Designing Production-Friendly Machine Learning 2021 VLDB -
Previous Page 1 / 1 Next

Frequent Co-authors

Co-authored at least 5 papers.

Co-author Shared Papers Rank Pagerank
Peter D. Bailis 15 132 0.34260431
Reynold Xin 15 254 0.21762263
Daniel Kang 12 321 0.18446789
Peter Kraft 10 786 0.089070328
Ion Stoica 7 135 0.33524546
Michael Armbrust 7 491 0.12710904
Ali Ghodsi 6 440 0.14079915
Christos Kozyrakis 6 1,171 0.062815972
Mike Stonebraker 5 7 1.0501788
Shoumik Palkar 5 1,608 0.047844277
Qian Li 5 1,732 0.045094032