| 23 |
Spark SQL: Relational Data Processing in Spark
|
2015 |
SIGMOD |
0.00055406774 |
| 459 |
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores
|
2020 |
VLDB |
0.00017856221 |
| 950 |
Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics
|
2021 |
CIDR |
0.00012895553 |
| 1,036 |
Fine-grained Partitioning for Aggressive Data Skipping
|
2014 |
SIGMOD |
0.00012377471 |
| 1,123 |
Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark
|
2018 |
SIGMOD |
0.0001193233 |
| 1,487 |
Photon: A Fast Query Engine for Lakehouse Systems
|
2022 |
SIGMOD |
0.00010521722 |
| 1,868 |
G-OLA: Generalized On-Line Aggregation for Interactive Analysis on Big Data
|
2015 |
SIGMOD |
9.4754064e-05 |
| 2,347 |
Filter Before You Parse: Faster Analytics on Raw Data with Sparser
|
2018 |
VLDB |
8.6012176e-05 |
| 3,337 |
The Composable Data Management System Manifesto
|
2023 |
VLDB |
7.4104865e-05 |
| 3,455 |
Scaling Spark in the Real World: Performance and Usability
|
2015 |
VLDB |
7.2884813e-05 |
| 4,287 |
VIVA: An End-to-End System for Interactive Video Analytics
|
2022 |
CIDR |
6.6869736e-05 |
| 4,332 |
Analyzing and Comparing Lakehouse Storage Systems
|
2023 |
CIDR |
6.6587744e-05 |
| 5,348 |
Adaptive and Robust Query Execution for Lakehouses at Scale
|
2024 |
VLDB |
6.1690434e-05 |
| 6,227 |
iOLAP: Managing Uncertainty for Efficient Incremental OLAP
|
2016 |
SIGMOD |
5.8452386e-05 |
| 6,636 |
Physical Visualization Design: Decoupling Interface and System Design
|
2025 |
SIGMOD |
5.7262507e-05 |
| 6,987 |
SparkR: Scaling R Programs with Spark
|
2016 |
SIGMOD |
5.6283496e-05 |
| 7,542 |
Foreign Keys Open the Door for Faster Incremental View Maintenance
|
2023 |
SIGMOD |
5.4986181e-05 |
| 7,582 |
Unity Catalog: Open and Universal Governance for the Lakehouse and Beyond
|
2025 |
SIGMOD |
5.4910839e-05 |
| 7,996 |
Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads
|
2025 |
SIGMOD |
5.4108483e-05 |
| 8,346 |
SparkCruise: Workload Optimization in Managed Spark Clusters at Microsoft
|
2021 |
VLDB |
5.3510511e-05 |
| 8,535 |
New Query Optimization Techniques in the Spark Engine of Azure Synapse
|
2022 |
VLDB |
5.320973e-05 |
| 8,771 |
Automatic Indexing in Oracle
|
2025 |
VLDB |
5.2806451e-05 |
| 9,282 |
Making Data Engineering Declarative
|
2023 |
CIDR |
5.2034134e-05 |
| 9,352 |
Bringing the Operational and Analytical Worlds Together with Lakebase
|
2025 |
VLDB |
5.1887972e-05 |
| 9,575 |
A Step Toward Deep Online Aggregation
|
2023 |
SIGMOD |
5.1571823e-05 |
| 9,832 |
[Demo] Low-latency Spark Queries on Updatable Data
|
2019 |
SIGMOD |
5.1245795e-05 |
| 9,988 |
Introduction to Spark 2.0 for Database Researchers
|
2016 |
SIGMOD |
5.0998353e-05 |
| 9,995 |
Enzyme Demo: Incremental View Maintenance for Data Engineering
|
2026 |
VLDB |
5.0979044e-05 |
| 10,079 |
Delta Sharing: An Open Protocol for Cross-Platform Data Sharing
|
2025 |
VLDB |
5.0830849e-05 |
| 10,103 |
Still Asking: How Good Are Query Optimizers, Really?
|
2025 |
VLDB |
5.0789354e-05 |
| 10,155 |
Photon: A High-Performance Query Engine for the Lakehouse
|
2022 |
CIDR |
5.0710314e-05 |
| 10,916 |
AutoLiquid: Autonomic Data Layout Optimization for the Databricks Lakehouse
|
2026 |
VLDB |
4.9793485e-05 |
| 10,927 |
A Decade of Apache Spark Structured Streaming: How We Evolved The Architecture To Meet Real-World Needs
|
2026 |
VLDB |
4.9793485e-05 |
| 10,941 |
Ultron: History-Based Query Optimization at Databricks
|
2026 |
VLDB |
4.9793485e-05 |
| 10,943 |
Lakebase: Serverless Postgres over Open Lake Storage
|
2026 |
VLDB |
4.9793485e-05 |
| 11,120 |
Ultraverse: An Efficient What-if Analysis Framework for Software Applications Interacting with Database Systems
|
2025 |
SIGMOD |
4.9793485e-05 |
| 11,607 |
A Flexible Forecasting Stack
|
2024 |
VLDB |
4.9793485e-05 |
| 11,873 |
Statistical Schema Learning using Occam's Razor
|
2022 |
SIGMOD |
4.9793485e-05 |
| 13,614 |
The Three Golden Ages of Database Engineering: From SIGMOD’85 to the Agentic Era
|
2026 |
VLDB |
- |
| 13,623 |
Blink Twice - Automatic Workload Pinning and Regression Detection for Versionless Apache Spark using Retries
|
2025 |
SIGMOD |
- |
| 13,785 |
Designing Production-Friendly Machine Learning
|
2021 |
VLDB |
- |
| 13,793 |
The Challenge of Building Effective Data Lakes
|
2020 |
SIGMOD |
- |