BigLake: BigQuery’s Evolution toward a Multi-Cloud Lakehouse
Summary: BigLake extends BigQuery to a multi-cloud lakehouse, unifying data lake and data warehouse workloads with governance on open formats. Innovations: Parquet/Iceberg treated as first-class formats with governance; Object tables enable AI/ML over unstructured data; Omni enables cross-cloud BigQuery deployment. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Justin Levandoski (Google)
- 2. Garrett Casto (Google)
- 3. Mingge Deng (Google)
- 4. Rushabh Desai (Google)
- 5. Pavan Edara (Google)
- 6. Thibaud Hottelier (Google)
- 7. Amir Hormati (Google)
- 8. Anoop Johnson (Google)
- 9. Jeff Johnson (Google)
- 10. Dawid Kurzyniec (Google)
- 11. Sam McVeety (Google)
- 12. Prem Ramanathan (Google)
- 13. Gaurav Saxena (Google)
- 14. Vidya Shanmugam (Google)
- 15. Yuri Volobuev (Google)
BibTeX Citation
@inproceedings{levandoski_sigmod24,
title = {{BigLake: BigQuery’s Evolution toward a Multi-Cloud Lakehouse}},
author = {Levandoski, Justin and Casto, Garrett and Deng, Mingge and Desai, Rushabh and Edara, Pavan and Hottelier, Thibaud and Hormati, Amir and Johnson, Anoop and Johnson, Jeff and Kurzyniec, Dawid and McVeety, Sam and Ramanathan, Prem and Saxena, Gaurav and Shanmugam, Vidya and Volobuev, Yuri},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3626246.3653388},
url = {https://dl.acm.org/doi/10.1145/3626246.3653388},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 9 of 9 citing papers.
| Rank | Citing Paper | Year | Venue | Pagerank |
|---|---|---|---|---|
| 9,127 | Databricks Lakeguard: Supporting Fine-grained Access Control and Multi-user Capabilities for Apache Spark Workloads | 2025 | SIGMOD | 5.3189314e-05 |
| 9,772 | SAP HANA Cloud: Data Management for Modern Enterprise Applications | 2025 | SIGMOD | 5.2209769e-05 |
| 10,057 | AnyBlox: A Framework for Self-Decoding Datasets | 2025 | VLDB | 5.1664022e-05 |
| 10,121 | End-to-End Declarative Data Analytics: Co-designing Engines, Interfaces, and Cloud Infrastructure | 2026 | CIDR | 5.093636e-05 |
| 10,536 | Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability | 2026 | VLDB | 5.093636e-05 |
| 10,690 | Dynamic Pruning for Recursive Joins | 2025 | SIGMOD | 5.093636e-05 |
| 10,768 | Intra-Query Runtime Elasticity for Cloud-Native Data Analysis | 2025 | SIGMOD | 5.093636e-05 |
| 10,997 | The HANA Native Query Engine for Lakehouse Systems | 2025 | VLDB | 5.093636e-05 |
| 11,006 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning Workloads | 2025 | VLDB | 5.093636e-05 |
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 7,684 | Deep Lake: a Lakehouse for Deep Learning | 2023 | CIDR |
| 2 | 10,536 | Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability | 2026 | VLDB |
| 3 | 10,997 | The HANA Native Query Engine for Lakehouse Systems | 2025 | VLDB |
| 4 | 13,479 | The Challenge of Building Effective Data Lakes | 2020 | SIGMOD |
| 5 | 11,742 | Pixels: Multiversion Wide Table Store for Data Lakes | 2020 | CIDR |
| 6 | 9,694 | Bringing the Operational and Analytical Worlds Together with Lakebase | 2025 | VLDB |
| 7 | 4,120 | Big Metadata: When Metadata is Big Data | 2021 | VLDB |
| 8 | 7,780 | Petabyte-Scale Row-Level Operations in Data Lakehouses | 2024 | VLDB |
| 9 | 5,983 | Adaptive and Robust Query Execution for Lakehouses at Scale | 2024 | VLDB |
| 10 | 1,138 | Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics | 2021 | CIDR |