BigLake: BigQuery’s Evolution toward a Multi-Cloud Lakehouse
Summary: BigLake extends BigQuery to a multi-cloud lakehouse, unifying data lake and data warehouse workloads with governance on open formats. Innovations: Parquet/Iceberg treated as first-class formats with governance; Object tables enable AI/ML over unstructured data; Omni enables cross-cloud BigQuery deployment. (summarized by gpt-5-nano on Feb 09 2026)
Incoming Non-self Citations Over Time
Authors
- 1. Justin Levandoski (Google)
- 2. Garrett Casto (Google)
- 3. Mingge Deng (Google)
- 4. Rushabh Desai (Google)
- 5. Pavan Edara (Google)
- 6. Thibaud Hottelier (Google)
- 7. Amir Hormati (Google)
- 8. Anoop Johnson (Google)
- 9. Jeff Johnson (Google)
- 10. Dawid Kurzyniec (Google)
- 11. Sam McVeety (Google)
- 12. Prem Ramanathan (Google)
- 13. Gaurav Saxena (Google)
- 14. Vidya Shanmugam (Google)
- 15. Yuri Volobuev (Google)
BibTeX Citation
@inproceedings{levandoski_sigmod24,
title = {{BigLake: BigQuery’s Evolution toward a Multi-Cloud Lakehouse}},
author = {Levandoski, Justin and Casto, Garrett and Deng, Mingge and Desai, Rushabh and Edara, Pavan and Hottelier, Thibaud and Hormati, Amir and Johnson, Anoop and Johnson, Jeff and Kurzyniec, Dawid and McVeety, Sam and Ramanathan, Prem and Saxena, Gaurav and Shanmugam, Vidya and Volobuev, Yuri},
series = {{SIGMOD} '24},
booktitle = {Proceedings of the {ACM} {SIGMOD} International Conference on Management of Data},
publisher = {Association for Computing Machinery},
doi = {10.1145/3626246.3653388},
url = {https://dl.acm.org/doi/10.1145/3626246.3653388},
year = {2024}
}
Incoming Citations (Sorted by Pagerank)
Showing 11 of 11 citing papers.
Previous
Page 1 / 1
Next
Outgoing Citations (Sorted by Pagerank)
Showing 13 of 13 cited papers.
Citations counted here include only citations to other VLDB/SIGMOD/CIDR/PODS papers in this database.
Previous
Page 1 / 1
Next
Semantically Similar Papers
| # | Overall Rank | Paper | Year | Venue |
|---|---|---|---|---|
| 1 | 10,226 | Active Data Lakes: Regaining Physical Data Independence Without Losing Interoperability | 2026 | VLDB |
| 2 | 9,684 | The HANA Native Query Engine for Lakehouse Systems | 2025 | VLDB |
| 3 | 13,793 | The Challenge of Building Effective Data Lakes | 2020 | SIGMOD |
| 4 | 12,045 | Pixels: Multiversion Wide Table Store for Data Lakes | 2020 | CIDR |
| 5 | 9,352 | Bringing the Operational and Analytical Worlds Together with Lakebase | 2025 | VLDB |
| 6 | 3,928 | Big Metadata: When Metadata is Big Data | 2021 | VLDB |
| 7 | 5,348 | Adaptive and Robust Query Execution for Lakehouses at Scale | 2024 | VLDB |
| 8 | 6,920 | Petabyte-Scale Row-Level Operations in Data Lakehouses | 2024 | VLDB |
| 9 | 950 | Lakehouse: A New Generation of Open Platforms that Unify Data Warehousing and Advanced Analytics | 2021 | CIDR |
| 10 | 10,943 | Lakebase: Serverless Postgres over Open Lake Storage | 2026 | VLDB |