Google Cloud unveils BigLake, a data lake storage engine based on its BigQuery data warehouse, to let customers analyze data across storage systems
Google Cloud plans to launch a new data lake storage engine based on its popular BigQuery data warehouse to help remove barriers preventing customers …
Context & Ripple Effects
BigLake is the next step in Google Cloud's long campaign to make BigQuery the analysis layer regardless of where data physically sits. In 2020 it shipped BigQuery Omni to query data across GCP and AWS, with Azure promised, and years earlier it graduated Cloud SQL, Cloud Bigtable, and Cloud Datastore into enterprise-grade services to broaden its database portfolio.
What changes with BigLake is scope: rather than reaching across clouds, the BigQuery engine now reaches across storage systems, targeting the barrier between data warehouses and data lakes that keeps customers from analyzing everything in one place.
First-order effects
- BigQuery customers can now point the same engine at data living in lakes and other storage systems instead of copying it into the warehouse, removing the migration step that previously gated analysis.
Second-order effects
- Rival warehouse vendors face pressure to open their engines to external lake storage the way BigQuery Omni already opened Google's engine to AWS, turning engine portability into a competitive requirement rather than a differentiator.
Third-order effects
- If engines rather than storage become the locus of lock-in, the industry splits into a storage layer customers keep portable and a query layer they consolidate around — the pattern Omni and BigLake together sketch out for Google Cloud.
The trend: Cloud data platforms are decoupling the analytics engine from storage, with Google Cloud extending BigQuery from multi-cloud queries (Omni) to multi-storage-system analysis (BigLake).