Google Cloud experienced a 15-hour service disruption last week when an electrical fault on the upstream utility grid triggered power and cooling failures at a datacenter serving three specialized services in the europe-west4-a zone. The outage affected Google Cloud VMware Engine (GCVE), NetApp Volumes, and Bare Metal Solutions (BMS). Google proactively shut down workloads to protect customer data from risks posed by high-temperature conditions, though the company has not disclosed whether backup generators were available or why they proved insufficient.
The incident exposed a critical architectural detail that many cloud customers may not realize: some Google Cloud managed services operate from single datacenters within availability zones, rather than being distributed across multiple facilities. This configuration differs from the standard cloud architecture where zones typically contain multiple datacenters with built-in redundancy. Google has not clarified whether it explicitly informs customers which services have single-datacenter dependencies or how these limitations affect the resilience guarantees typically associated with multi-zone deployments.
According to Google's incident report, the electrical fault originated on the utility grid upstream of the datacenter and disrupted both electrical distribution gear and cooling equipment. The company stated that the affected datacenter serves the three impacted services exclusively, explaining why the outage did not affect other Google Cloud services in the same zone. Google indicated its incident analysis remains ongoing and promised a final report detailing preventative measures, though the company has not yet responded to questions about its backup power infrastructure at the site.
Industry analysts note this architecture is not unique to Google Cloud. Forrester Principal Analyst Biswajeet Mahapatra explained that AWS, Azure, and Google all operate services requiring dedicated hardware or tightly coupled infrastructure that may not distribute across multiple facilities like standard compute and storage services. The challenge lies in transparency, as customers receive general guidance to use multiple zones and regions for resilience but rarely get visibility into single-datacenter dependencies for specific managed services. This information gap leads organizations to assume cloud abstractions provide more facility-level redundancy than actually exists for specialized offerings.
The outage follows a similar 2023 incident at Google Cloud's europe-west9-a region, where a water leak in a non-Google portion of the facility caused service disruptions. In that case, Google's Spanner data replication system failed to maintain availability when one building became unavailable. Security and infrastructure teams should verify the physical architecture underlying their critical cloud services, particularly for managed offerings that may require dedicated hardware. Organizations should request explicit documentation from cloud providers about single-datacenter dependencies and consider whether their disaster recovery plans account for facility-level failures affecting specialized services.
Source: https://www.theregister.com/off-prem/2026/07/21/google-cloud-outage-shows-its-still-hard-to-understand-hyperscalers-real-resilience-regimes/5275405


