Google Cloud: 3 ms voltage drop triggers a 15-hour outage
This article examines Google Cloud’s July 2026 europe-west4-a incident, where a three-millisecond voltage drop exposed failures in DRUPS backup power and cooling, disrupting VMware Engine, Bare Metal Solution, and NetApp Volumes for nearly 15 hours. We explore Google’s response and what teams can learn about facility redundancy, thermal monitoring, and dependency-ordered recovery.
Company and product
Google Cloud is Google’s cloud computing platform, providing infrastructure, storage, networking, databases, analytics, and managed application services to organizations worldwide.
The incident affected three services hosted in the europe-west4-a zone. Google Cloud VMware Engine provides managed VMware private clouds on Google Cloud infrastructure. Bare Metal Solution offers dedicated physical servers and storage for enterprise workloads that require direct hardware access. Google Cloud NetApp Volumes provides managed file storage for applications with enterprise performance and availability requirements.
Unlike services that run entirely on virtualized infrastructure, these products depend heavily on physical data center systems, including utility feeds, DRUPS backup power, transfer equipment, chilled-water cooling, network switches, storage appliances, and physical servers. A failure in shared power or cooling infrastructure can therefore disrupt several managed services at the same time, even when Google Cloud services in other zones remain operational.
What happened
At 16:24 PDT, an upstream electrical fault caused a three-millisecond voltage drop across utility feeds A and B.
Both protective breakers tripped and initiated transfers to DRUPS backup power. Feed B transferred successfully, but feed A’s DRUPS failed to carry the facility load because of electrical component failures.
The failure triggered an automatic transfer of three rows from feed A to feed B:
- Rows 1 and 2 transferred successfully.
- Row 3 failed to transfer.
- An overload-protection breaker tripped.
- Row 3 lost power.
Google attributed the failed Row 3 transfer to a load deployment discrepancy that remained under investigation in the final report.
At the same time, a chiller controller went offline and failed to signal the chilled-water distribution pumps to restart. Although the chillers remained powered, the lack of water circulation caused cooling system A to shut down.
The redundant cooling source was unavailable because of ongoing construction. Temperatures rose until the data hall crossed its safe operating threshold.
Google then shut down reachable systems to prevent equipment damage. The prolonged outage resulted not from the duration of the original voltage drop, but from the combined power failure, cooling loss, emergency shutdown, and controlled restart.
Timeline
- July 15, 16:24 PDT: A three-millisecond voltage drop affected both utility feeds.
- 16:24 PDT: Feed A’s DRUPS failed, and a chiller controller went offline.
- 16:30 PDT: VMware Engine power monitoring generated the first alert.
- 16:32 PDT: NetApp monitoring reported high temperatures and switch failures.
- 16:37 PDT: Rows 1 and 2 transferred to feed B; Row 3 tripped a breaker and lost power.
- 16:39 PDT: The NetApp Volumes customer-impact window began.
- 17:09 PDT: VMware Engine customer impact began.
- 17:32 PDT: All six affected NetApp clusters shut down because of high temperatures.
- 17:52 PDT: Bare Metal Solution monitoring detected rising temperatures.
- 18:29 PDT: Data hall temperatures reached 44°C.
- 18:37 PDT: Bare Metal Solution customer impact began.
- 19:55 PDT: Shutdown procedures finished for reachable servers and network devices.
- 21:05 PDT: Cooling recovered, and temperatures began returning to normal.
- July 16, 01:10 PDT: NetApp Volumes impact ended.
- 02:33 PDT: VMware Engine impact ended.
- 07:34 PDT: Bare Metal Solution impact ended.
Time to Detect (TTD): Approximately six minutes, from the 16:24 PDT fault to the first power alert at 16:30 PDT.
Time to Resolve (TTR): 14 hours and 55 minutes for the combined customer-impact window.
Who was affected?
- VMware Engine: Twenty-four private clouds belonging to 20 customers lost connectivity.
- Bare Metal Solution: Nine customers lost access to servers and storage appliances.
- NetApp Volumes: Standard, Premium, and Extreme customers lost access to storage volumes.
- Control-plane users: Creation of storage pools, volumes, and backups failed in the affected region.
- Single-zone deployments: Customers relying only on europe-west4-a had no in-zone alternative.
- Multi-regional deployments: Google advised customers to route traffic to another location.
Google did not publish the total number of affected NetApp Volumes customers.
How did Google respond?
Google’s immediate priority was protecting physical equipment and customer data.
Teams shut down servers, storage systems, private clouds, and network devices as temperatures exceeded safe levels. Facility engineers manually changed the chilled-water pumps from automatic to hand mode, restoring circulation.
An interim portable UPS was deployed to support the chiller controllers. After utility power returned, engineers verified Remote Power Panels, reset tripped breakers, and restored power to the affected racks.
A swing DRUPS unit was aligned with feed A to restore redundant backup capacity while the failed DRUPS underwent repair.
Google then restored infrastructure in dependency order:
- Network and routing.
- Storage clusters.
- Physical hosts.
- VMware private clouds.
- Customer workloads.
- Service health validation.
Google committed to follow-up actions with published ETAs:
- Investigating feed A’s DRUPS failure by August 2026.
- Resolving the Row 3 overload condition by August 2026.
- Reviewing cooling-control redundancy by September 2026.
- Adding earlier power and cooling alerts by August 2026.
- Improving automated thermal shutdown by October 2026.
- Updating BMS incident classifications and SLOs by September 2026.
- Reviewing NetApp recovery runbooks and introducing monthly drills by August 2026.
How did Google communicate?
Google used its Service Health dashboard to provide product-specific updates.
Initial messages warned about rising temperatures and explained that systems were being intentionally shut down for protection. Updates separated VMware Engine, Bare Metal Solution, and NetApp Volumes so customers could follow the recovery of their services.
Google stated that no workaround was available for single-zone deployments. Customers with multi-regional architectures were advised to route traffic elsewhere.
A preliminary report was published on July 17, followed by the final report on July 25.
Key learnings for other teams
- Test the complete power path: Validate DRUPS, breakers, transfers, panels, and deployed load together.
- Treat construction as reduced redundancy: Temporary maintenance conditions require additional safeguards or workload relocation.
- Monitor cooling as production infrastructure: Controllers, pumps, water flow, and temperature trends need high-priority alerts.
- Validate automatic restart behavior: Powered equipment can still fail when its control systems do not restart.
- Track actual row-level load: Transfer capacity must reflect real deployment rather than planned design.
- Detect thermal escalation earlier: Earlier thresholds provide more time for controlled shutdown.
- Automate emergency protection: Rapid temperature changes may outpace manual response.
- Design for multiple zones: Critical workloads should have tested failover outside one facility.
- Restore in dependency order: Power and cooling must stabilize before network, storage, compute, and workloads.
- Separate facility and service recovery: Restored utilities do not immediately restore customer environments.
Quick summary
On July 15, 2026, a three-millisecond voltage drop triggered a major Google Cloud outage in europe-west4-a. Feed B transferred successfully to DRUPS backup power, while feed A’s DRUPS failed because of electrical component problems. Rows 1 and 2 transferred to feed B, but Row 3 tripped an overload-protection breaker and lost power, which Google linked to a load deployment discrepancy still under investigation. The same disturbance took a chiller controller offline, preventing distribution pumps from restarting while redundant cooling was unavailable because of construction. Temperatures reached 44°C, forcing shutdowns across servers, storage, networking, and customer workloads. Google restored cooling manually, repaired the power path, aligned a swing DRUPS, and brought services back in dependency order. The combined impact lasted 14 hours and 55 minutes.
How ilert can help
Facility incidents require coordination across infrastructure, SRE, data center, and vendor teams.
- Alerting on facility conditions: Route power-feed, DRUPS, cooling, pump, temperature, and reachability signals to ilert.
- Grouping related alerts: Consolidate power, cooling, network, storage, and service symptoms into one incident.
- Escalating thermal emergencies: Use high-priority policies that immediately notify backup responders.
- Cross-team coordination: Use ilert ChatOps to bring data center, networking, storage, compute, SRE, and vendor teams into dedicated incident channels for real-time collaboration.
- Controlled recovery: Track cooling, power, network, storage, compute, and workload restoration in dependency order.
- Regional communication: Use ilert Status Pages to identify affected zones and explain available failover options.

