Postmortem Library

OpenAI: Regional identity failure disrupts ChatGPT and Codex

This article examines OpenAI’s July 2026 partial outage, where the loss of a regional identity database replica reduced capacity, weakened cache performance, and disrupted ChatGPT, Codex, Voice, Connectors, and file operations. We explore OpenAI’s mitigation and what teams can learn about capacity headroom, replica-loss testing, and automated cross-region failover.

Company and product

OpenAI develops artificial intelligence systems and products, including ChatGPT and Codex.

ChatGPT provides conversational and agent-based experiences across web, mobile, voice, files, and connected services. Codex supports software engineering tasks across repositories and development environments.

Both products depend on shared backend infrastructure, including identity, authentication, account management, storage, caching, and regional database services. A capacity failure in one shared dependency can therefore affect several customer-facing features simultaneously.

What happened

During cloud infrastructure maintenance, a regional database replica supporting OpenAI’s internal identity service became unavailable.

Database replicas normally provide redundancy and additional processing capacity. In this incident, the remaining capacity in the region could not carry the complete workload.

Identity-service requests began slowing down or failing. This reduced the effectiveness of internal caches, forcing more requests to reach the already constrained backend systems.

The incident followed this sequence:

  • A regional database replica became unavailable.
  • Remaining capacity could not support the workload.
  • Identity-service requests slowed or failed.
  • Cache effectiveness declined.
  • Dependent services generated additional backend traffic.
  • Regional infrastructure became increasingly overloaded.
  • Protective controls began rejecting requests.

OpenAI had failover mechanisms designed to move traffic away from unhealthy regional infrastructure. However, they did not automatically redirect enough traffic to other regions.

Because the identity service supported multiple products and features, the impact extended beyond conversations to Voice, Work Mode, Connectors, files, library functionality, and some Codex requests.

Timeline

  • Before 07:08 AM PDT — exact time not disclosed: Cloud maintenance caused a regional identity database replica to become unavailable.
  • 07:08 AM PDT: Customer impact began. A subset of users encountered errors across conversations and related features.
  • During the investigation: Engineers found that the remaining regional database capacity could not handle the workload.
  • During the incident: Database failures reduced cache effectiveness and increased backend traffic.
  • During the incident: Automated failover did not redirect enough traffic away from the affected region.
  • During the incident: Protective controls began rejecting requests to limit further overload.
  • 07:46 AM PDT: Engineers began manually redirecting traffic to healthy capacity in other regions.
  • After 07:46 AM PDT: Regional load decreased, and recovery began.
  • During recovery: The cloud provider paused further maintenance while stability was validated.
  • 08:05 AM PDT: OpenAI confirmed that affected services had fully recovered.

Time to Detect (TTD): Not publicly disclosed.

Time to Resolve (TTR): Approximately 57 minutes from the beginning of customer impact at 07:08 AM PDT to restoration at 08:05 AM PDT.

Who was affected?

  • A subset of ChatGPT users who could not load or continue conversations.
  • Users of ChatGPT Voice and Work Mode.
  • Users accessing Connectors, files, and library functionality.
  • Users whose Codex requests failed during the incident.

OpenAI classified the incident as a partial outage rather than a complete platform failure.

Its published availability metrics are aggregated across subscription tiers, models, and error types. Individual availability may vary depending on the subscription tier, model, and API features in use.

How did OpenAI respond?

OpenAI traced the elevated errors to insufficient database capacity following the loss of the regional replica.

At 07:46 AM PDT, engineers manually redirected requests away from the affected region and toward healthy capacity elsewhere. This reduced database pressure and allowed the identity service, caches, and dependent systems to stabilize.

The cloud provider paused further maintenance while OpenAI validated recovery and planned capacity rebalancing.

OpenAI identified three main prevention areas:

  • Strengthening regional capacity planning and verifying that services can tolerate the loss of a database replica.
  • Improving automatic failover and cross-region traffic redirection.
  • Hardening retries and circuit breakers through expanded failure testing and topology verification.

How did OpenAI communicate?

OpenAI communicated through its public status page.

Initial updates stated that some users were experiencing errors when loading or continuing conversations. Later updates identified additional impact to Voice and Work Mode and explained that mitigation work was underway.

OpenAI moved the incident into monitoring before confirming full recovery and later published a detailed write-up covering the impact, root cause, mitigation, and planned improvements.

Key learnings for other teams

  • Test replica loss under load: Remaining infrastructure must be proven capable of carrying the complete production workload.
  • Maintain capacity headroom: Regions operating near their limits may not tolerate maintenance or component loss.
  • Automate regional failover: Multi-region architecture offers limited protection when traffic movement requires manual intervention.
  • Monitor cache degradation: Declining hit rates can send additional traffic to a backend that is already overloaded.
  • Use load-aware retries: Backoff, retry budgets, circuit breakers, and load shedding should work together.
  • Test maintenance as failure: Planned work can remove capacity just as effectively as an unexpected fault.
  • Verify the live topology: Validate actual replica placement and capacity rather than relying only on architecture diagrams.
  • Protect identity services: Shared authentication and identity dependencies carry a large blast radius.

Quick summary

On July 19, 2026, OpenAI experienced a partial outage affecting ChatGPT conversations, Codex requests, Voice, Work Mode, Connectors, files, and library functionality. During cloud infrastructure maintenance, a regional database replica supporting an internal identity service became unavailable. Remaining capacity could not handle the full workload, causing database failures, reduced cache effectiveness, and increased backend pressure. Automated failover did not redirect enough traffic, so protective controls began rejecting requests. Engineers manually rerouted traffic at 07:46 AM PDT, and service was fully restored by 08:05 AM PDT. The incident demonstrated that redundancy is only effective when the remaining capacity has been tested under full production load.

How ilert can help

Shared infrastructure failures can generate alerts across many products at once.

  • Alerting on capacity risks: Route replica health, identity errors, regional utilization, cache hit rates, and rejected-request metrics to ilert.
  • Grouping related alerts: Consolidate signals from conversations, Voice, files, Connectors, and Codex into one ilert.
  • Automated escalation: Notify database, identity, SRE, and platform teams through multi-step escalation policies.
  • Coordinating failover: Use ilert’s incident management workspace to create a dedicated incident channel, assign responders, coordinate traffic redirection, and track database and cache recovery.
  • Targeted communication: Use ilert Status Pages to distinguish affected features from services that remain operational.
  • Supporting the postmortem: Build the review from alert history, responder activity, status updates, and timeline events.
Find more Postmortems:
Ready to elevate your incident management?
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Our Cookie Policy
We use cookies to improve your experience, analyze site traffic and for marketing. Learn more in our Privacy Policy.
Open Preferences
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.