AWS Outage Exposes Resilience Blind Spots

Four lessons midmarket IT leaders can apply to stress-test recovery plans.

There is no telling how good a disaster recovery plan is until the day it must carry a business through a real failure.

A brutal reminder came in May, when an AWS outage took giants like Coinbase and FanDuel offline for hours. While services have since been restored, and affected companies will receive a standard 10 percent service credit, the burden of lost revenue and fractured customer trust remains.

There are lessons midmarket IT leaders can learn from these outages — and about hardening their own IT estate.

[RELATED: Cloudflare Outage Lessons: 5 Ways The Midmarket Can Survive Massive Cloud Downtime]

Sort the Estate Before Paying For Cross-Region Recovery

Cross-region failover is expensive, and not every system justifies the cost. Your first move must be a strict triage: separate the critical workloads that cannot survive a regional collapse from the non-essential systems that can safely ride out an outage.

Most midmarket teams rely on multi-zone redundancy. While this handles a single data center outage, database provider SingleStore notes that surviving a total regional failure requires entirely different engineering. Yet, the Uptime Institute reports that many resilience assessments still ignore these massive external, systemic cloud risks.

Lesson 1 for midmarket IT leaders: Do not overspend trying to protect everything. Reserve budget and fully test cross-region or multi-provider failover strategies for mission-critical systems.

Rehearse Failover Before An Incident

An untested recovery plan is ultimately just a guess. Recent data highlights exactly how costly that blind confidence can be. According to Veeam’s 2026 Data Trust and Resilience Report, 90 percent of organizations felt certain they could recover from an incident within their target windows; only 28 percent actually recovered their data in full.

To bridge this gap between expectation and reality, experts recommend treating failover testing as a continuous discipline measured against realistic outage scenarios, rather than treating it as a superficial, once-a-year compliance check.

The necessity of this approach was proven during the May AWS outage, where Coinbase’s postmortem revealed that automated recovery completely broke down when their matching engine lacked automated cross-zone failover, and an AWS-managed Kafka (MSK) control plane defect failed, forcing engineers into a manual recovery path, which took hours to execute.

[RELATED: AWS Downtime Shakes Business World: ‘Outages Now Cascade Not Just Across Services, But Entire Economies,’ Says Cockroach Labs CEO]

Lesson 2: To avoid a similar fate, organizations must schedule regular failover drills tied to specific recovery objectives and run them until the execution stops surprising to the engineering team.

Translate Downtime Into Dollars Before The Budget Meeting

Resilience spending gets approved when it reads as avoided loss-the figures to make that case are public. The Uptime Institute’s Annual Outage Analysis report revealed that more than half of recent major outages cost over $100,000, with one in five crossing the $1 million mark.

Standard cloud contracts provide almost no shield against these losses; as typical service credits return a mere 10 percent of monthly compute spend on impacted instances, doing absolutely nothing to recoup lost revenue. Furthermore, Veeam’s research reports that roughly 40 percent of organizations hit by an incident suffer direct financial damage alongside extended downtime.

Lesson 3: At budget meetings, do not pitch abstract uptime. Instead, share the per-hour cost of downtime for each critical system. Doing so turns a vague request into a hard business-use case for the finance team.

[RELATED: After Massive Google Outage, Just How Resilient Is The Cloud? One Report Sheds Some Light]

Align Recovery Targets with the Business

Recovery targets define how much downtime and data loss the company can absorb, even though IT usually sets them in isolation. Under the shared responsibility model, the provider owns the infrastructure and the customer owns the recovery, which makes the call about acceptable loss a leadership decision well beyond procurement. The Veeam report exposed how wide that disconnect runs, finding that while 90 percent of leaders trusted their recovery times, only 69 percent said those targets were fully aligned with the business continuity goals they are meant to serve.

Lesson 4: There is therefore a need to fix and set recovery time and data-loss targets by business impact, so the systems that move revenue come back first.

Agreeing on these numbers with business leaders during a calm planning cycle keeps both ownership and expectations crystal clear. The good thing is that none of this needs a hyperscaler's budget or a standalone resilience team.

[RELATED: Midmarket Reacts, Recovers From CrowdStrike Outage]

A midmarket team that has sorted its critical systems, drilled its failover until it runs seamlessly, and priced an hour of downtime has done most of the preparation it needs for disaster recovery. The teams that handle it during the quiet stretches are the ones that recover fastest when the next zone goes dark.