6 minute read
- Most supposedly resilient systems are fragile by design because they ignore realistic failure modes and optimise for fantasy conditions.
- Treating resilience as paperwork, box‑ticking, or an afterthought creates a dangerous illusion of continuity that collapses under real stress. see why — jump to the section that argues this
- Leadership cultures that avoid confronting operational weaknesses and hide failures behind fake success metrics actively produce systemic fragility. see why — jump to the section that argues this
- Start by asking, for every critical component, “when, not if, this fails, how will we recover and how will customers experience it?” and bake those answers into roadmap and architecture.
- Regularly trigger real, end‑to‑end failure of critical infrastructure under production‑like conditions, and verify systems are actually usable, not just “back online,” before declaring success.
- Lean into painful failures early and often—intentionally expose weaknesses and iterate over the pain instead of relying on paperwork or simulations for disaster recovery.
Most systems are not resilient. They are fragile by design, propped up by a fantasy of “continuity” that vanishes the moment real pressure hits.
Spain’s national blackout. Portugal’s cascading failures. Oracle’s hospital cloud outage. Heathrow’s catastrophic shutdown. These were not accidents. They were not rare, unpredictable events. They were the inevitable consequences of bad engineering, shallow product thinking, and organisational self-delusion.
Resilience is not a checkbox. It is not a compliance exercise. It is not a hope and a prayer filed away in a disaster recovery plan. Resilience is hard. It is costly. It must be engineered, tested, and verified under real-world conditions, or it does not exist.
Bad Engineering
Real resilience assumes things will fail. Networks will fail. Authentication systems will fail. People will make mistakes. If your architecture does not assume failure at every level, you are not resilient; you are brittle.
Spain’s energy grid collapsed because it was optimised for efficiency, not survivability. No dynamic rerouting. No true load isolation. No meaningful observability. Their system was designed for perfect operating conditions that do not exist outside PowerPoint decks.
Oracle’s outage was even worse. Critical healthcare systems went offline because Oracle’s cloud infrastructure had no effective multi-region failover. Their architecture did not degrade gracefully; it fell over completely. That is not resilience. That is negligence at scale.
Bad Product and Continuity Thinking
Resilience is a product capability. If your product cannot survive failure, it is not a product. It is a liability.
Spain, Portugal, Oracle, all treated continuity as an afterthought. As long as the lights were on today, everything was declared fine. Until it was not.
Real product leadership demands harder questions: When, not if, this part fails, how will our system recover? How will our customers experience it? How fast can we restore service? How much risk are we carrying, and is that risk acceptable?
If those questions are not part of your roadmap, your architecture, and your operational strategy, you are not building resilience. You are building a house of cards.
Organisational Blindness
The real failure sits higher up the chain. Leadership failed to create a culture that prioritised operational survivability over operational fantasy.
I have lived through this firsthand. At Merrill Lynch, I participated in two major disaster recovery exercises. Both were declared “successful.” Both were complete failures.
Not a single system restored was actually usable. Systems were technically “back online”, but functionally, nothing worked. And the root cause was obvious: Active Directory, the system everything depended on for authentication, was never successfully recovered. Without it, every other “restored” system was dead weight.
Ironically, my application was successfully restored. We assumed it would have been usable, if Active Directory had been available. But we never found out. Two years running, the same critical dependency remained broken, and nobody was willing to call it what it was: systemic failure hidden behind fake success metrics.
Heathrow Airport offers another textbook case of organisational blindness disguised as resilience. When a fire broke out at one of their substations, they publicly blamed the disruption on their third-party power supplier. What they failed to mention was critical: Heathrow receives power from three independent substations, any one of which can fully power the airport alone.
The real problem was not the power supply it was a fluctuation in the power supply. It was Heathrow’s own disaster recovery system, designed to “protect” infrastructure by shutting everything down that detected that fluctuation and activated. The result? Heathrow’s entire IT backbone collapsed. It took the rest of the day to get basic systems running again, and much longer to recover from the cascading operational chaos.
Instead of owning the internal failure, leadership pointed fingers outward. It is the same story everywhere: an unwillingness to face the reality that their own fake resilience made the disaster worse.
Real Resilience: Iterating Over the Pain
Not every story ends in failure. There are organisations that do it right, and the difference is discipline.
Take Rackspace. During catastrophic floods in London, when almost every other datacentre in the city failed, Rackspace’s facility stayed operational. Their backup generators worked exactly as expected. While others blamed suppliers and scrambled for excuses, Rackspace quietly kept their customers online.
When asked why their systems worked when everyone else’s failed, the CEO simply held up a key.
It was the key to the power room.
Every month, without fail, he would walk down, unlock the main breaker, and physically pull it, shutting off external power. Not in theory. Not in a simulation. A real, full transfer to emergency backup power under real-world conditions.
Because of that brutal discipline, they did not hope their disaster recovery systems would work. They knew. They had tested it, again and again, under real conditions. They iterated over the pain.
And that is the lesson:
If something is hard, you must do it more often, not less.
If failure is painful, you must lean into it, not avoid it.
Only by living through controlled, intentional failures, early, often, and brutally, can you build true resilience.
You cannot wait until it matters. You cannot prepare only on paper. You must earn resilience by testing your systems, exposing your weaknesses, and getting punched in the face repeatedly until you are strong enough to survive the real thing.
Subscribe to Martin's articles
Writing since 2006
Articles, usually weekly on Mondays. One click to leave.
Resilience Is Built, Not Bought
You cannot buy resilience from a vendor. You cannot inherit it automatically because you deployed to “the cloud.” You cannot declare yourself resilient by writing it into your incident response plan.
Real resilience is built. It is designed in. It is iterated over. It is relentlessly tested. It is painful, slow, and expensive. But the alternative, the fragility we saw in Spain, Portugal, Oracle, and Heathrow, is far more costly.
If you are not engineering for failure, you are engineering for collapse.
Fragility is not an accident. It is a design choice.
Pretending otherwise only guarantees you will learn the hard way.
Enjoyed this? One click, no account.
Questions this answers
What does real system resilience look like and how is it actually built?
Real resilience assumes that components will fail and is intentionally engineered, tested, and verified under real-world conditions. It requires architectures that degrade gracefully, product leadership that treats resilience as a core capability, and organisational discipline to run painful, frequent, real tests rather than paper exercises or simulations. You cannot buy it from vendors or get it automatically from the cloud; you earn it through deliberate design, continuous iteration over failures, and brutal, recurring practice.
Why do supposedly resilient systems at places like Spain’s grid, Oracle, and Heathrow still fail catastrophically?
These systems fail because they are optimised for efficiency and appearance rather than survivability, with architectures that don’t assume failure and lack features like dynamic rerouting, true isolation, and effective failover. Organisationally, leadership treats continuity as a checkbox and engages in self-delusion, running exercises that declare success despite unusable systems and blaming external suppliers instead of acknowledging internal design and cultural failures.
How can organisations practically test and prove their disaster recovery and backup systems actually work?
Organisations must perform regular, real-world tests that intentionally trigger full failovers instead of relying on theoretical or simulated exercises. By repeatedly cutting primary power or dependencies under controlled conditions—like Rackspace’s CEO physically pulling the main breaker every month—they can validate that backup generators, failover paths, and dependent systems function correctly when truly needed.
Smart Classifications
Each classification [Concepts, Categories, & Tags] was assigned using AI-powered semantic analysis and scored across relevance, depth, and alignment. Final decisions? Still human. Always traceable. Hover to see how it applies.
