- Industry
- Cross-sector
- Capability
- Architecture & Resilience
- Decision
- Independent recovery design
- Outcome
- Primary vendor failure does not become total operational lockout
In July 2024, a faulty CrowdStrike Falcon sensor update crashed Windows systems around the world. Airlines grounded flights, hospitals lost access to systems and accreditation operations for the Paris Olympics were disrupted days before the Games opened.
This was a faulty security product update, not a cyberattack, yet the operational effect was severe.
Microsoft Entra ID has also experienced widespread authentication incidents, including multi-factor authentication failures that prevented users from reaching Microsoft 365, Teams and SharePoint across multiple regions.
Both vendors operate at enormous scale and are among the most capable suppliers in their markets. Their incidents demonstrate that even well-resourced platforms fail.
During one widespread identity outage, organisations that had placed every login, application and remote session behind a single service could not reach systems that were otherwise operating normally. Supplier reputation had been allowed to stand in for a recovery design.
The size of the supplier and the consultancy supporting it encouraged an assumption that resilience had already been covered. In practice, the brand on the contract had become the contingency plan.
Some clients using the same primary identity service retained essential access because they had already tested independent multi-factor authentication, break-glass accounts and alternative access paths. Their advantage was the recovery route available when the primary service failed.
A resilience review should therefore ask what remains available when a critical supplier service stops working, regardless of the name on the contract.
We do not design environments where one endpoint security product, one identity provider or one vendor-controlled path represents the only route to business operation. Not because those products are poor, but because no product should be the entire plan.
Break-glass accounts, an authentication path with a different failure mode and a tested recovery procedure are basic safeguards. They prevent trust in a supplier from becoming the entire resilience plan.
Engineering lessons
- Market leadership reduces some risks; it does not eliminate failure.
- Supplier scale and reputation are not substitutes for architectural resilience.
- Resilience requires an alternative that does not share the primary platform's failure mode.
- The strongest brand in the market is not a resilience strategy.
Read the engineering principles behind this work →
Confidentiality: Engineering Notes are based on real engagements. Client identities, timelines and identifying details may be changed to protect confidentiality. The engineering decisions and lessons remain representative of the work undertaken.
