CrowdStrike, the latest Global IT outage

by | Jul 19, 2024

What a day!

I write this in the middle of the most recent Global IT outage that has had a major effect on multiple Public and Corporate services, causing them to grind to a halt. It will still be some time before the effects are known, but it can easily be said – this is the most significant outage in a very long time.

UK’s BBC News website the moment they announced the statement from CrowdStrike’s CEO – George Kurtz.

The list is massive, from Public Health Services in the UK, multiple global airports, shipping companies, entertainment industry, even the BBC and Sky News had some issues with their services – one can say that this was indeed a catastrophic incident.

Early on in this event, several people picked up that this was a call to the industry to become more “Hybrid” in their cloud choices, as this was evidence of sticking all your eggs in one basket. Unfortunately, they had not quite got all the facts right at this time.

Col Chambers LinkedIn Post

While the news post on this LinkedIn post from Col Chambers, MD of Technology at IBM UK Defense, the original had pointed to an issue with Azure. It would make sense that in the early hours of the outage, Azure was being pointed at as being the issue – however, that is down to mathematical probabilities. There are more WInodws-based VM’s running on Microsoft’s Own Cloud than anywhere else! Couple that with the wide reach that CrowdStrike has, and the faster ephemeral nature of VMs being replaced in Azure, it was probably seen as a higher % of servers there would blue screen and go offline. That and Microsoft’s back-end systems would also be Windows-based.

The CrowdStrike Content Update

The issue was far from that, a security product – CrowdStrike was to blame. A security service that is designed to protect endpoints across your whole estate, pushed out, worldwide, a content update that had a flaw that caused specifically Windows Based Operating Systems to Blue Screen on reboot, requiring manual intervention on each server to recover. A process that will take the IT industry many days to recover from.

So why wouldn’t a #HybridCloud solution resolve this? It’s because CrowdStrike is a platform-agnostic tool, that will run on one of many Operating Systems, that sit squarely (for Virtual Machines), in the “Customer” end of the Cloud Service Provider’s (CSP’s) Shared Responsibility model.

AWS’s Shared Responsibility Model

If we were to host, a set of VM’s on both Azure and AWS that had the CrowdStrike tooling installed, and the update was pushed out – both “clouds” would have seen an outage. You are still in a critical situation which means your business is now offline, and recovering this would be difficult.

Hybrid Cloud? Probably, but it should be more.

In a hybrid cloud environment, being cloud-agnostic is not enough for reliability. Evaluate the full stack, from the cloud infrastructure to the operating system, and your code.

Quote from our LinkedIn

In any Well-Architected environment, outages should be far and few between, but that isn’t just on an infrastructure or a cloud level. The AWS Well-Architected Framework talks about 6 pillars that should underpin any workload, regardless of its location. Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimisation, Sustainability.

AWS Well-Architected Reviews
Find out more about our AWS Well-Architected Framework reviews

However, if we dive deeper into each pillar we can see that it doesn’t always reference the underlying cloud, or the infrastructure only. Operational Excellence will look at the cultural and operational elements of the business. Reliability will talk about your Disaster Recovery plans, both from a workload and a user support point of view.

So, how do we link this back to the CrowdStrike outage, and not being Hybrid Cloud? Well, Col Chambers was correct in part of his comment:

Throwing all your #cloud eggs in one basket is just lazy

All you have to do is remove the #cloud part and you are bang on. So many companies have put all their eggs into CrowdStrike and on top of that their trust. CrowdStrike’s agent installs itself at the lowest possible level of the operating system to work, by design to stop malware and other security incidents from happening. All it took was one bad content update, and suddenly we get Blue Screens of Death (BSOD) across the globe. On a Friday too!

This was never about a single cloud provider, this is the responsibility of the customer to ensure that they had completed enough due-diligence to ensure that incidents like this, couldn’t affect the business as widely as this.

My recommendation: Why not choose two security providers, and have them install different agents on different servers, so that if you have an incident like this, then only half your services are affected. You wouldn’t keep all your VM’s in a single Availability Zone and expect it to be resilient, you would spread the load across multiple, why can’t you do the same with security services too?

Ultimately there is always a risk factor to take in to account, it might not be high of a risk to you if your whole estate goes offline, but for some companies – Health Services, Shops, Flight Operators, and Banks – you probably want to have some sort of resiliency in place to prevent your whole business from going down. Build and test your DR strategies, and my second recommendation – add an incident like this to your list of risks, and build your DR/BCP plans to test this type of issue in the future!

The knock-on effect

This isn’t going to go away soon, far from it. As I have typed this paragraph out, the UK Government has just completed an emergency COBRA meeting, announcing they will be working closely with the sectors and industries affected by this issue. Root Cause Analysis (RCA) will be completed by CrowdStrike and we will soon find out how such a simple content update, not only caused the issue but why it caused such a wide-ranging outage. There will be legal battles going on for some time, and I can already predict that business practices will need to change, even the best practices we come to know today will be updated after this.

I wouldn’t be surprised if we do end up with a new best practice that suggests multiple security providers across your estate, and the need to review Disaster Recovery plans will be high on everyone’s list. However, we can prepare, and we can test, and we can make sure that these types of incidents truly, as Col Chambers said in his post, “things like this should be [in the] past”.

Image (c) Bethesda Softworks – Be prepared for anything!

How can we help.

Don’t miss out on the chance to optimise your organisation’s security, performance, and efficiency. Book your free AWS Well-Architected review today.

Featured Image Photo by NASA on Unsplash

See how we can help you on the next stage of your cloud journey.

Cube

Cloud
Platforms

Run your applications quickly and reliably

Cloud Data
and AI

Gain visibility and harness the power of AI

Cloud
Migration

Move your business data to the cloud

DevOps and
Automation

Security and automations to protect and optimise