In this talk from SRE Day London, I walk through a real incident from earlier in my career: reducing storage across a 142-server fleet, an automation script that skipped one critical step, and the ~700GB of production data that paid the price for it.
I cover:
- What actually went wrong, and why it was not really the script's fault
- How SRE practices (alerts, health checks, logs, metrics) turned a possible disaster into a contained incident
- Why disaster recovery and business continuity need to be the default for every system
- The bigger idea: automate the work, never automate the thinking
Whether you're in SRE, DevOps, or platform engineering, this talk is about building the judgment to know when automation needs a human check before it runs.
Watch the talk on YouTube ↗