Loading_
We let the platform close 71% of incidents without a human. Here is where it worked, where it failed embarrassingly, and the guard rails that made the difference.
State is the part of Terraform nobody teaches properly. These five production incidents — all real, all survivable — map to five habits.
Every network team follows the same four stages. Most stall at scripts-that-one-person-understands. The difference between stage two and stage four is not tooling.
Error budgets changed how we run services. The same idea, applied to detection engineering, changed how we run our SOC.
The technology for four-minute onboarding is a solved problem. The reason most projects deliver four-day onboarding anyway is upstream of IT.
Most patch schedules are calendar-driven waves. Ring-based deployment with health gates catches the bad patch while it can still only hurt forty machines.
One email a month
New posts, new automations and the incidents worth learning from. 28,000 engineers subscribe.
Book thirty minutes and watch the platform do what these posts describe, against a scenario from your own estate.