40% of incidents, gone
- Company
- Miro
- Role
- Senior Engineering Manager
- When
- Sep 2022 – Feb 2025
- Result
- incident count −40%
Every fast-growing company reaches the same place: the product is winning, the pager is losing. At Miro, incidents weren't rare events — they were a workload. Engineers planned their weeks around them.
The instinct in that situation is to add: more dashboards, more alerts, more runbooks, more on-call layers. I went the other way.
You don't fight incidents one by one. You find their species — and delete it.
Step one: taxonomy. We stopped treating incidents as unique snowflakes and started classifying them. It turned out a huge share came from a handful of repeating classes — the same failure wearing different costumes.
Step two: blameless reviews with teeth. A post-mortem that produces a document is theatre. A post-mortem that produces a shipped fix is engineering. Action items got owners, deadlines, and — this was the controversial part — priority over feature work when they targeted a recurring class.
Step three: subtraction. We deleted alerts nobody acted on. We set SLOs so that "is this bad?" had a numeric answer instead of a debate. Alert fatigue is a reliability problem, not an annoyance.
Incident count dropped by 40%. But the number I'm proudest of is quieter: the on-call rotation stopped being something engineers negotiated their way out of. When the pager is mostly silent, and meaningful when it isn't, reliability becomes culture instead of heroics.
Next story →