Dmitry Yackevich
MIRO · RELIABILITY · 2 MIN READ

40% of incidents, gone

−40%
incidents
Company
Miro
Role
Senior Engineering Manager
When
Sep 2022 – Feb 2025
Result
incident count −40%

Every fast-growing company reaches the same place: the product is winning, the pager is losing. At Miro, incidents weren't rare events — they were a workload. Engineers planned their weeks around them.

The instinct in that situation is to add: more dashboards, more alerts, more runbooks, more on-call layers. I went the other way.

You don't fight incidents one by one. You find their species — and delete it.

Step one: taxonomy. We stopped treating incidents as unique snowflakes and started classifying them. It turned out a huge share came from a handful of repeating classes — the same failure wearing different costumes.

Step two: blameless reviews with teeth. A post-mortem that produces a document is theatre. A post-mortem that produces a shipped fix is engineering. Action items got owners, deadlines, and — this was the controversial part — priority over feature work when they targeted a recurring class.

Step three: subtraction. We deleted alerts nobody acted on. We set SLOs so that "is this bad?" had a numeric answer instead of a debate. Alert fatigue is a reliability problem, not an annoyance.

Incident count dropped by 40%. But the number I'm proudest of is quieter: the on-call rotation stopped being something engineers negotiated their way out of. When the pager is mostly silent, and meaningful when it isn't, reliability becomes culture instead of heroics.

Next story →
Moving to AWS while millions were watching

Working on the same problems? Let's talk.

✉ Email me