Dmitry Yackevich
MANYCHAT · SRE · 2 MIN READ

Observability from day one

Day 1
visibility before incidents
Company
Manychat
Role
Head of Infrastructure
When
Feb 2026 – now
Status
in progress — SLOs first, Grafana stack

Most companies buy observability the way people buy smoke detectors — right after the fire. It gets bolted on after the outage, shaped by whatever that outage happened to be, and it grows into an expensive pile of dashboards nobody trusts.

At Manychat I got the rare version of the problem: building the infrastructure team from scratch, in 2026, for a platform where AI agents are becoming users of the system. Which means I got to do observability first, not after.

You can't run what you can't see. And you definitely can't run what a thousand AI agents are doing to what you can't see.

The order of operations matters. SLOs came before dashboards: decide what "healthy" means, numerically, then instrument toward it. Unified telemetry — metrics, logs, traces in one place, in partnership with Grafana — instead of five tools with five opinions. And a rule I insist on: the cost of observability is itself observed. A monitoring bill that grows faster than the infrastructure it monitors is not a monitoring strategy.

The AI era raises the stakes. When agents open PRs and touch production, "is it up?" stops being enough — you need SLOs on outputs: eval pass rates and regressions, paged like availability. That's the platform we're building, and it starts with being able to see.

This story is still being written. If you want to write part of it — we're hiring.

Next story →
The $10M nobody could see

Working on the same problems? Let's talk.

✉ Email me