Dmitry Yackevich
FLO HEALTH · MIGRATION · 2 MIN READ

Moving to AWS while millions were watching

99.9%
uptime, up from ~98%
Company
Flo Health
Role
Director of Engineering, Infrastructure
When
Oct 2018 – Sep 2022
Result
~98% → 99.9% uptime
In public
slides · talk video

In 2018, Flo ran on bare metal in a colocation facility, and the product was growing roughly 10× — a women's health app that millions of people checked every day. Downtime wasn't a metric. It was broken trust, at scale.

The servers were full, the racks were finite, and every capacity conversation ended with a delivery truck. We had to move to AWS. The question was how to move a living product without dropping it.

There is no such thing as a successful big-bang migration. There are only big bangs.

The rules we set: stateless services first, data moves are planned like surgery, every cutover gets a canary and a way back. One service at a time. A multi-account AWS landing zone ready before the first workload landed, so we weren't migrating into a mess.

Here's the part most migration stories skip: the move itself was the smaller half of the work. The bigger half was refusing to lift-and-shift our problems. Every service that moved got redesigned for availability — redundancy across zones, real change management, observability before alerts, not after.

Within the first year, uptime went from ~98% on hardware to 99.9% on AWS — roughly twenty times less downtime. And a few years later we did it all again, from EC2 to Kubernetes on EKS, while the plane was flying. I told both stories on stage as “2 Epic Migrations at Flo” (talk video, RU), and the field notes are in the essay.

Next story →
ISO 27001 without killing velocity

Working on the same problems? Let's talk.

✉ Email me