Moving to AWS while millions were watching
- Company
- Flo Health
- Role
- Director of Engineering, Infrastructure
- When
- Oct 2018 – Sep 2022
- Result
- ~98% → 99.9% uptime
- In public
- slides · talk video
In 2018, Flo ran on bare metal in a colocation facility, and the product was growing roughly 10× — a women's health app that millions of people checked every day. Downtime wasn't a metric. It was broken trust, at scale.
The servers were full, the racks were finite, and every capacity conversation ended with a delivery truck. We had to move to AWS. The question was how to move a living product without dropping it.
There is no such thing as a successful big-bang migration. There are only big bangs.
The rules we set: stateless services first, data moves are planned like surgery, every cutover gets a canary and a way back. One service at a time. A multi-account AWS landing zone ready before the first workload landed, so we weren't migrating into a mess.
Here's the part most migration stories skip: the move itself was the smaller half of the work. The bigger half was refusing to lift-and-shift our problems. Every service that moved got redesigned for availability — redundancy across zones, real change management, observability before alerts, not after.
Within the first year, uptime went from ~98% on hardware to 99.9% on AWS — roughly twenty times less downtime. And a few years later we did it all again, from EC2 to Kubernetes on EKS, while the plane was flying. I told both stories on stage as “2 Epic Migrations at Flo” (talk video, RU), and the field notes are in the essay.
Next story →