The Graviton math: when ARM pays for the migration
The Lizrd team · · 3 min read
Graviton — AWS’s ARM processors — is one of the few optimizations that improves price and performance at the same time. Same workload, ~20% cheaper, often faster. On paper it’s a no-brainer.
In practice it’s a migration, and migrations have a cost that the headline number ignores. The teams that get burned are the ones that treat “20% cheaper” as free. The teams that win are the ones that do the actual math: savings-at-stake versus migration effort, ranked, top-down.
Where the savings are real
The price/performance gain is genuine and it compounds with fleet size. A handful of large, always-on instances — an EKS node group, a fleet of application servers, a database running on a Graviton-supported engine — is where the dollars concentrate. Twenty percent of a big steady-state fleet is a serious number; twenty percent of a single small box is a rounding error that isn’t worth a sprint.
That’s the first filter: rank candidates by dollars-at-stake. You’ll find most of the savings live in a few fleets, and the long tail isn’t worth the risk.
Where the migration cost hides
The move from x86 to ARM is usually smooth, but “usually” is doing work in that sentence. The real effort lives in:
- Container images — you need
arm64builds. Multi-arch images make this painless if your pipeline already produces them, and annoying if it doesn’t. - Native dependencies — anything compiled: a Python wheel, a Node native addon, a Go cgo binary. Pure-interpreted code moves for free; native code needs an ARM build and a test.
- Third-party agents — observability, security, and sidecar vendors mostly ship ARM builds now, but “mostly” is the thing to verify before you commit a fleet.
None of this is hard. All of it is a reason the migration isn’t literally free — and why you sequence by effort as well as savings.
Migrate like you rightsize: one variable, reversible
The safe pattern is the familiar one. Change the instance family, keep everything else constant, and keep a rollback:
node_group {
- instance_types = ["m6i.2xlarge"] # x86
+ instance_types = ["m7g.2xlarge"] # Graviton
}
Shift a slice of the fleet, watch latency and error rates against the x86 baseline across a full business cycle, then roll forward. If the numbers regress, you revert one line. Big-bang cutovers are where confidence — and the on-call engineer’s weekend — goes to die.
The decision needs a clear picture
Every step above depends on knowing two things you probably can’t see at a glance: which fleets carry enough steady-state spend to be worth moving, and which ones have dependencies that make the move expensive. Guess and you either leave the big wins on the table or burn a sprint migrating a fleet that saved you $40.
Lizrd does that ranking for you — it reads your fleet cost and utilization, flags the Graviton-eligible EC2 and EKS candidates ordered by dollars-at-stake, and proposes the family change as an exact diff with the savings estimate attached. The Graviton math only works when you can see both sides of it.