🦎 Lizrd is in private beta. Request early access →

Autoscaling that saves money instead of hiding waste

The Lizrd team · · 3 min read

Autoscaling has a reputation as a cost control: it scales down when demand drops, so it saves money, right? Often it does the opposite. Autoscaling is a multiplier, and if you feed it over-provisioned inputs it faithfully multiplies your waste — spinning up more oversized replicas, packing them onto more nodes, all perfectly automated.

The problem is never the autoscaler. It’s that teams reach for it instead of getting the fundamentals right, when it only pays off on top of them.

The HPA multiplies whatever the pod is

The Horizontal Pod Autoscaler adds and removes replicas to hold a target metric. But each replica is a copy of the pod you defined — so if that pod requests 4 vCPU and uses 0.5, the HPA doesn’t fix the waste, it clones it. Ten replicas of a 4x-oversized pod is 40x the waste you started with, and it looks healthy because everything’s “scaling.”

Fix the pod first. Right requests and limits (the boring rightsizing work) are the precondition. Then horizontal scaling adds capacity that’s actually the right size.

Requests are the scaling signal, so get them honest

The HPA usually scales on CPU-as-a-percentage-of-request. That means your request is the reference point for every scaling decision — set it wrong and every threshold is wrong:

  resources:
    requests:
-     cpu: "2000m"   # scaling math is relative to this fiction
+     cpu: "500m"    # relative to reality

Inflate the request and the pod looks underutilized forever, so it never scales up when it should and you compensate with a permanently-high replica floor. The autoscaler was never given a chance to save you anything.

Scale the floor to zero where you can

The biggest autoscaling savings are at the bottom of the curve, not the top. A dev environment, an internal tool, a batch consumer with a quiet window — anything that can go to zero replicas when idle stops billing entirely during that idle. A minReplicas: 3 floor on a service that sees no traffic at night is paying for a peak that isn’t there. Match the floor to real demand, not to a comfort number.

VPA and HPA solve different halves

Horizontal scaling handles how many; the Vertical Pod Autoscaler handles how big. VPA’s real value for most teams is its recommendation mode — it watches actual usage and tells you what the requests should be, turning “what should I set?” from a guess into a measured answer. Use it to set the honest requests that make horizontal scaling work.

The inputs are the whole game

Every failure mode here is an input problem: a request set by guesswork, a floor set by fear, a pod nobody rightsized. The autoscaler just executes faithfully on whatever you gave it. So the leverage isn’t in the scaling policy — it’s in continuously knowing what your workloads actually use, and keeping the requests honest as that drifts.

That’s the signal Lizrd keeps live: it reads per-workload utilization against what you requested, flags the pods whose requests are inflating your replica counts and node bill, and proposes the corrected values as an exact diff you can apply to your cluster. Autoscaling saves money only when it’s scaling something the right size to begin with.

Keep reading

Stop reading about savings — find yours

Connect read-only and Lizrd surfaces your highest-impact fixes, with the exact change to make.