🦎 Lizrd is in private beta. Request early access →

Use case

Rightsize over-provisioned GPU instances

Find training and inference boxes running on far more GPU than they use, and get the exact smaller instance or accelerator to switch to — safely.

Cost Optimization $6,932/mo found
1
Rightsize Downsize the prod EKS node group

eks · prod-workers

$4,800/mo

High

2
Idle Stop the idle staging database

rds · analytics-staging

$1,180/mo

Medium

3
Rightsize Trim the over-provisioned checkout API

ecs · checkout-api

$640/mo

High

4
Orphaned Delete 14 unattached EBS volumes

ec2 · 14 × gp3

$312/mo

High

The problem

GPU instances are expensive and easy to over-size. A box gets provisioned for the peak of a training run, or a multi-GPU type is chosen “to be safe,” and then it runs for weeks at a fraction of its capacity. Because GPU hours are billed at a premium, a little over-provisioning is a lot of money — and CPU-oriented cost tools don’t even look at GPU utilization.

How Lizrd fixes it

Lizrd reads real GPU and GPU-memory utilization read-only (via the cloud’s GPU metrics), finds accelerators running well under what they’re paying for, and proposes the exact smaller instance type or accelerator — the right accelerator at the right size — with 30-day evidence, a confidence level, and the diff to apply. It works across EC2/SageMaker, Vertex AI, and Azure ML.

The outcome

Right-sized GPUs without a benchmarking project, every change reversible, and the savings proven back against your real bill once it ships.

“A four-GPU training box was sitting at 14% utilization for weeks. Lizrd showed the evidence and the one-line diff to drop it to a single GPU.”
— Staff ML engineer, fintech ~$5.4k/mo

More ways to save

Find this in your own cloud

Connect read-only and Lizrd surfaces the highest-impact fixes — with the exact change to make.