Meridian Fleet is a B2B SaaS platform for real-time fleet management, used by mid-market logistics operators running 200–2,000-vehicle fleets. The product ingests GPS and telemetry from in-vehicle devices, powers a live operations dashboard, scores driver behaviour, and offers route optimization on its premium tier. Customers are billed per active vehicle per month. ~45 engineers.
What the tech stack is for. Three things drive the architecture and, with it, the bill: (1) low-latency telemetry ingest — a vehicle's position must appear on the dashboard in under five seconds, which is both the core SLA and the top churn driver; (2) dashboard availability for 24/7 logistics operations; and (3) the route-optimization ML that differentiates the premium tier. Infrastructure is ~22% of cost of goods sold, so cost efficiency is gross margin — every $10K/mo recovered is roughly a point of margin.
Architecture, as it stands.
Before the findings below — the parts of this setup that are in good shape, and shouldn’t change.
Nine changes, ranked by monthly saving. Each lists the reason, the evidence, and the effort/risk. Expand for detail. None of these touches the customer-facing SLA; the first three are low-risk and capture most of the value.
Why. Compute and database usage is flat week-over-week, but everything is billed on-demand. A 1-year, no-upfront Compute Savings Plan sized to the observed p5 floor (not the peak) converts the predictable base load to a ~30% lower rate while leaving spikes on-demand. Start at ~70% coverage of the floor to stay conservative.
Why. 30-day CPU averages 19% and memory 34% on the prod API nodes, with peaks well within a node half the size. Halving the instance size keeps 2× headroom over observed peak. Pair with Karpenter (optimization 06 context) so the group also scales for genuine spikes instead of standing at peak 24/7.
Why. Staging (prod-sized) and the dev databases run around the clock but see no traffic outside working hours. A scheduled scale-to-zero / stop overnight and on weekends cuts their runtime by ~70% with zero production impact. The forgotten eu-west-1 staging environment (see hidden costs) is included here.
Why. This both saves money and removes the §04 High risk. Graviton (r6g) is ~20% cheaper at equal performance; a read replica offloads analytics/reporting reads off the writer, which then safely downsizes two sizes. Net: lower cost and the writer is no longer a single point of contention at month-end.
Why. A GPU inference endpoint has served zero requests in 21 days — the premium route-optimization feature was moved to a batch job and the real-time endpoint was never torn down. Delete it (or convert to a serverless/async endpoint that scales to zero if real-time is ever needed again).
Why. A large share of NAT-gateway processing is S3, ECR image pulls, and CloudWatch traffic routed out through NAT. Gateway endpoints (S3) and interface endpoints (ECR, CloudWatch, STS) keep that traffic on the AWS backbone — removing both the NAT data-processing charge and part of the cross-AZ transfer surfaced below.
Why. Brokers run at 28% CPU and well under half their storage throughput — the cluster is over-provisioned on brokers while the hot topic is under-partitioned (the §04 lag). Halving brokers and raising the positions topic from 6 → 18 partitions both cuts cost and resolves the consumer-lag bottleneck. Do the repartition first, validate lag, then remove brokers.
Why. 23 log groups never expire (years of debug logs retained), and per-vehicle custom metrics multiply cardinality. Set 30–90 day retention by group and aggregate the per-vehicle metrics to per-fleet dimensions. No observability lost that anyone uses.
Why. 1.2 TB of EBS volumes are unattached and a set of snapshots reference volumes that no longer exist — residue from past instance churn. Snapshot-then-delete the volumes; age out the orphan snapshots. Small, but free.
Spend that doesn't show up under a recognizable service line — it hides in "EC2-Other," duplicate environments, and usage-based services. ~$4,380/mo was not attributed to any resource before this review.
| Hidden cost | How it hides | Per month |
|---|---|---|
| Cross-AZ data transfer telemetry consumers ↔ brokers/DB across AZs | Billed as "EC2-Other," no resource tag | $2,300 |
| Forgotten eu-west-1 staging full env left after an EU migration | Spread across EC2/RDS/EKS in a second region | $1,100 |
| BigQuery full-table daily scan reporting job re-scans history nightly | On-demand bytes-scanned, no partition pruning | $900 |
| Idle load balancers 2 ALBs with zero healthy targets | Small per-ALB hourly charge, long forgotten | $44 |
| Unassociated Elastic IPs 4 EIPs not attached to anything | Charged only because they're unused | $36 |
Note: the cross-AZ and eu-west-1 items partly overlap with optimizations 03 and 06 — they are listed here to show where the money was hiding, not double-counted in the headline savings.
The rating weights each resource by spend, so the big, underused line items (prod compute, the idle GPU endpoint, the oversized writer) pull the overall score down hardest — which is exactly where the top three optimizations are aimed.
Where to start, and in what order. The numbers below map to the optimizations in §05. Four-fifths of the savings are low-risk quick wins you can capture this week without touching the customer-facing SLA; the structural work that follows pays for itself and clears the two High risks on the way.
● quick wins (low effort, no SLA impact) · ● structural (also fixes a High risk). The spend clusters top-left: the biggest savings are the easiest to capture.
Method. Meridian's AWS and GCP accounts were connected to Lizrd read-only. Lizrd built a knowledge graph of 437 resources, reconciled it against 30 days of billed spend and utilization telemetry, and surfaced the candidates above; each was then reviewed by hand for business context, risk, and sequencing. No changes were made to any environment — every recommendation is Meridian's to approve and implement, and Lizrd tracks realized savings as they land.
This is a sample deliverable for a fictional company. Want one for your infrastructure? Book a review at calendly.com/georg-lizrd/30min.
Produced with Lizrd · lizrd.ai · Confidential — Sample