Skip to content
Kernforce

Case study · anonymised

Making a savings number smaller

A European SaaS platform on AWS, two EKS clusters, a five-figure annual cloud bill. Automated detectors produced a headline savings figure. Reconciling them against each other removed roughly three quarters of it — and that is what made the remainder usable.

What the tools said

The estate was already instrumented: cost data by usage type, a metrics agent in both clusters, and a set of detectors producing findings across compute, storage, databases, networking and security. Summed naively, they produced a large and encouraging number.

It would not have survived the first technical review, for three separate reasons.

Three ways a savings figure inflates

1. The same money, counted from two directions

Most of the cluster nodes flagged as idle were managed by an autoscaler that sizes the fleet from pod requests. Their cost and the separately measured over-request of the workloads running on them are the same money, reached from two ends. Both appeared in the total. We folded them into one finding, with the measured request overhead as the low end of the range and the node-level estimate as the high end — and said explicitly what the gap between the two represents.

2. Advice the reader cannot carry out

The same findings recommended resizing those nodes. On an autoscaled cluster that is not an action anyone can take: the node type follows the requests, and a node deleted by hand returns within a minute. The recommendation was rewritten to name the lever that exists — cut the requests, let consolidation remove the node — and the finding was reclassified as evidence rather than as an independent saving.

3. A window that does not match its label

One family of findings reported a multi-month total under a monthly label, inflating those lines several-fold. It was caught by the oldest check there is: every figure in the report has to reconcile to a line on the bill for a named month. Those findings were rebuilt from the billing data directly, and the defect was fixed at the source rather than worked around.

What the client got

  • A range rather than a maximum, with the basis for each end stated — roughly a quarter of the naive total, and defensible line by line.
  • Every figure traceable to a bill line and a named month, so their engineers could check it without us.
  • A sequenced plan: what to do first, what it depends on, and what to do only after the fleet has stopped changing shape.
  • An explicit list of what we could not evaluate and what data would close each gap — including the parts of the estate we had no access to.
  • Two findings marked as costing nothing and worth doing anyway, because the risk was worse than the spend.

Why we tell it this way round

A large savings figure is easy to produce and expensive to defend. The moment someone senior asks an engineer to check it, every double-count and every unexecutable recommendation becomes a reason to distrust the whole exercise — including the findings that were real.

We would rather hand over a smaller number that holds. It is worth more, because it gets acted on.

Want the same review of your estate?

Fixed scope, fixed price, delivered as a report you keep — not a dashboard login you stop opening after a fortnight.