Skip to content
Kernforce

Kubernetes management

Day-2 operation of EKS, GKE and self-managed clusters: upgrades, node lifecycle, autoscaling, capacity and the boring reliability work that decides whether a platform is trusted.

Day one was the easy part

A cluster is straightforward to create and expensive to keep. Version support windows expire on someone else's schedule, node pools drift away from what the workloads actually need, autoscalers make decisions nobody reviews, and the person who set it up has moved on. Nothing is broken, exactly — it is just that no one can say what would happen if a zone went away.

What we do

  • Version upgrades planned against the provider's support window, rehearsed on non-production first, executed without a maintenance night where the architecture allows it.
  • Node lifecycle: capacity types, instance families, consolidation policy, and disruption budgets that reflect what the workload can actually tolerate.
  • Autoscaling that is reviewed rather than assumed — both the cluster autoscaler or Karpenter layer and the HPAs sitting on top of it.
  • Capacity and headroom measured rather than guessed, including the DaemonSet overhead every node pays before a single application pod schedules.
  • Reliability hygiene: probes, pod disruption budgets, topology spread, priority classes — the settings that only matter on the day they matter.

How the work goes

  1. 01

    Inventory

    Two weeks of measurement before any change: what runs, what it requests, what it actually uses, and where the cluster has no headroom left.

  2. 02

    Stabilise

    The findings that carry risk go first — version end-of-life, single points of failure, workloads with no disruption budget, nodes that cannot survive a zone loss.

  3. 03

    Operate

    A standing arrangement: upgrades, capacity reviews and an on-call path, with a monthly written summary of what changed and what it cost.

What you get

  • A written cluster inventory: versions, node groups, workloads, and the gaps with severities
  • An upgrade plan with dates and rollback points
  • Monthly report of changes, incidents and capacity trend

A good fit when

  • Teams running production Kubernetes without a dedicated platform person
  • Clusters that grew organically and now nobody wants to touch
  • Companies whose auditor has started asking about upgrade policy

Not us, and we will say so

  • Clusters we are not allowed to observe — we do not operate what we cannot measure
  • Teams looking for someone to take the blame rather than the work

Tell us what you are running.

A first call is a conversation. If this is not the service you need, we will point you at the one that is — or tell you that you do not need us yet.

Book a call

Related