Technology
Karpenter
Karpenter changes where the decisions live. Once it is running, the node fleet is a consequence of pod requests and NodePool constraints rather than something anyone chooses directly — which is excellent, and which invalidates most node-level advice written before it existed.
What we think, and why
Deleting a node is not a saving.
It comes back, because the pods that were on it still need somewhere to go. The fleet shrinks when the requests shrink or when consolidation is allowed to do its job, and not otherwise.
Broad requirements are a feature, not an oversight.
A NodePool that permits many instance families lets Karpenter find the cheap shape of the moment. Narrowing it to a familiar type is a common instinct and usually costs money for no reliability gain.
Consolidation needs permission to work.
An idle node that never goes away is usually held by something explicit: a policy that forbids disruption, a budget set to zero, an annotation, or a pod with no controller. That is a finding, not a mystery.
Burstable instances hide their failure mode.
A T-family node that cannot sustain its baseline does not report an error. It throttles, or it silently buys surplus credits that appear on the bill under a name nobody reads. Fit is measured against sustained capacity, not nominal vCPU.
What we do with it
- NodePool and EC2NodeClass design: requirements, capacity types, limits, expiry
- Consolidation and disruption policy tuned against what the workloads tolerate
- Diagnosing nodes that will not go away, and nodes that go away too eagerly
- Cost attribution per NodePool, so a pool's price is visible next to its purpose