Kubernetes operations
Understanding Karpenter Node Churn
A production investigation into consolidation, node churn, and the constraints that block disruption.
Professional case study
Production clusters investigated
Disruption blockers identified
Improvements tracked
Problem
Karpenter consolidation and node churn across production EKS clusters needed deeper analysis to understand disruption blockers and the effect of scheduling constraints.
Action
Led a deep dive into pod disruption budgets (PDBs), topology spread, disruption windows, and instance selection across production EKS clusters.
Outcome
Identified disruption blockers and translated the findings into tracked platform improvements.
engineering takeaways
Reusable patterns from the work.
These notes focus on the engineering judgment, tradeoffs, and patterns behind the work.
- Examined disruption controls alongside scheduling constraints.
- Included instance selection in the investigation of consolidation and churn.
- Converted operational findings into concrete follow-up work.
stack
KarpenterEKSPDBsTopology Spread
contact