Production incident analysis
When Rightsizing Meets Production Latency
Root-cause analysis of EKS CPU saturation that connected workload sizing, autoscaling, and .NET runtime behavior.
Professional case study
Root cause identified
Vendor collaboration
Reusable policy designed
Problem
A production EKS workload experienced CPU saturation at approximately 15 load per core. Understanding the failure required investigating the interaction between resource sizing, scaling behavior, and the application runtime.
Action
Led root-cause analysis and identified a structural interaction between workload rightsizing, Kubernetes requests, horizontal pod autoscaler (HPA) behavior, and .NET thread-pool exhaustion. Worked with the vendor on the findings.
Outcome
Designed a reusable rightsizing policy for latency-sensitive workloads to address the failure mechanism and reduce the risk of recurrence across critical workloads.
engineering takeaways
Reusable patterns from the work.
These notes focus on the engineering judgment, tradeoffs, and patterns behind the work.
- Investigated resource requests, autoscaling, and runtime behavior together.
- Used production findings to inform vendor collaboration.
- Designed a rightsizing policy around the needs of latency-sensitive workloads.
stack
contact