Production incident analysis

When Rightsizing Meets Production Latency

Root-cause analysis of EKS CPU saturation that connected workload sizing, autoscaling, and .NET runtime behavior.

Professional case study

Root cause identified

Vendor collaboration

Reusable policy designed

Problem

A production EKS workload experienced CPU saturation at approximately 15 load per core. Understanding the failure required investigating the interaction between resource sizing, scaling behavior, and the application runtime.

Action

Led root-cause analysis and identified a structural interaction between workload rightsizing, Kubernetes requests, horizontal pod autoscaler (HPA) behavior, and .NET thread-pool exhaustion. Worked with the vendor on the findings.

Outcome

Designed a reusable rightsizing policy for latency-sensitive workloads to address the failure mechanism and reduce the risk of recurrence across critical workloads.

engineering takeaways

Reusable patterns from the work.

These notes focus on the engineering judgment, tradeoffs, and patterns behind the work.

  • Investigated resource requests, autoscaling, and runtime behavior together.
  • Used production findings to inform vendor collaboration.
  • Designed a rightsizing policy around the needs of latency-sensitive workloads.

stack

EKSHPARightsizing.NETIncident Analysis

contact

Talk platform engineering, reliability, or developer tooling.