Platform release reliability
Protecting Shared Helm Chart Releases
Kyverno Chainsaw tests that validate a shared application chart through upgrades, canaries, and rollback, catching multiple regressions before publication.
Professional case study
Multiple regressions caught before release
Every PR + publication gate
Same test suite locally and in CI
Problem
Developers rely on an existing shared Helm chart to deploy images into EKS with configurable platform defaults for Argo Rollouts canaries, Datadog integration and monitoring, labels, KEDA autoscaling, and Istio configuration and HTTP routing. Faulty rollout and Kubernetes Service-selector configuration in patch releases had caused production outages. With developers advised to allow patch updates within major/minor version constraints, validating compatibility before publication was especially important.
Action
Implemented declarative, extensible end-to-end tests with Kyverno Chainsaw, supporting targeted runs and the same test suite on chart developers’ machines and in CI. Integrated the tests into every PR and as a release gate before new chart versions are published.
Outcome
Caught multiple regressions before shipping, particularly selector-label changes that could break API endpoints. Established repeatable validation of chart upgrades, canary progression, traffic, workload health, and rollback before publication.
engineering takeaways
Reusable patterns from the work.
These notes focus on the engineering judgment, tradeoffs, and patterns behind the work.
- Test version transitions as well as installations: the matrix starts from selected releases one and five patches back, the previous minor, and the previous major. Each run installs and checks its baseline, upgrades to the target, progresses through canaries, checks HTTP traffic and workload health, then rolls back to the baseline and checks again.
- Verify application behavior through HTTP probes and workload-health checks. This helps catch regressions such as selector-label changes that disrupt routing to application endpoints.
- Make failures reproducible locally with the same declarative suite used in CI. Support targeted test runs and extend the suite as new regression scenarios emerge.
stack
contact