Platform Engineer / Kubernetes / AWS / Developer Platforms

Platform Engineer making AWS and Kubernetes easier to debug, operate, and evolve.

At Just Eat Takeaway.com, I build network diagnostics and developer tooling, investigate production failures, and turn findings into safer platform workflows and informed migration decisions.

/usr/local/bin/lian
LZY

operator

Lian Zhen Yang

$ kubectl get impact

network diagnostics | quota analysis | DR mapping

$ helm status delivery

rollout observability, rollback fixes, KEDA validation

$ terraform plan platform

AWS, EKS, DNS, IAM, certificates, developer guardrails

case studies

Platform engineering case studies.

Faster network investigations, lessons from production incidents, automated release validation, and repeatable dependency discovery.

View all case studies

Developer tooling

AWS Network Path Tracing

A Go service that turns 30–60 minutes of manual AWS route investigation into minutes.

Problem
Tracing connectivity across accounts meant manually following VPC route tables, Transit Gateways, and Cloud WAN. A typical investigation took 30–60 minutes.
Outcome
Reduced typical manual investigations from 30–60 minutes to minutes with a repeatable tracing workflow.
GoAWS NetworkingTransit GatewayCloud WAN
Read case study

Production incident analysis

When Rightsizing Meets Production Latency

Root-cause analysis of EKS CPU saturation that connected workload sizing, autoscaling, and .NET runtime behavior.

Problem
A production EKS workload experienced CPU saturation at approximately 15 load per core. Understanding the failure required investigating the interaction between resource sizing, scaling behavior, and the application runtime.
Outcome
Designed a reusable rightsizing policy for latency-sensitive workloads to address the failure mechanism and reduce the risk of recurrence across critical workloads.
EKSHPARightsizing.NET
Read case study

Platform release reliability

Protecting Shared Helm Chart Releases

Kyverno Chainsaw tests that validate a shared application chart through upgrades, canaries, and rollback, catching multiple regressions before publication.

Problem
Developers rely on an existing shared Helm chart to deploy images into EKS with configurable platform defaults for Argo Rollouts canaries, Datadog integration and monitoring, labels, KEDA autoscaling, and Istio configuration and HTTP routing. Faulty rollout and Kubernetes Service-selector configuration in patch releases had caused production outages. With developers advised to allow patch updates within major/minor version constraints, validating compatibility before publication was especially important.
Outcome
Caught multiple regressions before shipping, particularly selector-label changes that could break API endpoints. Established repeatable validation of chart upgrades, canary progression, traffic, workload health, and rollback before publication.
Kyverno ChainsawHelmEKSArgo Rollouts
Read case study

Architecture and resilience

Disaster Recovery Dependency Mapping

Repeatable dependency discovery across 200+ components, later reused to map six applications in less than a day.

Problem
Disaster-recovery planning needed a dependency view across more than 200 platform and application components, with evidence spread across repositories, infrastructure, and AWS cost data.
Outcome
Mapped dependencies across 200+ components for disaster-recovery planning. Later reused the tooling to map dependencies for six applications in less than one day.
Disaster RecoveryC4StructurizrRepository Analysis
Read case study

capabilities

How I improve cloud platforms.

I connect production reliability, infrastructure automation, and developer experience to make everyday platform work easier.

Production Reliability

Root-cause analysis across workload sizing, autoscaling, and runtime behavior; Karpenter consolidation and disruption analysis.

EKSKubernetesHPAKarpenterDatadogPrometheus

AWS Networking & Infrastructure

Network path tracing across accounts, infrastructure automation, and validation of ingress hosts and certificates before deployment.

GoVPCTransit GatewayCloud WANTerraformTerragruntACM

Resilience & Migration

Repeatable dependency discovery, automated storage migration, and compute evaluations grounded in reliability, effort, and cost.

C4StructurizrAWS Cost AnalysisEBSEKS Auto Mode

Developer Platforms

Backstage workflows that automate EKS onboarding, service metadata, traceability, and documentation.

BackstageTypeScriptGoPythonReactNext.js

Deployment Guardrails

End-to-end Helm chart release gates, upgrade and rollback validation, pre-deployment checks, and clearer rollout signals.

HelmKyverno ChainsawHelmfileTerragruntArgo RolloutsKEDAACM

about

From process engineering to platform engineering.

I started in chemical engineering and data automation, where the work was always about understanding systems, finding bottlenecks, and making processes measurable. Platform engineering became the natural next step: the systems are cloud infrastructure and Kubernetes, the bottlenecks are developer friction and operational risk, and the best solutions are often small tools, clear guardrails, and documentation that lets teams move without waiting.