portfolio

Professional case studies and public projects.

How I investigate production systems, build reusable tooling, and evaluate platform changes, alongside selected personal projects.

platforminfraautomationdatapersonal

Developer tooling

AWS Network Path Tracing

A Go service that turns 30–60 minutes of manual AWS route investigation into minutes.

Problem
Tracing connectivity across accounts meant manually following VPC route tables, Transit Gateways, and Cloud WAN. A typical investigation took 30–60 minutes.
Outcome
Reduced typical manual investigations from 30–60 minutes to minutes with a repeatable tracing workflow.
GoAWS NetworkingTransit GatewayCloud WAN
Read case study

Production incident analysis

When Rightsizing Meets Production Latency

Root-cause analysis of EKS CPU saturation that connected workload sizing, autoscaling, and .NET runtime behavior.

Problem
A production EKS workload experienced CPU saturation at approximately 15 load per core. Understanding the failure required investigating the interaction between resource sizing, scaling behavior, and the application runtime.
Outcome
Designed a reusable rightsizing policy for latency-sensitive workloads to address the failure mechanism and reduce the risk of recurrence across critical workloads.
EKSHPARightsizing.NET
Read case study

Platform release reliability

Protecting Shared Helm Chart Releases

Kyverno Chainsaw tests that validate a shared application chart through upgrades, canaries, and rollback, catching multiple regressions before publication.

Problem
Developers rely on an existing shared Helm chart to deploy images into EKS with configurable platform defaults for Argo Rollouts canaries, Datadog integration and monitoring, labels, KEDA autoscaling, and Istio configuration and HTTP routing. Faulty rollout and Kubernetes Service-selector configuration in patch releases had caused production outages. With developers advised to allow patch updates within major/minor version constraints, validating compatibility before publication was especially important.
Outcome
Caught multiple regressions before shipping, particularly selector-label changes that could break API endpoints. Established repeatable validation of chart upgrades, canary progression, traffic, workload health, and rollback before publication.
Kyverno ChainsawHelmEKSArgo Rollouts
Read case study

Architecture and resilience

Disaster Recovery Dependency Mapping

Repeatable dependency discovery across 200+ components, later reused to map six applications in less than a day.

Problem
Disaster-recovery planning needed a dependency view across more than 200 platform and application components, with evidence spread across repositories, infrastructure, and AWS cost data.
Outcome
Mapped dependencies across 200+ components for disaster-recovery planning. Later reused the tooling to map dependencies for six applications in less than one day.
Disaster RecoveryC4StructurizrRepository Analysis
Read case study

Cost and technical judgment

EKS Auto Mode: The No-Go Decision

An evaluation of EKS Auto Mode for 10+ clusters and 1,000+ workloads, weighing migration feasibility, reliability, and compute cost.

Problem
Adopting EKS Auto Mode required evidence that networking, load balancing, storage, and controller migration could meet the needs of the existing platform at an acceptable cost.
Outcome
Recommended a no-go after identifying EBS migration and controller-coexistence blockers alongside a 10–12% compute-cost increase in this evaluation. The recommendation weighed reliability, migration effort, and cost.
AWS EKSEKS Auto ModeEBSCost Analysis
Read case study

Kubernetes operations

Understanding Karpenter Node Churn

A production investigation into consolidation, node churn, and the constraints that block disruption.

Problem
Karpenter consolidation and node churn across production EKS clusters needed deeper analysis to understand disruption blockers and the effect of scheduling constraints.
Outcome
Identified disruption blockers and translated the findings into tracked platform improvements.
KarpenterEKSPDBsTopology Spread
Read case study

Storage migration

Automating the Move from gp2 to gp3

An automated Kubernetes storage migration completed with minimal downtime and zero data loss.

Problem
All in-use gp2 EBS-backed Kubernetes storage needed to move to gp3 while preserving application data and minimizing downtime.
Outcome
Completed migration of all in-use gp2 storage to gp3 with minimal downtime and zero data loss, then enforced the new storage choice for future workloads.
KubernetesAWS EBSgp3PVCs
Read case study

Reliability engineering

Kubernetes Capacity Governance

Python tooling that scanned Kubernetes namespace quotas, HPA limits, and Karpenter capacity before scale risks turned into incidents.

Problem
Namespace quotas could silently block workloads from scaling to HPA maximums during peak traffic. Existing checks did not give teams an actionable view before the risk mattered.
Outcome
Warned 45+ teams before peak-traffic risk became incidents and validated quota suggestions against Karpenter capacity limits.
PythonKubernetesHPAResourceQuota
Read case study

Deployment reliability

Progressive Delivery Hardening

Improved rollout observability, rollback behavior, and autoscaling validation around Helm, Argo Rollouts, and KEDA.

Problem
Canary releases, rollback workflows, and autoscaling configurations had edge cases that could confuse developers or make operational signals harder to trust.
Outcome
Made progressive delivery workflows safer, reduced developer context switching, and improved confidence in rollout and autoscaling behavior.
HelmHelmfileArgo RolloutsKEDA
Read case study

projects

Selected personal projects.

These are public projects and experiments that show continued learning across web, infrastructure, data, and automation.

SubsGuard

May 2025 — Present

  • Built a privacy-first SaaS expense and subscription tracker for CSV imports, optional read-only US bank connections, recurring spend, budgets, shared expenses, and payback tracking.

  • Designed consent-first AI review workflows for transaction categorization, subscription grouping, transfer detection, split suggestions, and financial insights before data is saved.

  • Added shared-money workflows including one-off splits, saved people, always-split rules for recurring transactions, payback matching, and household-level settlement tracking.

CloudflareClerkStripeAI-assisted categorizationCSV workflows

Personal Website

February 2025 — Present

  • My website built with react/next.js for all things about me.

  • Templates and styling is done using tailwind css.

  • Hosted in Cloudflare

Route53TypeScriptCloudflare

HTTP Status Codes Learner with SMS

August 2025 — August 2025

  • A scheduled cloudflare worker that sends SMS to my personal number every x days of a random HTTP Status Code

  • To help me learn these status codes

TypeScriptCloudflare Worker

Investment Analysis Platform

October 2022 — Present

  • Developed a Python-based website (streamlit) for investment analysis using financial reporting API and SEC Database.

  • SEC Filings data is stored in MongoDB, retrievable via a built-in SEC Scraper.

  • Included a DCF calculator with custom inputs based on DCF parameters.

Python

NBA Data Scraper Package (PyPI)

October 2023 — January 2024

  • A tool built in Python to scrape a basketball website for shot data from NBA games (using BeautifulSoup).

  • Tested and published to PyPI using GitHub actions and poetry.

Python

NBA API

July 2023 — September 2023

  • Designed a database (MS SQL) to store NBA shots data (+4.5 million rows self-scraped) — shots taken, teams, players, positions, team arenas, and games played from 2000-2023.

  • Developed endpoints using FastAPI to access the database.

Python