Self-Healing LLM Inference Platform
It serves an LLM and heals itself: detect, diagnose, propose, review, reconcile.
Portfolio
Platform & DevOps Engineer crafting resilient cloud architectures.
Working across Kubernetes, AWS and GitOps. Currently focused on running LLM inference reliably, and on the automation that keeps a platform boring for the people who depend on it.
I have spent the last four years on the operations side of software, carrying on-call for a high-availability EKS platform at T-Mobile, migrating build systems that engineers wait on every day, and writing the tooling that makes a cluster answer questions about itself.
The pattern across my work is automating what does not need a human, and keeping a human where judgment matters. That is why my self-healing platform lets a model pick an action but never write the YAML, and why my pipelines verify signatures at admission rather than trusting the build that produced them.
Lately that has pulled me toward AI infrastructure: GPU scheduling, inference SLOs, and what "healthy" even means for a model server. It is the same reliability problem with a more expensive failure mode.
It serves an LLM and heals itself: detect, diagnose, propose, review, reconcile.
Verifies AI-generated code by executing it, not by having another model read the diff. Runs the pre-change and post-change versions of a function in isolated cloud sandboxes and compares actual outputs.
Sandbox environments in three minutes instead of three days.
No long-lived cloud keys, no unsigned images, no unchecked manifests.
Indiana University Bloomington
UST (Client: T-Mobile)
Tata Motors
Master of Science, Computer Science · GPA 3.9 / 4.0
Bachelor of Technology, Computer Engineering · GPA 3.65 / 4.0
I am open to DevOps, Site Reliability and Platform Engineering roles. The fastest way to reach me is email. I read everything.
Download résumé (PDF)