Hritik Munde.

Platform & DevOps Engineer crafting resilient cloud architectures.

Working across Kubernetes, AWS and GitOps. Currently focused on running LLM inference reliably, and on the automation that keeps a platform boring for the people who depend on it.

Sunnyvale, CAMS Computer Science, Indiana University

Reliability is a design decision, not a rota.

I have spent the last four years on the operations side of software, carrying on-call for a high-availability EKS platform at T-Mobile, migrating build systems that engineers wait on every day, and writing the tooling that makes a cluster answer questions about itself.

The pattern across my work is automating what does not need a human, and keeping a human where judgment matters. That is why my self-healing platform lets a model pick an action but never write the YAML, and why my pipelines verify signatures at admission rather than trusting the build that produced them.

Lately that has pulled me toward AI infrastructure: GPU scheduling, inference SLOs, and what "healthy" even means for a model server. It is the same reliability problem with a more expensive failure mode.

DevOps Engineer, Indiana University Bloomington
Sunnyvale, California
RHCSA , Certified
HashiCorp Terraform Associate , Certified
CKA , In progress, Jul 2026

Things I built, and why they work the way they do.

Self-Healing LLM Inference Platform

It serves an LLM and heals itself: detect, diagnose, propose, review, reconcile.

AWS EKSvLLMNVIDIA T4AWS BedrockPrometheusAlertmanagerGrafanaArgoCDPython
Under 1 min
detection to pull request
100%
of injected config faults resolved
4 modules, 8 manifests
Terraform and GitOps
Under 20 min
full rebuild from code

Alibi

Verifies AI-generated code by executing it, not by having another model read the diff. Runs the pre-change and post-change versions of a function in isolated cloud sandboxes and compares actual outputs.

CodexModalClaude-MemPython

Shipyard | Internal Developer Platform

Sandbox environments in three minutes instead of three days.

TerraformKubernetes (EKS)BackstageCrossplaneArgoCDGitOps

Zero-Trust CI/CD Pipeline

No long-lived cloud keys, no unsigned images, no unchecked manifests.

GitHub ActionsOIDCCosignOPA GatekeeperTrivySyft SBOMKubernetes
Repository link coming soon

Where I have done the work.

Aug 2025 to May 2026
Bloomington, IN

DevOps Engineer (Graduate Assistant)

Indiana University Bloomington

  • Provisioned Linux lab environments on AWS EC2 with Ansible automation for 200+ students per semester, cutting environment setup time by 70% and removing per-machine manual steps.
  • Automated grading-environment teardown and rebuild with Bash and cron against instance inventories, reducing recurring course operations effort by 60% each semester.
Jul 2022 to Jul 2024
Pune, India

DevOps Engineer

UST (Client: T-Mobile)

  • Migrated legacy Jenkins build and test jobs to GitHub Actions with parallel test execution across distributed runners and dependency caching, cutting average pipeline runtime by 65%.
  • Developed internal tooling in Go with client-go for orphaned-resource sweeps and namespace audits across clusters, replacing manual kubectl checks and saving 6+ engineer-hours weekly.
  • Standardized monitoring and logging on Prometheus, Grafana and Loki with SLO alert rules tied to runbooks, dropping incident detection from roughly 30 minutes to under 5.
  • Carried weekly on-call for a high-availability EKS platform, troubleshooting deploy failures and node incidents, cutting repeat incidents by 35% through postmortem fixes.
Aug 2021 to Mar 2022
Pune, India

Software Engineer Intern

Tata Motors

  • Engineered Java REST backend services for an internal workflow platform, adding response caching and pagination that cut data-retrieval latency by 40% on the heaviest queries.
  • Diagnosed slow SQL queries with execution plans to root-cause latency, restructuring joins and indexes across reporting services and improving dashboard load times by 25%.

The toolkit, grouped by what it is for.

PythonGoBash / ShellJavaSQLRust
Linux / Unix (RHEL, Debian/Ubuntu)systemdDNS, TLS, load balancingDistributed systems troubleshooting
Kubernetes (EKS)DockerHelmECR
GitHub ActionsJenkinsArgoCDGitOpsTerraformAnsibleBackstageCrossplane
PrometheusGrafanaLokiOpenTelemetryAlertmanagerSLOs & runbooks
OIDCRBAC & IRSACosignOPA GatekeeperTrivySyft SBOMPatch automation
vLLMGPU model serving (NVIDIA)AWS BedrockAIOps remediationInference SLOs
AWS (EC2, VPC, IAM, Lambda, CloudWatch)Azure (AKS)

Academics.

Aug 2024 to May 2026

Indiana University Bloomington

Master of Science, Computer Science · GPA 3.9 / 4.0

Aug 2018 to May 2022

Pune University

Bachelor of Technology, Computer Engineering · GPA 3.65 / 4.0

Let's talk about your infrastructure.

I am open to DevOps, Site Reliability and Platform Engineering roles. The fastest way to reach me is email. I read everything.

Download résumé (PDF)