Self-Healing LLM Inference Platform

It serves an LLM and heals itself: detect, diagnose, propose, review, reconcile.

AWS EKSvLLMNVIDIA T4AWS BedrockPrometheusAlertmanagerGrafanaArgoCDPython
View repository →

Problem

GPU-backed model serving fails in ways that page a human at 3am for a fix that is usually mechanical: a replica that never came back, a node taint, or a bad rollout. The diagnosis is the slow part, not the edit.

Architecture

Architecture diagram of the Self-Healing LLM Inference Platform showing vLLM on an EKS GPU node group, Prometheus and Alertmanager observability, a Python remediation agent calling AWS Bedrock, and ArgoCD reconciling merged manifests from GitHub
System architecture: serving path, observability, remediation, and GitOps reconcile.
Sequence diagram of the self-healing loop across vLLM, Prometheus, Alertmanager, the agent, the Kubernetes API, AWS Bedrock, GitHub, the engineer, and ArgoCD
The healing loop: alert fires, evidence gathered, Bedrock diagnoses, human merges, ArgoCD syncs.

Approach

vLLM serves Qwen2.5-0.5B-Instruct on an EKS GPU node group behind an OpenAI-compatible API, with Prometheus scraping /metrics and Alertmanager routing a firing VLLMDown to a Python remediation agent. The agent gathers evidence with read-only RBAC (pod phase plus the last 50 log lines), asks AWS Bedrock for a structured {root_cause, action, reason}, then applies a deterministic manifest edit and opens a fix PR. ArgoCD auto-syncs once a human merges.

Walkthrough of a real incident

  1. fault injected: --max-model-len 999999 committed to vllm.yaml
  2. vLLM enters CrashLoopBackOff, exit code 1
  3. VLLMCrashLooping alert fires, routed to remediation-agent
  4. agent reads logs from the crashed pod via the Kubernetes API
  5. Bedrock diagnosis: max_model_len 999999 exceeds model max position embeddings 32768
  6. action selected: lower_max_model_len
  7. PR opened: "agent: fix VLLMDown (lower max-model-len 999999 -> 4096)"

Evidence

EVIDENCE / ONE REAL INCIDENT, END TO END

ArgoCD application graph for llm-platform showing the vLLM pod in CrashLoopBackOff while the remediation-agent pod is healthy and sync status is OK.
ArgoCD: the vLLM pod in CrashLoopBackOff while the remediation agent stays healthy. Sync OK on the commit that injected the fault.
Prometheus platform.rules page showing the VLLMDown and VLLMCrashLooping alerting rules, both firing.
The detection rules. VLLMDown watches available replicas, VLLMCrashLooping watches container waiting reason.
Alertmanager alerts view showing VLLMCrashLooping and VLLMDown routed to the platform-webhook remediation-agent receiver.
Alertmanager routing the firing alerts to the remediation-agent webhook receiver.
Kubernetes pod detail for vllm-d56d5c496-mqbz2 in CrashLoopBackOff, last terminated with exit code 1.
The evidence the agent reads: CrashLoopBackOff, last terminated with exit code 1.
GitHub pull request #8 authored by the SRE agent, describing the root cause, selected action lower_max_model_len, and the Bedrock reasoning.
PR #8, authored by the agent. Root cause, selected action, and reasoning, diagnosed via AWS Bedrock.

Outcome

Remediation PRs open in under a minute, and unreviewed infrastructure edits stay at zero. The safety boundary is the design. The model only picks one action from a vetted four-action set. It never writes raw YAML. The edit itself is deterministic code, a pull request is the human gate, and Bedrock runs outside the failure domain it diagnoses, so the healer stays up when the patient doesn't.

Hard-won lessons

Lesson 01

A Kubernetes Service named `vllm` causes the cluster to auto-inject a VLLM_PORT environment variable in host:port form, which collides with vLLM's own VLLM_PORT config and crashes the engine at startup. The fix is setting VLLM_PORT explicitly on the container so the injected value is overridden.

Lesson 02

The terraform-aws-eks module v20 silently ignores a top-level disk_size attribute on a node group. Node disk must be configured through block_device_mappings instead, otherwise nodes come up with the default volume and the model download fills the disk.

Lesson 03

An NVIDIA T4 is compute capability 7.5 and cannot run bfloat16, so vLLM must be started with --dtype half.