βš™οΈ AI Operations
Β· 3 min read
Last updated on

Kubernetes CrashLoopBackOff for AI Workloads: Models, GPUs and Probes


CrashLoopBackOff means a container repeatedly exits and Kubernetes is delaying the next restart. It is not the root cause. For an AI workload, the real failure is usually visible in previous logs, the termination reason, events or model-server startup sequence.

First five commands

kubectl get pod POD -n NAMESPACE -o wide
kubectl describe pod POD -n NAMESPACE
kubectl logs POD -n NAMESPACE --all-containers
kubectl logs POD -n NAMESPACE --all-containers --previous
kubectl get events -n NAMESPACE --sort-by=.lastTimestamp

--previous is critical: the current container may not yet show the error that killed its predecessor.

Inspect termination state:

kubectl get pod POD -n NAMESPACE \
  -o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.lastState.terminated.reason}{"\t"}{.lastState.terminated.exitCode}{"\n"}{end}'

Do not repeatedly delete the pod before capturing logs and events.

Triage by evidence

EvidenceLikely direction
OOMKilled, often exit 137Memory limit, host pressure or model size
Missing model/config fileVolume, mount path, image or init-container failure
Authentication exceptionMissing Secret, ConfigMap or service identity
Probe failuresStartup, liveness or readiness timing
CUDA/device errorGPU scheduling, device plugin, driver or runtime mismatch
Process exits with code 0A batch command runs under a controller expecting a service
Connection refusedDependency ordering or unavailable backend

Exit codes are clues, not universal diagnoses. Pair them with termination reason and application logs.

Model loading exceeds memory

An inference container may load weights, allocate KV cache and initialize CUDA before opening its port. If that exceeds the limit, Kubernetes restarts it before readiness succeeds.

kubectl describe pod POD -n NAMESPACE
kubectl top pod POD -n NAMESPACE --containers

Fix the workload rather than merely raising a limit:

  • choose an appropriate model or quantization;
  • reduce concurrency, context length or cache allocation;
  • separate download/conversion jobs from serving;
  • set requests and limits from observed startup and steady state;
  • confirm node capacity and eviction pressure.

Use the Kubernetes OOMKilled guide and AI workload memory guide for deeper diagnosis.

Models or caches are missing

Check mounts, claims and init containers:

kubectl get pod POD -n NAMESPACE -o yaml
kubectl get pvc -n NAMESPACE
kubectl logs POD -n NAMESPACE -c INIT_CONTAINER --previous

Verify the mount path, storage permissions, free bytes and inodes, incomplete downloads and concurrent access to shared writable caches. Host capacity problems are covered in Linux disk full on AI servers.

GPU discovery and runtime failures

Requesting a GPU is only one part of the path. The node needs compatible hardware, drivers and a device plugin; the image and model runtime must support the exposed device.

kubectl describe node NODE
kubectl get pod POD -n NAMESPACE -o jsonpath='{.spec.nodeName}{"\n"}'
kubectl describe pod POD -n NAMESPACE

Check for unsatisfied GPU requests, driver/library mismatches, unsupported architecture or a process that assumes a GPU exists when none was assigned.

Probes kill a slow-starting model server

Large models may need minutes to load. A liveness probe that starts too early can create the crash loop. Use a startup probe for initialization, readiness for accepting traffic and liveness only for a process that cannot recover without restart.

startupProbe:
  httpGet:
    path: /health/startup
    port: 8080
  periodSeconds: 10
  failureThreshold: 60

readinessProbe:
  httpGet:
    path: /health/ready
    port: 8080
  periodSeconds: 5

livenessProbe:
  httpGet:
    path: /health/live
    port: 8080
  periodSeconds: 20

Tune from measured startup behavior. Do not make liveness depend on an external model provider; an upstream outage should not restart every pod.

Secrets and configuration

List names and referencesβ€”not values:

kubectl describe pod POD -n NAMESPACE
kubectl get secret -n NAMESPACE
kubectl get configmap -n NAMESPACE

Common failures include a renamed provider key, missing namespace-specific Secret, invalid endpoint or service account without dependency access. Validate configuration before starting and return a clear fatal error without printing credentials.

Debug and verify safely

When the image has no shell or crashes immediately, use Kubernetes ephemeral-container debugging where supported rather than replacing the production command blindly.

After the fix:

kubectl rollout status deployment/DEPLOYMENT -n NAMESPACE
kubectl get pods -n NAMESPACE -w
kubectl get pod POD -n NAMESPACE -o jsonpath='{.status.containerStatuses[*].restartCount}'

Prevention

  • Test images with production-like model and memory settings.
  • Separate startup, readiness and liveness semantics.
  • Emit structured startup stages without secrets.
  • Alert on restart rate and termination reason.
  • Use immutable image tags and controlled model revisions.
  • Budget disk, RAM and GPU capacity for rollouts.
  • Run evaluations and smoke checks before shifting traffic.

Connect this procedure to AI Kubernetes troubleshooting, Docker vs Kubernetes for AI applications and the AI Operations hub.

Primary documentation