Kubernetes CrashLoopBackOff for AI Workloads: Models, GPUs and Probes
CrashLoopBackOff means a container repeatedly exits and Kubernetes is delaying the next restart. It is not the root cause. For an AI workload, the real failure is usually visible in previous logs, the termination reason, events or model-server startup sequence.
First five commands
kubectl get pod POD -n NAMESPACE -o wide
kubectl describe pod POD -n NAMESPACE
kubectl logs POD -n NAMESPACE --all-containers
kubectl logs POD -n NAMESPACE --all-containers --previous
kubectl get events -n NAMESPACE --sort-by=.lastTimestamp
--previous is critical: the current container may not yet show the error that killed its predecessor.
Inspect termination state:
kubectl get pod POD -n NAMESPACE \
-o jsonpath='{range .status.containerStatuses[*]}{.name}{"\t"}{.lastState.terminated.reason}{"\t"}{.lastState.terminated.exitCode}{"\n"}{end}'
Do not repeatedly delete the pod before capturing logs and events.
Triage by evidence
| Evidence | Likely direction |
|---|---|
OOMKilled, often exit 137 | Memory limit, host pressure or model size |
| Missing model/config file | Volume, mount path, image or init-container failure |
| Authentication exception | Missing Secret, ConfigMap or service identity |
| Probe failures | Startup, liveness or readiness timing |
| CUDA/device error | GPU scheduling, device plugin, driver or runtime mismatch |
| Process exits with code 0 | A batch command runs under a controller expecting a service |
| Connection refused | Dependency ordering or unavailable backend |
Exit codes are clues, not universal diagnoses. Pair them with termination reason and application logs.
Model loading exceeds memory
An inference container may load weights, allocate KV cache and initialize CUDA before opening its port. If that exceeds the limit, Kubernetes restarts it before readiness succeeds.
kubectl describe pod POD -n NAMESPACE
kubectl top pod POD -n NAMESPACE --containers
Fix the workload rather than merely raising a limit:
- choose an appropriate model or quantization;
- reduce concurrency, context length or cache allocation;
- separate download/conversion jobs from serving;
- set requests and limits from observed startup and steady state;
- confirm node capacity and eviction pressure.
Use the Kubernetes OOMKilled guide and AI workload memory guide for deeper diagnosis.
Models or caches are missing
Check mounts, claims and init containers:
kubectl get pod POD -n NAMESPACE -o yaml
kubectl get pvc -n NAMESPACE
kubectl logs POD -n NAMESPACE -c INIT_CONTAINER --previous
Verify the mount path, storage permissions, free bytes and inodes, incomplete downloads and concurrent access to shared writable caches. Host capacity problems are covered in Linux disk full on AI servers.
GPU discovery and runtime failures
Requesting a GPU is only one part of the path. The node needs compatible hardware, drivers and a device plugin; the image and model runtime must support the exposed device.
kubectl describe node NODE
kubectl get pod POD -n NAMESPACE -o jsonpath='{.spec.nodeName}{"\n"}'
kubectl describe pod POD -n NAMESPACE
Check for unsatisfied GPU requests, driver/library mismatches, unsupported architecture or a process that assumes a GPU exists when none was assigned.
Probes kill a slow-starting model server
Large models may need minutes to load. A liveness probe that starts too early can create the crash loop. Use a startup probe for initialization, readiness for accepting traffic and liveness only for a process that cannot recover without restart.
startupProbe:
httpGet:
path: /health/startup
port: 8080
periodSeconds: 10
failureThreshold: 60
readinessProbe:
httpGet:
path: /health/ready
port: 8080
periodSeconds: 5
livenessProbe:
httpGet:
path: /health/live
port: 8080
periodSeconds: 20
Tune from measured startup behavior. Do not make liveness depend on an external model provider; an upstream outage should not restart every pod.
Secrets and configuration
List names and referencesβnot values:
kubectl describe pod POD -n NAMESPACE
kubectl get secret -n NAMESPACE
kubectl get configmap -n NAMESPACE
Common failures include a renamed provider key, missing namespace-specific Secret, invalid endpoint or service account without dependency access. Validate configuration before starting and return a clear fatal error without printing credentials.
Debug and verify safely
When the image has no shell or crashes immediately, use Kubernetes ephemeral-container debugging where supported rather than replacing the production command blindly.
After the fix:
kubectl rollout status deployment/DEPLOYMENT -n NAMESPACE
kubectl get pods -n NAMESPACE -w
kubectl get pod POD -n NAMESPACE -o jsonpath='{.status.containerStatuses[*].restartCount}'
Prevention
- Test images with production-like model and memory settings.
- Separate startup, readiness and liveness semantics.
- Emit structured startup stages without secrets.
- Alert on restart rate and termination reason.
- Use immutable image tags and controlled model revisions.
- Budget disk, RAM and GPU capacity for rollouts.
- Run evaluations and smoke checks before shifting traffic.
Connect this procedure to AI Kubernetes troubleshooting, Docker vs Kubernetes for AI applications and the AI Operations hub.