βš™οΈ AI Operations
Β· 2 min read
Last updated on

Kubernetes OOMKilled for AI Workloads: Models, Memory and GPU Resources


State:      Terminated
Reason:     OOMKilled
Exit Code:  137

OOMKilled normally means the container exceeded its memory limit or the node exhausted memory. For AI workloads, model weights, runtime overhead, tokenization, KV cache, batch size, context length, and concurrent requests can all raise RAM use sharply.

Confirm the event

kubectl describe pod POD_NAME -n NAMESPACE
kubectl get pod POD_NAME -n NAMESPACE -o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'
kubectl top pod POD_NAME -n NAMESPACE --containers

Historical peaks may be missing after termination, so use your metrics system to inspect working set, limits, restarts, request concurrency, context length, and model-loading events.

RAM is not VRAM

Kubernetes memory requests and limits generally describe host RAM. GPU memory is managed through the accelerator, runtime, and model server. A pod can have free host RAM and still fail inside CUDA because VRAM is exhaustedβ€”or be OOMKilled by the kernel while GPU memory is available.

Identify the actual failure source before changing resources.

Estimate the working set

Account for:

  • model weights and quantization;
  • runtime and framework allocations;
  • CPU-side loading and preprocessing;
  • KV cache per active sequence;
  • maximum context length;
  • batch size and concurrent requests;
  • temporary buffers and fragmentation.

Measure representative prompts and concurrency. A model that starts successfully can still fail under long-context traffic.

Correct requests and limits

resources:
  requests:
    cpu: "2"
    memory: "12Gi"
  limits:
    memory: "16Gi"
    nvidia.com/gpu: "1"

Requests influence scheduling; limits constrain the container. Do not copy these example values into production. Set them from measured workload peaks plus an explicit safety margin.

Removing the limit can move the failure from one pod to the entire node. Raising it without matching node capacity can leave pods Pending or increase eviction pressure.

Reduce memory safely

  • Use a smaller or appropriately quantized model.
  • Bound context length, batch size, and concurrency.
  • Queue excess work instead of admitting every request.
  • Load only required models per worker.
  • Stream results without retaining unnecessary buffers.
  • Separate model serving from API and retrieval processes.

Changing quantization or context can change quality and behavior; validate it as a model change, not only an infrastructure fix.

Prevent recurrence

Alert on memory working set versus limit, OOM events, restarts, queue depth, active sequences, context distribution, latency, and GPU memory. Use readiness only after model load and remove an unhealthy replica before routing more requests to it.

See Docker vs Kubernetes for AI applications, Pending GPU pods, and AI Kubernetes troubleshooting. This incident belongs in the broader AI Operations capacity plan.