βš™οΈ AI Operations
Β· 2 min read
Last updated on

Kubernetes Pod Evictions in AI Systems: Resource Pressure and Model Services


An evicted pod was removed because the node or cluster needed to reclaim resources. AI model services amplify common causes: large model caches consume disk, inference uses substantial memory, and verbose prompt or token logging can fill ephemeral storage.

Confirm the pressure signal

kubectl describe pod POD_NAME -n NAMESPACE
kubectl describe node NODE_NAME
kubectl get events -n NAMESPACE --sort-by=.lastTimestamp
kubectl top node NODE_NAME

Look for MemoryPressure, DiskPressure, PIDPressure, or ephemeral-storage messages. OOMKilled is a container memory-limit event; eviction is generally a node-level resource-management decision. The remedies differ.

Memory pressure

Set realistic requests so the scheduler understands the working set. A model server with tiny requests but large actual consumption can overcommit a node and displace other workloads. Bound concurrency, context length, and loaded models, and isolate memory-heavy services when necessary.

Disk and ephemeral-storage pressure

Container images, model downloads, caches, temporary media, and logs can exhaust node storage. Declare ephemeral-storage requests and limits, rotate logs, move durable artifacts to object storage, and manage model caches explicitly.

Do not repeatedly download large weights into an unbounded temporary directory on every restart.

Priority and disruption policy

Priority classes influence which pods survive pressure; they do not create capacity. Use them only with a documented workload hierarchy. PodDisruptionBudgets help with voluntary disruptions but do not prevent every pressure eviction.

For inference, maintain enough ready replicas or queue capacity that losing one pod does not drop accepted work. Agent and tool jobs must be idempotent because recovery may redeliver them.

Stabilization sequence

  1. Identify the exact node condition.
  2. Protect or drain accepted work where possible.
  3. Stop the source of uncontrolled memory, disk, logs, or processes.
  4. Correct requests, limits, cache retention, and placement.
  5. Add capacity only after the growth mechanism is understood.
  6. Verify recovery under representative load.

Deleting evicted pods only removes historical objects; it does not fix the node pressure that caused them.

Prevention

Monitor node allocatable capacity, working set, ephemeral storage, filesystem usage, image and model cache growth, evictions, restarts, ready replicas, queue depth, and request latency. Correlate infrastructure events with model versions and deployments.

Use Docker vs Kubernetes for AI applications for platform boundaries, OOMKilled guidance for container-limit failures, and AI Operations for reliability design.