βš™οΈ AI Operations
Β· 2 min read
Last updated on

Too Many Open Files in AI Services: Diagnose FD Exhaustion


Too many open files, EMFILE or ENFILE means a process or the system cannot allocate another file descriptor. Sockets, pipes, files, event streams and watchers all consume descriptors. AI services can exhaust them through high concurrency, leaked model streams, tool subprocesses or connection pools.

Raising a limit may restore capacity, but it does not fix a leak.

Measure limits and live use

ulimit -Sn
ulimit -Hn
cat /proc/PID/limits | grep -i 'open files'
ls /proc/PID/fd | wc -l
lsof -p PID | head
cat /proc/sys/fs/file-nr

Identify descriptor types:

lsof -p PID | awk '{print $5}' | sort | uniq -c | sort -nr | head
ss -s

Compare use over time. A count that rises after each request and never returns suggests a leak; a count that tracks legitimate concurrency suggests a capacity or pooling problem.

AI workload failure patterns

  • streaming responses remain open after client cancellation;
  • provider SDK connections are recreated instead of pooled;
  • agent tool subprocess pipes are not closed;
  • uploaded files or model shards remain open;
  • vector database or Redis connections leak;
  • file watchers run inside a production image;
  • retries multiply concurrent sockets;
  • sidecars and proxies have lower limits than the application.

Trace one request through the AI Application Architecture to find which boundary owns closure and cancellation.

Fix the lifecycle first

Use structured cleanup (finally, context managers or equivalent), bounded connection pools and cancellation propagation. Test disconnects, provider timeouts and agent-tool failures; happy-path tests rarely expose descriptor leaks.

For long-lived streams, define idle timeouts and close both upstream and downstream resources when either side ends. Monitor active streams and sockets separately from completed requests.

Set limits at the real runtime boundary

Shell ulimit changes do not automatically configure systemd services or containers. Inspect and set the limit where the process starts, then restart through the normal deployment process.

For systemd, use an explicit LimitNOFILE override. For containers and Kubernetes, confirm node, runtime and process limits rather than assuming a pod resource limit controls file descriptors.

Increase limits only after estimating legitimate peak concurrency and host capacity. Very high limits can postpone failure until memory, ports or an upstream service becomes the next bottleneck.

Validate under failure conditions

Load-test a representative request pattern in an authorized environment, including:

  • client cancellation during model streaming;
  • provider timeout and reset;
  • repeated agent tool failures;
  • rollout overlap between replicas;
  • connection-pool saturation.

Verify descriptor counts plateau and return toward baseline. Alert on utilization relative to the process limit, not only the terminal EMFILE error.

Use AI Operations for production monitoring, connection timeout troubleshooting for the next network layer and AI Testing & Evaluation to preserve the failure as a regression scenario.

The correct outcome is bounded resource use with a limit sized for expected trafficβ€”not an unlimited process that can consume the host.