πŸ—οΈ AI Application Architecture
Β· 4 min read
Last updated on

Rate Limiting AI APIs: Tokens, Requests, Quotas, and Cost Control


β€œ100 requests per minute” is not enough for an AI application. One request may classify a sentence; another may send a huge context, invoke tools, and stream thousands of output tokens. They consume different provider capacity and create very different costs.

Production rate limiting usually combines requests, tokens, concurrency, tenant quotas, and spend.

Five limits that solve different problems

LimitProtectsExample
Requests per intervalAPI and control-plane capacity60 requests/minute per user
Input/output tokensModel quota and cost200,000 tokens/day per workspace
Concurrent generationsRuntime capacity and latency4 active generations per user
Job queue depthBackground workers100 queued jobs per tenant
Monetary budgetBusiness exposure$50 model spend/month per workspace

Use the smallest set that corresponds to real failure modes. Showing a nominal request limit while leaving token and concurrency use unlimited provides false reassurance.

Separate provider limits from product limits

Providers may expose requests per minute, tokens per minute, concurrent requests, or account-level quotas. Your product needs its own limits per user, tenant, feature, and model tier.

The application limit should normally remain below the upstream ceiling, leaving capacity for retries and internal work. If several services share one provider account, coordinate centrally through an AI gateway.

Token bucket for burstable request limits

A token bucket refills continuously. Each allowed request consumes capacity; short bursts are possible without allowing unlimited sustained traffic.

class TokenBucket {
  constructor(capacity, refillPerSecond) {
    this.capacity = capacity;
    this.tokens = capacity;
    this.refillPerSecond = refillPerSecond;
    this.updatedAt = Date.now();
  }

  take(cost = 1) {
    const now = Date.now();
    const elapsed = (now - this.updatedAt) / 1000;
    this.tokens = Math.min(
      this.capacity,
      this.tokens + elapsed * this.refillPerSecond
    );
    this.updatedAt = now;
    if (this.tokens < cost) return false;
    this.tokens -= cost;
    return true;
  }
}

The in-memory example explains the algorithm but is not a distributed production limiter. Multiple instances need shared atomic state or an infrastructure service.

Charge by estimated work

Not every request must cost one unit. Estimate input size before dispatch:

request cost = base unit
             + input token units
             + requested output allowance
             + premium model multiplier

Reconcile estimated usage with provider-reported usage after completion. Estimates protect capacity before execution; actual usage supports billing and quota accounting afterwards.

Do not trust a client-supplied token count. Compute or validate it server-side.

Concurrency limits matter most for latency

Ten simultaneous generations can overload a small worker or exhaust provider concurrency even when the per-minute request count looks safe.

Use semaphores or queues per tenant and globally. Interactive traffic may receive reserved capacity while batch work waits. Limit tool execution separately when tools hit databases, browsers, or third-party APIs.

Queue background work

When capacity is unavailable, choose intentionally between:

  • rejecting with 429 and a retry hint;
  • queueing a durable job;
  • degrading to a cheaper model;
  • reducing output limits;
  • asking the user to upgrade.

Do not keep unlimited requests waiting inside web processes. Long AI workflows belong in durable queues with visible status. See Webhook Architecture for AI Workflows.

Return useful 429 responses

{
  "error": {
    "code": "tenant_token_quota_exceeded",
    "message": "This workspace has reached its daily AI token allowance.",
    "retryable": true,
    "retryAfterSeconds": 1800,
    "requestId": "req_01J..."
  }
}

Use the standard Retry-After header where appropriate. Do not reveal other tenants’ traffic or internal provider account limits.

Clients should apply backoff and jitter rather than retrying immediately. Coordinate this with AI API failure handling.

Cost controls are not ordinary rate limits

Token counts approximate provider cost only when model and pricing tier are known. Track normalized usage and estimated money separately.

Useful guardrails include:

  • per-request maximum input and output tokens;
  • model allowlists by plan;
  • daily tenant soft warnings;
  • monthly hard budgets;
  • unusually rapid spend alerts;
  • maximum agent steps and tool calls;
  • approval before expensive jobs.

Cached input, batch pricing, reasoning tokens, media generation, and provider promotions can complicate estimates. Keep the calculation versioned and reconcile against provider invoices.

Multi-tenant fairness

One large customer should not starve every smaller tenant. Combine:

  • per-user limits;
  • per-tenant limits;
  • global provider capacity;
  • feature-specific queues;
  • plan-based entitlements.

Make quota state visible so users can understand remaining capacity. Hidden throttling feels like random failure.

Failure behavior

Rate-limit storage can fail too. Decide which endpoints fail closed and which can degrade safely. Never silently disable budget enforcement for expensive autonomous agents because Redis is unavailable.

When provider headers report remaining quota, treat them as useful observationsβ€”not the only source of truth. Several workers may consume the same quota concurrently.

Observe the limiter

Track:

  • allowed and rejected work by dimension;
  • queue wait time;
  • current concurrency;
  • tokens and cost per tenant;
  • provider 429 rate;
  • retries caused by throttling;
  • unused reserved capacity;
  • budget overrides.

Production checklist

  • Limit requests, tokens, concurrency, and money according to actual risks.
  • Separate product entitlements from upstream provider quotas.
  • Use atomic shared state for distributed enforcement.
  • Queue work that can wait; reject work that cannot.
  • Return stable error codes and retry guidance.
  • Bound agent steps and tool calls.
  • Reconcile estimated usage with provider data.
  • Test limiter-store failure and provider throttling.

Rate limiting belongs inside AI Application Architecture and should share tenant identity with Authentication for AI Applications. For a practical implementation, continue with Rate Limiting AI API Requests.