Rate Limiting AI APIs: Tokens, Requests, Quotas, and Cost Control
β100 requests per minuteβ is not enough for an AI application. One request may classify a sentence; another may send a huge context, invoke tools, and stream thousands of output tokens. They consume different provider capacity and create very different costs.
Production rate limiting usually combines requests, tokens, concurrency, tenant quotas, and spend.
Five limits that solve different problems
| Limit | Protects | Example |
|---|---|---|
| Requests per interval | API and control-plane capacity | 60 requests/minute per user |
| Input/output tokens | Model quota and cost | 200,000 tokens/day per workspace |
| Concurrent generations | Runtime capacity and latency | 4 active generations per user |
| Job queue depth | Background workers | 100 queued jobs per tenant |
| Monetary budget | Business exposure | $50 model spend/month per workspace |
Use the smallest set that corresponds to real failure modes. Showing a nominal request limit while leaving token and concurrency use unlimited provides false reassurance.
Separate provider limits from product limits
Providers may expose requests per minute, tokens per minute, concurrent requests, or account-level quotas. Your product needs its own limits per user, tenant, feature, and model tier.
The application limit should normally remain below the upstream ceiling, leaving capacity for retries and internal work. If several services share one provider account, coordinate centrally through an AI gateway.
Token bucket for burstable request limits
A token bucket refills continuously. Each allowed request consumes capacity; short bursts are possible without allowing unlimited sustained traffic.
class TokenBucket {
constructor(capacity, refillPerSecond) {
this.capacity = capacity;
this.tokens = capacity;
this.refillPerSecond = refillPerSecond;
this.updatedAt = Date.now();
}
take(cost = 1) {
const now = Date.now();
const elapsed = (now - this.updatedAt) / 1000;
this.tokens = Math.min(
this.capacity,
this.tokens + elapsed * this.refillPerSecond
);
this.updatedAt = now;
if (this.tokens < cost) return false;
this.tokens -= cost;
return true;
}
}
The in-memory example explains the algorithm but is not a distributed production limiter. Multiple instances need shared atomic state or an infrastructure service.
Charge by estimated work
Not every request must cost one unit. Estimate input size before dispatch:
request cost = base unit
+ input token units
+ requested output allowance
+ premium model multiplier
Reconcile estimated usage with provider-reported usage after completion. Estimates protect capacity before execution; actual usage supports billing and quota accounting afterwards.
Do not trust a client-supplied token count. Compute or validate it server-side.
Concurrency limits matter most for latency
Ten simultaneous generations can overload a small worker or exhaust provider concurrency even when the per-minute request count looks safe.
Use semaphores or queues per tenant and globally. Interactive traffic may receive reserved capacity while batch work waits. Limit tool execution separately when tools hit databases, browsers, or third-party APIs.
Queue background work
When capacity is unavailable, choose intentionally between:
- rejecting with
429and a retry hint; - queueing a durable job;
- degrading to a cheaper model;
- reducing output limits;
- asking the user to upgrade.
Do not keep unlimited requests waiting inside web processes. Long AI workflows belong in durable queues with visible status. See Webhook Architecture for AI Workflows.
Return useful 429 responses
{
"error": {
"code": "tenant_token_quota_exceeded",
"message": "This workspace has reached its daily AI token allowance.",
"retryable": true,
"retryAfterSeconds": 1800,
"requestId": "req_01J..."
}
}
Use the standard Retry-After header where appropriate. Do not reveal other tenantsβ traffic or internal provider account limits.
Clients should apply backoff and jitter rather than retrying immediately. Coordinate this with AI API failure handling.
Cost controls are not ordinary rate limits
Token counts approximate provider cost only when model and pricing tier are known. Track normalized usage and estimated money separately.
Useful guardrails include:
- per-request maximum input and output tokens;
- model allowlists by plan;
- daily tenant soft warnings;
- monthly hard budgets;
- unusually rapid spend alerts;
- maximum agent steps and tool calls;
- approval before expensive jobs.
Cached input, batch pricing, reasoning tokens, media generation, and provider promotions can complicate estimates. Keep the calculation versioned and reconcile against provider invoices.
Multi-tenant fairness
One large customer should not starve every smaller tenant. Combine:
- per-user limits;
- per-tenant limits;
- global provider capacity;
- feature-specific queues;
- plan-based entitlements.
Make quota state visible so users can understand remaining capacity. Hidden throttling feels like random failure.
Failure behavior
Rate-limit storage can fail too. Decide which endpoints fail closed and which can degrade safely. Never silently disable budget enforcement for expensive autonomous agents because Redis is unavailable.
When provider headers report remaining quota, treat them as useful observationsβnot the only source of truth. Several workers may consume the same quota concurrently.
Observe the limiter
Track:
- allowed and rejected work by dimension;
- queue wait time;
- current concurrency;
- tokens and cost per tenant;
- provider
429rate; - retries caused by throttling;
- unused reserved capacity;
- budget overrides.
Production checklist
- Limit requests, tokens, concurrency, and money according to actual risks.
- Separate product entitlements from upstream provider quotas.
- Use atomic shared state for distributed enforcement.
- Queue work that can wait; reject work that cannot.
- Return stable error codes and retry guidance.
- Bound agent steps and tool calls.
- Reconcile estimated usage with provider data.
- Test limiter-store failure and provider throttling.
Rate limiting belongs inside AI Application Architecture and should share tenant identity with Authentication for AI Applications. For a practical implementation, continue with Rate Limiting AI API Requests.