AI API Design: Streaming, Tool Calls, Jobs and Reliable Model Responses
An AI API is not an ordinary CRUD API with a prompt field added. A response may stream for a minute, request a tool, fail after consuming tokens, or complete asynchronously. Models and providers also change faster than most client integrations.
Good AI API design gives clients a stable contract without pretending model behavior is deterministic. This guide focuses on the boundaries an application owns: requests, events, jobs, errors, tool execution, usage, and versioning.
Start with an application contract
Do not expose a providerβs complete request body as your public contract. Define the tasks and capabilities your product supports:
{
"task": "support_reply",
"input": { "ticket_id": "t_123", "message": "..." },
"response_mode": "stream",
"metadata": { "request_id": "client_456" }
}
Your backend can translate this into OpenAI, Anthropic, or another provider format. That boundary lets you change models, enforce policy, and validate input without forcing every client to migrate.
Use provider-specific escape hatches sparingly. If clients depend on every upstream field, your gateway is only a proxy.
Separate synchronous, streaming, and job APIs
Use a synchronous response for short, bounded work. Use streaming when partial output improves the experience. Use an asynchronous job for long-running agents, media generation, batch work, or tasks that must survive a disconnected client.
POST /v1/responses -> completed response
POST /v1/streams -> server-sent event stream
POST /v1/jobs -> 202 + job ID
GET /v1/jobs/{id} -> current state
POST /v1/jobs/{id}/cancel -> cancellation request
Do not hold an HTTP connection open for work that may take many minutes. For job completion, combine polling with reliable webhooks.
Design streaming as an event protocol
A stream is more than fragments of text. Give every event a type and sequence number:
event: response.started
data: {"response_id":"resp_123","sequence":0}
event: output.delta
data: {"response_id":"resp_123","sequence":1,"text":"Hello"}
event: response.completed
data: {"response_id":"resp_123","sequence":2,"usage":{"input_tokens":81,"output_tokens":24}}
Document ordering, heartbeats, cancellation, reconnect behavior, and terminal events. Clients should not infer completion from a closed socket alone. If resumption is unsupported, say so explicitly.
Treat tool calls as proposed actions
A model-produced tool call is untrusted input, not authorization. Return a stable tool name and schema-validated arguments, then let application policy decide whether it may run.
{
"type": "tool_call",
"id": "call_789",
"name": "create_refund",
"arguments": { "order_id": "o_42", "amount": 25 }
}
The executor should authenticate the actor, validate arguments against live data, enforce scopes and limits, and require confirmation for consequential actions. Record the user, agent run, policy, tool, and result. See authentication for AI applications and AI Security.
Structured output needs semantic validation
JSON Schema can enforce types and required fields. It cannot prove a price is current, a citation supports a claim, or an identifier belongs to the signed-in tenant.
Validate in layers:
- Parse and schema-check the model output.
- Apply business rules and authorization.
- Resolve identifiers against trusted systems.
- Reject, repair, or request review when confidence is insufficient.
Never execute generated SQL, shell commands, URLs, or tool arguments merely because they are valid JSON.
Use a stable error envelope
Clients need to distinguish invalid input, policy rejection, provider throttling, timeouts, and internal failure.
{
"error": {
"code": "provider_rate_limited",
"message": "The model is temporarily unavailable.",
"retryable": true,
"request_id": "req_abc"
}
}
Do not leak raw provider errors, prompts, credentials, or internal topology. Use appropriate HTTP status codes, but keep a machine-readable application code because different upstream failures may map to the same status. The AI API failure guide covers retry and fallback policy.
Make unsafe retries impossible
Generation requests can consume money even when a client never receives the response. Tool actions can create duplicate side effects. Accept an idempotency key for operations that clients may safely retry, persist the result, and bind the key to the authenticated caller and request hash.
Read idempotency for AI workflows before adding automatic retries to agent actions.
Return usage without promising exact billing
Return the provider, actual model, latency, finish reason, and available usage fields. Keep estimated cost separate from authoritative billing:
{
"model": "capable",
"provider_model": "provider/model-version",
"usage": { "input_tokens": 840, "output_tokens": 216 },
"estimated_cost": { "amount": "0.0041", "currency": "USD" }
}
Pricing can vary by provider, cache behavior, region, tier, or promotion. Version your pricing data and disclose that estimates may differ from invoices. Enforce quotas at the gateway; see rate limiting AI APIs and building an AI gateway.
Version capabilities, not provider fashions
Prefer stable application concepts such as tools, structured_output, or reasoning_level over provider-specific flags. Publish a capability matrix when aliases differ. Add fields compatibly, provide deprecation dates, and test old clients against new model adapters.
Changing the model behind an alias can still change output quality or safety. Treat major model migrations as product changes even when the JSON schema stays identical.
Observability belongs in the contract
Assign a request ID at entry and carry it through gateway, provider, tool calls, queues, and webhooks. Track:
- tenant and application task;
- provider and resolved model;
- queue, first-token, and total latency;
- retries and fallbacks;
- input and output usage;
- tool calls and policy decisions;
- cancellation and terminal state.
Avoid logging full prompts and outputs by default. If content logging is necessary, define access, redaction, retention, deletion, and consent.
Production checklist
- The public contract describes application tasks, not arbitrary provider payloads.
- Sync, streaming, and asynchronous lifecycles are explicit.
- Every stream and job has terminal states and cancellation behavior.
- Tool calls pass schema, authorization, and policy checks.
- Errors are stable, redacted, and marked retryable only when safe.
- Retried side effects use idempotency keys.
- Usage and cost estimates identify their limits.
- Model aliases and capability changes are versioned and tested.
- Request IDs connect all components without exposing sensitive content.
This contract becomes the center of an AI application architecture: authentication at the edge, policy in the gateway, durable state for jobs, and explicit control over model behavior.