🆕 Updated July 30, 2026: GPT-5.6 Luna drops to $0.20/$1.20 (-80%) and Terra to $2/$12 (-20%). Luna is now the cheapest capable model by a huge margin. See GPT-5.6 price drop analysis.
July 17, 2026: Kimi K3 at $3/$15 (with 90% cache hits effectively ~$0.30 input) and Meta Muse Spark 1.1 at $1.25/$4.25 join the mid-tier war. K3 hits 88.3% Terminal-Bench (near Sol). See K3 pricing breakdown and Muse Spark 1.1 guide.
AI API pricing has dropped 85% since GPT-4 launched in 2023. Frontier model input now costs under $3/1M tokens. But prices vary wildly between providers, and the cheapest model isn’t always the cheapest for your use case. For the full set of pricing breakdowns and cost-optimization guides, see our APIs & pricing hub.
Here’s every major provider’s pricing in one place.
Frontier models
🆕 July 27, 2026: Claude Opus 5 launched at $5/$25 — same price as Opus 4.8, double the performance, within 0.5% of Fable 5 on coding.
June 10, 2026: Claude Fable 5 launched — Anthropic’s most powerful model ever at $10/$50 per million tokens. It’s a new premium tier above Opus, scoring 95% on SWE-bench Verified. See our full pricing analysis.
| Model | Input/1M | Output/1M | Context | Provider |
|---|---|---|---|---|
| Claude Fable 5 🆕 | $10.00 | $50.00 | 1M | Anthropic |
| GPT-5.4 | $2.50 | $15.00 | 1M | OpenAI |
| GPT-5.4 Pro | $30.00 | $180.00 | 1M | OpenAI |
| Claude Opus 5 | $5.00 | $25.00 | 1M | Anthropic |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | Anthropic |
| Claude Opus 4.6 | $15.00 | $75.00 | 1M | Anthropic |
| Claude Sonnet 4 | $3.00 | $15.00 | 200K | Anthropic |
| Grok 4.7 | $2.00 below 200K input | $6.00 below 200K input | 500K | xAI |
| Gemini 3.6 Flash 🆕 | $1.50 | $7.50 | 1M | |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1M | |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M+ | |
| Gemini 2.5 Pro | $1.25 | $5.00 | 1M |
Mid-tier models (best value)
| Model | Input/1M | Output/1M | Context | Provider |
|---|---|---|---|---|
| GPT-5.6 Terra 🆕 | $2.00 | $12.00 | 1M | OpenAI |
| GPT-5.4 Mini | $0.75 | $4.50 | 400K | OpenAI |
| Gemini 3.5 Flash-Lite 🆕 | $0.30 | $2.50 | 1M | |
| Claude Haiku 4 | $0.25 | $1.25 | 200K | Anthropic |
| Gemini 2.5 Flash | $0.15 | $0.60 | 1M | |
| Mistral Large 2 | $2.00 | $6.00 | 128K | Mistral |
| Mistral Small | $0.10 | $0.30 | 32K | Mistral |
Budget models (cheapest)
| Model | Input/1M | Output/1M | Context | Provider |
|---|---|---|---|---|
| GPT-5.6 Luna 🆕 | $0.20 | $1.20 | 1M | OpenAI |
| DeepSeek V4-Flash | $0.14 | $0.28 | 1M | DeepSeek |
| DeepSeek V4-Pro | $1.74 | $3.48 | 1M | DeepSeek |
| DeepSeek Chat | $0.14 | $0.28 | 128K | DeepSeek |
| DeepSeek Reasoner | $0.55 | $2.19 | 128K | DeepSeek |
| GPT-5 Mini | $0.25 | $2.00 | 128K | OpenAI |
| Qwen 3.6 Plus | Free tier | Free tier | 1M | Alibaba |
| Qwen 3.6 Flash | $0.25 | $1.00 | 1M | Alibaba |
| Qwen 3.6 Max Preview | $1.50 | $6.00 | 1M | Alibaba |
| GLM-5.1 | Free tier | Free tier | 128K | Z.ai |
Open-source via hosting providers
New (April 24, 2026): DeepSeek V4-Flash at $0.14/$0.28 per 1M tokens is now the cheapest frontier-class model available — scoring 79.0% on SWE-bench with 1M context. V4-Pro at $1.74/$3.48 scores 80.6%. Both MIT licensed. See V4 Flash: cheapest frontier model.
| Model | Provider | Input/1M | Output/1M |
|---|---|---|---|
| Llama 4 Maverick | Together AI | ~$0.49 | ~$0.49 |
| Llama 4 Scout | Together AI | ~$0.10 | ~$0.10 |
| Qwen 3.5 27B | Fireworks | ~$0.20 | ~$0.20 |
| Any model | Ollama (local) | $0 | $0 |
| Any model | OpenRouter | Varies | Varies |
Tired of API costs? Self-hosting can be cheaper long-term. See how to self-host an LLM for €4.99/month or deploy on Vultr with $250 free credits.
Cost per common task
What does a typical developer task actually cost?
| Task | Tokens (approx) | GPT-5.4 | GPT-5.4 Mini | DeepSeek | Local |
|---|---|---|---|---|---|
| Code review (1 file) | 3K in / 1K out | $0.02 | $0.007 | $0.001 | $0 |
| Bug fix | 5K in / 2K out | $0.04 | $0.01 | $0.002 | $0 |
| Full feature | 10K in / 5K out | $0.10 | $0.03 | $0.005 | $0 |
| Codebase analysis | 50K in / 5K out | $0.20 | $0.06 | $0.01 | $0 |
| 8-hour coding session | 500K in / 100K out | $2.75 | $0.83 | $0.10 | $0 |
Monthly cost estimates
| Usage level | GPT-5.4 | GPT-5.4 Mini | DeepSeek | Local |
|---|---|---|---|---|
| Light (10 tasks/day) | $6-12 | $2-4 | $0.30-0.60 | $0 |
| Medium (50 tasks/day) | $30-60 | $10-20 | $1.50-3 | $0 |
| Heavy (200 tasks/day) | $120-240 | $40-80 | $6-12 | $0 |
The smart approach: model routing
Don’t use one model for everything. Route by task complexity:
async def route(message, complexity):
if complexity == "simple":
return await call("deepseek-chat", message) # $0.001
elif complexity == "medium":
return await call("gpt-5.4-mini", message) # $0.01
else:
return await call("claude-sonnet-4", message) # $0.04
This gives you frontier quality when you need it and near-zero cost when you don’t. See our model routing guide for implementation details.
Batch API pricing: the discount most budgets miss
Three of the four major providers sell a second, cheaper lane for work that doesn’t need a real-time response. It’s easy to miss because it’s not on the main pricing table — you have to opt into a separate batch endpoint.
| Provider | Batch discount | Notes |
|---|---|---|
| OpenAI | 50% off sync rate | Fixed 24-hour completion window, often faster |
| Anthropic | 50% off sync rate | ≤24h window, most batches finish in under an hour |
| Google (Gemini) | 50% off sync rate | Stacks with introductory pricing on Gemini 3.6/3.7 Flash |
| xAI (Grok) | 20% off, 4 older models only | Flagship Grok 4.6 has no batch discount at all |
Verified directly against each vendor’s own pricing documentation: Google’s Gemini API pricing page states “Batch API (50% cost reduction),” OpenAI’s Batch API guide states a “50% cost discount compared to synchronous APIs,” and Anthropic’s Message Batches documentation prices all batch usage at 50% of standard rates. For Claude Sonnet 5 specifically, that puts the batch rate at $1/$5 per million input/output tokens against the $2/$10 standard rate.
How batch differs from standard API pricing: same model, same output quality — you give up streaming and a real-time response in exchange for the discount. OpenAI’s batch jobs complete within a fixed 24-hour window; Anthropic’s typically finish in under an hour but can take up to 24. Anthropic’s batch discount also stacks with prompt caching, so a high-cache-hit batch workload pays 50% off an already-discounted cached-read rate.
When to use batch processing: anything without a human waiting on the response. Bulk classification, tagging, embeddings, nightly eval runs, and scheduled report generation are the clearest fits — paying the synchronous rate for this kind of work is a standing 2x overspend. Interactive chat, coding-agent tool loops, and anything using streaming should stay on the synchronous API; batch requests explicitly reject stream: true.
Practical cost-saving implication: if you’re running any kind of bulk or offline job — dataset labeling, log analysis, content generation queues — check whether your provider’s batch endpoint applies before assuming the standard per-token rate is your floor. On a large enough job, that’s the difference between a $200 bill and a $100 one for identical output.
| Provider | What’s free | Limitation |
|---|---|---|
| Gemini CLI | Full CLI access | Rate limited |
| Qwen 3.6 Plus | API access | Rate limited |
| GLM-5.1 | Z.ai free tier | Quota limited |
| OpenRouter | Several free models | Model-dependent |
| Ollama | Unlimited local | Hardware-dependent |
See our free AI coding tier review for real-world testing of each.
Price trends
Input token costs have dropped ~85% since 2023. The trend continues as competition intensifies between OpenAI, Anthropic, Google, and Chinese providers. Expect another 30-50% drop by end of 2026.
The biggest driver: open-source models (GLM-5.1, Qwen 3.6, Llama 4) matching proprietary quality at zero cost, forcing paid providers to lower prices.
Related: AI Coding Tools Pricing · FinOps for AI · AI Agent Cost Management · Monitor AI API Spending · OpenRouter Complete Guide · Cheapest AI Coding Setup · Tested Every Free AI Tier
-
Cloud Hosting Pricing Compared — Railway vs Cloudways vs Hetzner vs Vultr vs RunPod (2026)
-
Qwen 3.7 Flash vs Gemini 3.6 Flash: $0.03 vs $1.50 Per Million Tokens
-
Best AI APIs for Startups — Free Tiers and Pricing Compared (2026)
-
Grok 4.5 Pricing and Cursor Integration: What Developers Need to Know
For real-time pricing updates and a cost calculator, see APIpulse which tracks 90+ models across all providers.