🤖 AI Tools
· 7 min read
Last updated on

AI API Pricing Compared: Every Provider in One Table (2026)


🆕 Updated July 30, 2026: GPT-5.6 Luna drops to $0.20/$1.20 (-80%) and Terra to $2/$12 (-20%). Luna is now the cheapest capable model by a huge margin. See GPT-5.6 price drop analysis.

July 17, 2026: Kimi K3 at $3/$15 (with 90% cache hits effectively ~$0.30 input) and Meta Muse Spark 1.1 at $1.25/$4.25 join the mid-tier war. K3 hits 88.3% Terminal-Bench (near Sol). See K3 pricing breakdown and Muse Spark 1.1 guide.

AI API pricing has dropped 85% since GPT-4 launched in 2023. Frontier model input now costs under $3/1M tokens. But prices vary wildly between providers, and the cheapest model isn’t always the cheapest for your use case. For the full set of pricing breakdowns and cost-optimization guides, see our APIs & pricing hub.

Here’s every major provider’s pricing in one place.

Frontier models

🆕 July 27, 2026: Claude Opus 5 launched at $5/$25 — same price as Opus 4.8, double the performance, within 0.5% of Fable 5 on coding.

June 10, 2026: Claude Fable 5 launched — Anthropic’s most powerful model ever at $10/$50 per million tokens. It’s a new premium tier above Opus, scoring 95% on SWE-bench Verified. See our full pricing analysis.

ModelInput/1MOutput/1MContextProvider
Claude Fable 5 🆕$10.00$50.001MAnthropic
GPT-5.4$2.50$15.001MOpenAI
GPT-5.4 Pro$30.00$180.001MOpenAI
Claude Opus 5$5.00$25.001MAnthropic
Claude Opus 4.8$5.00$25.001MAnthropic
Claude Opus 4.6$15.00$75.001MAnthropic
Claude Sonnet 4$3.00$15.00200KAnthropic
Grok 4.7$2.00 below 200K input$6.00 below 200K input500KxAI
Gemini 3.6 Flash 🆕$1.50$7.501MGoogle
Gemini 3.5 Flash$1.50$9.001MGoogle
Gemini 3.1 Pro$2.00$12.001M+Google
Gemini 2.5 Pro$1.25$5.001MGoogle

Mid-tier models (best value)

ModelInput/1MOutput/1MContextProvider
GPT-5.6 Terra 🆕$2.00$12.001MOpenAI
GPT-5.4 Mini$0.75$4.50400KOpenAI
Gemini 3.5 Flash-Lite 🆕$0.30$2.501MGoogle
Claude Haiku 4$0.25$1.25200KAnthropic
Gemini 2.5 Flash$0.15$0.601MGoogle
Mistral Large 2$2.00$6.00128KMistral
Mistral Small$0.10$0.3032KMistral

Budget models (cheapest)

ModelInput/1MOutput/1MContextProvider
GPT-5.6 Luna 🆕$0.20$1.201MOpenAI
DeepSeek V4-Flash$0.14$0.281MDeepSeek
DeepSeek V4-Pro$1.74$3.481MDeepSeek
DeepSeek Chat$0.14$0.28128KDeepSeek
DeepSeek Reasoner$0.55$2.19128KDeepSeek
GPT-5 Mini$0.25$2.00128KOpenAI
Qwen 3.6 PlusFree tierFree tier1MAlibaba
Qwen 3.6 Flash$0.25$1.001MAlibaba
Qwen 3.6 Max Preview$1.50$6.001MAlibaba
GLM-5.1Free tierFree tier128KZ.ai

Open-source via hosting providers

New (April 24, 2026): DeepSeek V4-Flash at $0.14/$0.28 per 1M tokens is now the cheapest frontier-class model available — scoring 79.0% on SWE-bench with 1M context. V4-Pro at $1.74/$3.48 scores 80.6%. Both MIT licensed. See V4 Flash: cheapest frontier model.

ModelProviderInput/1MOutput/1M
Llama 4 MaverickTogether AI~$0.49~$0.49
Llama 4 ScoutTogether AI~$0.10~$0.10
Qwen 3.5 27BFireworks~$0.20~$0.20
Any modelOllama (local)$0$0
Any modelOpenRouterVariesVaries

Tired of API costs? Self-hosting can be cheaper long-term. See how to self-host an LLM for €4.99/month or deploy on Vultr with $250 free credits.

Cost per common task

What does a typical developer task actually cost?

TaskTokens (approx)GPT-5.4GPT-5.4 MiniDeepSeekLocal
Code review (1 file)3K in / 1K out$0.02$0.007$0.001$0
Bug fix5K in / 2K out$0.04$0.01$0.002$0
Full feature10K in / 5K out$0.10$0.03$0.005$0
Codebase analysis50K in / 5K out$0.20$0.06$0.01$0
8-hour coding session500K in / 100K out$2.75$0.83$0.10$0

Monthly cost estimates

Usage levelGPT-5.4GPT-5.4 MiniDeepSeekLocal
Light (10 tasks/day)$6-12$2-4$0.30-0.60$0
Medium (50 tasks/day)$30-60$10-20$1.50-3$0
Heavy (200 tasks/day)$120-240$40-80$6-12$0

The smart approach: model routing

Don’t use one model for everything. Route by task complexity:

async def route(message, complexity):
    if complexity == "simple":
        return await call("deepseek-chat", message)      # $0.001
    elif complexity == "medium":
        return await call("gpt-5.4-mini", message)       # $0.01
    else:
        return await call("claude-sonnet-4", message)     # $0.04

This gives you frontier quality when you need it and near-zero cost when you don’t. See our model routing guide for implementation details.

Batch API pricing: the discount most budgets miss

Three of the four major providers sell a second, cheaper lane for work that doesn’t need a real-time response. It’s easy to miss because it’s not on the main pricing table — you have to opt into a separate batch endpoint.

ProviderBatch discountNotes
OpenAI50% off sync rateFixed 24-hour completion window, often faster
Anthropic50% off sync rate≤24h window, most batches finish in under an hour
Google (Gemini)50% off sync rateStacks with introductory pricing on Gemini 3.6/3.7 Flash
xAI (Grok)20% off, 4 older models onlyFlagship Grok 4.6 has no batch discount at all

Verified directly against each vendor’s own pricing documentation: Google’s Gemini API pricing page states “Batch API (50% cost reduction),” OpenAI’s Batch API guide states a “50% cost discount compared to synchronous APIs,” and Anthropic’s Message Batches documentation prices all batch usage at 50% of standard rates. For Claude Sonnet 5 specifically, that puts the batch rate at $1/$5 per million input/output tokens against the $2/$10 standard rate.

How batch differs from standard API pricing: same model, same output quality — you give up streaming and a real-time response in exchange for the discount. OpenAI’s batch jobs complete within a fixed 24-hour window; Anthropic’s typically finish in under an hour but can take up to 24. Anthropic’s batch discount also stacks with prompt caching, so a high-cache-hit batch workload pays 50% off an already-discounted cached-read rate.

When to use batch processing: anything without a human waiting on the response. Bulk classification, tagging, embeddings, nightly eval runs, and scheduled report generation are the clearest fits — paying the synchronous rate for this kind of work is a standing 2x overspend. Interactive chat, coding-agent tool loops, and anything using streaming should stay on the synchronous API; batch requests explicitly reject stream: true.

Practical cost-saving implication: if you’re running any kind of bulk or offline job — dataset labeling, log analysis, content generation queues — check whether your provider’s batch endpoint applies before assuming the standard per-token rate is your floor. On a large enough job, that’s the difference between a $200 bill and a $100 one for identical output.

ProviderWhat’s freeLimitation
Gemini CLIFull CLI accessRate limited
Qwen 3.6 PlusAPI accessRate limited
GLM-5.1Z.ai free tierQuota limited
OpenRouterSeveral free modelsModel-dependent
OllamaUnlimited localHardware-dependent

See our free AI coding tier review for real-world testing of each.

Input token costs have dropped ~85% since 2023. The trend continues as competition intensifies between OpenAI, Anthropic, Google, and Chinese providers. Expect another 30-50% drop by end of 2026.

The biggest driver: open-source models (GLM-5.1, Qwen 3.6, Llama 4) matching proprietary quality at zero cost, forcing paid providers to lower prices.

Related: AI Coding Tools Pricing · FinOps for AI · AI Agent Cost Management · Monitor AI API Spending · OpenRouter Complete Guide · Cheapest AI Coding Setup · Tested Every Free AI Tier

For real-time pricing updates and a cost calculator, see APIpulse which tracks 90+ models across all providers.