πŸ€– AI Tools
Β· 6 min read
Last updated on

Claude Sonnet 5 Complete Guide: Pricing, API, and Benchmarks (2026)


Anthropic released Claude Sonnet 5 on June 30, 2026. It brought a 1 million token context window, adaptive thinking, stronger agentic coding, and lower token pricing to the Sonnet tier. Its API model ID is claude-sonnet-5.

Successor notice: Claude Sonnet 5.5 is the newer Sonnet generation. It keeps the same headline token prices but changes thinking and tool behavior, so it is not a model-ID-only migration. See Sonnet 5.5 vs Sonnet 5 before switching a production integration.

This page remains the version-specific owner for Sonnet 5. It documents the model as released, rather than turning into a rolling page for whichever Sonnet is newest.

Claude Sonnet 5 quick specs

SpecClaude Sonnet 5
API model IDclaude-sonnet-5
Release dateJune 30, 2026
Context window1,000,000 tokens
Maximum output128,000 tokens
Input price$2 per 1M tokens
Output price$10 per 1M tokens
Cache read$0.20 per 1M tokens
Standard 5-minute cache write$2.50 per 1M tokens
ThinkingAdaptive, with thinking disable support
Default effortHigh
Input and outputText and images to text

Sonnet 5 launched with $2/$10 pricing that was initially described as introductory. Anthropic later cancelled the planned increase to $3/$15, leaving $2 input and $10 output as the standard rate.

What Sonnet 5 was built for

Sonnet 5 was Anthropic’s value model for coding, agents, computer use, and everyday professional work. It narrowed the capability gap between Sonnet and the more expensive Opus tier while remaining practical for repeated development tasks.

The main use cases were:

  • repository-scale code analysis
  • multi-step bug fixing
  • terminal and browser agents
  • long-document analysis
  • tool-calling applications
  • high-volume professional workflows

Anthropic described Sonnet 5 as more agentic than earlier Sonnet models. That description referred to planning, sustained tool use, and verification behavior. It did not remove the need for runtime permissions, tests, or human review.

Benchmarks and what they mean

Anthropic reported 63.2% on SWE-bench Pro and 81.2% on OSWorld for Sonnet 5 at launch. Those results were first-party evaluations and depend on the published harness, effort setting, and agent scaffolding.

They showed that Sonnet 5 could approach more expensive models on some coding and computer-use workloads. They did not prove that the model was universally better or cheaper for every repository. A model that uses more tokens or retries can have a higher completed-task cost even when its token rate is lower.

Keep historical Sonnet 5 results attached to this model. Sonnet 5.5 has different benchmark results and runtime behavior, even though both models share $2/$10 token pricing.

Pricing and prompt caching

The standard Sonnet 5 API rates are:

Billing itemPrice per 1M tokens
Input$2.00
Cache read$0.20
5-minute cache write$2.50
Output$10.00

Prompt caching is useful for stable system instructions, large codebase context, and repeated evaluation inputs. Cache pricing only helps when the prefix remains reusable. Frequently changing tool output and conversation history still create new input.

Read Claude Sonnet 5 pricing explained for the original cost model and AI API pricing compared for cross-provider rates.

The tokenizer change

Sonnet 5 introduced an updated tokenizer. Anthropic documented that the same source text could produce approximately 30% more tokens, with the exact change depending on content. That meant teams migrating from Sonnet 4.6 needed to recount prompts and revisit output budgets.

The lower token rates partly offset the increased token count, but they did not make every workload exactly the same price. Repository content, generated code, and long agent traces can all change the real bill.

Thinking and effort controls

Sonnet 5 uses adaptive thinking and supports effort controls. It defaults to high effort on the Claude Platform. Unlike Sonnet 5.5, Sonnet 5 accepts disabled thinking at any effort level.

Higher effort can improve difficult tasks but usually increases latency and token use. Use the Sonnet 5 effort-level guide to choose an explicit setting rather than leaving cost behavior accidental.

Applications using manual extended-thinking token budgets from older Claude integrations had to migrate to adaptive thinking. Non-default temperature, top_p, or top_k values return an error on Sonnet 5, so older sampling code also needed review.

Tool use and coding agents

Sonnet 5 supports tool use, vision, PDF processing, Files API workflows, prompt caching, and batch processing. For coding agents it can inspect repositories, call terminals, edit files, and operate browsers when the surrounding product exposes those tools.

The model does not grant itself permissions. Claude Code, cloud platforms, and custom agents each determine available tools, filesystem access, network access, and credentials. The AI Security hub and agent sandboxing guide cover those controls.

Setup guides remain version-specific:

Provider availability

Sonnet 5 shipped through the Claude API and supported cloud channels, including Amazon Bedrock, Google Cloud, and Microsoft Foundry. Provider model identifiers, regions, billing, and feature timing can differ from Anthropic’s direct API.

Third-party routers are another separate layer. A provider may expose Sonnet 5 with its own routing, limits, or availability lifecycle. Always verify the provider-specific model record rather than assuming Anthropic-direct behavior.

Sonnet 5 versus Sonnet 5.5

Sonnet 5.5 is the successor, but several core dimensions stay the same: $2/$10 token prices, a 1M context window, and up to 128K standard output. The newer model adds different adaptive-thinking behavior, changed tool semantics, newer computer-use tooling on some providers, and stronger vendor-reported performance.

The model IDs are different:

Sonnet 5:   claude-sonnet-5
Sonnet 5.5: claude-sonnet-5-5

Sonnet 5.5 requires migration work for applications that disable thinking, force a specific tool, parse text between tool calls, use the older computer-use tool, or use Sonnet 5 as an advisor model. That is why production users should follow the dedicated migration comparison.

When should you keep Sonnet 5 temporarily?

Remain on Sonnet 5 long enough to validate an integration when:

  • your client assumes the first response block is text
  • your UI renders text between tool calls
  • you force tool selection with tool_choice
  • you rely on the older computer-use tool version
  • you switch models within preserved-thinking conversations
  • you have not recalibrated latency and task-cost budgets

This is a temporary compatibility reason, not a claim that Sonnet 5 is generally the better model. New integrations should normally evaluate Sonnet 5.5 first.

My take

Sonnet 5 was a major value release because it paired frontier-style agent behavior with $2/$10 token pricing. It remains relevant for existing integrations and historical comparisons, but it should no longer be presented as the current Sonnet generation.

Teams should preserve Sonnet 5 as a fallback while they test Sonnet 5.5’s changed thinking and tool behavior. Once those differences are handled, the newer model is the logical default candidate.

Frequently asked questions

What is the Claude Sonnet 5 API model ID?

The direct Claude API model ID is claude-sonnet-5.

How much does Sonnet 5 cost?

It costs $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 and standard 5-minute cache writes cost $2.50 per million tokens.

Does Sonnet 5 have a 1M context window?

Yes. Its standard context window is 1 million tokens, with up to 128,000 output tokens.

Is Sonnet 5 the newest Sonnet model?

No. Claude Sonnet 5.5 is the newer generation and has a different API model ID and migration requirements.

Is Sonnet 5.5 a drop-in replacement?

Not for every integration. Several thinking, tool-use, response-block, computer-use, and advisor behaviors changed. Retest the full agent loop before switching production traffic.