๐Ÿค– AI Tools
ยท 5 min read

Claude Sonnet 5.5 vs Sonnet 5: Pricing, API and Migration


Claude Sonnet 5.5 succeeds Claude Sonnet 5 without changing the headline $2 input and $10 output token rates. The newer model is Anthropicโ€™s current Sonnet choice for coding and agent workflows, but it changes enough thinking and tool behavior that production teams should not treat migration as a string replacement.

This comparison separates what stayed the same from what needs new tests.

Quick verdict

Choose Sonnet 5.5 for new integrations and for existing systems that pass the migration checklist. Keep Sonnet 5 temporarily when an agent depends on older thinking controls, forced tool selection, preserved message history, an older computer-use tool or an advisor configuration that Sonnet 5.5 no longer accepts.

Sonnet 5 remains a valid historical/version-specific model owner. It is not the current generation.

Sonnet 5.5 vs Sonnet 5 specs

FeatureSonnet 5.5Sonnet 5
API model IDclaude-sonnet-5-5claude-sonnet-5
Release dateSeptember 28, 2026June 30, 2026
Context window1,000,0001,000,000
Maximum output128,000128,000
Input price$2/M$2/M
Cache read$0.20/M$0.20/M
5-minute cache write$2.50/M$2.50/M
Output price$10/M$10/M
ThinkingAdaptive, changed disable behaviorAdaptive, can disable thinking
Default effortHighHigh
PositionCurrent Sonnet generationPrevious Sonnet generation

The specifications explain why the migration looks simple. The runtime behavior explains why it is not.

Pricing: the same rates, not necessarily the same bill

Both models cost $2 per million input tokens and $10 per million output tokens. Cache reads are $0.20 and standard 5-minute cache writes are $2.50 per million tokens. Sonnet 5.5 also documents a $4 per million 1-hour cache write and 50% batch discounts.

Anthropic says Sonnet 5.5 is more than 30% faster and can cost up to 30% less per task than Sonnet 5. Those are vendor workload claims, not guaranteed savings. A real bill depends on token use, cache hits, tool-call count, retries and reviewer effort.

Use a fixed evaluation dataset and measure cost per successful task. The LLM regression-testing guide explains how to preserve those cases across model changes.

Thinking behavior is the main compatibility boundary

Sonnet 5 permits disabled thinking. Sonnet 5.5 interprets thinking-off configurations through between_tools behavior, which can reason between tool calls at high effort or below. That mode is not available at xhigh or max, and effort cannot change midway through the preserved conversation.

Applications that expose a simple thinking toggle should not assume it maps to identical behavior. Test latency, block ordering, token accounting and UI rendering.

Both models reject non-default temperature, top_p and top_k settings. Remove inherited sampling controls rather than assuming a client silently ignores them.

Tool-use changes

Four tool boundaries deserve explicit tests:

  1. Forced tools: Sonnet 5.5 can error when a specific tool is forced. Build a recovery path instead of relying on unconditional tool selection.
  2. Thinking blocks: text between tool calls can appear in thinking blocks. Render and log those blocks according to the API contract.
  3. Computer use: Claude API and Google Cloud integrations must move from the old computer_20251124 tool to the current toolset. AWS has a different compatibility boundary.
  4. Advisor models: Sonnet 5 is not accepted as an advisor model in the newer configuration, alongside Opus 4.7 and Opus 4.8.

These changes matter most in coding agents, browser agents and workflows that preserve long message histories. The AI Application Architecture hub covers retries, idempotency and reliable tool orchestration.

Benchmarks and performance

Anthropicโ€™s launch table reports 70.6% for Sonnet 5.5 on Terminal-Bench 4 at the stated setting, compared with 10.3% for Sonnet 5. The vendor also reports speed and cost-per-task improvements.

The gap is large enough to justify evaluation, but not enough to skip it. Benchmark scaffolding, effort, tools and scoring rules affect the result. A repository-specific coding agent can fail for reasons a public benchmark does not test, including credentials, flaky tests and unsupported tool output.

Treat Anthropicโ€™s figures as first-party evidence of the intended improvement. Do not present them as independently reproduced results.

Context and long-document work

Both models support a 1M-token context window and up to 128K output. The larger number does not eliminate context engineering. Long prompts raise latency and can bury critical instructions.

Use prompt caching for stable prefixes, retrieve only relevant repository or document sections, and test long-context accuracy separately. The best long-context models guide compares architectural tradeoffs without assuming every maximum window is equally useful.

Provider availability

Sonnet 5.5 is available through Anthropic and supported cloud platforms including AWS, Google Cloud and Microsoft Foundry. Provider identifiers and tool versions differ. Sonnet 5 also shipped through these channels, but lifecycle timing is provider-specific.

GitHub Copilot availability is another separate surface. Sonnet 5.5 is generally available for eligible paid Copilot plans through a gradual rollout, with administrator policy controls for managed accounts. That does not imply an Anthropic API entitlement.

Migration checklist

  • inventory every place that sets the model ID
  • remove unsupported custom sampling values
  • test thinking-off and between_tools behavior
  • exercise forced-tool errors and fallbacks
  • inspect text and thinking blocks between tool calls
  • update the computer-use tool by provider
  • replace invalid advisor configurations
  • run structured-output and schema tests
  • compare latency, retries and successful-task cost
  • retain Sonnet 5 as a temporary rollback target

The Claude Sonnet 5.5 guide provides the full current model reference, while the Sonnet 5 guide preserves historical specifications and integrations.

When should you remain on Sonnet 5?

Keep Sonnet 5 temporarily when compatibility work is incomplete or a provider has not exposed the required Sonnet 5.5 feature in your region. Do not stay solely because list pricing is familiar: the prices are already the same.

Use a controlled rollout. Route a small share of representative traffic to Sonnet 5.5, compare quality and operational metrics, then expand only after error handling and review workflows remain stable. The canary-deployment guide provides a production pattern.

My take

Sonnet 5.5 is the sensible default candidate, but the word candidate matters. Anthropic kept the pricing and context envelope stable while changing the agent runtime. That is good for economics and less comfortable for compatibility. Teams that test the whole workflow should benefit. Teams that only change the model ID can create subtle failures.

Specifications and migration boundaries were checked against Anthropicโ€™s Sonnet 5.5 model overview and prompting guidance.

Frequently asked questions

Is Sonnet 5.5 more expensive than Sonnet 5?

No. Both use $2 per million input tokens and $10 per million output tokens at standard rates.

Is Sonnet 5.5 a drop-in replacement?

No. Thinking, tool choice, response blocks, computer-use tooling and advisor compatibility can require changes.

Do both models have 1M context?

Yes. Both support a 1 million token context window and up to 128,000 output tokens.

Should a new project use Sonnet 5?

Usually no. A new integration should start by evaluating Sonnet 5.5 unless a provider or compatibility requirement blocks it.

Should the Sonnet 5 URL redirect to Sonnet 5.5?

No. Sonnet 5 has its own version-specific API, history, benchmarks and migration intent, so both pages should remain available.