Claude Sonnet 4 vs Claude 3.5 Sonnet: Full Comparison

Claude Sonnet 4 outperforms Claude 3.5 Sonnet on nearly every published benchmark, most notably jumping to 72.7% on SWE-bench Verified versus 49% for the upgraded Claude 3.5 Sonnet, while both models now share the same retired status on Anthropic’s API. Claude 3.5 Sonnet launched in June 2024 and was retired on October 28, 2025.
Claude Sonnet 4 launched in May 2025 and was retired on June 15, 2026. For a related comparison, see AI Comparison.
This guide compares both models on features, performance, and pricing for historical and migration reference, and points to the active model Anthropic recommends in their place today.
Quick Comparison Table
Claude Sonnet 4 and Claude 3.5 Sonnet differ most in coding performance, context window, and current API availability. The table below summarizes the specifications each model shipped with while active. For a related comparison, see Best AI Chatbot.
| Specification | Claude 3.5 Sonnet | Claude Sonnet 4 |
|---|---|---|
| Release date | June 20, 2024 (updated Oct 22, 2024) | May 22, 2025 |
| API status (Sept 2026) | Retired | Retired |
| Retirement date | October 28, 2025 | June 15, 2026 |
| Context window | 200K tokens | 200K tokens (1M tokens in beta) |
| Max output tokens | 8,192 tokens | 64,000 tokens |
| Knowledge cutoff | April 2024 | January 2025 |
| SWE-bench Verified | 49% | 72.7% |
| Input price | $3 per million tokens | $3 per million tokens |
| Output price | $15 per million tokens | $15 per million tokens |
| Recommended replacement | Claude Sonnet 4.6 | Claude Sonnet 4.6 |
What Is Claude 3.5 Sonnet?
Claude 3.5 Sonnet is Anthropic’s mid-tier language model released on June 20, 2024, and upgraded on October 22, 2024. The original version introduced Artifacts, a dedicated window for viewing generated code, documents, and website designs alongside the chat. The October 2024 upgrade added a computer use beta, letting Claude control a desktop by viewing screenshots and issuing mouse and keyboard commands. For a related comparison, see Claude Code vs Cursor.
Claude 3.5 Sonnet ran on a 200K-token context window with a 4,096-token output limit by default, extendable to 8,192 tokens through a beta API header. Anthropic’s own benchmarks showed the upgraded version solving 64% of an internal agentic coding evaluation, up from 38% for Claude 3 Opus. Anthropic retired both Claude 3.5 Sonnet model IDs on October 28, 2025.
What Is Claude Sonnet 4?
Claude Sonnet 4 is Anthropic’s coding-focused Claude model released on May 22, 2025, as a successor to Claude 3.7 Sonnet. It scored 72.7% on SWE-bench Verified, a state-of-the-art result at launch, and cut agentic navigation errors from around 20% to near zero compared to earlier Sonnet versions. For a related comparison, see Claude vs ChatGPT vs Gemini.
Claude Sonnet 4 shipped with the same 200K-token context window as Claude 3.5 Sonnet, then gained a 1-million-token beta context window on August 12, 2025, for developers on API usage Tier 4 and above. It also added parallel tool execution, memory file creation for long tasks, and extended thinking output up to 64,000 tokens.
Anthropic retired claude-sonnet-4-20250514 on June 15, 2026, directing traffic to Claude Sonnet 4.6. For a related comparison, see Claude Code vs OpenAI Codex.
Feature Comparison
Claude Sonnet 4 extends nearly every feature Claude 3.5 Sonnet introduced, adding a larger output window, parallel tool use, and a beta 1-million-token context option. Both models support vision input, Artifacts, and computer use, but Claude Sonnet 4 executes these capabilities with materially fewer errors.
Claude Sonnet 4 covers 5 capability upgrades over Claude 3.5 Sonnet worth noting directly:
- Output length — 64,000 tokens versus 8,192 tokens, useful for long code files or full reports in a single response.
- Context ceiling — a 1-million-token beta tier versus a fixed 200K-token maximum.
- Tool use — parallel tool execution versus sequential-only tool calls.
- Reasoning control — extended thinking with adjustable effort versus no dedicated reasoning mode.
- Agentic reliability — near-zero file navigation errors versus roughly 20% navigation error rates.
Both models remain identical on one point: neither ever exceeded a 200K-token context window at standard API pricing, since Claude Sonnet 4’s 1-million-token tier requires the higher-cost long-context pricing described below.
Performance Comparison
Claude Sonnet 4 beats Claude 3.5 Sonnet on every shared coding and reasoning benchmark Anthropic published. The gap is widest on SWE-bench Verified, a real-world software engineering benchmark, where Claude Sonnet 4 scores 72.7% against 49% for the upgraded Claude 3.5 Sonnet and 33% for the original June 2024 release.
| Benchmark | Claude 3.5 Sonnet | Claude Sonnet 4 |
|---|---|---|
| SWE-bench Verified | 49% | 72.7% |
| GPQA Diamond | Not disclosed at launch | 70.0% |
| MMLU | Sets new benchmark per Anthropic | 85.4% |
| AIME (math) | Not disclosed at launch | 33.1% |
| Agentic coding (internal eval) | 64% | Not directly comparable metric |
Anthropic did not publish GPQA Diamond or AIME scores for Claude 3.5 Sonnet at its original release, so those rows reflect Claude Sonnet 4’s later, more standardized benchmark suite. The comparable coding delta between the two models is large enough that Anthropic marketed Claude Sonnet 4 primarily as a coding and agentic-workflow upgrade rather than a general chat improvement. For a related comparison, see Best AI Models.
Pricing
Claude Sonnet 4 and Claude 3.5 Sonnet cost exactly the same at standard context lengths: $3 per million input tokens and $15 per million output tokens. Anthropic kept pricing flat across this generation jump so the SWE-bench and reasoning gains came at no extra cost per token. For a related comparison, see Gemini 2 5 Pro vs Claude 3 7 Sonnet.
Claude Sonnet 4 introduced one pricing difference: requests using its 1-million-token beta context window cost more once a prompt exceeds 200K tokens. Anthropic priced those long-context requests at $6 per million input tokens and $22.50 per million output tokens, roughly double the standard rate.
Claude 3.5 Sonnet had no equivalent long-context tier, since it never supported beyond 200K tokens. For a related comparison, see Claude 3.5 Sonnet vs Claude 3 Opus.
Both prices are now historical, since Claude retired both models from the API. Their recommended replacement, Claude Sonnet 4.6, launched February 17, 2026, at the same $3/$15 per million token rate as Claude Sonnet 4, while the newer Claude Sonnet 5 launched June 30, 2026, at a lower $2 per million input tokens and $10 per million output tokens.
Pros and Cons
Claude 3.5 Sonnet’s advantage was cost-effective speed for everyday writing and light coding tasks, while Claude Sonnet 4’s advantage was reliability on complex, multi-file coding work. Neither model can be provisioned for new production traffic today, since both are retired.
Claude 3.5 Sonnet carried 3 consistent advantages while active:
- Fast response times for chat, summarization, and short-form writing tasks.
- Lower output cost per response due to its 8,192-token output ceiling limiting runaway generations.
- Simpler integration, since it never required the Tier 4 access Claude Sonnet 4’s 1M-context beta needed.
Claude 3.5 Sonnet’s main drawbacks were a 33-49% SWE-bench Verified score well behind Claude Sonnet 4, a fixed 200K-token context ceiling, and no extended thinking mode for multi-step reasoning.
Claude Sonnet 4 carried 3 consistent advantages while active:
- Materially stronger coding accuracy, at 72.7% on SWE-bench Verified.
- A 64,000-token output ceiling, 8 times larger than Claude 3.5 Sonnet’s default.
- Extended thinking and parallel tool use for agentic, multi-step tasks.
Claude Sonnet 4’s main drawback was the added complexity and cost of its long-context tier, plus a shorter total support window before Anthropic retired it in June 2026.
User Reviews
Developer sentiment favored Claude Sonnet 4 for production coding work, while some early testers reported confusion over the model’s self-reported knowledge cutoff. Community reports on GitHub and developer forums documented Claude Sonnet 4 occasionally describing itself as having an April 2024 cutoff, matching Claude 3.5 Sonnet’s, even though its actual training data extended through January 2025. For a related comparison, see Claude Code vs GitHub Copilot.
Coding-tool communities such as Cursor and Zed users tracked Claude 3.5 Sonnet’s rollout of the 8,192-token output limit closely in 2024, since the earlier 4,096-token default cut off longer code generations mid-file. Anthropic’s own release notes cited a drop in agentic navigation errors from roughly 20% to near zero for Claude Sonnet 4, a change developers on real codebases echoed in practice. Overall sentiment across both launches treated Claude Sonnet 4 as a substantial coding upgrade rather than an incremental release.
Use Cases
Claude 3.5 Sonnet suited high-volume, latency-sensitive tasks like chat support and short content drafts, while Claude Sonnet 4 suited coding agents, long documents, and multi-step workflows. Both models served the same broad chatbot and API use cases, differing mainly in how much complexity each could reliably handle in one pass.
Claude 3.5 Sonnet was the better fit for 4 recurring workloads while active:
- Customer support chat — fast, low-cost responses for straightforward queries.
- Short-form content — social copy, email drafts, and summaries under a few thousand words.
- Basic code review — single-file suggestions and bug spotting.
- Vision tasks — chart reading and document transcription from images.
Claude Sonnet 4 was the better fit for coding agents like Claude Code, long-document analysis using its 1M-token beta context, and autonomous multi-step workflows that needed parallel tool calls. Teams building agentic products today should evaluate Claude Sonnet 4.6 or Claude Sonnet 5 for these same use cases, since Claude Sonnet 4 no longer accepts new API traffic.
Final Recommendation
Neither Claude 3.5 Sonnet nor Claude Sonnet 4 is a valid choice for new projects, since Anthropic retired both from the Claude API. The comparison below reflects which model made sense for which workload while both were still active, useful for understanding legacy code, old benchmarks, or migration history.
Choose Claude 3.5 Sonnet if you needed:
– The lowest-latency responses for simple chat and writing tasks.
– A smaller, more predictable output size per request.
– Basic vision and Artifacts support without agentic complexity.
Choose Claude Sonnet 4 if you needed:
– SWE-bench-grade coding accuracy for real software engineering tasks.
– A context window scaling past 200K tokens for large codebases or documents.
– Extended thinking and parallel tool use for autonomous agents.
For any project starting now, the equivalent decision point sits between Claude Sonnet 4.6 and Claude Sonnet 5, both active on the Claude API as of September 2026.
Alternatives
Claude Sonnet 4.6 is Anthropic’s direct, currently active replacement for both Claude 3.5 Sonnet and Claude Sonnet 4, scoring 79.6% on SWE-bench Verified at the same $3/$15 per million token pricing as Claude Sonnet 4. Anthropic points every retired Sonnet migration path, from claude-3-5-sonnet-20241022 to claude-sonnet-4-20250514, to this same model ID.
Claude Sonnet 5, released June 30, 2026, is Anthropic’s newer flagship Sonnet, priced lower at $2 per million input tokens and $10 per million output tokens with a standard 1-million-token context window. For reasoning-heavy work beyond what Sonnet-tier models handle, Claude Opus remains Anthropic’s higher-cost, higher-capability tier. Outside the Claude lineup, teams comparing options typically also evaluate ChatGPT and Gemini for the same coding and agentic workloads.
Frequently Asked Questions
Is Claude 3.5 Sonnet still available in 2026?
No. Anthropic deprecated Claude 3.5 Sonnet on August 13, 2025, and fully retired it from the Claude API, Claude Platform on AWS, and Microsoft Foundry on October 28, 2025. API requests to either claude-3-5-sonnet-20240620 or claude-3-5-sonnet-20241022 now return an error.
Can I still use Claude Sonnet 4 through the API?
No. Anthropic deprecated Claude Sonnet 4 on April 14, 2026, and retired the model ID claude-sonnet-4-20250514 on June 15, 2026. Requests to that model ID fail, and Anthropic’s migration guide points existing integrations to Claude Sonnet 4.6.
What replaced Claude Sonnet 4?
Claude Sonnet 4.6 is the model Anthropic officially recommends as the replacement for Claude Sonnet 4, listed directly in Anthropic’s model deprecation table. Claude Sonnet 4.6 launched February 17, 2026, scores 79.6% on SWE-bench Verified, and keeps the same $3/$15 per million token pricing Claude Sonnet 4 used.
Was Claude Sonnet 4 better than Claude 3.5 Sonnet for coding?
Yes. Claude Sonnet 4 scored 72.7% on SWE-bench Verified compared to 49% for the upgraded Claude 3.5 Sonnet, a 23.7 percentage point gap. Anthropic also reported agentic file-navigation errors dropping from roughly 20% under earlier Sonnet versions to near zero under Claude Sonnet 4.
Why do my API calls to claude-3-5-sonnet or claude-sonnet-4 fail?
Both model IDs are past their retirement dates, so Anthropic’s API rejects any request naming them. Claude 3.5 Sonnet retired October 28, 2025, and Claude Sonnet 4 retired June 15, 2026. Update the model parameter to claude-sonnet-4-6 or claude-sonnet-5 to restore functionality.
What is the context window difference between Claude Sonnet 4 and Claude 3.5 Sonnet?
Claude 3.5 Sonnet capped out at a fixed 200K-token context window. Claude Sonnet 4 launched with the same 200K-token limit, then added a 1-million-token beta context window on August 12, 2025, for API accounts on Tier 4 or higher, at roughly double the standard per-token price.
Should I migrate straight to Claude Sonnet 5?
For new projects with no existing Claude Sonnet 4 integration, Claude Sonnet 5 is a reasonable default, since it costs less per token and ships with a standard 1-million-token context window. For an existing Claude Sonnet 4 codebase, Anthropic’s own migration guide points to Claude Sonnet 4.6 first, since it is the officially listed replacement model ID.
Final Verdict
Claude Sonnet 4 was the stronger model of the two by a wide margin, particularly for coding, matching a 72.7% SWE-bench Verified score against Claude 3.5 Sonnet’s 49% at identical per-token pricing. Both models share one practical fact that matters more than any benchmark today: Anthropic retired both from the Claude API, with Claude 3.5 Sonnet gone since October 28, 2025, and Claude Sonnet 4 gone since June 15, 2026. Any team still referencing either model ID migrates to Claude Sonnet 4.6 or Claude Sonnet 5, both active and priced competitively as of September 2026.