
- Claude Sonnet 5 launched June 30, 2026 at an introductory $2 per million input tokens and $10 per million output through August 31 — then $3 and $15, still well below Opus 4.8’s $5 and $25.
- It scores 63.2% on SWE-bench Pro (ahead of GPT-5.5 at 58.6% and Gemini 3.5 Flash at 55.1%) and 81.2% on the OSWorld-Verified computer-use test.
- Sonnet 5 is now the default model on Free and Pro plans and ships with an adjustable effort dial that can reach Opus 4.8 performance on some tasks.
- Cyber safeguards are on by default; on a Mozilla-built Firefox exploit test, Sonnet 5 built a working exploit 0.0% of the time.
For most of the past year, Anthropic’s strongest agentic performance lived in its priciest tier. On June 30, 2026, that calculation shifted. Claude Sonnet 5 arrived at $2 per million input tokens — roughly 60% below Opus 4.8’s $5 — while landing within striking distance of that flagship on coding, tool use, and computer-use benchmarks. In other words, the company just made its own most expensive model harder to justify for a large slice of everyday work.
A Cheaper Model That Behaves Like the Expensive One
The pricing that reframes the lineup
At launch, Sonnet 5 carries introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026. After that it settles at $3 and $15. For comparison, Opus 4.8 — Anthropic’s general-purpose flagship — costs $5 per million input and $25 per million output. That gap matters most in agentic workloads, where a single task can chain dozens of tool calls and burn through large token budgets. One caveat sits in the fine print: Sonnet 5 uses an updated tokenizer, so the same text can map to between 1.0 and 1.35 times as many tokens. Anthropic says the introductory pricing is set so the move from Sonnet 4.6 is roughly cost-neutral.
Benchmarks: near-Opus, ahead of rivals
On published evaluations, Sonnet 5 clears its predecessor across the board. It posts 63.2% on SWE-bench Pro — the harder variant of the coding benchmark — versus 58.6% for GPT-5.5 and 55.1% for Gemini 3.5 Flash. On OSWorld-Verified, which measures an agent driving a real computer, it reaches 81.2%, edging past the updated 78.5% mark for Sonnet 4.6. It also ships with a 1M-token context window. Anthropic frames the model as sitting just below Opus 4.8 on most rows, while nearly matching or nudging ahead on reasoning-with-tools and knowledge-work tasks.
Trend Insight — The headline is not a new capability ceiling — Opus 4.8 still tops most charts. It is the price at which near-frontier agentic work becomes routine. When the default plan model can carry multi-step engineering tasks, the economics of automation change for every team, not just well-funded labs.
The Effort Dial Is the Real Story
Sonnet 5’s most consequential feature may be control rather than raw score. The model exposes adjustable effort levels, letting developers trade cost for capability on a per-task basis. Anthropic’s own cost-performance curves show Sonnet 5 covering a far wider range than Sonnet 4.6 did: substantially better cost efficiency at medium effort, and higher-effort settings that match Opus 4.8 on some agentic search and computer-use tasks. Practically, that means a team can dial Sonnet 5 down for cheap bulk work and up for the few tasks that genuinely need flagship-level reasoning — instead of paying Opus rates for everything.
What early partners are seeing
Early-access testers emphasized follow-through. Cursor co-founder Sualeh Asif said agents built on Sonnet 5 stay on plan, follow their conventions, and ship clean multi-step changes at an efficient cost. Lovable co-founder Fabian Hedin said the model gets more done with less and stressed safety, adding that a model that knows when to say no is just as important as one that knows how to build. In one test, an engineer asked Sonnet 5 to investigate a bug; unprompted, it wrote a reproducing test, implemented the fix, then stashed the change to confirm the bug returned without it — all in a single pass. Another tester handed it a two-part job, updating Salesforce account tiers and sending a launch announcement to enterprise contacts, and the model finished end to end where earlier versions would stall halfway.
Trend Insight — Autonomy that checks its own work is the quiet unlock here. A model that writes a failing test before fixing a bug behaves less like autocomplete and more like a junior engineer — exactly the workflow that lets small teams take on work that used to need more hands.
Safer by Default, With Cyber Guardrails On
Anthropic’s pre-deployment evaluations put Sonnet 5 ahead of Sonnet 4.6 on safety. It is better at refusing malicious requests and resisting hijack attempts in prompt-injection attacks, and it shows lower rates of hallucination and sycophancy. On an automated behavioral audit spanning a wide range of misaligned behaviors, Sonnet 5 scored safer overall than 4.6 — though still somewhat higher than the more capable Opus 4.8 and Mythos Preview. On cybersecurity, the limits are deliberate: the model was not trained for offensive cyber work and performs far worse than Opus at it. On a Firefox 147 exploit test built with Mozilla, neither Sonnet model produced a working exploit — both scored 0.0% — though Sonnet 5 showed a slightly higher partial-success rate, which Anthropic attributes to general intelligence gains rather than cyber training. Because Sonnet 5 is modestly stronger here, it launched with real-time cyber safeguards enabled by default, the same ones used on Opus 4.7 and 4.8.
Trend Insight — Shipping a cheaper, more agentic model with guardrails on by default is a notable posture. It signals that Anthropic expects Sonnet 5 to be deployed widely and autonomously — and is pricing in the risk that broad access, not raw capability, is where most real-world misuse would come from.
What It Means for Teams Shipping With AI
For product and engineering leaders, the practical takeaway is a re-drawn default. Until now, getting reliable multi-step agent behavior often meant reaching for a flagship-priced model. Sonnet 5 pushes that behavior down into the tier most teams already pay for — and makes it the out-of-the-box choice on Free and Pro plans. A sensible pattern emerging from launch feedback: run Sonnet 5 at low-to-medium effort for the bulk of automation and coding, reserve high effort for the hard tasks, and keep Opus 4.8 for the narrow set of problems that still need it, including sensitive cybersecurity work. For a lean team, that is the difference between automation being a line-item you ration and one you can apply broadly.
Trend Insight — The competitive signal points outward as much as inward. By pricing near-Opus agentic performance at Sonnet rates, Anthropic is pressuring rival mid-tier models on the exact axis buyers now care about — cost per completed task, not cost per token.
Related
- Coding Skill Wasn’t the Edge. Domain Expertise Was.
- AI’s Next Crisis Isn’t Compute — It’s $570 Billion in Debt
- Ford Bet on AI Over Its Engineers. Then This Happened.
- Tokenmaxxing Is Fading: Why More AI Isn’t Better AI
- AI Oversight Startups Just Raised Big. Here’s Why.
Sources
- Anthropic — Introducing Claude Sonnet 5 (June 30, 2026)
- MarkTechPost — Claude Sonnet 5 vs Sonnet 4.6 vs Opus 4.8: Benchmarks and Pricing (June 30, 2026)
- TechCrunch — AI coverage
AI Biz Insider · AI Trends EN · aibizinsider.com
댓글 남기기