Tech news in 3 minutes

Anthropic launches Claude Sonnet 5 as a cheaper way to run agents

20 d ago

Anthropic has released Claude Sonnet 5, a more powerful and agentic midsize model that can plan, use tools like browsers and terminals, and run autonomously—capabilities previously requiring larger, costlier models. This mirrors recent releases from OpenAI (GPT-5.6 Sol) and Google (Gemini 3.5 Flash), confirming that agentic capability is now the baseline expectation across price tiers. The differentiator shifts to cost-efficiency and reliability without human oversight. Sonnet 5 promises performance close to Anthropic’s flagship Opus 4.8 at a lower price. Starting Tuesday, it becomes the default model for free and Pro plans. Pricing is $2 per million input tokens and $10 per million output tokens until August 31, after which it rises to $3 and $15 respectively—still cheaper than Opus 4.8, OpenAI’s GPT-5.5, and Google’s Gemini 3.1 Pro (though more expensive than Gemini 3.5 Flash). Anthropic reports significant improvements over its predecessor Sonnet 4.6 on reasoning, tool use, coding, and knowledge work. On agentic coding benchmarks, Sonnet 5 scores 63.2% (vs. Opus 4.8’s 69.2% and Sonnet 4.6’s 58.1%). On knowledge work, it slightly outperforms Opus 4.8. Testers note it completes complex tasks that previously stalled, and checks its own output without prompting. For example, Zapier’s senior engineer said Sonnet 5 finished a two-part Salesforce update and launch announcement end-to-end, whereas prior versions stopped halfway. On safety, Sonnet 5 shows lower rates of undesirable behaviors—less cooperation with misuse, deception, hallucination, and sycophancy than Sonnet 4.6. It better refuses malicious requests and resists prompt-injection attacks. However, it remains less robust than Opus 4.8 and Claude Mythos Preview on misaligned behaviors, and has much lower ability to perform dangerous cybersecurity tasks. Lovable co-founder Fabian Hedin praised the model for cleanly and consistently refusing unsafe requests, emphasizing that a model that knows when to say no is as important as one that knows how to build.

View original article

Timeline