Claude Sonnet 4.6 Outperforms Pricier Opus Model in User Preference, Expands Capabilities


Anthropic has released Claude Sonnet 4.6, an updated model that users preferred over its predecessor, Sonnet 4.5, in 70% of early coding tests. Notably, users also favored Sonnet 4.6 over the more expensive flagship model, Opus 4.5, in 59% of cases. The new Sonnet version features upgrades across six key areas: coding, computer operations, long-context reasoning, agent planning, knowledge work, and design. It also introduces a 1 million token context window in beta, while maintaining its pricing at $3 per million input tokens and $15 per million output tokens.
Performance Benchmarks and General Utility
Sonnet 4.6 demonstrated strong performance across 15 evaluations, leading or ranking near the top in most categories. These include agent tool use, large-scale tool calling, office tasks, and financial analysis. While Opus models, particularly Opus 4.6, retain their position as the gold standard for tasks requiring the deepest reasoning, such as code refactoring and multi-agent coordination, Sonnet 4.6 is now considered sufficient for most daily tasks at a significantly lower cost.

Data analyst reviewing complex charts and code on multiple monitors in a modern office.
Databricks testing indicated that Sonnet 4.6's performance in enterprise document understanding tasks (OfficeQA) is now on par with Opus 4.6. Replit’s assessment highlighted Sonnet 4.6’s "amazing" value for money, noting its performance improves with increasing task difficulty.
Advancements in Computer Operations
Anthropic's computer operation model, first introduced in October 2024, has evolved from an experimental stage to near human-level capability. Its score on OSWorld, a benchmark for AI computer operations, has increased from 14.9% to 72.5% over 16 months. OSWorld evaluates a model's ability to interact with real software like Chrome, LibreOffice, and VS Code on a simulated computer, mimicking human interaction through screen observation, mouse clicks, and keyboard input, rather than relying on special API interfaces.
Early users reported Sonnet 4.6 performing at near human-level in scenarios such as operating complex spreadsheets, filling out multi-step web forms, and integrating information across multiple browser tabs. Pace Insurance achieved a 94% accuracy rate in its internal benchmarks, identifying Sonnet 4.6 as the strongest computer operation model it had tested. Convey's evaluation similarly found it significantly superior to other models.

Human hand precisely operating a complex spreadsheet on a computer screen.
Despite these advancements, the 72.5% success rate indicates that approximately 30% of tasks may still fail, suggesting that critical business processes are not yet fully ready for complete automation. However, the rapid progress, marked by a five-fold increase in 16 months, is notable. Anthropic also confirmed that Sonnet 4.6 has significantly improved its resistance to prompt injection attacks compared to Sonnet 4.5. The capability of AI to operate computers like humans is making the automation of legacy systems without APIs increasingly feasible.
Expanded Context Window and Long-Context Reasoning
Sonnet 4.6's context window has been expanded to 1 million tokens in beta, equivalent to a large codebase, an extensive contract, or dozens of research papers. This larger window is complemented by improved reasoning capabilities within long contexts, addressing a common limitation where models might "remember but cannot reason."
The Vending-Bench Arena test provided an example of Sonnet 4.6's long-context planning. In this simulation, the model devised a strategy to heavily invest in production capacity during the initial 10 months, incurring higher expenditures than competitors, before pivoting to focus on profitability in the final stage. This approach led to significant outperformance, demonstrating the model's ability to utilize long contexts for strategic planning.
Combined with the beta Context Compaction feature, which automatically summarizes older content as conversations approach their limit, the effective usable context can extend beyond 1 million tokens. This enables applications such as large codebase analysis, lengthy contract reviews, and comprehensive literature reviews, tasks that previously required manual segmentation.

Abstract digital library with flowing data lines converging into a glowing core, symbolizing expanded context.
Enhanced Coding Capabilities
Coding remains a primary use case for the Sonnet series, and Sonnet 4.6 introduces fundamental changes beyond minor speed or accuracy improvements. User feedback highlights several key enhancements: better context reading before code modification, integration of shared logic instead of copy-pasting, reduced over-engineering and "laziness," improved adherence to instructions, fewer hallucinations, and enhanced consistency in multi-step tasks. These improvements are particularly evident in long conversations, where previous AI coding tools often struggled with "forgetting" or "going rogue."
Companies utilizing Sonnet 4.6 for coding have reported positive results. GitHub noted excellent performance in complex code fixes, especially in large codebase searches, with high resolution rates and stability. Cursor described it as a "comprehensive and significant improvement" over Sonnet 4.5 for long-span tasks and difficult problems. Bolt found it delivered cutting-edge results in complex application building and bug fixing, becoming a preferred choice for deep codebase work previously requiring more expensive models. Rakuten reported that Sonnet 4.6 produced the best iOS code they had tested, demonstrating superior architecture, compliance, first-time success, and proactive use of modern tools. Cognition stated that Sonnet 4.6 "powerfully narrowed the gap with Opus" in bug detection, enabling more parallel review procedures and bug catching without increased costs.

Software developer coding on multiple monitors in a focused, professional setting.
API and Tool Ecosystem Updates
Anthropic has also updated its API tools, with Web Search Dynamic Filtering being a notable addition. This feature addresses the token-intensive nature of web searches, where models typically initiate queries, retrieve results, obtain full HTML, and then reason. Dynamic filtering allows Claude to automatically generate code to filter and process search results, retaining only relevant content. This approach involves the model filtering with code before reasoning, rather than reasoning directly over massive HTML.
Results from this feature include:
BrowseComp benchmark: Sonnet 4.6 improved from 33.3% to 46.6%, and Opus 4.6 from 45.3% to 61.6%.
DeepsearchQA: Sonnet improved from 52.6% to 59.4%, and Opus from 69.8% to 77.3%.
Overall, average accuracy increased by 11%, while token consumption was reduced by 24%.
Quora/Poe's evaluation indicated that Opus 4.6 with dynamic filtering achieved the highest accuracy in internal assessments, with the model behaving like a "true researcher" by parsing, filtering, and cross-referencing results with Python. While the data on reduced token consumption applies to Sonnet 4.6, token costs for Opus 4.6 reportedly increased, with Anthropic recommending developers test with their own queries.
Five officially released tools include Code Execution, Memory Function, Programmatic Tool Calling, Tool Search, and Tool Usage Examples. For financial users, the Excel plugin now supports MCP connectors, enabling Claude to directly call data sources like S&P Global, LSEG, Daloopa, PitchBook, Moody's, and FactSet within Excel.

Modern API interface on a tablet screen, with subtle connections to financial data sources.
Recommendations for Users
For free users, the default model has been upgraded to Sonnet 4.6, which includes file creation, connectors, skills, and context compression features. Pro and Team users are advised to use Sonnet 4.6 for daily tasks, as its performance is now near Opus level in most scenarios. Switching to Opus 4.6 is recommended only for tasks requiring code refactoring, multi-agent coordination, or absolute precision.
Developers can access the new model via the claude-sonnet-4-6 API, maintaining the same pricing as Sonnet 4.5. Experimenting with different thought intensity settings is suggested, as Sonnet 4.6 performs strongly even with extended thinking off.
Enterprise users can leverage the combination of Computer Use and MCP connectors to automate legacy systems lacking APIs. Pace Insurance's 94% accuracy rate in this area demonstrates the potential for AI to directly operate such systems.
While Opus 4.6 remains the preferred choice for the deepest reasoning tasks, Sonnet 4.6 offers a sufficient and more cost-effective solution for most users, demonstrating that a lower price no longer implies weaker performance.
Stay Ahead of the AI Curve
Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.
No spam, unsubscribe at any time.