Anthropic Releases Claude Opus 4.6 Alongside OpenAI's GPT-5.3 Codex
Anthropic has launched its new large language model, Claude Opus 4.6, introducing enhanced capabilities and product updates. The release occurred shortly after OpenAI unveiled its GPT-5.3 Codex model, marking a significant simultaneous development from two leading AI developers.
Screenshot or graphic of Anthropic's Claude Opus 4.6 interface or announcement graphic.
Claude Opus 4.6 Innovations
Claude Opus 4.6 features several key advancements, including improved benchmarks and new product functionalities. Anthropic also introduced "Agent Teams" and updated its Excel and PowerPoint plugins.
Digital dashboard showing AI performance benchmarks like 65.4% and 72.7%.
Performance benchmarks for Opus 4.6 indicate strong performance across various evaluations. On Terminal-Bench 2.0, which assesses programming in a terminal environment, Opus 4.6 scored 65.4%. The OSWorld evaluation, measuring an AI's ability to operate a computer, showed Opus 4.6 achieving 72.7%, an increase from Opus 4.5's 66.3%. For online information retrieval, Opus 4.6 scored 84.0% on BrowseComp. The GDPval-AA evaluation, which measures performance in real-world professional tasks, gave Opus 4.6 an Elo score of 1606. Additionally, Opus 4.6 scored 68.8% on ARC AGI 2, an assessment of fluid intelligence.
Beyond benchmarks, Opus 4.6 introduces a 1 million token context window, a substantial increase from the previous 200,000 tokens. This expansion aims to improve the model's ability to process extensive documents and codebases. Tests on the MRCR v2 (needle-in-a-haystack) benchmark showed Opus 4.6 achieving 76% accuracy with 8 hidden "needles" in a 1 million token context. The output limit has also doubled to 128K tokens.
Abstract visual of expanding digital data points, symbolizing a 1 million token context window.
A new feature, Context Compaction, allows Claude to automatically summarize older conversation content to manage context window limits, enabling longer, uninterrupted tasks. Adaptive Thinking and Effort Control features offer more granular control over the model's processing depth, balancing speed, cost, and output quality.
The Agent Teams update in Claude Code allows multiple AI agents to collaborate on complex tasks, such as code reviews, by assigning specialized roles and facilitating direct communication between agents. This differs from subagents, which operate within a single session and report to a main agent.
Abstract glowing figures representing AI agents collaborating on a holographic code projection.
Anthropic has also integrated Claude Opus 4.6 into Excel and PowerPoint. The Excel plugin now supports pivot table editing, chart modification, and financial-grade formatting. The PowerPoint integration enables Claude to create and edit presentations while adhering to existing layouts and templates. The API pricing for Claude Opus 4.6 remains $5 per million input tokens and $25 per million output tokens, with additional pricing for contexts exceeding 200,000 tokens.
OpenAI Unveils GPT-5.3 Codex

Screenshot or graphic of OpenAI's GPT-5.3 Codex interface or announcement graphic.
OpenAI's GPT-5.3 Codex was released concurrently with Claude Opus 4.6. A notable aspect of GPT-5.3 Codex's development is its reported role in its own creation. OpenAI stated that early versions of the model were used by its Codex team to debug training processes, manage deployments, and evaluate test results.
In benchmark comparisons, direct evaluation between GPT-5.3 Codex and Claude Opus 4.6 is challenging due to differing evaluation methodologies. However, on the shared Terminal-Bench 2.0, GPT-5.3 Codex scored 77.3%, surpassing Claude Opus 4.6's 65.4%.
For the OSWorld evaluation, Claude Opus 4.6 reported 72.7% on the original OSWorld, while GPT-5.3 Codex reported 64.7% on OSWorld-Verified. OSWorld-Verified is a refactored version designed to address issues in the original, generally considered more stringent.
The GDPval evaluation, which assesses performance on real-world knowledge tasks, also presented comparability issues due to different scoring methods. GPT-5.3 Codex's "GDPval wins or ties: 70.9%" was based on human expert review, while Claude Opus 4.6's "GDPval-AA Elo: 1606" used an independent evaluation agency's framework.
On SWE-bench, which tests an AI's ability to fix GitHub issues, Claude Opus 4.6 scored 80.8% on SWE-bench Verified (a 500-problem, Python-only subset), while GPT-5.3 Codex scored 56.8% on SWE-bench Pro Public (a 731-problem, multi-language benchmark considered more complex).
Modern racing game interface on a monitor, with subtle AI development indicators.
OpenAI demonstrated GPT-5.3 Codex's capabilities by showcasing two complete games, a racing game and a diving game, that were autonomously developed by the model using a "develop web game" skill and iterative prompts.
A new feature in GPT-5.3 Codex allows users to interact with the model while it is working, enabling real-time intervention and direction adjustments. OpenAI also indicated that GPT-5.3 Codex requires less than half the tokens of its predecessor, GPT-5.2 Codex, for the same tasks and operates over 25% faster per token.
Dynamic light trails and energy flows symbolizing rapid AI advancement and interaction.
The simultaneous releases from Anthropic and OpenAI highlight a period of rapid advancement in AI, particularly in coding and agentic capabilities.
Stay Ahead of the AI Curve
Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.
No spam, unsubscribe at any time.