Chinese Team Feeling AI Ranks Second Globally in Agentic AI Benchmark Terminal-Bench 2.0

Feeling AI, a Chinese startup, has secured second place globally in the authoritative Terminal-Bench 2.0 ranking with its CodeBrain-1, placing it just behind OpenAI's latest flagship model. This achievement highlights China's advanced capabilities in Agentic AI for complex task planning and autonomous coding.

Two glowing digital entities, blue and orange, representing competing AI models in a global landscape.
The global AI landscape has recently seen significant developments, with Anthropic introducing Claude Opus 4.6 and OpenAI countering with GPT-5.3-Codex. This competition signals a shift in the large model race from a focus on benchmark scores to practical application and autonomous workflows in real-world business scenarios.
Both OpenAI and Anthropic used Terminal-Bench 2.0 to demonstrate their models' capabilities. Opus 4.6 achieved a 65.4% success rate in Agentic Terminal Coding Tasks. OpenAI's combination of 5.3-Codex and Simple Codex reached a 77.3% success rate, with a reported 75.1% on the benchmark, which the company stated represents a peak in global coding performance. NVIDIA Chief Scientist Jim Fan has described the real terminal environment as AI's "devil's training ground," emphasizing its role in measuring a model's engineering capabilities through self-evolution in a closed-loop environment.

Developer's hands typing on a mechanical keyboard, with complex code on a monitor in the background.
In this competitive environment, Feeling AI's CodeBrain-1, built on the GPT-5.3-Codex base model, achieved a 72.9% success rate (70.3% on the benchmark), making it the only Chinese entrant in the top ten.
Feeling AI's Recent Advancements
Five days prior to the Terminal-Bench 2.0 announcement, Feeling AI released MemBrain1.0, which achieved new state-of-the-art (SOTA) results in memory benchmarks such as LoCoMo, LongMemEval, and PersonaMem-v2. MemBrain1.0 surpassed existing memory systems and full-context models like MemOS, Zep, and EverMemOS. It also demonstrated over 300% improvement in the most challenging levels of KnowMeBench Level III evaluations, marking a significant advance in Agentic Memory.
The development of robust memory capabilities and hierarchical memory systems is enabling a paradigm shift in Agentic AI, moving beyond core model capabilities to enhance user experience. Following MemBrain 1.0, Feeling AI introduced CodeBrain-1, an "evolving brain" designed with dynamic planning and strategy adjustment features. This system rapidly ascended to second place globally on Terminal-Bench 2.0, trailing only OpenAI's Simple Codex, which partners with GPT-5.3-Codex.

Hierarchical AI architecture diagram showing InteractBrain, InteractSkill, and InteractRender components.
Feeling AI has consistently emphasized that dynamic interaction is crucial for world models to achieve Artificial General Intelligence (AGI). Its cross-modal hierarchical architecture includes three core components: InteractBrain for understanding, memory, and planning; InteractSkill for capability execution; and InteractRender for rendering and presentation. Both MemBrain and CodeBrain are integral to the InteractBrain layer, focusing on deep understanding and long-term planning in complex dynamic interaction scenarios. The company suggests these developments are the result of deliberate strategic planning, focusing on complex "dynamic interaction" scenarios for both Agentic Memory and task planning/execution.
OpenAI has described Simple Codex as "the optimal solution for long-term software engineering tasks." The integration of models with agent frameworks is emerging as a standard for the commercial deployment of large models. Agentic Memory's capabilities are expected to become part of future agent frameworks, acting as an external memory system to enhance models through systematic capabilities. The ability of Chinese frameworks to control leading global models positions them as key intelligent hubs in the AI era, influencing the engineering standards for future large models.
CodeBrain-1's Dynamic Planning and Strategy Adjustment
The official Terminal-Bench evaluation website shows CodeBrain-1 positioned second, behind OpenAI's Simple Codex (GPT-5.3-Codex), and ahead of Factory's Droid, which uses Anthropic's Claude Opus 4.6. Other notable entities on the list include Warp, Coder, Google, and Princeton.
Terminal-Bench encompasses a variety of tasks, including complex system operations and extensive coding tasks within a real terminal environment. CodeBrain-1 focuses on ensuring code correctness and execution. Its technical implementation emphasizes two key aspects for successful and efficient task completion:
Useful Context Searching: CodeBrain-1 prioritizes "truly useful" context, recognizing that in complex tasks, relevance is more important than quantity to mitigate large language model (LLM) hallucination. It uses Language Server Protocol (LSP) functionalities to efficiently retrieve relevant information based on current task requirements and existing Code Base indexes, aiding the code generation process. For instance, when planning a game bot task, CodeBrain-1 uses LSP Search to accurately obtain method signatures, documentation, and usage examples for functions like move_to(target) and do(action), reducing retrieval overhead and context interference.
Validation Feedback: CodeBrain-1 converts failures into actionable information. It efficiently identifies issues from LSP Diagnostics and supplements error-related code and documentation, thereby shortening the generate-validate loop. For example, if a type error occurs when calling on(observation, exec), LSP provides details such as argument type mismatch, caller examples, documentation for the erroneous parameter, and usage of the exec parameter.
The team evaluated CodeBrain-1 on a subset of 47 Python-only tasks from Terminal-Bench, where it demonstrated consistent completion capabilities, including efficient retrieval of associated code and documentation, and faster problem localization during code inspection and validation.

Abstract representation of data validation and error feedback in a digital system, with red errors and green corrections.
CodeBrain-1 also demonstrated efficiency in token consumption. When both base models used Claude Opus 4.6, CodeBrain-1 reduced total token consumption for successful Python tasks by over 15% compared to Claude Code, according to Anthropic's technical documentation.
CodeBrain-1's performance on Terminal-Bench 2.0 is attributed to its end-to-end task execution in a real command-line interface (CLI) environment and its ability to dynamically adjust plans and strategies. By optimizing task execution logic and error feedback, it significantly improves the model's operational success rate in real terminal environments.
CodeBrain-1's approach involves generating executable programs based on "intelligence" within defined constraints and continuously adjusting based on feedback. These plans and strategies can operate at individual and group levels. For individuals, roles can adapt schedules, behaviors, and interactions based on goals, memories, and observations. For groups, shared memories can be formed, and overall planning and response rules can be adjusted based on external conditions.
To illustrate its capabilities, CodeBrain-1 was tested in game scenarios:
Case 1: Real-time driving of game bots In open-world games, CodeBrain-1 can act as a game companion, interpreting natural language commands like "build me a house" or "make a pickaxe." It then plans actions such as collecting resources, clearing workspaces, and crafting, generating and executing action scripts to achieve goals, thereby enhancing the player's experience.
Case 2: Tactical evolution driven by collective memory In "search, attack, and retreat" games, CodeBrain-1 can enable opposing groups to develop "collective memory" of player habits. During map construction and deployment, the system adjusts its strategy, for example, by distributing forces in player hotspots. Behavioral rules can also be superimposed for immersion, such as bots exclaiming "Got you!" or "Misjudgment!" Additionally, simple squad combat strategies, like front-line charge and rear-line cover, can be configured, with behaviors dynamically generated by group strategies.

AI-controlled game bots building structures and executing tactical maneuvers in a futuristic open-world game.
The Significance of Terminal-Bench 2.0
Terminal-Bench, an open-source benchmark developed by Stanford University and Laude Institute, is considered a "gold standard" for evaluating AI agents' end-to-end execution capabilities in real command-line interface (CLI) environments. Its rigor stems from:
Closed-loop practical environment: AI agents must perform compilation, debugging, training, and deployment in an isolated Docker container, mimicking human experts in a real Linux ecosystem.
High-pressure long-term tasks: The benchmark includes 89 deep scenarios spanning software engineering and scientific computing, requiring high logical span and eliminating simple pattern matching.
Zero-tolerance verification: A 0/1 judgment criterion is applied, where only the production of expected deliverables, such as fixed code or running services, counts as a pass.
2.0's "ceiling" effect: The upgraded 2.0 version significantly increased the difficulty, with top global models generally struggling to exceed a 65% solution rate. CodeBrain-1's second-place debut underscores its value in this challenging environment.

Digital maze with a short, green, efficient path and longer, blue, convoluted paths, symbolizing AI strategy adjustment.
While top models like the GPT series possess strong reasoning chains, they can sometimes lead to lengthy execution paths due to "overthinking." CodeBrain-1 functions as an executive brain that continuously adjusts plans and strategies. It acts as a "dispatching hub" and "efficiency calibrator," guiding the model for rapid responses in routine operations and activating deep thinking only for critical errors. This precise control over the base model is a key factor for commercial implementation.
Robust closed-loop error recovery, efficient task decomposition, and accurate environmental perception are vital for AGI's commercialization. Powerful agents are considered essential for models to achieve practical application. Sam Altman's statement following the GPT-5.3-Codex release, describing Codex's evolution into an "all-around agent" for computer operations, supports this trend. OpenAI envisions models and frameworks evolving into deeply integrated "intelligent family packages."
Despite the presence of major industry players, vertical industries offer significant commercial opportunities for effective engineering frameworks. System-level agent frameworks and lean developer productivity tools, being closer to the user, hold potential for growth. Feeling AI's ability to integrate with OpenAI's models and achieve leading results demonstrates its engineering responsiveness and the increasing prominence of Chinese AI teams in global engineering collaboration.

Abstract network of glowing lines and nodes, with a central bright hub representing Chinese AI teams connected to other major AI players.
Securing second place globally on Terminal-Bench 2.0, a benchmark known for its real-world environments and long-term evolution, is symbolically significant. It indicates that Chinese startup teams are contributing to the transition of agents from "conversational toys" to "productivity tools," establishing a leading position in redefining workflows. Within the ecosystems built by OpenAI and Anthropic, Chinese teams are emerging as "framework definers," showcasing a distinct and resilient approach to AI innovation. The competition in the commercial implementation of models, following the initial phase of global base model development, is expected to intensify.
Stay Ahead of the AI Curve
Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.
No spam, unsubscribe at any time.