OpenAI Introduces GPT-5.3-Codex-Spark for Real-Time Code Generation at 1,000 Tokens Per Second


OpenAI has released GPT-5.3-Codex-Spark, a new model designed for rapid code generation, capable of producing over 1,000 tokens per second. The company describes this as its first model specifically engineered for real-time programming, emphasizing its speed.
Accelerated Code Generation
The primary feature of GPT-5.3-Codex-Spark is its generation speed, which aims to reduce latency in coding workflows. OpenAI demonstrated the model's ability to provide near-instant responses, with code appearing almost immediately after input. This speed is intended to eliminate wait times often associated with code generation.
Hardware and Architectural Optimizations
To achieve this performance, GPT-5.3-Codex-Spark operates on Cerebras' Wafer Scale Engine 3, a hardware platform designed for low-latency processing. OpenAI also re-engineered the model's underlying architecture, incorporating persistent WebSocket connections. This change reportedly reduced round-trip overhead by 80% and increased the speed of the first character's appearance by 50%. Despite its focus on speed, OpenAI states the model maintains strong performance in programming ability benchmarks.

Close-up of Cerebras Wafer Scale Engine 3, a large, square chip with dense circuitry.
Performance and Application
In evaluations such as SWE-Bench Pro and Terminal-Bench 2.0, which assess agent software engineering capabilities, GPT-5.3-Codex-Spark demonstrated strong performance while significantly reducing task completion times compared to GPT-5.3-Codex.
The model is intended for applications requiring real-time interaction, such as collaborative coding environments. It allows users to make immediate changes to logic or reshape interfaces rapidly. OpenAI suggests this experience is akin to pair programming.

Two software engineers collaboratively coding on a shared screen, symbolizing pair programming.
GPT-5.3-Codex-Spark features a 128k context window and currently supports text-based interactions. It also incorporates safety measures. The model is available to ChatGPT Pro users through the Codex App, CLI, and VS Code plugin. OpenAI indicated that future programming paradigms might involve both deep, extended model processing and real-time, interactive model assistance, with Spark addressing the latter.
Stay Ahead of the AI Curve
Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.
No spam, unsubscribe at any time.