Taalas Unveils AI Chip Running Llama 3.1 at 17,000 Tokens Per Second


Taalas, a Toronto-based startup, has introduced an AI chip that reportedly runs the Llama 3.1 8B model at 17,000 tokens per second. This performance significantly outpaces existing solutions, with the company claiming its HC1 chip is nearly 10 times faster than Cerebras and almost 50 times faster than Nvidia's B200 for the same model.

Detailed view of a high-performance AI chip on an anti-static mat in a lab setting.
The company, founded less than three years ago, has adopted a distinct approach to AI hardware. Instead of relying on liquid cooling, expensive High Bandwidth Memory (HBM), or general-purpose computing, Taalas directly integrates large language models into the chip's physical structure. This design choice aims to reduce complexity and improve efficiency.
Novel Architecture for AI Acceleration
Taalas's HC1 chip is designed with a "brutal aesthetic," where the model's weights are permanently welded into the silicon at the time of manufacture. This means each transistor on the chip corresponds to a specific weight of the Llama 3.1 8B model. Matrix multiplication is handled directly by electrical currents within the physical circuit, eliminating the need for software scheduling.

Microscopic view of silicon pathways, some glowing and solidified, representing hardwired AI model weights.
This specialized architecture allows the HC1 to operate without complex storage hierarchies, resulting in a significantly lower cost and power consumption compared to traditional solutions. Taalas states that the HC1's cost is about 1/20th, and its power consumption is 1/10th, of conventional systems. A cluster of ten HC1 cards reportedly requires only 2.5 kilowatts and can be air-cooled.
The company has launched an experience website, chatjimmy.ai, to demonstrate the chip's speed. Users have noted the near-instantaneous response times, describing the AI's output as appearing "as if it had been premeditated."
Trade-offs and Industry Perspectives
While the HC1's speed is a notable achievement, its design introduces specific limitations. The chip is hardwired to run only the Llama 3.1 8B model and cannot be fine-tuned, upgraded, or adapted to other models. This raises questions about its long-term utility in a rapidly evolving AI landscape where models frequently update. Critics suggest that such a specialized chip could quickly become obsolete if newer, more capable models emerge.

Digital hourglass with glowing data bits, symbolizing rapid technological obsolescence in AI.
However, some optimists view this approach as a potential future direction for AI, particularly for applications requiring extremely low latency. They suggest that such high-speed token output might be more relevant for inter-agent communication than for human interaction.
Divergent Paths in Chip Design
Ljubisa Bajic, CEO of Taalas and a former core architect at AMD and Nvidia, previously founded Tenstorrent, an AI chip company. Jim Keller, a prominent chip architect, joined Tenstorrent as CTO and later became CEO, advocating for general-purpose, programmable platforms. Bajic's move to establish Taalas reflects a departure from this general-purpose philosophy, pursuing an extreme specialization in exchange for maximal performance and efficiency.
Portrait of Ljubisa Bajic, CEO of Taalas and former AMD and Nvidia architect.
The concept of creating application-specific integrated circuits (ASICs) for AI models has drawn comparisons to the human brain's efficiency. Some observers on social media platforms have highlighted the brain's precise and low-power operation as a form of "hardware solidification." They argue that for numerous specialized tasks—such as voice assistants, automated data labeling, or robotic navigation—a highly efficient, fixed-function chip could be more practical and cost-effective than a general-purpose AI.
This divergence in AI hardware development suggests a future where AI systems could polarize. One path leads to large, versatile, and expensive cloud-based general-purpose AI, while the other focuses on highly specialized, low-cost, and ultra-fast chips embedded in everyday devices, offering near-zero latency for specific functions.
Stay Ahead of the AI Curve
Join 50,000+ subscribers getting the latest AI tools, trends, and tutorials delivered to their inbox weekly.
No spam, unsubscribe at any time.