AMD buys Taalas to hard-wire AI models onto silicon
AMD has agreed to acquire Canadian startup Taalas, a move that could radically accelerate AI inference by hard-coding specific models directly into specialized silicon chips.

AMD is expanding its hardware capabilities by acquiring Taalas, a Toronto-based startup founded in 2023 that emerged from stealth in February. Taalas specializes in an unconventional hardware design where an AI model's architecture and trained parameters are permanently embedded directly into the silicon. While this approach locks the resulting chip to one specific model, it yields unprecedented processing speeds.
The startup's approach has already demonstrated remarkable performance. A demo chip hard-wired with Meta's Llama 3.1-8B model achieved processing speeds of over 16,000 tokens per second per user. This rate significantly outpaces traditional general-purpose hardware. The potential of this hard-coded silicon architecture has also attracted other industry giants, with reports indicating that Google is developing a similar dedicated chip for its Gemini models.
AMD plans to integrate this hard-wired technology into its broader accelerator roadmap. The company intends to offer these specialized chips alongside its existing Instinct graphics processing units (GPUs) as part of a comprehensive, system-level solution. Vamsi Boppana, senior vice president of AMD's artificial intelligence division, noted that the acquisition strengthens the chipmaker's overall AI portfolio, while Taalas co-founder Ljubisa Bajic stated that AMD provides the necessary scale to commercialize the technology. The transaction remains subject to standard regulatory approvals.
For AI practitioners, this acquisition signals a shift toward hyper-specialized infrastructure. While traditional GPUs offer the flexibility to run and fine-tune various models, they introduce latency. Hard-wired chips offer a trade-off: developers lose the ability to update or change the underlying model, but they gain massive throughput and efficiency for high-volume, static workloads. This makes the technology highly attractive for enterprises running mature, unchanging models at massive scale where operational costs and speed are the primary bottlenecks.
This is our own summary of reporting by The Decoder



