CourionAI
EN
Newsletter
← All news
ai-basics 3 min read

AMD Buys Taalas, a Startup That Bakes an Entire AI Model Into the Chip

Taalas etches a model's architecture and weights directly into silicon. That makes it extremely fast and completely inflexible. AMD wants the technology alongside its Instinct GPUs.

A silicon wafer seen from above with a dense circuit labyrinth etched permanently into it and one single route through it picked out in coral

AMD is buying Taalas, a small Canadian chip startup with an unusual idea: instead of building a general chip that can run any AI model, build a chip that can run exactly one, by etching that model directly into the hardware. The deal was announced Friday and is subject to the usual regulatory approvals.

A bit of vocabulary first, because it makes the rest make sense. When an AI model answers you, that is called inference, as opposed to training, which is the expensive one-off process of building the model in the first place. Inference is what happens millions of times a day, so it is where most of the running cost sits. Normally a GPU loads the model’s weights, the billions of numbers that encode what the model learned, out of memory and does the maths. Shuffling those numbers back and forth is a big part of what makes it slow and power-hungry.

Taalas skips that step by putting the architecture and the trained weights into the chip itself. The company came out of stealth in February 2026 with a demo chip running Llama 3.1-8B at more than 16,000 tokens per second per user, many times faster than competing inference hardware. A token is roughly a word fragment, so that number means text appearing far faster than anyone can read it. The catch is right there in the design: that chip runs Llama 3.1-8B and nothing else, forever. New model, new silicon.

That trade is less strange than it sounds. Chip design and manufacturing take many months, and until recently a model would be obsolete before its chip arrived. Now that top models are staying in service longer and the same handful of them serve enormous volumes of traffic, locking one into hardware starts to pencil out. Google is reportedly working on the same trick with a chip that bakes in Gemini’s architecture. AMD says it will fold the technology into its accelerator roadmap and offer it alongside its Instinct GPUs as a system-level product. AMD AI chief Vamsi Boppana called it a strengthening of the portfolio, and Taalas co-founder Ljubisa Bajic said AMD brings the scale the startup lacked.

What this means for you: nothing you can buy, and nothing that changes your chatbot this week. What it changes over the next few years is price. The chips underneath AI services are quietly getting cheaper and more efficient per answer, and that is the main reason the free tiers you use keep getting more generous rather than less. It is also a small hint about which direction “AI on your own device” is going: not a general-purpose brain in your phone, but specific models fixed into specific hardware. A fair caveat: this is an acquisition, not a product. Regulators still have to clear it, and nothing ships tomorrow.

Sources

Source: https://ir.amd.com/news-events/press-releases/detail/1296/amd-acquires-taalas-to-advance-compute-solutions-for-rapidly-growing-ai-inference-market

Next story

TikTok Owner ByteDance Is Reportedly Training China's Largest AI Model Yet

Three insiders told the Financial Times that ByteDance has a model with up to ten trillion parameters in pretraining, three times the size of Kimi K3.

An enormous lattice dome of connected nodes under construction, dwarfing the cranes and scaffolding around it