CourionAI
EN
Newsletter
← All news
hardware 2 min read

AMD and Cerebras team up to make AI answers come back faster

Two chip companies are splitting the work of running an AI model between them, aiming for lower latency and better energy efficiency. Here is the plain-English version.

Risograph illustration of a tall server rack and a round silicon wafer joined by a streak of light, sharing one task

AMD and Cerebras announced a partnership this week to run AI models faster by splitting the job between two very different kinds of chip. It sounds like deep hardware plumbing, but the goal is something you feel every day: the wait between hitting enter and the answer starting to appear.

Here is the idea in plain terms. When an AI model responds to you, it does two jobs. First it reads and digests your prompt, along with any documents or history attached to it. Then it writes the answer back one piece at a time. Those two jobs have different appetites. Reading a long prompt is a bulk task that rewards raw throughput; writing the answer word by word rewards raw speed per step. AMD’s “Helios” racks, built on its EPYC processors and Instinct GPUs, will handle the reading part, where they can chew through lots of big requests at once. Cerebras, whose claim to fame is a chip the size of a dinner plate (a “wafer-scale engine”, essentially one giant processor instead of many small ones wired together), takes over the writing part, spitting out tokens with very low delay. AMD and Cerebras say splitting the work this way delivers up to five times better energy efficiency. The combined service is expected to arrive first through Cerebras Cloud in the second half of this year.

So what is really going on here? The AI industry has spent years racing to build bigger models. The new race is about serving them well: fast, cheap, and without burning absurd amounts of power. “Disaggregated” setups like this one, where you use the right chip for each part of the task instead of one chip for everything, are a big part of that answer. It is also a notable moment for Cerebras, which has long been the unusual outsider in a Nvidia-dominated world, and for AMD, which recently signed a separate inference deal with Anthropic. Worth a caveat: these are performance claims from the companies themselves, and the joint product is not out yet, so the real-world numbers still have to land.

What this means for you: You will never see an AMD or Cerebras chip, and you do not need to. But when the assistant you use replies more quickly, or a provider drops its prices, this is the kind of behind-the-scenes work that makes it possible. Faster, cheaper serving is what turns an impressive demo into something you actually want to use all day. For most people, nothing changes today; it is just a good sign of where the plumbing is heading.

Sources

Source: https://www.cerebras.ai/press-release/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference

Next story

China's new rules for AI companions just took effect. Here is what they do

New Chinese rules on AI 'companion' apps came into force last week, banning romantic chatbots for minors and forcing dependency safeguards. A look at what changed and why it is being watched worldwide.

Risograph illustration of a heart-shaped speech bubble held behind a low barrier gate beside a small clock