CourionAI
EN
Newsletter
← All news
ai-basics 2 min read

Nvidia's Next AI Chips Reportedly Squeeze 10x More Work From the Same Power

CoreWeave says Nvidia's new Vera Rubin systems produced about 10 times more tokens per megawatt than last generation's Blackwell in an early benchmark. Here is what 'tokens per megawatt' means and why it decides whether AI can keep growing.

A lightning bolt powering a server rack beside a small arrow and a much larger arrow, and a speedometer gauge swung high next to a sage leaf

Cloud provider CoreWeave published early benchmark numbers for Nvidia’s next-generation AI system, the Vera Rubin NVL72, and the headline is striking: about 10 times more “tokens per megawatt” than the previous Blackwell generation. In the test, running a DeepSeek reasoning model, a Blackwell rack reportedly produced around 80,000 tokens per second at 150 megawatts, while the Vera Rubin rack hit roughly 800,000 tokens per second at the same power draw.

Let’s unpack that, because two bits of jargon are doing the heavy lifting. A “token” is a chunk of text an AI model reads or writes, very roughly a word or part of a word; when a chatbot answers you, it is producing tokens. A “megawatt” is just a unit of electrical power, the kind used to size data centres and power plants. So “tokens per megawatt” is a fancy way of asking: how much AI work do you get for each unit of electricity? Ten times more means either far more output from the same power, or the same output from a fraction of it.

For the last two years the AI story has quietly been an energy story. Training and running big models eats enormous amounts of electricity, and power, not chips, is increasingly the thing operators cannot get enough of. That is why “performance per watt” has become the number Nvidia keeps putting front and centre: it is what decides whether a company can serve more users profitably, and whether AI’s growth runs into the wall of the electric grid. A jump this size, if it holds up, buys the industry some breathing room. Vera Rubin racks are reportedly ramping up at CoreWeave, Google Cloud, Microsoft Azure and Oracle.

What this means for you: For most people, nothing today, and that’s fine. You will not buy one of these racks. But the knock-on effects reach you anyway. Cheaper, more efficient inference (the industry term for running a trained model to answer requests) is what makes it viable for companies to give away powerful AI features, run bigger models behind your favourite apps, or lower prices. It also softens one of the more uncomfortable criticisms of the AI boom, its power appetite, though it does not erase it, since more efficiency often just invites more usage.

Worth keeping expectations grounded: this is a single vendor-partner benchmark, and CoreWeave is both an Nvidia cloud partner and an Nvidia investee. The published result does not include enough setup detail for an outside lab to reproduce it independently. Impressive, plausibly real, but a marketing debut rather than a peer-reviewed measurement.

Sources

Source: https://wf.coreweave.com/blog/nvidia-vera-rubin-nvl72-on-coreweave-10x-more-tokens-per-megawatt-than-blackwell

Next story

Samsung Is in Talks to Back Mistral, Pushing Europe's AI Champion Toward a €20 Billion Valuation

Samsung is reportedly weighing around a €1 billion investment in French AI lab Mistral, part of a round that would value the company near €20 billion. Here is why a big Korean electronics maker wants a piece of Europe's answer to OpenAI.

Two tall towers built from stacked coins joined by a bridge of coins, with a swirling wind gust and a faint European coastline in the background