CourionAI
EN
Newsletter
← All news
nvidia 3 min read

Nvidia's New Rack Is 10x, 30x or 67x Better, Depending Entirely on Who Is Counting

Fresh numbers for the Vera Rubin NVL72 landed this week from Nvidia and from SemiAnalysis, and the multipliers range wildly. The spread is not dishonesty, it is a lesson in how efficiency claims are built.

Three rulers of different lengths and scales propped against the same tall block

New performance figures for Nvidia’s Vera Rubin NVL72, the rack sized system that will run a large share of the world’s AI over the next couple of years, arrived this week. Nvidia says it delivers up to 30 times the throughput per megawatt of the GB300 NVL72 it replaces, and up to 35 times lower cost per million tokens, measured on an agentic coding workload built on DeepSeek V4-Pro. The analysts at SemiAnalysis published their own look at the same hardware under the headline of 67 times better performance per dollar. Back in July, CoreWeave measured 10 times the tokens per megawatt against an older Blackwell rack on a different model. Nvidia’s own chief executive has previously talked about 3 times in public.

Throughput per megawatt is the number that matters most right now, because AI data centres are limited by available electricity rather than by money or floor space. Ten, thirty and sixty seven are very different numbers, and none of them is a lie. They differ because each is measured at a different setting. Nvidia’s 30x figure applies at 160 tokens per second per user, meaning the speed each individual user or agent experiences while the model works. Push that responsiveness target up or down and the ratio moves, sometimes a lot, because a chip that can serve many slow users at once behaves differently from one asked to serve a few very fast ones. CoreWeave compared against an older generation, on a different model, in July. And Nvidia’s numbers were measured by Nvidia, on a benchmark suite from SemiAnalysis, with the review by SemiAnalysis still pending.

What is behind this

There is a useful habit buried here. When a hardware or model claim arrives as a single multiplier, three questions usually collapse it into something meaningful: compared with what, at what setting, and measured by whom. In this case, compared with the immediately previous generation or one two steps back, at which responsiveness target, and by the vendor or by an outsider. The underlying improvement is real, everyone who has measured this hardware agrees the gain is large, and for AI providers that translates directly into cheaper inference, which is the running cost of answering your questions. It just does not translate into one number.

What this means for you: Indirectly, this is why the price of AI keeps falling. Hardware that produces several times more text per unit of electricity is the main reason a request that cost cents two years ago now costs fractions of a cent, and that flows through to what you pay for a subscription or an API call. Directly, the transferable thing is the habit. The next time a product page tells you something is “10x faster”, ask compared with what, at what setting, and who measured it. That single question separates a real improvement from a well chosen chart in almost every AI announcement you will read this year.

Sources

Source: https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/

Next story

The Company Behind the GLM Models Just Raised 5 Billion Dollars in a Day

Z.AI, formerly Zhipu, raised about 5 billion dollars through a discounted share placement and zero coupon convertible bonds. Sixty percent goes into the next generation of GLM models and the computers to train them.

An arcing fountain of coins pouring down into one wide container already overflowing