CourionAI
EN
Newsletter
← All news
models 3 min read

A Diffusion Model Now Writes 1,100 Words a Second, and It Costs Almost Nothing

Inception Labs released Mercury 2.5, a language model built the way image generators are built rather than the way ChatGPT is. It produces 1,107 tokens per second on ordinary Nvidia cards at 20 cents per million input tokens.

A scattered cloud of printed grain resolving all at once into rows of crisp speed lines, with one slow chain of beads crawling along the bottom edge

Inception Labs released Mercury 2.5 on Monday, and the number everyone is quoting is 1,107 tokens per second. Tokens are the small chunks of text a model reads and writes, roughly three quarters of a word each, so that is somewhere around 800 words per second coming out of a single model on widely available Nvidia hardware. For comparison, most chat models you use today feel fast at 50 to 150 tokens per second.

The company says Mercury 2.5 is about 40 percent more capable than Mercury 2 on its internal measures, and that it lands in the same quality band as the cheap, fast tiers of the big labs: GPT-5.6 Luna on its low setting, Gemini 3.5 Flash-Lite, Claude Haiku 4.5. It ships with a 260,000 token context window, which is how much text the model can hold in mind at once, plus tunable reasoning effort, parallel tool calls, and JSON output that follows a schema you define. Standard pricing is 0.20 dollars per million input tokens and 0.75 per million output, with an 80 percent launch discount that drops those to 0.04 and 0.15. Inception also claims sub-170ms time to first token for voice pipelines, which is the delay before the first word comes back.

What is behind this

Almost every language model you have used works left to right. It writes one token, looks at everything it has written, then writes the next one. That sequence is why generation has a speed ceiling: the model literally cannot start token 200 before token 199 exists.

A diffusion language model works the way image generators like Stable Diffusion work. It starts with a rough, noisy draft of the whole answer and then refines it in passes, sharpening many tokens at the same time. Because the refinement happens in parallel, the same GPU produces far more text per second. The tradeoff has historically been quality, since a model that commits to a rough shape early has a harder time with long chains of careful reasoning. Mercury 2.5 is Inception’s argument that the gap has closed enough to matter for everyday work, and by their own account it is the largest diffusion language model anyone has trained. That claim has not been independently checked yet, which is worth keeping in mind.

What this means for you: If you are just curious about AI, nothing changes today, but this is a useful thing to know: speed and price are not fixed properties of AI, they depend on how the model is built, and someone just found a shortcut. If you build things, the interesting case is anything where waiting is the problem rather than brilliance is: autocomplete in an editor, a voice assistant that has to answer before the pause gets awkward, bulk cleanup of a few thousand documents. At 15 cents per million output tokens during the launch discount, jobs that were too expensive to run over a whole archive suddenly are not. A fair caveat: run your own comparison before you switch anything important, because “comparable to the cheap tier of a frontier model” is a claim from the vendor, not a verdict from the field.

Sources

Source: https://www.inceptionlabs.ai/blog/introducing-mercury-2-5

Next story

Meta Stops Grading Engineers on How Much AI They Use

After staff raced each other up an internal token leaderboard, Meta told engineers that AI usage dashboards and token counts will no longer count in performance reviews. The episode has a name now: tokenmaxxing.

A cranking mechanical counter machine spilling drifts of identical tokens across the floor beside a gauge with a snapped needle