CourionAI
EN
Newsletter
← All news
models 3 min read

Sakana's New Model Is Not a Model, It's a Dispatcher for Other Models

Fugu Max and Fugu Ultra v2, released on 11 September, do not answer your question themselves. They pick which of a pool of open and specialist models should, and at 2 dollars per million input tokens Fugu Max undercuts the frontier tier by half.

An empty conductor's podium with a fan of batons, facing a semicircle of empty music stands

Sakana AI released two new systems on 11 September, Fugu Max and Fugu Ultra v2, and the interesting part is what they are not. Neither is a single large model that read the internet and now answers your questions. Each is an orchestrator: a model whose actual skill is deciding which other model should handle the job in front of it, then handing the work over, sometimes to several models in sequence, sometimes to another copy of itself. The pool it routes to is mostly open weight and specialist models, meaning models whose files anyone can download and run.

The pricing is the pitch. Fugu Max costs 2 dollars per million input tokens and 6 per million output, roughly 40 to 60 percent below Sonnet 5 and GPT-5.6 Terra, and Sakana says it ranks best on six of ten benchmarks it reports, including Terminal Bench 2.1 and GPQA Diamond. Fugu Ultra v2 aims higher and costs more, 5 dollars in and 30 out, with longer prompts billed at a higher rate, and reports 48.3 on Chartography and 74.3 on DeepSWE. Sakana makes a point of noting that it reaches those numbers without Fable 5, Fable 5.1 or GPT-6 Astra anywhere in its pool, so the results come from cheaper parts assembled well. Both sit behind an API that speaks the same format as OpenAI’s, so switching is mostly a change of address. One line in the fine print deserves attention: the routing itself consumes tokens, and they are billed.

What is behind this

The unspoken assumption of the last three years has been that if you want a better answer, you buy a bigger model. Routing questions this. Most requests are not hard. Summarising an email, rewriting a paragraph, fixing an obvious bug: a small open model handles all of it perfectly well for a fraction of the cost, and only a minority of requests genuinely need the expensive reasoning model. A router that can tell the two apart captures most of the quality for a fraction of the money. The catch is that a router is one more thing that can be wrong, in a way that is annoying to debug, because an answer that came out mediocre might be a bad answer or simply a bad choice of who answered. And the benchmark numbers here are Sakana’s own, on benchmarks Sakana selected.

What this means for you: If you use AI through a chat window, this changes nothing you can see, but it is a good thing to understand, because routing is quietly becoming how the big assistants work too. That is why the same product sometimes feels brilliant and sometimes feels lazy: you are not always talking to the same model. If you pay for API usage, Fugu Max is worth an honest test on your own workload, since a 40 to 60 percent cut on a bill you already pay is real money. Two caveats before you move anything important: run your own comparison rather than trusting the launch chart, and add the orchestration tokens into your maths before you decide it is cheaper.

Sources

Source: https://sakana.ai/fugu-max-release/

Next story

DeepMind Published a Prediction for Every Single Letter You Could Change in Human DNA

AlphaGenome Atlas is a one petabyte dataset covering all 9 billion possible single-letter DNA variants, precomputed so researchers can look up an answer instead of running a model. It is free for non-commercial research.

A wall of library card catalogue drawers, two of them pulled open and unspooling into a twisting double helix ribbon