CourionAI
EN
Newsletter
← All news
open-source 3 min read

Ling 3.0 Flash Is the Smartest Open Model of Its Size, and It Stopped Making Things Up

Ant Group's inclusionAI released Ling 3.0 Flash under the MIT license. The headline number is not the intelligence score, it is the hallucination rate falling from 97 to 44 percent.

One small lantern burning brightly beside a wide open padlock, with three much larger unlit lanterns behind it

Ling 3.0 Flash, released this week by inclusionAI, the AI arm of Ant Group, is currently the highest-scoring open model under 124 billion total parameters. On the Artificial Analysis Intelligence Index it reaches 38 points, a large jump over its predecessor and level with Qwen3.6 27B while using far fewer active parameters. It is released under the MIT license, with weights on Hugging Face and API access through inclusionAI and DeepInfra.

The more interesting number is elsewhere. On AA Omniscience, a test of whether a model admits when it does not know something, Ling’s hallucination rate fell from 97 percent to 44 percent. A hallucination is when a model states something false with full confidence, and the test measures how often it invents an answer instead of declining. Going from “almost always makes something up” to “makes something up on fewer than half of the questions it cannot answer” is a bigger practical improvement than any benchmark point. It still means roughly two in five, so this is progress, not a solved problem.

Two more things worth knowing. Ling 3.0 Flash is cheaper per token than every comparably capable model, though it burns through more tokens on complex tasks, so the per-token advantage narrows on real work. Even measured per completed task, it stays cheaper than Qwen3.6 27B. And it is not the leader overall: DeepSeek V4 Flash sits at 52 on the same index, well ahead.

What is actually going on here

Two patterns are visible here at once. The first is that “small” open models keep eating the ground the big ones stood on, and the gap between an open model you can download and a frontier model you rent is now measured in months rather than years. The second is that labs have started competing on refusal rather than only on capability. For years the incentive ran the other way: a model that answers everything looks more useful in a demo than one that says “I don’t know”. As these models move into work where a confident wrong answer costs real money, being honest about uncertainty becomes a feature people will pay for. Fewer active parameters, incidentally, means fewer calculations per word, which is why a model can be large in total but cheap and fast to run.

What this means for you: for most people this is not a model you will use directly, and that is fine. What it changes is the menu. Anyone building a product now has a genuinely capable model they can download, run on their own hardware, and never send customer data over the internet with, under a licence that permits commercial use. That is the whole argument for open models, and it is getting easier to make each month. If you do run models locally, this one is worth downloading. If you do not, watch the hallucination number, because “how often does it admit it does not know” is becoming a better question than “how smart is it”.

Sources

Source: https://huggingface.co/inclusionAI/Ling-3.0-flash

Next story

Mistral OCR 4.1 Can Now Point at the Exact Paragraph It Read

The French lab's document reader now returns paragraph-level coordinates, labels each block by type, and scores its own confidence. Boring on paper, and exactly what makes AI document answers checkable.

Stacked document pages covered in abstract squiggle lines, each block framed by a rectangle, one page lifted and glowing