CourionAI
EN
Newsletter
← All news
models 2 min read

Are Models Getting Worse at Facts on Purpose?

A widely shared post argues that labs are trading factual knowledge for reasoning efficiency. The benchmark numbers behind the claim are real, even if the intent is harder to prove.

A half-empty brain-shaped bookshelf with books flying out toward an open filing cabinet

A post by Walter van der Giessen titled “Models Are Getting Dumber on Purpose” spent Sunday on the Hacker News front page, and the argument is worth understanding even if you disagree with the framing. His claim: frontier labs are deliberately squeezing raw factual knowledge out of their models in order to buy reasoning ability and speed.

The numbers he points at are striking. GLM-5.2 scores 99.2 percent on AIME 2026, a hard maths competition benchmark, using 40 billion active parameters. Qwen 3.5 hits 91.3 percent with 17 billion active. Active parameters means the portion of the model actually switched on for any given word, which in a mixture-of-experts design is much smaller than the total. Meanwhile SimpleQA, which just asks short factual questions and checks the answers, has a leader board topped by Gemini 2.5 Pro at around 53 percent. Models that can nearly ace olympiad maths still get half of simple factual questions wrong.

Whether that is on purpose is the contested part. There is a much duller explanation available: memorised facts take up parameters, parameters cost money at every single request, and a model tuned for reasoning benchmarks will naturally end up with less trivia inside it. Nobody has to sit in a meeting and decide to make the model forget things. Either way the practical result is the same, and van der Giessen’s conclusion follows from it: if the model cannot be trusted to remember, the sensible architecture is a smaller model plus retrieval, where the facts live in a document store the model looks things up in rather than inside its weights.

The wider point is one that keeps getting rediscovered. A benchmark score tells you what a model was optimised for, not what it is good at. A number that goes up because a lab pushed on it is not the same as general improvement, and the things nobody benchmarks quietly get worse. Factual recall is one of those things, because it is unglamorous and hard to make a chart out of.

What this means for you: treat confident factual answers with the same suspicion you would treat a confident stranger in a pub, no matter how clever the model seems at reasoning. The practical fix is simple and free: give the model the source. Paste in the document, switch on web search, upload the PDF. Answers grounded in text you supplied are dramatically more reliable than answers pulled from memory, and this is the single habit that most improves everyday results. For anyone building something, this is a nudge toward retrieval over bigger models: a small local model with a good document index will often beat a frontier model guessing from memory, and it costs a fraction as much.

Sources

Source: https://w4g1.dev/blog/models-are-getting-dumber-on-purpose

Next story

Stripe Is Buying OpenRouter for More Than 7 Billion Dollars

The payments company wants the layer that decides which AI model answers your request, and then bills you for it. OpenRouter's valuation jumped more than fivefold in three months.

Many highway lanes narrowing into a single tollbooth arch with cargo shapes passing through