CourionAI
EN
Newsletter
← All news
ai-detection 2 min read

AI Could Write Like a Human. Safety Training Is What Gives It Away

Pangram's CTO argues that language models are detectable not because they lack ability but because post-training narrows how they express themselves. The effect has a name: mode collapse.

Rows of identical typewriters funneling into one narrow bottleneck shape

Here is a claim worth sitting with: language models could in principle write with all the variety of human beings, and the reason they do not is a side effect of how we train them to behave. That is the argument Bradley Emi, CTO of the AI text detector Pangram, made in a post on 20 August, and it reframes what AI detection is actually detecting.

The mechanism has a name: mode collapse. During post-training, the stage after a model has read the internet where it learns rules of behaviour, a model is taught to avoid dangerous outputs, decline certain requests and stay within particular bounds. That training does not only remove the bad options. It narrows the whole range, so the model settles into one preferred way of phrasing things instead of spreading its choices across the many ways a person might say the same sentence. The opposite state, where probability is distributed more evenly across phrasings, is called mode coverage. Emi’s supporting evidence is the exception that proves the rule: base models, the raw versions before post-training, write with enough variety that Pangram’s detector does not flag them. The same goes for narrow fine-tunes trained only on Hemingway or on one subreddit, and for broken, incoherent output.

If that is right, it explains a lot of things that otherwise look like coincidence, including the em dash habit that Pew’s study of the web found rising in step with AI use. It also sets up a genuine tension. Every safety guardrail a lab adds makes the model’s writing a little more uniform, and therefore a little easier to detect. You could read that as a happy accident, since detectability is useful. You could equally read it as evidence that we are flattening these systems in ways nobody deliberately chose. Emi notes one important limit: all of this applies only to text without a watermark, the invisible statistical signature Anthropic began embedding in Claude’s output earlier this month. Watermarks will likely keep working regardless of how varied a model’s writing is.

What this means for you: if you have wondered why AI-written text has a “sound,” now you know it is not a limitation of the technology but a consequence of the training that makes it safe to use. Practically, this is a reason not to trust detectors blindly in either direction. A person who writes in a clean, careful, slightly formal register will trip them, and someone using an unfiltered model may not. If you are ever on the receiving end of a detector’s verdict, at school or at work, that asymmetry is the thing to point out.

Sources

Source: https://pangram.substack.com/p/no-llms-dont-just-mimic-human-text

Next story

Amazon Just Made Its AI Assistant Free on Fire TV, and Quietly Dropped a 19.99 Dollar Fee

Alexa+ is rolling out to all compatible US Fire TV devices at no cost, with no Prime membership required. The upgrade is automatic. The catch is that the free tier stops at the television.

A boxy retro television on a low stand with sound ripples around it and a paper price tag drifting away