CourionAI
EN
Newsletter
← All news
ai-detection 2 min read

Pangram says its new detector is wrong once every 24,000 documents. That number deserves a closer look

Pangram 4 claims a 0.0041 percent false positive rate and can spot AI text through 13 humanizer tools. The company's user base grew 44-fold in a year, which is the part that matters.

Risograph illustration of a fine mesh sieve sorting falling paper sheets into two piles with one sheet caught wrongly

Pangram has released Pangram 4, a new model for detecting AI-generated text. The company says it is six times larger than its predecessor, with 14 times fewer false positives and six times fewer missed AI texts. In Pangram’s own benchmarks, it correctly identifies 99.66 percent of AI text and wrongly flags human writing 0.0041 percent of the time, which the company frames as roughly one error per 24,000 documents.

There is more in the release than the headline figure. Pangram 4 claims it can distinguish lightly AI-polished text from fully AI-generated text, a distinction that matters enormously in practice and that earlier detectors handled badly. It also says it resists “humanizer” tools, services that rewrite AI output specifically to evade detection, catching the AI component across 13 common humanizers 98.83 percent of the time. The API now costs $0.05 per 100 words, a two to tenfold increase depending on document length, with image scanning included on all plans. Pangram 3 stays available until 30 September 2026.

The business numbers are arguably the bigger story. According to the New York Times, annual revenue has grown 35 times year over year, and monthly users jumped from 2,700 in June 2025 to 120,000 in June 2026. Detection has gone from a niche academic worry to a product category with real money in it, which is why Substack’s decision to add Pangram tools last week caused the reaction it did.

Now the caveat, and it is a real one. Every number above comes from the company selling the product. Independent evaluation of detectors has historically been unkind: a Stanford study found GPT detectors systematically biased against non-native English writers, whose more formulaic sentence construction reads as machine-like. And a false positive rate of 0.0041 percent sounds negligible until you notice what it means at scale. Run 120,000 documents a month and you produce roughly five wrongly accused people every month, each of whom has no way to prove a negative. In a classroom or a disciplinary hearing, the aggregate accuracy is not what matters. What matters is the cost of being the one it got wrong.

What this means for you: If you write, especially in English as a second language, know that these tools are now widely deployed and that being flagged is not the same as being guilty. Keep drafts, version history and notes, they are the only real defence. If you are on the other side, a teacher, an editor, a platform, treat a detector score as one signal that starts a conversation, never as evidence that ends one. And treat vendor-reported accuracy as a marketing claim until someone independent checks it.

Sources

Source: https://www.pangram.com/blog/introducing-pangram-4

Next story

Another Big Four firm caught publishing reports with sources that do not exist

GPTZero found fabricated citations in four PwC Middle East reports. One is 84 percent likely to be entirely AI-generated. KPMG, Deloitte and EY got there first.

Risograph illustration of a stack of bound reports with footnote threads dangling into empty air, attached to nothing