CourionAI
EN
Newsletter
← All news
security 3 min read

Anthropic Put Its Most Restricted Model to Work Scanning Code for Bugs. A Human Still Has to Approve Every Fix

Claude Security now runs on Mythos 5, the model Anthropic keeps off general release because it is too good at cyber tasks. It is in public beta for Enterprise customers, and partners protecting hospitals and banks get access too.

A shield with a magnifying glass scanning rows of code lines and a small bug icon

Anthropic has moved its code-scanning tool, Claude Security, onto Mythos 5. That is notable mostly because of which model it is: Mythos is Anthropic’s most capable model and it is deliberately not broadly available, specifically because it is so strong at cyber tasks. The company announced the change on 21 August.

The tool itself is straightforward. It reads a codebase, looks for vulnerabilities, and suggests patches. Every finding comes with a CWE category, which is an industry-standard way of classifying a type of software flaw so that different tools describe the same problem the same way, plus a severity rating and a proposed fix. It is in public beta for Enterprise customers, and scans count as ordinary token usage rather than a separate product.

The important detail sits at the end: a human still has to sign off on every patch. Anthropic is not letting the model commit changes to code on its own.

The asymmetry this is meant to fix

Attackers have had access to capable AI for a while now, and they face no approval process, no compliance review and no consequences for a bad suggestion. Defenders face all three. That gap is the reason a lab would take its most tightly held model and point it specifically at defence. Anthropic is also plugging Mythos 5 into partner security products that protect hospitals, utilities and banks, where end users never touch the model directly and only see the output, such as a suggested patch.

That structure is the actual policy here. Rather than releasing a powerful cyber model and hoping it gets used well, Anthropic keeps it behind a gate and hands it out through defensive channels only. Whether this holds is an open question. Any capability strong enough to find a vulnerability is strong enough to exploit one, and a gate is only as good as the people it lets through. But it is a coherent attempt at a genuinely hard problem, and more thought-through than simply publishing the weights and hoping.

Worth noting the timing. This lands the same week researchers demonstrated an unpatched Grok attack that hid encrypted instructions inside a web page, and days after Microsoft shipped a fix for a one-click Copilot data-theft flaw it had known about for eight months. AI security is currently a story about defenders catching up.

What this means for you: If you do not write code, this is background, but useful background: the AI security story is not only about attacks, and some of the defensive work is real. If you do write code, or you run a small company that does, the pattern to copy is the approval step. Let a model find and propose, never let it apply. That single rule is what separates a helpful scanner from a new attack surface, and it costs you nothing to enforce.

Sources

Source: https://claude.com/blog/bringing-claude-mythos-5-to-more-defenders

Next story

Anthropic Kept Your Data for 30 Days to Hunt Attackers. After Enterprise Pushback, It Will Sit in Your Cloud Instead

The 30-day window stays, but the data moves to the customer's own cloud. Anthropic spent months building the system with more than 100 customers from regulated industries, and admits the original rule was a business risk.

A data folder moving from one cloud into a second cloud shaped like a building