CourionAI
EN
Newsletter
← All news
ai-basics 2 min read

OpenAI and Anthropic are arguing about a benchmark, and the argument is more useful than the scores

GPT-5.6 Sol scored 7.8 percent on ARC-AGI-3 in the official setup and 38.3 percent in OpenAI's own. The gap explains why benchmark numbers are so hard to compare.

Two identical measuring gauges side by side, their needles pointing to wildly different readings

Last week Anthropic’s Claude Opus 5 posted a record 30.2 percent on ARC-AGI-3, a puzzle benchmark built to test whether a model can work out rules it has never seen before. On July 30, OpenAI answered: run GPT-5.6 Sol with two extra settings, it says, and the model hits 38.3 percent. In the official test setup, the same model scored 7.8 percent.

That is not a typo. The gap comes from two features in OpenAI’s Responses API. “Retained Reasoning” keeps the model’s chain of thought, its internal working-out, alive between steps instead of throwing it away. “Compaction” summarises older context rather than simply cutting it off when the window fills up. Strip both away, as the official harness does, and the model loses its train of thought after every move. ARC Prize replied that its environment is deliberately standardised so comparisons stay fair. Then co-founder Francois Chollet conceded ground: setups built specifically to game the benchmark are off limits, he said, but general API settings available to every customer are fair play. In effect, ARC Prize’s own score had put OpenAI at a disadvantage, possibly by testing through an older API that lacked features Anthropic’s already had.

What is behind it

This is worth understanding even if you never read a benchmark again. A benchmark never measures a model alone. It measures the model plus the scaffolding around it: how memory is handled, how context is trimmed, how many attempts it gets, how much compute it is allowed to burn. Change the scaffolding and the number moves by a factor of five, as it just did here. Neither side is being dishonest. They are measuring different things and both calling it “the model’s score”. Chollet’s own resolution is the sensible one: any setup is acceptable as long as the settings and the cost are reported clearly.

What this means for you: Treat single benchmark numbers as advertising, not measurement. When a lab says its model beats another by some margin, the question to ask is what setup produced the number and what it cost to run. If you are choosing a model for real work, your own small test on your own actual tasks will tell you more than any leaderboard. For everyone else, the practical read is simply that these two models are close, and the marketing gap between them is wider than the capability gap.

Sources

Source: https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

Next story

Hidden white text in a Word file can hijack Copilot, and the file it produces carries the trick onward

A researcher showed that invisible instructions in a document can make Microsoft 365 Copilot alter figures and copy the instructions into the new file. Microsoft's mitigations have not closed it.

A blank sheet of paper under raking light revealing hidden marks, duplicating into an endless chain behind it