Two Benchmark Labs Measured GPT-6 Astra and Reached Opposite Conclusions
Epoch AI puts Astra clearly in first place across 50 tests. Artificial Analysis rates it level with its predecessor and behind Claude Fable 5.1. The one number both sides find remarkable is on ARC-AGI-3.
A week after OpenAI shipped GPT-6 Astra, two independent evaluation labs have published overall scores for it, and they disagree. Epoch AI, which rolls up more than 50 individual benchmarks, ranks Astra first out of 267 models with 169 points. Artificial Analysis, which tests knowledge, coding, and text comprehension, gives it 61 points, exactly level with its predecessor GPT-5.6 Sol and behind Claude Fable 5.1 at 66.
Both are careful, well-regarded outfits. They are simply not measuring the same thing, and the detail underneath explains why. Epoch’s individual numbers show Astra leading on maths, knowledge, and puzzles. Fable 5.1 holds the top mark on almost every coding test. Epoch has so far recorded only one coding score for Astra, from a run at medium reasoning effort. So the two headline numbers are largely a question of what got weighted.
A few concrete figures from the same reporting: Astra’s hallucination rate on the AA-Omniscience test drops from 92 to 51 percent, but it loses around 80 Elo points on GDPval-AA v2 and slips on banking support and long-context reasoning. It costs two and a half times more per unit of text than Sol, yet on coding tasks it matches Claude Fable 5’s score at less than half the cost per task, because it uses roughly a third of the steps.
What is behind this
The result everyone is talking about is ARC-AGI-3, which drops a model into unfamiliar little game worlds and explains nothing. It has to work out the rules by trying things. Astra scored 62.7 percent on the neutral test harness, against 7.8 percent for Sol and 30.2 percent for Claude Opus 5. More striking than the score is the efficiency: on 96 percent of levels, Astra needed fewer moves than the median human tester who solved that level.
It got there by inventing its own shorthand, an algebra-like notation for recording objects, coordinates, and plans, because anything it does not write into its own notes it forgets. ARC Prize founder François Chollet calls this “on the fly symbolic world modeling” and notes that behaviour like this used to require an elaborate external scaffold. Now it is in the model.
Two caveats matter. OpenAI reported 99.9 percent on the same test using its own harness, and ARC Prize says only the 62.7 percent figure allows a fair comparison between vendors. This is the same harness dispute that produced the eye-catching Nvidia AVO result in August. And Chollet is explicit that none of this proves general intelligence: the games are short, closed, and tiny next to real tasks. What he does say is that progress arrived about twice as fast as he expected, and he is moving his AGI forecast earlier than 2030.
What this means for you: Treat single headline benchmark numbers as marketing until you know what went into them. When you are picking a model, the only ranking that matters is your own: take three tasks you actually do, run them on two models, and compare. That takes twenty minutes and beats any leaderboard.
Sources
ChatGPT Is Back Up to 55.5 Percent of Chatbot Web Traffic, but It Held 73 Percent a Year Ago
Similarweb's August figures show ChatGPT regaining share from 52.7 percent three months ago while Gemini slips back to 25.6 percent. Claude grew from 1.9 to 9.3 percent in a year. Apps are not counted.