CourionAI
EN
Newsletter
← All news
ai-basics 2 min read

An OpenAI researcher quit after eight months, betting that better data matters more than bigger models

Andrew Ho left OpenAI to build specialised training datasets and expects labs to spend over $100 billion on data collection. A Cambridge researcher sees the same pattern from the outside.

A geological cross section where one narrow peak spikes far above surrounding flat plateaus

Andrew Ho has left OpenAI after eight months, and his reason is more interesting than the departure. He thinks large language models generalise badly, meaning they are much worse at transferring a skill to a new situation than their headline scores suggest. Even in coding, the most heavily funded area of all, he says performance is inconsistent. His diagnosis is a shortage of the right training data, and he is starting a company to sell it.

Most economically valuable work, Ho argues, barely exists in the data these models learn from. “Most work is highly contextual and not easily encoded into a gradable environment,” he wrote, adding that even when you can watch a human do a job well, it is hard to tell whether the paths they did not take would have worked too. He expects AI labs to spend more than $100 billion on targeted data collection in the coming years. His first products target scientific work: datasets for bioinformatics analysis, where he says even GPT-5.6 Sol succeeds only about 30 percent of the time, and datasets for everyday lab tasks like judging photos of experiments. Chemistry, materials science and healthcare are meant to follow. Ho is also openly sceptical of frontier lab valuations, arguing they stay unprofitable because they must keep spending to stay ahead of cheaper rivals like Qwen and Kimi.

What is behind it

Cambridge researcher Adam Hunt describes the same shape from outside the labs. His claim is that recent models are not getting broader, they are getting spikier: coding and hard maths keep improving while language quality and simple logic stagnate or slip. The reason he gives is mechanical. Reinforcement learning, the training method that turned reasoning models into what they are, needs a clear signal for right and wrong. Code compiles or it does not. Maths proofs check out or they do not. For most human work, no such signal exists, so that training simply cannot be applied. Hunt puts his own confidence at about 40 percent, which is refreshingly honest for a public prediction, and DeepMind’s Tom Zahavy makes a related structural argument in a paper titled “LLMs can’t jump”: models handle deduction and induction well but fail at inventing an explanation nothing in their training has hinted at.

What this means for you: For most people, nothing changes today. But it is a useful correction to a common assumption. The models are not improving evenly across the board, so the fact that AI got dramatically better at your colleague’s coding job says almost nothing about whether it got better at yours. If a tool disappoints you at something that feels easy while acing something that looks hard, you are not doing it wrong. That unevenness is the current state of the technology, and it explains why hands-on testing beats reading release notes.

Sources

Source: https://x.com/andrewho03/status/2082615798011744270

Next story

Amazon quietly stops developing most of its own AI models

Nova Premier, Omni, the Reel video model and the Canvas image generator go into keep-the-lights-on mode. Resources shift to a new Frontier Model Research group under Pieter Abbeel.

Risograph illustration of a row of factory machines draped in dust sheets with only the last one still running