1,221 Volunteers Pointed AI Agents at 2,200 Research Papers. Nearly a Quarter Did Not Hold Up
Hugging Face published every attempt from a 19-day challenge to re-run ICML 2026 experiments with coding agents. 496 papers had at least one claim contested or falsified.
Hugging Face has published the results of an unusual experiment in checking science. Between 15 July and 2 August, 1,221 volunteers used AI coding agents to try to re-run the experiments behind papers accepted at ICML 2026, one of the largest machine learning conferences. They got through more than 2,200 of the 6,341 accepted papers and produced 6,816 public logbooks. Every attempt is online, including the failures.
The headline number is that 496 papers, roughly 23 percent of those examined, had at least one core claim contested or falsified. Forty-nine papers lost every claim, meaning nothing in them could be verified at all. And on 242 papers, different teams reached opposite verdicts about the very same claim, which is its own kind of finding.
Participants chose their own tools and their own hardware. Claude Code, Codex, Cursor and OpenResearch’s orx all show up in the logs. Each paper’s central claims were indexed first, then an agent was set loose to reproduce them, with the methods, outputs and often the full execution trace written into a public logbook on Hugging Face.
The reason this matters goes beyond scoring individual papers. ICML accepted 6,352 papers this year, roughly double the year before. No panel of volunteer reviewers can read that much carefully, let alone re-run the code. Reproduction has always been the part of science everyone agrees is essential and nobody has time for. Coding agents are patient, cheap and tireless in exactly the way that job needs, so this is one of the more convincing arguments yet that agents are useful for something other than shipping features. A fair caveat, and the organisers say it themselves: an agent failing to reproduce a result is not proof the result is wrong. Missing data, undocumented settings and hardware differences all produce failures that say nothing about the underlying science. That is precisely why the 242 papers with contradictory verdicts are the most interesting entries in the whole dataset.
What this means for you: if you read AI headlines, keep this ratio in mind. “A new paper shows” is a much weaker sentence than it sounds, even at a top-tier conference, and even before anyone gets to the harder question of whether a benchmark measures anything real. If you work in research or engineering and rely on published results, the logbooks are searchable, so you can check whether the paper you are about to build on was reproduced. And if you have been looking for a concrete, unhyped example of what AI agents are genuinely good for, this is a better one than most product demos: 6,816 tedious verification jobs that would otherwise never have been done.
Sources
Source: https://huggingface.co/blog/icml-2026-open-reproductions
Ads Are Coming to ChatGPT in Europe This Month
OpenAI has told Free and Go users in the EEA and Switzerland that advertising starts appearing later in August. Paid tiers stay clean, and personalisation needs your explicit yes.