Give a Frontier Model Only a Bibliography and Ask for the Idea. It Gets There 3 to 15 Percent of the Time
A new benchmark called Reconstruction strips away everything except a paper's reference list and asks models to recover its core finding. Seven frontier models scored between 3 and 15 percent. A multi-agent tournament reached 42.
A benchmark posted to arXiv on 17 August sets language models a task that sounds almost easy and turns out not to be. Take a research paper. Delete everything except the list of works it cited before publication. Hand a model that list and ask it to recover the paper’s core idea. Across 643 papers in six scientific fields, seven frontier models managed match rates between 3 and 15 percent. The benchmark, from a team led by Shaolong Chen and Yanlin Fei, is called Reconstruction.
The design is the interesting part. A bibliography points at a problem space, but the creative jump from existing literature to a new hypothesis is exactly what the paper itself supplies, so that jump is the only route to a correct answer. To stop models from simply remembering the paper, the authors built in three defences: no cited reference may postdate the paper’s own knowledge horizon, author names and identifying metadata are stripped from references, and each bibliography is frozen so nothing leaks in at test time. An independent model acts as judge, matching proposed hypotheses against the real ones.
What makes the low scores meaningful is how tightly they cluster. If the task were being solved by reasoning, you would expect stronger models to pull ahead. Instead the best barely reached 15 percent and the rest bunched below, which rules out explanations about prompting or model size and points at something structural. One route did work better: a multi-agent pipeline that collects hypotheses from several models, has them critique each other, then eliminates weaker candidates through a bracket-style tournament reached 23 to 42 percent, roughly 2.4 times the best single model. No external search was involved, so the improvement came purely from diversity and argument. The authors note this is a rough computational stand-in for peer review, which is not a coincidence: it is the same error-correction mechanism scientists use.
Some caveats belong here. This is a preprint that has not been through peer review. The tournament approach needs many model calls and coordination infrastructure, and the paper does not detail the compute cost. And 42 percent, while much better, still leaves most ideas unrecovered.
What this means for you: there is now a sizeable industry selling language models as hypothesis-generation engines, and many of those claims rest on tests where the model could see full paper text, author names or post-publication signals. Scores earned under those conditions do not tell you much about the inferential step. If you are evaluating an AI science tool, the question to ask is whether the evaluation was contamination-resistant, meaning whether the answer could have been retrieved rather than reasoned toward. For everyone else, this is a useful piece of ballast against the current mood. The same month brought real results on mathematical proofs and protein design, and those are not diminished by this. But summarising the literature brilliantly and inventing the next idea are different jobs, and current models are visibly much better at the first.
Sources
Source: https://arxiv.org/abs/2608.16645
Anthropic Is Reportedly Paying 7 Billion Dollars for Software That Makes Chips Go Further
Anthropic and Decart are exchanging advanced drafts of a 7 billion dollar acquisition, mostly in shares, with signing targeted for early September. Decart claims models can run up to 8 times faster on the same hardware.