CourionAI
EN
Newsletter
← Glossary Term

Reconstruction benchmark

A 2026 test that gives a model only a paper's reference list and asks it to recover the paper's core idea, with frontier models scoring 3 to 15 percent.

Reconstruction, posted to arXiv in August 2026 by a team led by Shaolong Chen and Yanlin Fei, tests whether a language model can do the specific thing that defines scientific discovery: form a genuinely new hypothesis from prior work. The input is deliberately thin. The model gets the list of papers a research paper cited before publication, and nothing else. No full text, no author names, no clues from after publication.

Across 643 papers in six fields, seven frontier models recovered the core idea between 3 and 15 percent of the time. The narrow spread matters as much as the low numbers, because if the task were being solved by reasoning you would expect stronger models to pull clearly ahead. A multi-agent pipeline that had several models propose ideas, critique each other and eliminate the weakest through a tournament reached 23 to 42 percent, which suggests the useful shape for AI-assisted discovery is a structure resembling peer review rather than one very large model. The paper is a preprint and has not been peer reviewed.