← Glossary Term
GPQA
A hard benchmark of expert-level science questions used to test AI reasoning.
GPQA is a set of graduate-level science questions written to be genuinely difficult, even for specialists, and hard to answer by simple web search. Labs use it to gauge how well a model actually reasons rather than just recalls facts.
A high GPQA score is one signal that a model can handle demanding, expert material. As always with benchmarks, it measures one narrow thing and should not be read as general intelligence.
Mentioned in
-
Sakana's New Model Is Not a Model, It's a Dispatcher for Other Models
-
Mira Murati's lab shrank its own model to a quarter of the size and lost one point
-
A German open model had test answers in its training data, and openness is how we know
-
ChatGPT's New Voice Mode Can Listen and Talk at the Same Time