← Glossary Term
GDPval
An OpenAI benchmark that tests models on realistic office work rather than puzzles.
GDPval measures how well a model handles the kind of tasks people are actually paid to do: drafting documents, working through spreadsheets, pulling together research, multi-step jobs with messy inputs. It was built by OpenAI as a counterweight to benchmarks made of exam questions, which reward a narrow kind of cleverness.
Results are reported as Elo points rather than a percentage, so the useful reading is the gap between models rather than the absolute number. Since OpenAI both created the benchmark and competes on it, it is worth treating the scores as one input among several rather than a neutral verdict.
Mentioned in
-
OpenAI's New Ultrafast Mode Runs Its Best Model 14 Times Faster
-
Nvidia's New Free Model Is Not the Smartest. It Is Just Very, Very Fast
-
Alibaba's newest model scores higher and guesses more: hallucination rate jumps from 23 to 40 percent
-
DeepSeek updated its cheap model and it now runs neck and neck with OpenAI's cheap model