← Glossary Term
Terminal-Bench
A test of whether an AI model can get real work done from a command line.
Terminal-Bench puts a model in front of a terminal, the plain text window developers use to run commands, and gives it jobs that only count as finished if the machine ends up in the right state. Installing something, fixing a broken build, making a test suite pass. There is no partial credit for a plausible-looking answer.
That is why the scores show up in every agent announcement. Chat benchmarks reward good writing about a task; this one rewards actually completing it across many steps without a human nudging things back on track. Scores here have climbed fast in 2026, which is a decent proxy for how much unsupervised work an assistant can be trusted with.
Mentioned in
-
Grok 4.7 Thinks Longer for the Same Money, and Wins an Odd Benchmark
-
Sakana's New Model Is Not a Model, It's a Dispatcher for Other Models
-
Anthropic's Fable 5.1 Costs the Same, Except for the Part Agents Use Most
-
DeepSeek Opens Up a 305 Billion Parameter Model That Can Finally See
-
Z.ai's GLM-5.3-Flash Gets Near Opus at a Tenth of the Price
-
GLM-5.3 Got Much Better at Breaking Software, So Z.ai Is Holding the Weights Back
-
Qwen3.8-27B Is Out, and It Runs on One Gaming Graphics Card
-
Meta gave its AI agent a second agent whose only job is remembering things
-
A Small Open Coding Model Just Beat Rivals Ten Times Its Size, and You Can Run It Yourself
-
Grok 4.5 Arrives at a Third of the Price, and That Might Matter More Than Benchmarks