← Glossary Term
Terminal-Bench
A test of whether an AI model can get real work done from a command line.
Terminal-Bench puts a model in front of a terminal, the plain text window developers use to run commands, and gives it jobs that only count as finished if the machine ends up in the right state. Installing something, fixing a broken build, making a test suite pass. There is no partial credit for a plausible-looking answer.
That is why the scores show up in every agent announcement. Chat benchmarks reward good writing about a task; this one rewards actually completing it across many steps without a human nudging things back on track. Scores here have climbed fast in 2026, which is a decent proxy for how much unsupervised work an assistant can be trusted with.
Mentioned in
-
GLM-5.3 Got Much Better at Breaking Software, So Z.ai Is Holding the Weights Back
-
Qwen3.8-27B Is Out, and It Runs on One Gaming Graphics Card
-
Meta gave its AI agent a second agent whose only job is remembering things
-
A Small Open Coding Model Just Beat Rivals Ten Times Its Size, and You Can Run It Yourself
-
Grok 4.5 Arrives at a Third of the Price, and That Might Matter More Than Benchmarks