CourionAI
EN
Newsletter
← Glossary Term

Terminal-Bench

A test of whether an AI model can get real work done from a command line.

Terminal-Bench puts a model in front of a terminal, the plain text window developers use to run commands, and gives it jobs that only count as finished if the machine ends up in the right state. Installing something, fixing a broken build, making a test suite pass. There is no partial credit for a plausible-looking answer.

That is why the scores show up in every agent announcement. Chat benchmarks reward good writing about a task; this one rewards actually completing it across many steps without a human nudging things back on track. Scores here have climbed fast in 2026, which is a decent proxy for how much unsupervised work an assistant can be trusted with.