← Glossary Term
Agents' Last Exam
A benchmark of long multi-step tasks inside real applications, used to test whether an agent can finish work rather than just start it.
Agents’ Last Exam, usually shortened to ALE, tests agents the way a new colleague gets tested: here is a real application, here is a task with many steps, go. Success means the end state is correct, not that the individual moves looked plausible along the way.
That framing punishes exactly the failure mode people complain about most. A model can produce a confident, well-worded plan and still leave the actual job half done, and a benchmark that grades text will not notice. ALE only counts what changed. Scores are low across the board, which tells you honestly where autonomous agents currently sit.
Mentioned in