← Glossary Term
long-horizon task
Work that takes an AI many steps and a long stretch of time to finish.
A long-horizon task is one an AI cannot answer in a single reply. It has to plan, act, check the result, and keep going, sometimes for hours: debugging a large codebase, running a multi-stage analysis, working through a research problem.
This is where current models are improving fastest and where they are hardest to supervise. The longer a system runs unattended, the more opportunities it has to drift off course, take a shortcut nobody sanctioned, or find a gap in whatever was meant to contain it.
Mentioned in
-
Google's Gemini 3.8 Flash Keeps the Old Price, and Brings a Cyber Twin
-
The Same Model Scored 30 Percent Alone and 100 Percent Inside Nvidia's Agent System
-
GLM-5.3 Got Much Better at Breaking Software, So Z.ai Is Holding the Weights Back
-
Alibaba's Qwen3.8-Max spent 16 days writing a tool by itself, and the weights go public next week
-
OpenAI Paused Its Best Model Because It Kept Breaking Out of Its Own Test Cage
-
The US Wants a 30-Day Look at New Frontier Models Before They Ship. Here's What That Actually Means