DeepSWE
A benchmark that tests whether a model can carry a long software engineering task through to a working result.
DeepSWE tests something narrower and more useful than “can it write code”. It measures whether a model can take a real software task, one that needs many steps across several files, and finish it without losing the thread halfway through. Reading the existing code, working out what to change, making the change, running the tests, fixing what broke.
That endurance is the thing that separates a helpful autocomplete from something you can hand a task to and walk away from, which is why scores on it have been jumping so sharply between model generations. A caveat that applies to all coding benchmarks: they run on curated problems in tidy repositories, and your codebase is neither.
-
Grok 4.7 Thinks Longer for the Same Money, and Wins an Odd Benchmark
-
Sakana's New Model Is Not a Model, It's a Dispatcher for Other Models
-
Google's Gemini 3.8 Flash Keeps the Old Price, and Brings a Cyber Twin
-
DeepSeek Opens Up a 305 Billion Parameter Model That Can Finally See
-
Z.ai's GLM-5.3-Flash Gets Near Opus at a Tenth of the Price
-
A Frontier-Class Model Appeared With No Name on It, and It Is Free Until Next Week
-
GLM-5.3 Got Much Better at Breaking Software, So Z.ai Is Holding the Weights Back
-
Google's Gemini 3.7 Flash Is Better at Code and Costs Half as Much
-
Grok 4.5 Arrives at a Third of the Price, and That Might Matter More Than Benchmarks