← Glossary Term
SWE-Bench
A widely used test of how well AI models fix real software bugs.
SWE-Bench is a benchmark that measures whether an AI can resolve real issues from real software projects, not toy puzzles. Scores on it are the most-quoted numbers in any coding-model announcement, and the harder SWE-Bench Pro variant arrived once models began saturating the original.
Benchmarks are imperfect, models can be tuned to the test, but SWE-Bench remains the quickest honest signal of coding ability.
Mentioned in
-
Qwen3.8-27B Is Out, and It Runs on One Gaming Graphics Card
-
Meta Put a 30B Model on Your Laptop, and Zuckerberg Used the Launch to Pick a Fight
-
Why Databricks Just Made a Chinese Open-Source Model Its Daily Coding Engine
-
OpenAI Says a Third of a Popular AI Coding Test Is Broken
-
Anthropic's Answer to Fable 5's Price: Let It Manage Cheaper Models
-
MiniMax M3: A Top-Tier AI Model You Can Download and Run Yourself