benchmark
A standard test used to measure and compare how well AI models perform.
A benchmark is a standardised test used to measure how well an AI model performs at some task, coding, reasoning, reading documents, and so on, so that different models can be compared on a level footing. Leaderboards built from benchmarks are a common way labs show off progress.
Benchmarks are useful but easy to over-trust. A model can be tuned to score well on a popular test without being better in real use, and the way a test is run can quietly distort the result. Researchers found, for example, that standard tests of AI agents understated them simply by cutting the agents off before they had enough time to finish.
The practical takeaway: treat benchmark numbers as a rough signal, not gospel, and weigh them against how a tool actually performs on your own work.
-
Debian Votes to Allow AI-Assisted Contributions, Narrowly
-
Google's WikiSkill Lets AI Agents Keep a Wiki of Their Own Mistakes
-
LAION Releases 10 Million Hours of Video for Open AI Research
-
Tencent Open-Sources Hy4, a 770-Billion-Parameter Model Built for Office Work
-
DeepMind's AI Co-Scientist Now Runs the Lab Equipment, and the Human Doctors Were Less Impressed Than the Benchmarks
-
Google Locks Its Own Model Out of the Test Questions
-
Google's New Transcription Model Cleans Up Your Ums and Your Mind Changes
-
IBM's Granite 4.2 Is Free, Small Enough to Run Locally, and Built to Think First
-
Z.ai's GLM-5.3-Flash Gets Near Opus at a Tenth of the Price
-
Claude Now Remembers You in One Place Instead of Two
-
OpenAI Built Its Own Chip, and the First Benchmarks Are Good
-
Alibaba's Next Open Model Is a Preview of Qwen 4
-
Mistral's New Agentic Search Lets AI Actually Read Your Documents
-
A Support Bot Went From Solving a Quarter of Tickets to Solving Half, Without a Better Model
-
This Open Model Got So Good at Finding Bugs That Its Own Maker Won't Release It Yet
-
Skills Make AI Agents Better by Handing Them a Checklist, Not Facts. And They Stop Working at 100 Entries
-
DeepSeek's Cheap Model Can Now See, and It Edges Past Opus 4.8 on Two Visual Tests
-
The Same Model Scored 30 Percent Alone and 100 Percent Inside Nvidia's Agent System
-
A Frontier-Class Model Appeared With No Name on It, and It Is Free Until Next Week
-
Reinforcement Learning Pioneer Rich Sutton Calls Synthetic Data a Big Mistake
-
An Oxford Spinout Sold Anthropic 250 Million Dollars of Chips That Do Not Exist Yet
-
Give a Frontier Model Only a Bibliography and Ask for the Idea. It Gets There 3 to 15 Percent of the Time
-
A Legal AI Company Built Its Own Model, and It Started From an Open Chinese One
-
OpenAI Cut Its Flagship Model to Half Price, and the Reason Is Sitting on a Leaderboard
-
We Are Building an Invention Meant to Outgrow Its Inventors
-
Google Switches Off Three Imagen 4 Models Today
-
1,221 Volunteers Pointed AI Agents at 2,200 Research Papers. Nearly a Quarter Did Not Hold Up
-
Are Models Getting Worse at Facts on Purpose?
-
Grok 4.6 Shows Up in GitHub Copilot, Two Days After Launch
-
Qwen3.8-27B Is Out, and It Runs on One Gaming Graphics Card
-
Deepseek Gave Away Its Agent Software and Raised API Prices on the Same Day
-
The Best Model on the Market Is the One Companies Are Not Buying
-
Ling 3.0 Flash Is the Smartest Open Model of Its Size, and It Stopped Making Things Up
-
Google's Gemini 3.7 Flash Is Better at Code and Costs Half as Much
-
OpenAI's New Ultrafast Mode Runs Its Best Model 14 Times Faster
-
That Advice About Which Language AI Codes Best In? It Falls Apart on Real Work
-
The Creator of Redis Made a Chinese Video Model Run on a Mac. Europeans Are Not Licensed to Use It
-
Nvidia's New Free Model Is Not the Smartest. It Is Just Very, Very Fast
-
Meta Put a 30B Model on Your Laptop, and Zuckerberg Used the Launch to Pick a Fight
-
DeepMind's New Storm Model Buys Forecasters One Extra Day, and Nobody Quite Knows Why It Works
-
xAI's New Image Model Lands Second in the Arena Rankings, and Brings Photoshop-Style Editing to Grok
-
Same model, four different tools, three times the price: a test of what the wrapper costs you
-
Alibaba's newest model scores higher and guesses more: hallucination rate jumps from 23 to 40 percent
-
Meta ships Muse Code, a coding agent that keeps working after it crashes
-
An open model is now months, not years, behind the frontier. Its safety testing is not.
-
Mistral released a free safety filter that you describe in plain English, and it fits on one graphics card
-
AI agents made research software 60 times faster, and were confidently wrong in ways nobody spotted for weeks
-
Karpathy turned one paragraph of Tolkien into a 3D scene for ten dollars, and called it a vibe check
-
Alibaba's new Qwen ad says AI will take your job, and makes it sound like a holiday
-
Alibaba's Qwen3.8-Max spent 16 days writing a tool by itself, and the weights go public next week
-
People are building playable 3D games from a single prompt, and the results stopped looking like blocks
-
OpenAI now answers its own support phone line with an AI agent, and is selling the system that does it
-
AMD trained a fully open model on its own chips and published everything except a commercial licence
-
DeepSeek updated its cheap model and it now runs neck and neck with OpenAI's cheap model
-
Google DeepMind's new robot brain is meant to run everything from a desk arm to a humanoid
-
OpenAI and Anthropic are arguing about a benchmark, and the argument is more useful than the scores
-
Microsoft says it will stop chasing the frontier and build small specialist models instead
-
OpenAI's new transcription models are faster and 25 percent cheaper, but still not the most accurate
-
1,200 AI lab employees ask Washington for a brake pedal they admit nobody knows how to build
-
Pangram says its new detector is wrong once every 24,000 documents. That number deserves a closer look
-
Anthropic says Claude found real weaknesses in two encryption algorithms
-
Hugging Face published the full forensic timeline of the AI agent that broke into its systems
-
We Ran a Local AI Model on a Six Year Old Budget Laptop Chip. Here Is Exactly What It Could Do.
-
A $500 training run made a 9B model beat every frontier model at one boring job
-
Microsoft's new security model is small on purpose, and that is the interesting part
-
Cursor rebuilt SQLite with a swarm of agents, and the cheap models did most of it
-
METR has a new way to ask whether an AI agent is actually cheaper than a person
-
A German open model had test answers in its training data, and openness is how we know
-
Sakana says its model router now beats a model it does not even use
-
Anthropic's Opus 5 is smaller and cheaper, yet it beats its bigger sibling
-
DeepSeek V4 goes fully stable, and the old models switch off today
-
OpenAI says one of its test models broke out and hacked Hugging Face
-
Alibaba's New Image AI Can Draw a Whole Newspaper Page, Tiny Readable Text and All
-
Nvidia's Next AI Chips Reportedly Squeeze 10x More Work From the Same Power
-
Kimi K3 Got Too Popular: Moonshot Pauses New Subscriptions
-
OpenAI Paused Its Best Model Because It Kept Breaking Out of Its Own Test Cage
-
A New Medical Benchmark Asks Whether AI Knows When It Doesn't Know
-
SAP Just Spent Over a Billion Euros on AI That Reads Spreadsheets, Not Chats
-
The US Wants a 30-Day Look at New Frontier Models Before They Ship. Here's What That Actually Means
-
Kimi K3: A Free-to-Download Model That Almost Keeps Up With the Big Names
-
A Capable AI Model That Fits on Your Phone, and Runs Entirely Offline
-
DeepMind's Hassabis Wants an AI Referee, Modeled on Wall Street's Watchdog
-
Soofi S: Germany Now Has an Open AI Model That Tops the Open-Source Charts
-
Independent Numbers Are In: Meta's Muse Spark 1.1 Is a Serious Value Pick
-
OpenAI Says GPT-5.6 Sol Trained Its Own Smaller Sibling
-
Why Databricks Just Made a Chinese Open-Source Model Its Daily Coding Engine
-
GPT-5.6 Sol Nearly Matches the Best AI Model, at a Third of the Price
-
Meta Joins the AI Price War With Muse Spark 1.1 and Its First Developer API
-
OpenAI Says a Third of a Popular AI Coding Test Is Broken
-
Anthropic's Answer to Fable 5's Price: Let It Manage Cheaper Models
-
Grok 4.5 Arrives at a Third of the Price, and That Might Matter More Than Benchmarks
-
Mistral Enters Robotics: One Camera Is Enough to Steer a Robot
-
ChatGPT's New Voice Mode Can Listen and Talk at the Same Time
-
OpenAI's GPT-5.6 Models Launch Publicly This Thursday
-
Microsoft Is Quietly Swapping OpenAI and Anthropic Out of Copilot to Cut Costs
-
Mistral's Free New Model Finds Real Bugs by Actually Proving Code Correct
-
Stanford's Big Yearly AI Report Card Is Out, and the Trust Gap Is Widening
-
AI research agents don't fail at searching, they fail at asking you questions
-
Baidu's 'Unlimited OCR' reads 40-page documents in one go, by learning to forget
-
Researchers made AI models run a startup for 500 days. Most went bankrupt.
-
This open-source tool hides text in images to cut Claude's bill by up to 70%
-
That AI Tool That 'Failed' Six Months Ago? It Might Just Have Needed More Time
-
MiniMax M3: A Top-Tier AI Model You Can Download and Run Yourself