inference
Running a trained AI model to get an answer, as opposed to training it in the first place.
Inference is the act of using a trained AI model, feeding it a prompt and getting a response. It is distinct from training, which is the far more expensive, one-time process of building the model’s weights in the first place.
The distinction is useful for understanding AI costs. Training a frontier model can require months of work and enormous compute, but that cost is paid once. Inference is paid every time someone uses the model, which is why providers meter it, usually by the token, and why efficiency improvements that make inference faster or cheaper matter so much.
It is also why running smaller models locally is practical: even a capable laptop can handle inference for many tasks, even though it could never have trained the model itself.
-
Tencent Open-Sources Hy4, a 770-Billion-Parameter Model Built for Office Work
-
OpenAI Built Its Own Chip, and the First Benchmarks Are Good
-
Anthropic Hired the Man Who Built Google's AI Chips, and It Is Not Subtle About Why
-
An Oxford Spinout Sold Anthropic 250 Million Dollars of Chips That Do Not Exist Yet
-
Watching Its Own Model Now Costs OpenAI 20 Percent Extra, and It Paused a Training Run to Do It
-
Anthropic Is Reportedly Paying 7 Billion Dollars for Software That Makes Chips Go Further
-
A Chip Startup Doubled Its Value in a Month by Being Deliberately Inflexible
-
We Are Building an Invention Meant to Outgrow Its Inventors
-
The Most Valuable AI Skill Is Not Prompting. It Is Verification.
-
Someone Else Is Paying for Your AI, Just Not the Part You Think
-
Backflip AI Turns 3D Scans Into Editable CAD Models, a Job That Usually Eats Hours
-
AMD Buys Taalas, a Startup That Bakes an Entire AI Model Into the Chip
-
Cloudflare wants to give your AI agent a wallet, and you the spending limit
-
Someone got DeepSeek V4 Flash running in production on a single AMD card and published every patch it took
-
AMD trained a fully open model on its own chips and published everything except a commercial licence
-
Amazon quietly stops developing most of its own AI models
-
AMD and Cerebras team up to make AI answers come back faster
-
Nvidia's Next AI Chips Reportedly Squeeze 10x More Work From the Same Power
-
Kimi K3 Got Too Popular: Moonshot Pauses New Subscriptions
-
A Hack Revealed What an AI Music Generator Was Trained On, and It's Two Million YouTube Clips
-
DeepSeek Wants Fresh Billions Weeks After Its First Round, Cheap AI Is Expensive
-
DeepSeek Is Designing Its Own AI Chip, and Taking Outside Money for the First Time