reinforcement learning
A training method where a model learns by being rewarded for good answers instead of being shown correct ones.
Reinforcement learning is a way of training a model by trial, error and reward. Instead of showing it thousands of correct answers to copy, you let it attempt something, score the attempt, and nudge it toward whatever earned a higher score. It is roughly how you would teach a dog a trick, except the dog is a few hundred billion numbers and the treat is a mathematical signal. Most of the recent jump in reasoning model quality came from this, not from bigger models.
The catch is that it needs a reliable way to tell good from bad. Code either compiles or it does not, a maths proof either checks out or it does not, so those areas improved fast. For work where nobody can automatically score the result, like writing a tactful email or judging a business decision, the method has much less to grip on. That is one reason models feel uneven: they were trained hard where scoring was easy.
-
Watching Its Own Model Now Costs OpenAI 20 Percent Extra, and It Paused a Training Run to Do It
-
OpenAI Froze Its Own Training for Two Weeks, and Now Watches Its Models Like a Hawk
-
Google Turned a Finished Model Into a Much Faster One, and Published the Recipe
-
Alibaba's Qwen3.8-Max spent 16 days writing a tool by itself, and the weights go public next week
-
An OpenAI researcher quit after eight months, betting that better data matters more than bigger models
-
A $500 training run made a 9B model beat every frontier model at one boring job
-
A Turing Award Winner Bets Against Deep Learning With His New Startup Oak Lab
-
Mistral Enters Robotics: One Camera Is Enough to Steer a Robot