Training data
The pile of text, images and code a model learns from, and the part labs are least willing to describe.
Training data is everything a model was shown while it was being built: web pages, books, code repositories, images, transcripts. The model does not store these files. It adjusts millions of internal numbers, its model weights, until it gets good at predicting what comes next in material like that. Everything a model appears to know traces back to what it was trained on, which is also why its blind spots and biases tend to mirror its sources.
This is the least transparent layer of modern AI. Most labs will tell you how big a model is and how it scores on benchmarks, but not what went into it, partly for competitive reasons and partly because the copyright questions around scraped material are still moving through courts. Projects such as OpenWALDO are trying to change that by building corpora where every document carries a licence and an origin.
-
Real People Are Reading Real ChatGPT Conversations, and Most Users Have No Idea
-
Chatbots Told a Third of Simulated Patients Their Sleep Apnoea Could Wait
-
A 2.5 Billion Parameter Model That Beats Bigger Ones, and Runs on Your Laptop
-
The US Justice Department Tells a Court That Training AI on Copyrighted Books Is Fair Use
-
K2 Horizon Ships Six Open Models, and the Training Data With Them
-
Meta's Cheapest Model Tier Costs 21 Times Less, and Trains on Your Prompts
-
DeepSeek Opens Up a 305 Billion Parameter Model That Can Finally See
-
LAION Releases 10 Million Hours of Video for Open AI Research
-
Google Locks Its Own Model Out of the Test Questions
-
Amazon Is Closing Mechanical Turk, the Human Workforce It Once Sold as AI
-
This Robot Learns a Ten-Minute Task From Watching One Video
-
X Shuts Down Nitter, and the Open Web Gets a Little Smaller
-
Anthropic Opened a Free School for Learning Its Own AI, 355 Lessons and No Credit Card
-
Adobe Firefly Can Now Make the Music, the Voiceover and the Sound Effects
-
This Robot Learns a New Task From One Short Video, No Training Required
-
Reinforcement Learning Pioneer Rich Sutton Calls Synthetic Data a Big Mistake
-
Hollywood and TikTok's Owner Signed a Truce on AI Video, but Not the Part That Matters Most
-
We Are Building an Invention Meant to Outgrow Its Inventors
-
Debian Is Voting on Whether AI May Touch Its Code
-
Rocky Linux's Founder Wants to Open the One Part of AI Nobody Opens: the Training Data
-
Small Models Read Old Books Better Than Big Ones, and It Costs Two Dollars per Thousand Pages
-
An open model is now months, not years, behind the frontier. Its safety testing is not.
-
Karpathy turned one paragraph of Tolkien into a 3D scene for ten dollars, and called it a vibe check
-
Mira Murati's lab shrank its own model to a quarter of the size and lost one point
-
An OpenAI researcher quit after eight months, betting that better data matters more than bigger models
-
Google's Lyria 3.5 lets you fix one part of an AI song without redoing the whole thing
-
Computing's biggest professional body is debating whether to open its library to AI
-
An Indian court just called AI training 'private use', and publishers should read the small print
-
Who Is Who: Black Forest Labs, the German Lab Inside Photoshop and Canva
-
Robots are running out of training data, so one startup is recording brain waves
-
Not everyone believes OpenAI's rogue agent story, and the doubts are worth hearing
-
A German open model had test answers in its training data, and openness is how we know
-
A Judge Just Approved Anthropic's $1.5 Billion Payout to Authors, the Biggest Copyright Deal Yet
-
A Hack Revealed What an AI Music Generator Was Trained On, and It's Two Million YouTube Clips
-
Same Question, Different Answer: Claude Is Warmer in Hindi, Stricter in Russian
-
Google's SensorFM Wants to Be One AI Brain for All Your Wearable's Health Data
-
Why Databricks Just Made a Chinese Open-Source Model Its Daily Coding Engine
-
MiniMax Reportedly Plans a 2.7 Trillion Parameter Open-Source Model
-
Stanford's Big Yearly AI Report Card Is Out, and the Trust Gap Is Widening