Training data
The pile of text, images and code a model learns from, and the part labs are least willing to describe.
Training data is everything a model was shown while it was being built: web pages, books, code repositories, images, transcripts. The model does not store these files. It adjusts millions of internal numbers, its model weights, until it gets good at predicting what comes next in material like that. Everything a model appears to know traces back to what it was trained on, which is also why its blind spots and biases tend to mirror its sources.
This is the least transparent layer of modern AI. Most labs will tell you how big a model is and how it scores on benchmarks, but not what went into it, partly for competitive reasons and partly because the copyright questions around scraped material are still moving through courts. Projects such as OpenWALDO are trying to change that by building corpora where every document carries a licence and an origin.
-
Hollywood and TikTok's Owner Signed a Truce on AI Video, but Not the Part That Matters Most
-
We Are Building an Invention Meant to Outgrow Its Inventors
-
Debian Is Voting on Whether AI May Touch Its Code
-
Rocky Linux's Founder Wants to Open the One Part of AI Nobody Opens: the Training Data
-
Small Models Read Old Books Better Than Big Ones, and It Costs Two Dollars per Thousand Pages
-
An open model is now months, not years, behind the frontier. Its safety testing is not.
-
Karpathy turned one paragraph of Tolkien into a 3D scene for ten dollars, and called it a vibe check
-
Mira Murati's lab shrank its own model to a quarter of the size and lost one point
-
An OpenAI researcher quit after eight months, betting that better data matters more than bigger models
-
Google's Lyria 3.5 lets you fix one part of an AI song without redoing the whole thing
-
Computing's biggest professional body is debating whether to open its library to AI
-
An Indian court just called AI training 'private use', and publishers should read the small print
-
Who Is Who: Black Forest Labs, the German Lab Inside Photoshop and Canva
-
Robots are running out of training data, so one startup is recording brain waves
-
Not everyone believes OpenAI's rogue agent story, and the doubts are worth hearing
-
A German open model had test answers in its training data, and openness is how we know
-
A Judge Just Approved Anthropic's $1.5 Billion Payout to Authors, the Biggest Copyright Deal Yet
-
A Hack Revealed What an AI Music Generator Was Trained On, and It's Two Million YouTube Clips
-
Same Question, Different Answer: Claude Is Warmer in Hindi, Stricter in Russian
-
Google's SensorFM Wants to Be One AI Brain for All Your Wearable's Health Data
-
Why Databricks Just Made a Chinese Open-Source Model Its Daily Coding Engine
-
MiniMax Reportedly Plans a 2.7 Trillion Parameter Open-Source Model
-
Stanford's Big Yearly AI Report Card Is Out, and the Trust Gap Is Widening