Big Video Dataset
LAION's open video training collection from August 2026, covering roughly 10 million hours of footage with machine-written captions for picture and sound.
The Big Video Dataset is an open training collection published by the non-profit LAION. The team started from 1.3 billion video links found in Common Crawl, downloaded 80 million of them totalling around 10 million hours, and cut those into 55 million clips plus 300 million still images. Most of the material comes from YouTube and most of it is in English, which tells you something useful about the bias any model trained on it will inherit.
What makes it valuable for research is the labelling. Rather than depending on whatever description a human happened to write, the team generated captions for both the visuals and the audio automatically, so a model can learn which pictures go with which words and which sounds all at once. It is released for non-commercial research only, resting on the same German research exemption LAION relies on for its image datasets.