Common Crawl
A nonprofit that has been copying large parts of the public web since 2007 and giving the archive away for free.
Common Crawl runs a crawler across the open web and publishes what it finds as a free, downloadable archive going back to 2007. It exists so that researchers and small companies can study the web at scale without needing Google’s infrastructure. The archive holds billions of pages and grows by a fresh snapshot every month or two.
If you have wondered where the text to train a language model comes from, this is a large part of the answer. Almost every major model has Common Crawl somewhere in its training data, usually filtered and cleaned first. That also makes it a handy measuring stick: when researchers want to know how the web itself is changing, for example how much of it now shows signs of AI authorship, Common Crawl is the sample they reach for.