Common Crawl gets a lot of opt-out requests.
We respect publisher rights, even if they don't possess time machines go fix their robots.txt files from 20 years ago. It is what it is. We endeavor to be polite, however.
commoncrawl.org/blog/common-cr…
Our ongoing opt-out registery /
August 2026 Crawl Archive Now Available
We are pleased to announce that the crawl archive for August 2026 is now available, containing 2.14 billion web pages or 360 TiB of uncompressed content.
Fascinating analysis by @pewresearch - How Much of the Internet Is Written With AI?
"To explore this question, we used the Common Crawl web archive to collect almost half a million English-language webpages from the past five years. We then ran the text of those pages through
AI-authored content has become increasingly common online, but where is it coming from?
This growth has been driven almost entirely by commercial (.com) websites.
#AIContent#AIAuthorship
Announcing the First Stable Release of CC-Downloader
Over a year ago we released cc-downloader, an experimental tool to politely download Common Crawl data. Today we're releasing its first stable version, with a Rust library and Python bindings.