1. X
  2. Common Crawl Foundation
Log inSign up
Common Crawl Foundation
1,479 posts
Common Crawl Foundation profile banner
@CommonCrawl

Common Crawl Foundation

@CommonCrawl
Common Crawl is a non-profit foundation dedicated to the Open Web.
San Francisco, CA
commoncrawl.org
Joined February 2010
1,611
Following
7,818
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @CommonCrawl
    Common Crawl Foundation
    @CommonCrawl
    8h
    Thom Vaughan of Common Crawl, speaking from the Paris Open Source AI Summit recently. opensourceaisummit.eu
    Image
    00:00
  • @CommonCrawl
    Common Crawl Foundation
    @CommonCrawl
    Aug 26
    Common Crawl gets a lot of opt-out requests. We respect publisher rights, even if they don't possess time machines go fix their robots.txt files from 20 years ago. It is what it is. We endeavor to be polite, however. commoncrawl.org/blog/common-cr… Our ongoing opt-out registery /
    Image
    Common Crawl - Blog - Common Crawl Foundation Opt-Out Registry
    From commoncrawl.org
  • @CommonCrawl
    Common Crawl Foundation
    @CommonCrawl
    Aug 25
    August 2026 Crawl Archive Now Available We are pleased to announce that the crawl archive for August 2026 is now available, containing 2.14 billion web pages or 360 TiB of uncompressed content.
    Image
  • @CommonCrawl
    Common Crawl Foundation
    @CommonCrawl
    Aug 22
    Fascinating analysis by @pewresearch - How Much of the Internet Is Written With AI? "To explore this question, we used the Common Crawl web archive to collect almost half a million English-language webpages from the past five years. We then ran the text of those pages through
    @pewresearch
    Pew Research Center
    @pewresearch
    Aug 21
    AI-authored content has become increasingly common online, but where is it coming from? This growth has been driven almost entirely by commercial (.com) websites. #AIContent #AIAuthorship
    Chart showing AI authorship is less common in .edu and .gov domains
  • @CommonCrawl
    Common Crawl Foundation
    @CommonCrawl
    Aug 10
    Announcing the First Stable Release of CC-Downloader Over a year ago we released cc-downloader, an experimental tool to politely download Common Crawl data. Today we're releasing its first stable version, with a Rust library and Python bindings.
    Image
Advertisement
Advertisement