Notes from HTTP Workshop Basel and IETF 126 Vienna
Two weeks in Basel and Vienna, at the HTTP Workshop and IETF 126. Protocol adoption measured across the whole web, and an attempt to define what "machine readable" actually means.
Common Crawl is a non-profit foundation dedicated to the Open Web.
- July 2026 Crawl Archive Now Available The crawl archive for July 2026 is now available. The data was crawled between July 7th and July 25th, and contains 2.14 billion web pages (or 364.01 TiB of uncompressed content). We also announce some improvements and changes.
- Common Crawl Joins Project Tapestry Common Crawl has joined Project Tapestry, a global initiative led by the AI Alliance to advance open, sovereign AI. We will contribute our expertise in responsible web data, multilingual coverage and culturally informed AI development. 📷
- Measuring Crawled Coverage of a Website in Common Crawl How can we measure how many pages we’ve crawled from a particular website? The answer is a lot more complicated than you might think.
- Congrats on [email protected] getting 5000 Github stars on his project GROBID! A machine learning software for extracting information from scholarly documents grobid.readthedocs.io/en/latest/

