Common Crawl is a US nonprofit that has collected the public web since 2008 and publishes the data free for anyone to use. Its crawler, the program that visits pages automatically, is called CCBot. It runs roughly one crawl a month and saves a few billion pages each time.
Researchers and companies use this data as a common starting point for AI training sets. Most of GPT-3's training text came from Common Crawl, and many datasets today are still built on it.
This tool queries Common Crawl's public index and shows which crawls saved your pages and, when a crawl didn't, the reason why.
Each crawl records what happened when CCBot visited your address. Every result is one of these.
- 200Saved
- CCBot fetched the page and stored its HTML, so the page can end up in datasets built from Common Crawl. Open the saved copy to see exactly what CCBot received. If it's nearly empty, the page probably builds its text with JavaScript.
- 301 →Redirect
- CCBot visited the address and was sent to another one (a 301, 302 and so on). Only the target page gets saved, so check that the target is the URL you meant and that it was saved too. Redirects from http to https or to and from www are common and usually fine.
- 403 / 429Error or block
- CCBot reached the address but got an error, so nothing was saved. A 403 (forbidden), 429 (too many requests) or 503 usually means a firewall, CDN or security plugin stopped the bot. A 404 means the page didn't exist at the time.
- Disallow: /Blocked by robots.txt
- Your robots.txt told CCBot not to fetch the page, so it didn't. That's fine if you're opting out of AI training on purpose. If you want AI models to know about you, remove the Disallow rule for CCBot.
- —No record
- None of the crawls checked has any record of the address. New sites and pages with few links often haven't been found yet. Publishing a sitemap and earning links from other sites helps. You can also widen the check to the whole site or another period.