Umoren.ai
Free · Check AI training data

Is your site inCommon Crawl's record?

Common Crawl is a nonprofit that has archived the public web every month since 2008, and many AI training datasets start from it. Enter a page or site to see which crawls saved it, and why the bot left without it when it didn't: robots.txt, a firewall, or a redirect.

What to checkWhat to check
PeriodPeriod

How it works

Step 01

Enter an address

Type a page or a domain and we check the crawls from the past 12 months, or any period of up to a year since 2008. With a free account you can check a whole site or include its subdomains.

Step 02

Read the record

The checker asks Common Crawl's index about each crawl, newest first, and collects what it saved, redirected or failed on.

Step 03

See why

When a page is missing, it reads the robots.txt CCBot saw and the block page it was shown, then explains the cause and the fix.

What is Common Crawl?

Common Crawl is a US nonprofit that has collected the public web since 2008 and publishes the data free for anyone to use. Its crawler, the program that visits pages automatically, is called CCBot. It runs roughly one crawl a month and saves a few billion pages each time.

Researchers and companies use this data as a common starting point for AI training sets. Most of GPT-3's training text came from Common Crawl, and many datasets today are still built on it.

This tool queries Common Crawl's public index and shows which crawls saved your pages and, when a crawl didn't, the reason why.

What the results mean

Each crawl records what happened when CCBot visited your address. Every result is one of these.

200Saved
CCBot fetched the page and stored its HTML, so the page can end up in datasets built from Common Crawl. Open the saved copy to see exactly what CCBot received. If it's nearly empty, the page probably builds its text with JavaScript.
301 →Redirect
CCBot visited the address and was sent to another one (a 301, 302 and so on). Only the target page gets saved, so check that the target is the URL you meant and that it was saved too. Redirects from http to https or to and from www are common and usually fine.
403 / 429Error or block
CCBot reached the address but got an error, so nothing was saved. A 403 (forbidden), 429 (too many requests) or 503 usually means a firewall, CDN or security plugin stopped the bot. A 404 means the page didn't exist at the time.
Disallow: /Blocked by robots.txt
Your robots.txt told CCBot not to fetch the page, so it didn't. That's fine if you're opting out of AI training on purpose. If you want AI models to know about you, remove the Disallow rule for CCBot.
—No record
None of the crawls checked has any record of the address. New sites and pages with few links often haven't been found yet. Publishing a sitemap and earning links from other sites helps. You can also widen the check to the whole site or another period.

Why Common Crawl matters

For AI search optimization, the starting point is making sure AI models know about you.

1

Where AI training data starts

Many LLM training sets for GPT, Claude, Gemini and others start from Common Crawl. If you're not in it, you're less likely to be in them.

2

Find blocks you didn't mean

Cloudflare's AI bot blocking or a security plugin often turns CCBot away without the site owner knowing.

3

Catch JavaScript-only pages

CCBot doesn't run JavaScript. If the saved HTML is nearly empty, models can't learn what your page says.

FAQ

Q

What is Common Crawl?

A

Common Crawl is a US nonprofit that has collected the public web since 2008 and publishes the data free for anyone to use. Its bot is called CCBot. It runs a new crawl roughly once a month, each holding a few billion pages. The data is hosted on AWS and is one of the most common raw materials for research and AI training datasets.

Q

Why isn't my page in it?

A

There are three usual reasons. First, CCBot may not have found the page yet. It discovers pages through links and sitemaps, so new sites and pages with few links can take several crawls to appear, and no crawl saves the whole web, so even a known site won't have every page saved every time. Second, your robots.txt may block CCBot. Third, a bot blocker at your CDN, host or security plugin may be turning it away. This tool checks the second and third for you.

Q

What does a redirect or error result mean?

A

Both mean CCBot reached the address but didn't save a page there. For a redirect (301, 302 and so on), only the target gets saved, so check that the target is the right URL and that it was saved. A 403 (forbidden), 429 (too many requests) or 503 usually means a firewall or rate limit stopped the bot: on Cloudflare, look at AI Crawl Control and Bot Fight Mode; on WordPress, look at your security plugin. A 404 means the page didn't exist at the time.

Q

Why is the saved copy nearly empty?

A

CCBot saves the first HTML your server sends and doesn't run JavaScript. Pages that draw their text in the browser with React, Vue and the like are saved as near-empty shells, and datasets built from Common Crawl see them that way too. Server-side rendering (SSR) or static generation (SSG) puts the text in the HTML so it gets saved.

Q

Does being in Common Crawl mean an AI model learned my page?

A

Not necessarily. AI labs heavily filter Common Crawl before training, dropping duplicates, low-quality pages and languages they don't need. A new crawl also takes months to a year or more to reach a released model. And AI tools that search the web live, like ChatGPT search and Perplexity, use their own crawlers. Being in Common Crawl is one prerequisite for AI models knowing about you, not a guarantee.

Q

Does the checker store what I check?

A

Partly, yes. To spare Common Crawl's servers and keep checks fast, we cache the results we fetch from Common Crawl (which are public data) on our server. If you're logged in, your results are saved to your account history so you can come back to them. We also log the address checked and a summary of the result to improve the service and prevent abuse. We never visit or fetch anything from your own site.

Q

How often does Common Crawl crawl?

A

In recent years a new crawl has come out roughly every month. The early years (about 2008 to 2012) had only a few large crawls covering several years each, so month-by-month comparisons start in 2013. By default this tool checks the crawls from the past 12 months.

Q

How do I allow CCBot?

A

In robots.txt, remove any "User-agent: CCBot" group with "Disallow: /". A blanket "User-agent: *" disallow applies to CCBot too. Even with a clean robots.txt, Cloudflare's AI bot blocking or Bot Fight Mode, a WordPress security plugin or your host's firewall can still turn CCBot away. If the results name a blocker, start with that setting.

Q

Can I check older periods?

A

Yes. Choose "Choose a period" to pick any span of up to a year from the crawls published since 2008. It's handy for seeing whether a redesign changed how CCBot was treated, or when a block started.

Q

How does this relate to LLMO?

A

LLMO (AI search optimization) starts with making sure AI models know accurate facts about you, and being in Common Crawl is one prerequisite. umoren.ai helps with everything from crawl checks to improving how you show up in AI answers.

Common Crawl Checker: Is Your Site in AI Training Data? | umoren.ai