janibert1/robots-check

See what your site's robots.txt actually serves -- including whether Cloudflare's AI Crawl Control is silently overriding it.

1

stars

1

commits

JavaScript

primary language

Sep 9, 2026

updated

github.com/janibert1/robots-check#readme
cli
cloudflare
devtools
github-action
nodejs
robots-txt
seo

README

robots-check

See what your site's robots.txt actually serves — including whether Cloudflare's "AI Crawl Control" feature is silently appending its own rules on top of yours.

Why this exists

While deploying a few small sites behind Cloudflare, we noticed our own robots.txt files weren't what was actually being served. Cloudflare injects a managed block (look for # BEGIN Cloudflare Managed content) that adds its own AI-crawler-blocking rules — real search engines are usually left alone, but plenty of site owners have no idea this is happening at all, and it's easy to assume the file in your repo/origin is the one being served when it isn't.

robots-check fetches the live file over HTTPS and tells you, in plain terms:

  • whether it's Cloudflare-managed (and where to check/adjust it)
  • whether real search engines (Googlebot, Bingbot, etc.) are actually allowed to crawl
  • which AI training/crawling bots (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, and others) are blocked vs. allowed
  • whether a sitemap is declared at all

No API keys, no signup, no dependencies — just Node's built-in https module.

Usage

npx github:janibert1/robots-check example.com

Or clone it and run directly:

git clone https://github.com/janibert1/robots-check
cd robots-check
node robots-check.js example.com

Requires Node 18+.

Example output

example.com — 47 non-blank lines

⚠ This robots.txt is (at least partly) Cloudflare-managed.
  Cloudflare's "AI Crawl Control" feature injects its own block into what's
  actually served — your origin server's own robots.txt file may say
  something different from what's shown below. Check your Cloudflare
  dashboard under Bots → AI Crawl Control if this doesn't match your intent.

Real search engines:
  (any not listed above default to the same as User-agent: * → allowed)

AI training/crawling bots:
  gptbot                         blocked
  google-extended                blocked
  ccbot                          blocked
  ...

Sitemap: https://example.com/sitemap.xml

What it doesn't do

It can't see your origin server's own robots.txt file directly — only what's actually served over HTTPS to a real request, which is what matters for crawling anyway. If Cloudflare (or another CDN/WAF) sits in front of your site, that's the version to trust regardless of what's in your repo.

License

MIT

Contributors

janibert1

1 commits

janibert1/robots-check

See what your site's robots.txt actually serves -- including whether Cloudflare's AI Crawl Control is silently overriding it.

1

stars

1

commits

JavaScript

primary language

Sep 9, 2026

updated

github.com/janibert1/robots-check#readme
cli
cloudflare
devtools
github-action
nodejs
robots-txt
seo

README

robots-check

See what your site's robots.txt actually serves — including whether Cloudflare's "AI Crawl Control" feature is silently appending its own rules on top of yours.

Why this exists

While deploying a few small sites behind Cloudflare, we noticed our own robots.txt files weren't what was actually being served. Cloudflare injects a managed block (look for # BEGIN Cloudflare Managed content) that adds its own AI-crawler-blocking rules — real search engines are usually left alone, but plenty of site owners have no idea this is happening at all, and it's easy to assume the file in your repo/origin is the one being served when it isn't.

robots-check fetches the live file over HTTPS and tells you, in plain terms:

  • whether it's Cloudflare-managed (and where to check/adjust it)
  • whether real search engines (Googlebot, Bingbot, etc.) are actually allowed to crawl
  • which AI training/crawling bots (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, and others) are blocked vs. allowed
  • whether a sitemap is declared at all

No API keys, no signup, no dependencies — just Node's built-in https module.

Usage

npx github:janibert1/robots-check example.com

Or clone it and run directly:

git clone https://github.com/janibert1/robots-check
cd robots-check
node robots-check.js example.com

Requires Node 18+.

Example output

example.com — 47 non-blank lines

⚠ This robots.txt is (at least partly) Cloudflare-managed.
  Cloudflare's "AI Crawl Control" feature injects its own block into what's
  actually served — your origin server's own robots.txt file may say
  something different from what's shown below. Check your Cloudflare
  dashboard under Bots → AI Crawl Control if this doesn't match your intent.

Real search engines:
  (any not listed above default to the same as User-agent: * → allowed)

AI training/crawling bots:
  gptbot                         blocked
  google-extended                blocked
  ccbot                          blocked
  ...

Sitemap: https://example.com/sitemap.xml

What it doesn't do

It can't see your origin server's own robots.txt file directly — only what's actually served over HTTPS to a real request, which is what matters for crawling anyway. If Cloudflare (or another CDN/WAF) sits in front of your site, that's the version to trust regardless of what's in your repo.

License

MIT

Contributors

janibert1

1 commits

Languages

JavaScript

100.0%