Know exactly what search engines and AI crawlers can reach.
Googlebot, GPTBot, ClaudeBot and PerplexityBot don't necessarily see the same site. A robots.txt rule written for one purpose can block a crawler nobody meant to block. This checks access per bot, keeps the sitemap in sync with what's actually live, and flags the gap.
You are here · the workflow
The technical surface · this page covers
Part of Technical Suite's six-stage loop. See the full system →
One rule blocks a bot nobody meant to block.
- A robots.txt rule written to block a staging subdomain ends up matching a path pattern on the live site too, and nobody notices for months.
- Googlebot is allowed everywhere, but GPTBot or ClaudeBot is quietly blocked on an entire section — docs, blog, product pages — because nobody checked access bot by bot.
- The sitemap still lists pages that were deleted six months ago, and is missing pages that shipped last week.
- Nobody finds out until a section stops appearing in AI answers, and by then it's unclear how long it's been invisible.
What needs to happen: access needs checking per bot, on every crawl, and the sitemap needs to stay in sync with the site rather than being generated once and forgotten.
How SEORCE does it: every crawl checks robots.txt matching and actual access separately for each major crawler, compares the sitemap against what's actually live, and surfaces exactly which rule is causing an unintended block. Whether to block a bot on purpose stays a decision for your team — SEORCE shows what's affected, not a change made on its own.
Search engines and AI agents don't necessarily see your site the same way.
GPTBot and OAI-SearchBot are checked separately — GPTBot relates to OpenAI's training-related crawling, OAI-SearchBot is the crawler behind ChatGPT Search inclusion. A block on one doesn't mean a block on the other.
What the sitemap says. What the site actually has.
Inside crawlability.
Crawlability is the gate everything else depends on.
What does Crawlability check?
Crawlability checks sitemap accuracy and per-bot access during the continuous crawl — whether Googlebot, Bingbot, GPTBot, ClaudeBot, PerplexityBot and other crawlers can each reach a page, checked separately rather than assumed to be the same for every bot. It compares the sitemap against what's actually live, flagging pages missing from it and stale entries pointing at pages that no longer exist. Whether to intentionally block a bot stays a decision for your team; SEORCE surfaces exactly which robots.txt rule is causing an unintended block and what it affects.
Key facts
- Checks access separately per bot, not just whether a page is crawlable in general
- Covers Googlebot, Bingbot, and AI crawlers including GPTBot, ClaudeBot and PerplexityBot
- Compares the sitemap against the live crawl for coverage and staleness
- Surfaces the exact robots.txt rule behind an unintended block
- Doesn't change robots.txt automatically — blocking a bot on purpose stays a human decision
Frequently asked about crawlability.
Yes. Access is checked per bot — Googlebot, GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and others — because a robots.txt rule can allow one and block another without anyone intending that.
They are OpenAI's two separate crawlers. GPTBot relates to OpenAI's training-related crawling; OAI-SearchBot is the crawler used for inclusion in ChatGPT Search. A robots.txt rule can allow one and block the other, so they're checked independently.
Yes, and it's one of the more common findings. A rule written for one purpose can end up matching a path pattern that also covers a section of the live site, and nobody notices unless something checks per-bot access directly.
Yes. The sitemap is compared against what the crawl actually finds, flagging pages missing from the sitemap and sitemap entries that no longer resolve to a live page.
Whether to block a bot on purpose is always a decision for a person — SEORCE surfaces exactly which rule is causing an unintended block and what it affects, rather than changing robots.txt on its own.
See who can actually reach your site.
Googlebot, GPTBot, ClaudeBot and more, checked separately, alongside a full sitemap comparison.