VellumUp
PricingBlogIntegrationsComponents
Sign inRegister
Back to tools

AI Crawler Access Checker

Fetch any site's robots.txt and check whether GPTBot, ClaudeBot, PerplexityBot, and other named AI crawlers are allowed in, split by what each one is actually used for.

About this tool

This tool fetches the robots.txt file at the root of the address you enter, and checks it against a curated list of publicly documented AI crawler user-agents from OpenAI, Anthropic, Perplexity, Google, Common Crawl, Meta, ByteDance, Apple, Amazon, and Cohere. Each crawler is checked the same way a real crawler resolves the file: a group naming that exact user-agent applies if one exists, otherwise the general User-agent: * group applies, otherwise everything is allowed by default.

The crawlers are split into four groups because "AI crawler" is not one thing to allow or block together. Training crawlers, like GPTBot or ClaudeBot, scrape pages to train a foundation model, blocking them says nothing about whether a site can still be found through that company's chat product, since training and live retrieval are separate crawlers for every vendor listed here. Live-answer crawlers, like ChatGPT-User or Perplexity-User, fetch a page in real time when a person actually asks the assistant to open a link, blocking one of these can make a specific page invisible to that exact request. AI-search crawlers, like PerplexityBot or OAI-SearchBot, index pages ahead of time for that company's own AI-powered search feature, functioning much closer to a traditional search engine crawler. Ad-check crawlers, like OpenAI's OAI-AdsBot, fetch a landing page to verify it meets ad quality and policy requirements before an ad can run, a separate concern from training or answering.

As with any robots.txt rule, this is a voluntary convention: it tells a well-behaved crawler what it should not request, it does not technically prevent a page from being fetched by a crawler that chooses to ignore the file.

Explore more free tools

Robots.txt Tester

Fetch and read any site's robots.txt, and check whether a specific URL is blocked for a given crawler.

XML Sitemap Validator

Fetch and validate any XML sitemap: malformed XML, invalid or duplicate URLs, and bad field values.

Sitemap vs. Robots.txt Conflict Checker

Find URLs listed in your sitemap that your own robots.txt then blocks.

Frequently asked questions

It depends on which ones and why. Blocking a training crawler like GPTBot keeps that specific company from using your content to train future models, but does nothing to keep your site out of that company's live chat answers, since a separate crawler handles that. Blocking a live-answer or AI-search crawler instead can make your pages invisible to that product's actual users. There is no single right answer for every site, it is a real tradeoff between training-data control and AI-driven visibility.
No. GPTBot is OpenAI's training crawler. ChatGPT-User and OAI-SearchBot are the separate crawlers involved when ChatGPT actually browses a link or searches the web on a user's behalf. Blocking only GPTBot leaves both of those untouched.
Google-Extended is not a separate crawler with its own IP ranges, it is a control token that, when disallowed, tells Google not to use a page's content to train Gemini and Vertex AI models. It has no effect on regular Google Search crawling or ranking, that is still controlled entirely by the standard Googlebot rules.
Access for a specific AI crawler is evaluated the same way as any other crawler: by the rules in its matching user-agent group, or the general group if it has none. Checking the root path (/) shows whether that crawler has any access to the site at all under the current rules. A site can still carve out per-section exceptions, for that level of detail the general Robots.txt Tester tool on this site lets you check any specific path and any specific user-agent.
It means the robots.txt file has no group written specifically for that crawler's name, so it falls back to whatever the general User-agent: * group says. If that general group disallows the path being checked, every unnamed crawler, AI or otherwise, is blocked the same way, not just the one shown.
It covers the crawlers with clear, current public documentation from their own companies as of this tool's last update. New AI crawlers are introduced periodically, and a crawler with no confirmed, documented purpose is deliberately left out rather than guessed at, since a wrong classification here would be worse than a missing one.
No. The address is used only to fetch robots.txt for that one check, and nothing is saved. Only public web addresses can be checked: requests to private or internal network addresses are refused.

Want SEO and GEO content
in your brand voice, automatically?

Try it risk-free — a content plan, keyword research, and up to 5 articles on us.

Say hello

support@vellumup.com

Follow us

Product

PricingFree ToolsShopify App StoreWix App MarketWordPress Pluginnpx vellumup-init

Company

BlogIntegrationsComponentsAffiliates

Integrations

DocumentationWordPressShopifyWebflowWixWebhooksNext.js

Legal

Terms of ServicePrivacy PolicyAccessibility
VellumUp

© 2026 VellumUp. All rights reserved.