VellumUp
PricingBlogIntegrationsComponents
Sign inRegister
Back to tools

Sitemap vs. Robots.txt Conflict Checker

Find URLs listed in your sitemap that your own robots.txt then blocks. This silent contradiction wastes crawl submissions and confuses search engines.

Enter a domain to check its sitemap.xml against its robots.txt, or a specific sitemap address.

About this tool

This tool fetches a site's sitemap and its robots.txt at the same time, then checks every URL the sitemap lists against the rules robots.txt sets for general crawlers. Listing a URL in a sitemap is an active submission for crawling and indexing, while a matching Disallow rule in robots.txt tells crawlers not to request that same URL at all. Both files usually update independently, on their own schedules or through different tools, so a page removed or restricted after the fact can end up disallowed while still sitting in the sitemap, submitted for a crawl that will never happen. A missing robots.txt is treated as no restrictions at all, matching the standard, documented behavior crawlers themselves follow.

If you have already seen this exact conflict without running this tool, it is likely because Google Search Console flagged it for you directly: this specific situation shows up in the Pages report under the label "Submitted URL blocked by robots.txt," one of the more common warnings site owners run into there. This tool is a way to find and understand that same conflict proactively, on your own schedule, rather than waiting for Search Console to surface it after Google has already tried and failed to crawl the page.

Explore more free tools

Content Freshness Gap Finder

Fetch a sitemap and see which pages are fresh, aging, or stale, oldest first.

Robots.txt Tester

Fetch and read any site's robots.txt, and check whether a specific URL is blocked for a given crawler.

XML Sitemap Validator

Fetch and validate any XML sitemap: malformed XML, invalid or duplicate URLs, and bad field values.

Frequently asked questions

The two files are usually maintained separately, sometimes by different tools or different people, and often on different schedules. A common cause is restricting a section of a site in robots.txt after its pages were already added to the sitemap, or an auto-generated sitemap that lists every page regardless of what robots.txt currently disallows.
Search engines generally will not request the page's content at all because robots.txt tells them not to, so the sitemap submission for that URL is effectively wasted. In some cases the URL can still appear in search results with no real description, since the page was never allowed to be read, which is a confusing and usually unwanted outcome.
The rules for the general User-agent: * group, which applies to crawlers not specifically named elsewhere in the file. A sitemap is a general submission meant for any search engine that reads it, so the general-purpose rule set is the relevant one to compare it against, not any single named crawler's own override.
There are two valid fixes, and which one is correct depends on intent: if the page should not be crawled, remove its URL from the sitemap so it stops being actively submitted; if the page should be crawled, update the robots.txt rule blocking it so it is no longer disallowed. Leaving both as they are keeps the contradiction in place.
No. When a site has no robots.txt file at all, the standard, documented behavior is that nothing is disallowed, so every sitemap URL is checked against an effectively empty rule set and no conflicts are found on that basis.
No. The address is used only to fetch that site's sitemap and robots.txt for that one check, and nothing is saved. Only public web addresses can be checked: requests to private or internal network addresses are refused.
Yes. That is Google Search Console's own name for precisely this conflict, a URL your sitemap submitted that your robots.txt then disallows. After fixing it here, whichever way you choose, use the URL Inspection tool in Search Console on the affected address and click Validate Fix to prompt Google to recrawl and re-evaluate it, rather than waiting for the next scheduled crawl.

Want SEO and GEO content
in your brand voice, automatically?

Try it risk-free — a content plan, keyword research, and up to 5 articles on us.

Say hello

support@vellumup.com

Follow us

Product

PricingFree ToolsShopify App StoreWix App MarketWordPress Pluginnpx vellumup-init

Company

BlogIntegrationsComponentsAffiliates

Integrations

DocumentationWordPressShopifyWebflowWixWebhooksNext.js

Legal

Terms of ServicePrivacy PolicyAccessibility
VellumUp

© 2026 VellumUp. All rights reserved.