Features / Robots.txt and sitemap audit
Robots.txt and sitemap audit: your rules and files on one screen
SEOSage is a desktop app for Mac and Windows. Its robots.txt and sitemap audit reads your robots.txt and your XML sitemaps and shows the rules, the sitemap files, how many URLs each one lists, and how that list compares with the pages your last crawl found.
Jump to: The screen · What it checks · Examples · How it works · Questions
See the robots.txt and sitemap audit in 20 seconds
See your robots.txt rules and every sitemap file with its URL count on one screen, and how your sitemap compares with the pages your last crawl found.
Read the video transcript
Your rules and your sitemap. See your robots.txt rules, and every sitemap file with its URL count. After a crawl, see which pages are missing from your sitemap, and which sitemap pages were never crawled. Start your fourteen-day free trial at seosage.co.
See your rules and your sitemap files
Open robots.txt in a site's menu, under Insights. SEOSage fetches the files when you open it, and Refresh does it again. The files do not need a crawl, but the coverage part does. A sitemap should list the pages you want search engines to index, and the coverage counts show where it and your crawl disagree.
robots.txt found · HTTP 200 52 URLs in sitemap · 2 sitemap files
What the robots.txt screen shows
- robots.txt status: Found or Not found, with the HTTP status, and a link to view the raw file
- Declared sitemaps: The Sitemap lines in robots.txt. If there are none, SEOSage says so and checks the default /sitemap.xml
- Crawl rules: The Allow and Disallow rules for each user agent. With no rules, it says the whole site is allowed
- Sitemap files: How many URLs the sitemap lists and how many sitemap files there are, with the status and URL count of each file. Index files are marked
- Coverage compared with the last crawl: Three counts: indexable pages in the sitemap, crawled but missing from the sitemap, and in the sitemap but not crawled. Each list can be copied
What it does not do
The screen shows what your files contain. It does not edit them, and it does not test one address against your rules. Coverage needs a crawl, because it compares the sitemap with the pages the crawl found.
The crawl checks for a crawl delay and for blocked CSS or JavaScript in the rules for all bots and for Googlebot. In the Issues list, Page not in sitemap and Sitemap page not crawled each show up to 500 pages, and Sitemap page not crawled is not reported when a crawl was cut short.
Problems a crawl reports for these files
These show in the Issues list after a crawl:
- robots.txt blocks CSS or JavaScript. A warning. Disallow rules hide files Google needs to display your pages.
- robots.txt has no sitemap line. A notice. Add a Sitemap line so search engines find your XML sitemap.
- Sitemap is not accessible. A warning. The sitemap did not answer, or answered with an error.
- Sitemap has a syntax error. An error. The XML cannot be read, so search engines may ignore the whole file.
- Sitemap has too many URLs. An error. A sitemap may list up to 50,000 URLs. Split it into several sitemaps.
- Page not in sitemap. A notice. The crawl found the page, but the sitemap does not list it.
- Noindex page in sitemap. An error. A page with a noindex tag is listed in the XML sitemap.
How it works in three steps
- Open robots.txt.It is in the site menu, under Insights. SEOSage fetches robots.txt and your sitemaps.
- Check the rules and the files.Look at the Disallow rules, the declared sitemaps and the status of each sitemap file.
- Crawl, then check coverage.After a crawl, the screen lists the pages that are missing from your sitemap and the sitemap pages the crawl did not reach.
Related: robots.txt tester · sitemap checker · all sitemap checks.
Questions about this check
Do I need to crawl first?
Not for the files. They are fetched when you open the screen. You need a crawl for Coverage, which compares the sitemap with the pages a crawl found.
What if my robots.txt has no Sitemap line?
The screen says so and checks the default /sitemap.xml. A crawl also reports robots.txt has no sitemap line, as a notice.
Which problems does a crawl report?
Among others: robots.txt blocks CSS or JavaScript, robots.txt sets a crawl delay, a sitemap that is not accessible, in the wrong format or with a syntax error, a sitemap over 50,000 URLs or 50 MB, URLs outside the sitemap's folder, missing or invalid lastmod dates, a page in more than one sitemap, and pages that are in the sitemap but broken, redirected, noindex or non-canonical.
What about a very large sitemap?
If a sitemap is very large, the screen reads the first URLs and tells you how many it read. A crawl flags a sitemap over 50,000 URLs or 50 MB.
How is it different from the robots.txt tester and the sitemap checker on this site?
The robots.txt tester says whether one address is allowed, and the sitemap checker counts the URLs and tests a sample. The app screen shows both files in full and compares the sitemap with your crawl.
More SEOSage features
Every check and tool in the app, each with its own page.
SEO crawler Crawl settings Page audit Change tracking Compare crawls Rank tracking Search Console report Google Analytics report Core Web Vitals Backup and restore Local site audit Custom extraction Internal linking tool Orphan pages Link strength Canonical tag checker Title tag width checker URL structure checker Security headers check Keyword cannibalization Search Console quick wins Competitor analysis Weekly SEO to-do list Fix and re-check White-label reports Ask Claude All features
Try it on your own site
Download SEOSage for Mac or Windows. The 14-day trial needs a card to start and charges nothing for 14 days. Then $19 a month or $149 a year.