Direct answer: define the URLs and host you intend to audit, choose a link-discovery (Spider) or supplied-URL (List) crawl, configure limits and exclusions, inspect technical findings, compare the crawl with your XML sitemap, and validate high-impact conclusions in Google Search Console. A crawler reports what its own requests found; it cannot prove Google crawled or indexed every page.
What a crawler audit can—and cannot—tell you
A crawler fetches pages in a controlled way, follows links or processes a URL list, and records responses, directives, canonicals, internal links and other extracted data. That makes it useful for finding patterns across templates and sections.
It does not reproduce Google’s complete crawl history or index decisions. Keep two evidence columns in your audit: crawler-observed (for example, a 404 returned to your crawler) and Google-observed (for example, a URL Inspection result). This prevents a local crawl result from being presented as proof of search visibility.
1. Define scope before starting
Choose the property
- Write down the canonical host, such as
https://example.com, and decide whether subdomains, staging hosts or alternate protocols are included. - List the sections that matter to the business and the templates that generate them.
- Collect important URL sources: navigation, internal links, XML sitemaps, analytics exports and known landing pages.
Choose Spider or List mode
In Screaming Frog SEO Spider, a normal Spider crawl starts from a homepage and discovers URLs through same-subdomain HTML hyperlinks. List mode accepts a pasted or uploaded set of known URLs. Use Spider mode to evaluate discoverability and internal architecture; use List mode when the question is “what happens to these exact URLs?” or when crawling a private inventory.
Recommended Free Tools
Prevent URL explosions
Before pressing Start, exclude patterns that add no audit value: faceted filters, calendar parameters, session IDs, infinite search results and tracking parameters. Set a sensible crawl depth, URL limit and request rate for the site’s capacity. There is no universal safe limit; the correct values depend on the site, server and audit purpose.
2. Configure and run the crawl
- Open the crawler and select the target mode.
- Enter the homepage for Spider mode, or paste/upload the complete URL set for List mode.
- Configure robots.txt handling, user agent, JavaScript rendering, authentication (if required), URL exclusions and crawl limits.
- Start the crawl and watch progress for stalled requests, repeated parameter patterns, server errors and unexpected hosts.
- Export the full URL list and issue reports when the crawl finishes; save the configuration so a later comparison uses the same settings.
Screaming Frog’s documentation describes real-time progress, review of directives and canonicals, and technical SEO auditing. Treat every flagged item as an investigation lead: inspect representative URLs and determine whether the pattern is intentional.
3. Read the core technical signals
Status codes and transport
Group URLs by 2xx, 3xx, 4xx and 5xx responses. A single broken URL is different from a template that returns errors across thousands of pages. Check redirect chains, loops, HTTP-to-HTTPS behavior and whether redirects end at the intended canonical URL.
Indexability directives
Record meta robots and X-Robots-Tag values, canonical targets, response content type and whether the page is accessible to your crawler. A canonical is a hint, not a guarantee that Google will select that URL.
Internal links and orphan candidates
Review pages with no internal links, shallowly linked important pages and links that point to redirects or errors. An “orphan” candidate is a page present in another source (such as a sitemap or analytics export) but absent from the crawl’s link graph; verify it before labeling it orphaned.
Rank #2
JavaScript-dependent content
Compare an HTML-only crawl with a rendered crawl when navigation, links or main content are injected by JavaScript. Note which configuration produced each finding. A missing link in an HTML-only crawl may be a rendering limitation rather than a production defect.
Robots.txt is not an index-exclusion control
Google defines robots.txt as instructions about which URLs crawlers may access, primarily for managing crawl traffic or avoiding unimportant and similar URLs. Google also warns that robots.txt is not a dependable way to keep a URL out of Search: a blocked URL can still be indexed if other pages link to it.
If the goal is exclusion from search results, use a noindex directive (while allowing Google to fetch the page) or password protection. A robots-blocked URL may prevent your crawler from seeing the page’s meta tags and content, so document that evidence gap rather than assuming the page is noindex.
4. Compare the crawl with your XML sitemap
Export sitemap URLs and compare them with crawled URLs and your list of important pages. Investigate both directions:
- In sitemap, not crawled: the URL may be blocked, disconnected from internal links, redirected, broken, or outside the crawler’s configured scope.
- Important and crawled, not in sitemap: decide whether it should be declared in the sitemap, especially if it is a canonical, indexable landing page.
- In sitemap but non-indexable: remove it or correct the directive, redirect or canonical depending on the intended outcome.
- Duplicate or parameter variants: establish which URL is canonical and prevent unnecessary variants from being submitted.
Screaming Frog documents XML sitemap analysis for missing, non-indexable and orphan-page investigations. Google says a sitemap is an important way to tell Google about URLs, but inclusion in a sitemap does not guarantee immediate crawling or indexing.
Rank #3
5. Validate findings in Google tools
Search Console Crawl Stats
Use the Crawl Stats report for Googlebot request history, response groups and host status. It answers a Google-specific question that your desktop crawler cannot: how Googlebot has been interacting with the property.
URL Inspection
Inspect representative URLs from every important issue pattern. Check indexed status, canonical selection, crawl information and the rendered result where available. Include both a failing example and a control URL that behaves as intended.
Recrawl requests
A request to validate a fix or ask Google to recrawl is only a request. Google states that it does not guarantee immediate crawling or inclusion in results. Keep the issue open until subsequent evidence shows the desired state.
6. Prioritize and report issues
Use one record per finding with these fields:
| Field | What to record |
|---|---|
| Sample URL | A reproducible affected page and a control URL |
| Pattern | Template, directory, status, directive or link behavior |
| Scope | Estimated affected URL set, with the method used to count it |
| Evidence | Crawler settings, response data, screenshots or Search Console result |
| Recommendation | The specific code, configuration or content change |
| Owner and validation | Responsible team, test condition and follow-up tool |
Prioritize sitewide template and directive errors before isolated low-impact defects, but confirm scope before assigning business impact. A crawler warning is not automatically a ranking problem.
Screenshot evidence without manual browser work
When an audit finding depends on what a visitor sees—such as a consent banner obscuring content, a redirect landing page or a broken responsive layout—capture representative URLs as evidence. ScreenshotNeo is a website screenshot API and MCP server; it removes cookie banners, newsletter popups and chat widgets before capture, so visual checks are easier to compare.
Or skip the browser setup
Use one request after you have an API key. See the full parameter reference in the ScreenshotNeo documentation.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallcurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are not billed, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Performance, reliability and cost controls
- Run large audits during a maintenance window agreed with the site owner, and throttle requests to avoid unnecessary load.
- Use exclusions and a URL limit to prevent parameter traps from consuming time and bandwidth.
- Save crawl exports and settings; repeat with identical configuration before comparing changes.
- Separate transient network failures from repeatable HTTP errors by retrying a small sample.
- For authenticated or blocked areas, document credentials, headers and the exact access boundary; do not infer public indexability from a private crawl.
- Capture timestamps, user agent and rendering mode so another analyst can reproduce the result.
Troubleshooting common crawl problems
The crawl stops or returns many timeouts
Reduce concurrency and request rate, check server logs and test whether a firewall or bot-management system is challenging the crawler. Retry a small URL sample before restarting the full audit.
Important pages are missing
Check that the pages are linked, the correct subdomain and protocol were entered, exclusions are not too broad, and JavaScript rendering is enabled when links are client-rendered. Compare against the sitemap and a known URL list.
The crawler cannot see a blocked page
Review robots.txt and access controls. A robots rule may be preventing inspection; it does not establish that Google has excluded the URL. Use Search Console and URL Inspection for Google-specific evidence.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Everything appears indexable, but Google does not show the page
Inspect the URL in Search Console, confirm the selected canonical and review coverage or indexing reasons. A crawler’s successful fetch only proves that its own request succeeded.
Sitemap and crawl counts disagree
Normalize protocol, host, trailing slash and parameter handling, then classify redirects, blocked URLs, duplicates and non-indexable pages. Do not treat raw count differences as errors until those categories are explained.
Best Value
- Features Over 160 Latin Songs
- Arranged for C Instruments
- Standard Notation
- 48 Pages
A repeatable audit checklist
- Write the host, sections, URL sources and audit question.
- Select Spider or List mode and record the configuration.
- Set exclusions, rendering, rate and limits before crawling.
- Export URLs, status codes, directives, canonicals and internal-link data.
- Group findings by template and verify representative examples.
- Compare crawled URLs with the XML sitemap and important-page inventory.
- Validate high-impact patterns in Crawl Stats and URL Inspection.
- Assign an owner, fix, evidence standard and follow-up date.
- Repeat with the saved configuration and compare only like-for-like crawls.
Frequently Asked Questions
Should every URL found in a crawl be added to the XML sitemap?
No. Include the URLs you want search engines to discover and consider, then investigate why other discovered URLs are duplicates, redirects, blocked, non-indexable or otherwise unsuitable.
Can a crawler prove that Google indexed a page?
No. Confirm Google’s state with Search Console URL Inspection and related reports; the crawler can only report what its own configured requests observed.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsWhen is List mode better than Spider mode?
Use List mode when you already have a defined URL inventory and need to test that set without relying on internal link discovery.
The Bottom Line
A reliable website audit combines a scoped, reproducible crawl with sitemap comparison and Google-specific validation. Use the crawler to find patterns; use Search Console to establish what Google actually saw and indexed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

