October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

Create a robots.txt File Safely: Rules, Testing, and SEO Limits

Create a minimal UTF-8 robots.txt file at the correct site root, test its rules against real URLs, and use noindex or authentication for goals robots.txt cannot achieve.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To create a robots.txt file, save a small UTF-8 text file named robots.txt at the root of the exact website origin you want to control, then test its rules against real URLs. Use it to guide compliant crawlers away from selected paths—not to secure private content or guarantee that a URL disappears from Google Search.

What robots.txt does—and what it cannot do

Google describes robots.txt as a file that tells search engine crawlers which URLs they can access. Its rules are crawler requests, not access controls: the Internet Engineering Task Force states in RFC 9309, “These rules are not a form of access authorization.” A bot may ignore the instructions, and a URL blocked from crawling may still appear in search results if it is discovered elsewhere.

As an Amazon Associate I earn from qualifying purchases.

That makes robots.txt useful for managing crawler access and requests to URL areas that do not need crawling. It is not a dependable way to remove a page from search or hide information. For search exclusion, leave the page crawlable so the crawler can see a noindex directive. For private content, require authentication.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Method Crawler access Search visibility Use it for
Disallow in robots.txt Asks compliant crawlers not to fetch matching paths Does not guarantee exclusion; a discovered URL can remain visible Managing crawl access and requests
noindex The crawler must be able to fetch the page to read the directive Requests that the page not be included in search results Pages that should be crawlable but excluded from search
Authentication or password protection Stops public crawlers from retrieving protected content Keeps the content unavailable to public search crawlers Private or restricted information

Where to put robots.txt

Publish the lowercase filename /robots.txt at the top-level path for the relevant service. For example, it belongs at https://www.example.com/robots.txt, not https://www.example.com/blog/robots.txt. RFC 9309 specifies UTF-8 text served as text/plain.

#1 Best Overall
TP-Link Deco 7 BE23 Dual-Band BE3600 WiFi 7 Mesh Wi-Fi Router
  • 𝐍𝐞𝐱𝐭-𝐆𝐞𝐧 𝐖𝐢-𝐅𝐢 𝟕 𝐰𝐢𝐭𝐡 𝟒-𝐒𝐭𝐫𝐞𝐚𝐦 𝐃𝐮𝐚𝐥-𝐁𝐚𝐧𝐝 𝐮𝐩 𝐭𝐨 𝟑.𝟔 𝐆?𝐩𝐬 - Designed with the latest Wi-Fi 7 technology, featuring Multi-Link Operation (MLO), Multi-RUs, and 4K-QAM, The Deco 7 BE23 delivers full speeds of up to 2882 Mbps on the 5GHz band, 688 Mbps on the 2.4GHz band with 4 streams and achieve optimized performance on latest WiFi 7 laptops and devices, like the iPhone 16 Pro, and Samsung Galaxy S24 Ultra.
  • 𝐖𝐢𝐝𝐞 𝐂𝐨𝐯𝐞𝐫𝐚𝐠𝐞 𝐰𝐢𝐭𝐡 𝐒𝐭𝐫𝐨𝐧𝐠 𝐂𝐨𝐧𝐧𝐞𝐜𝐭𝐢𝐨𝐧 - Enjoy seamless max Wi-Fi coverage up to 2,500 sq. ft (1-Pack) and 150 devices without compromising performance. 4x high-gain antennas per node and 4x high-power FEMs deliver far-reaching, reliable signals for remote workers, gamers, students, and more.
  • 𝐔𝐥𝐭𝐫𝐚-𝐅𝐚𝐬𝐭 𝟐.𝟓 𝐆𝐛𝐩𝐬 𝐖𝐢𝐫𝐞𝐝 𝐏𝐞𝐫𝐟𝐨𝐫𝐦𝐚𝐧𝐜𝐞 - Each Deco 7 BE23 unit is equipped with two 2.5 Gbps WAN/LAN ports, offering warp-speed connectivity for high-performance wired devices. Integrate with a multi-gig modem for gigplus internet.
  • 𝐎𝐮𝐫 𝐂𝐲𝐛𝐞𝐫𝐬𝐞𝐜𝐮𝐫𝐢𝐭𝐲 𝐂𝐨𝐦𝐦𝐢𝐭𝐦𝐞𝐧𝐭 - TP-Link is a signatory of the U.S. Cybersecurity and Infrastructure Security Agency’s (CISA) Secure-by-Design pledge. This device is designed, built, and maintained, with advanced security as a core requirement.
  • 𝐒𝐭𝐫𝐨𝐧𝐠𝐞𝐫, 𝐌𝐨𝐫𝐞 𝐑𝐞𝐥𝐢𝐚𝐛𝐥𝐞 𝐁𝐚𝐜𝐤𝐡𝐚𝐮𝐥 - The Deco 7 BE23 enhances stability with simultaneous wireless and wired backhaul, leveraging Wi-Fi 7 MLO for stronger, more stable connections.

Rules apply only to the protocol, host, and port where the file is published. That means https://example.com, https://www.example.com, and http://example.com have separate robots.txt scope. Check the exact origin used by the URLs you want to manage.

If you use a CMS or hosted website platform, check its official documentation first: it may generate robots.txt or provide search-visibility settings instead of requiring direct file editing.

How to write the rules

Rules are grouped by crawler. Each group starts with a User-agent line, followed by directives that apply to that crawler. Paths in Allow and Disallow are relative to the URL root. Anything not disallowed is allowed by default.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
User-agent: *
Disallow: /private-preview/

Sitemap: https://www.example.com/sitemap.xml

In this example, the rule asks crawlers that follow the wildcard group not to fetch paths beginning with /private-preview/. The Sitemap line advertises a fully qualified sitemap URL. It does not change which paths are allowed or blocked, and it does not guarantee that listed URLs will be indexed.

Use specific paths and check overlaps

Google supports * and $ wildcards in path values. Rules that overlap can produce unexpected results, so compare them against actual URL paths before publishing. Under RFC 9309, the most specific matching Allow or Disallow rule governs; if equivalent rules tie, Allow takes precedence.

Do not copy another site’s rules without checking its URL structure and intended crawler groups. Google, Bing, and other crawlers may not support exactly the same directives.

Use crawler names deliberately

User-agent: * applies to crawlers covered by that group, but Google says it does not cover AdsBot crawlers. If you need a rule for an AdsBot, name the relevant crawler explicitly. Google supports User-agent, Allow, Disallow, and Sitemap; it does not support crawl-delay.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to create and publish the file

  1. Set the goal. Decide which URL paths should not be fetched by compliant crawlers. Do not use robots.txt to protect secrets or guarantee deindexing.
  2. Check the existing setup. Open the exact origin’s /robots.txt and confirm whether your CMS or hosting platform already manages it.
  3. Inventory affected URLs. Review real pages and resources that match the paths you plan to block. Keep pages and assets accessible when crawlers need them to understand or render a page.
  4. Draft the smallest workable file. Put each directive on its own line, group it under the intended User-agent, and use root-relative paths. Add a sitemap line only if you have a valid sitemap URL, including its protocol and host.
  5. Save as plain UTF-8 text. Use the exact lowercase filename robots.txt.
  6. Publish at the origin root. The file should load at the relevant host’s /robots.txt address.
  7. Verify and monitor. Confirm the public file loads, test important allowed and blocked URLs, and review crawler or indexing reports after the change.

How to test robots.txt

First, visit the public /robots.txt URL for the origin you intend to control. Confirm that it returns the file you edited rather than an old, missing, or platform-generated version. Then test representative URLs—especially pages that should remain crawlable and paths that should be blocked—using Search Console or a compatible local parser.

Google recommends keeping CSS and JavaScript resources accessible when their absence would impair Google’s understanding of a page. Include those resources in your checks if your rules affect their directories.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common mistakes and how to avoid them

  • Expecting Disallow to remove a page from Google. A blocked URL can remain in results if Google discovers it elsewhere. Use a crawlable noindex directive for search exclusion.
  • Putting the file in a subdirectory. Place it at the top-level path of the relevant origin, not under a site section.
  • Assuming one file covers every hostname or protocol. Check each origin whose URLs need different crawler instructions.
  • Blocking resources needed to render a page. Review CSS, JavaScript, and other required assets before blocking their paths.
  • Treating robots.txt as a secret or security measure. The file is public, and its paths can reveal URL patterns. Use authentication for private material.
  • Adding unsupported directives. Google does not support crawl-delay; check each crawler’s current documentation before relying on nonstandard records.
  • Assuming a change takes effect immediately. Crawlers may cache robots.txt. RFC 9309 says crawlers should not use a cached file for more than 24 hours unless it is unreachable, but that is a standard recommendation, not a guarantee of when a specific crawler will refresh.

Why a robots.txt change may not appear to work

If the file does not load, check the exact protocol, host, and port, then verify that the path is the origin root and that the server is returning the current text file. If a URL’s behavior differs from what you expect, compare its full path with the relevant user-agent group and look for a more specific overlapping rule.

Also allow for crawler caching: RFC 9309 distinguishes an unavailable robots.txt response from a network or server failure, and its prescribed handling differs. Google documents its own behavior for robots.txt errors and retrieval in its robots.txt specification. Check that guidance when troubleshooting Googlebot rather than assuming every crawler handles failures the same way.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to put in robots.txt for SEO

There is no universal set of paths that every site should block. Start with a specific crawl-management need, inspect the site’s actual URL patterns, and add only rules whose effects you can verify. A sitemap reference can help crawlers discover its location, but it does not override access rules or promise indexing. Robots.txt is a crawl-control tool, not an SEO ranking shortcut.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.