To create a robots.txt file, save a small UTF-8 text file named robots.txt at the root of the exact website origin you want to control, then test its rules against real URLs. Use it to guide compliant crawlers away from selected paths—not to secure private content or guarantee that a URL disappears from Google Search.
What robots.txt does—and what it cannot do
Google describes robots.txt as a file that tells search engine crawlers which URLs they can access. Its rules are crawler requests, not access controls: the Internet Engineering Task Force states in RFC 9309, “These rules are not a form of access authorization.” A bot may ignore the instructions, and a URL blocked from crawling may still appear in search results if it is discovered elsewhere.
As an Amazon Associate I earn from qualifying purchases.
That makes robots.txt useful for managing crawler access and requests to URL areas that do not need crawling. It is not a dependable way to remove a page from search or hide information. For search exclusion, leave the page crawlable so the crawler can see a noindex directive. For private content, require authentication.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →| Method | Crawler access | Search visibility | Use it for |
|---|---|---|---|
Disallow in robots.txt |
Asks compliant crawlers not to fetch matching paths | Does not guarantee exclusion; a discovered URL can remain visible | Managing crawl access and requests |
noindex |
The crawler must be able to fetch the page to read the directive | Requests that the page not be included in search results | Pages that should be crawlable but excluded from search |
| Authentication or password protection | Stops public crawlers from retrieving protected content | Keeps the content unavailable to public search crawlers | Private or restricted information |
Where to put robots.txt
Publish the lowercase filename /robots.txt at the top-level path for the relevant service. For example, it belongs at https://www.example.com/robots.txt, not https://www.example.com/blog/robots.txt. RFC 9309 specifies UTF-8 text served as text/plain.
#1 Best Overall
- 𝐍𝐞𝐱𝐭-𝐆𝐞𝐧 𝐖𝐢-𝐅𝐢 𝟕 𝐰𝐢𝐭𝐡 𝟒-𝐒𝐭𝐫𝐞𝐚𝐦 𝐃𝐮𝐚𝐥-𝐁𝐚𝐧𝐝 𝐮𝐩 𝐭𝐨 𝟑.𝟔 𝐆?𝐩𝐬 - Designed with the latest Wi-Fi 7 technology, featuring Multi-Link Operation (MLO), Multi-RUs, and 4K-QAM, The Deco 7 BE23 delivers full speeds of up to 2882 Mbps on the 5GHz band, 688 Mbps on the 2.4GHz band with 4 streams and achieve optimized performance on latest WiFi 7 laptops and devices, like the iPhone 16 Pro, and Samsung Galaxy S24 Ultra.
- 𝐖𝐢𝐝𝐞 𝐂𝐨𝐯𝐞𝐫𝐚𝐠𝐞 𝐰𝐢𝐭𝐡 𝐒𝐭𝐫𝐨𝐧𝐠 𝐂𝐨𝐧𝐧𝐞𝐜𝐭𝐢𝐨𝐧 - Enjoy seamless max Wi-Fi coverage up to 2,500 sq. ft (1-Pack) and 150 devices without compromising performance. 4x high-gain antennas per node and 4x high-power FEMs deliver far-reaching, reliable signals for remote workers, gamers, students, and more.
- 𝐔𝐥𝐭𝐫𝐚-𝐅𝐚𝐬𝐭 𝟐.𝟓 𝐆𝐛𝐩𝐬 𝐖𝐢𝐫𝐞𝐝 𝐏𝐞𝐫𝐟𝐨𝐫𝐦𝐚𝐧𝐜𝐞 - Each Deco 7 BE23 unit is equipped with two 2.5 Gbps WAN/LAN ports, offering warp-speed connectivity for high-performance wired devices. Integrate with a multi-gig modem for gigplus internet.
- 𝐎𝐮𝐫 𝐂𝐲𝐛𝐞𝐫𝐬𝐞𝐜𝐮𝐫𝐢𝐭𝐲 𝐂𝐨𝐦𝐦𝐢𝐭𝐦𝐞𝐧𝐭 - TP-Link is a signatory of the U.S. Cybersecurity and Infrastructure Security Agency’s (CISA) Secure-by-Design pledge. This device is designed, built, and maintained, with advanced security as a core requirement.
- 𝐒𝐭𝐫𝐨𝐧𝐠𝐞𝐫, 𝐌𝐨𝐫𝐞 𝐑𝐞𝐥𝐢𝐚𝐛𝐥𝐞 𝐁𝐚𝐜𝐤𝐡𝐚𝐮𝐥 - The Deco 7 BE23 enhances stability with simultaneous wireless and wired backhaul, leveraging Wi-Fi 7 MLO for stronger, more stable connections.
Rules apply only to the protocol, host, and port where the file is published. That means https://example.com, https://www.example.com, and http://example.com have separate robots.txt scope. Check the exact origin used by the URLs you want to manage.
If you use a CMS or hosted website platform, check its official documentation first: it may generate robots.txt or provide search-visibility settings instead of requiring direct file editing.
How to write the rules
Rules are grouped by crawler. Each group starts with a User-agent line, followed by directives that apply to that crawler. Paths in Allow and Disallow are relative to the URL root. Anything not disallowed is allowed by default.
User-agent: *
Disallow: /private-preview/
Sitemap: https://www.example.com/sitemap.xml
In this example, the rule asks crawlers that follow the wildcard group not to fetch paths beginning with /private-preview/. The Sitemap line advertises a fully qualified sitemap URL. It does not change which paths are allowed or blocked, and it does not guarantee that listed URLs will be indexed.
Rank #2
Use specific paths and check overlaps
Google supports * and $ wildcards in path values. Rules that overlap can produce unexpected results, so compare them against actual URL paths before publishing. Under RFC 9309, the most specific matching Allow or Disallow rule governs; if equivalent rules tie, Allow takes precedence.
Do not copy another site’s rules without checking its URL structure and intended crawler groups. Google, Bing, and other crawlers may not support exactly the same directives.
Use crawler names deliberately
User-agent: * applies to crawlers covered by that group, but Google says it does not cover AdsBot crawlers. If you need a rule for an AdsBot, name the relevant crawler explicitly. Google supports User-agent, Allow, Disallow, and Sitemap; it does not support crawl-delay.
How to create and publish the file
- Set the goal. Decide which URL paths should not be fetched by compliant crawlers. Do not use robots.txt to protect secrets or guarantee deindexing.
- Check the existing setup. Open the exact origin’s
/robots.txtand confirm whether your CMS or hosting platform already manages it. - Inventory affected URLs. Review real pages and resources that match the paths you plan to block. Keep pages and assets accessible when crawlers need them to understand or render a page.
- Draft the smallest workable file. Put each directive on its own line, group it under the intended
User-agent, and use root-relative paths. Add a sitemap line only if you have a valid sitemap URL, including its protocol and host. - Save as plain UTF-8 text. Use the exact lowercase filename
robots.txt. - Publish at the origin root. The file should load at the relevant host’s
/robots.txtaddress. - Verify and monitor. Confirm the public file loads, test important allowed and blocked URLs, and review crawler or indexing reports after the change.
How to test robots.txt
First, visit the public /robots.txt URL for the origin you intend to control. Confirm that it returns the file you edited rather than an old, missing, or platform-generated version. Then test representative URLs—especially pages that should remain crawlable and paths that should be blocked—using Search Console or a compatible local parser.
Rank #3
Google recommends keeping CSS and JavaScript resources accessible when their absence would impair Google’s understanding of a page. Include those resources in your checks if your rules affect their directories.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common mistakes and how to avoid them
- Expecting Disallow to remove a page from Google. A blocked URL can remain in results if Google discovers it elsewhere. Use a crawlable
noindexdirective for search exclusion. - Putting the file in a subdirectory. Place it at the top-level path of the relevant origin, not under a site section.
- Assuming one file covers every hostname or protocol. Check each origin whose URLs need different crawler instructions.
- Blocking resources needed to render a page. Review CSS, JavaScript, and other required assets before blocking their paths.
- Treating robots.txt as a secret or security measure. The file is public, and its paths can reveal URL patterns. Use authentication for private material.
- Adding unsupported directives. Google does not support
crawl-delay; check each crawler’s current documentation before relying on nonstandard records. - Assuming a change takes effect immediately. Crawlers may cache robots.txt. RFC 9309 says crawlers should not use a cached file for more than 24 hours unless it is unreachable, but that is a standard recommendation, not a guarantee of when a specific crawler will refresh.
Why a robots.txt change may not appear to work
If the file does not load, check the exact protocol, host, and port, then verify that the path is the origin root and that the server is returning the current text file. If a URL’s behavior differs from what you expect, compare its full path with the relevant user-agent group and look for a more specific overlapping rule.
Also allow for crawler caching: RFC 9309 distinguishes an unavailable robots.txt response from a network or server failure, and its prescribed handling differs. Google documents its own behavior for robots.txt errors and retrieval in its robots.txt specification. Check that guidance when troubleshooting Googlebot rather than assuming every crawler handles failures the same way.
Free tools Windows power users keep installed
One-click scans. No signup required.
What to put in robots.txt for SEO
There is no universal set of paths that every site should block. Start with a specific crawl-management need, inspect the site’s actual URL patterns, and add only rules whose effects you can verify. A sitemap reference can help crawlers discover its location, but it does not override access rules or promise indexing. Robots.txt is a crawl-control tool, not an SEO ranking shortcut.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




