Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a stable, minimal User-Agent that truthfully identifies your crawler, set it explicitly in your HTTP client, check robots.txt before requesting pages, and provide an operator contact when appropriate. Do not impersonate Chrome or Firefox to evade access controls: a changed User-Agent cannot fix excessive request rates, missing authentication, JavaScript-only delivery, or a site policy that disallows automation.

What a User-Agent is in a scraping request

A User-Agent is an HTTP request header sent by the client program that initiates a request. RFC 9110 says a user agent should send a User-Agent field on each request unless it has been specifically configured not to. Servers use the value to identify the originating software and may use it when tailoring a response.

For a crawler, the header is an identification label, not a password and not a permission slip. A useful value tells an operator what your program is and how to contact its owner:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
User-Agent: catalog-crawler/1.0 (+https://example.com/crawler-info)

The product token is catalog-crawler; 1.0 is its version; the parenthesized URL points to information about the crawler. Keep that value consistent across requests so a site operator can recognize your traffic.

Choose a truthful, minimal User-Agent

Identify your software, not a browser you are not running

RFC 9110 defines product identifiers with optional versions and advises senders to limit identifiers to information necessary to identify the product. A scraper written with Python Requests should not copy a current Chrome string and claim to be Chrome. Using another implementation’s product tokens to declare compatibility defeats the purpose of identification and can create a misleading fingerprint.

Use a name you control, such as catalog-crawler/1.0, price-monitor/2.3, or research-bot/1.0. Update the version when the crawler changes in a way that matters to operators, but do not rotate the value on every request.

Add contact information when a robot could create unwanted traffic

RFC 9110 says a robotic user agent should send a valid From header so the responsible operator can be contacted if the robot sends excessive, unwanted, or invalid requests. The contact can be an email address or another maintained channel:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
User-Agent: catalog-crawler/1.0 (+https://example.com/crawler-info)
From: [email protected]

Only publish an address that is monitored. A dead mailbox is not useful contactability.

Avoid unnecessary device and platform detail

Long values that enumerate operating systems, devices, extensions, or implementation details add latency and can increase fingerprinting risk. They rarely help a data collector. Include only the product name, a meaningful version, and an information or contact URL.

Check robots.txt before crawling

Robots exclusion rules are published instructions for crawlers. RFC 9309 describes matching a crawler product token in the User-Agent header to the applicable User-agent group in robots.txt. Use the same product token in both places.

  1. Request https://target.example/robots.txt before beginning a crawl.
  2. Find the group whose User-agent token matches your product identifier. If there is no matching group, use the wildcard group.
  3. Apply that group’s Allow and Disallow rules to every URL you plan to fetch.
  4. Follow any published crawl-delay guidance. Treat it as an operational limit, not a suggestion to ignore.
  5. Review the site’s terms, authentication requirements, copyright restrictions, and applicable law as well.

A matching token does not grant permission to access private or authenticated content. It only determines which published crawler rules apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Set a User-Agent in Python Requests

Requests accepts custom headers through the headers dictionary. Header values must be strings or byte strings. This complete example identifies the crawler, supplies a contact, applies a timeout, and raises an exception for an unsuccessful HTTP response:

import requests

url = "https://example.org/data"
headers = {
    "User-Agent": "catalog-crawler/1.0 (+https://example.com/crawler-info)",
    "From": "[email protected]",
}

response = requests.get(url, headers=headers, timeout=20)
response.raise_for_status()
print(response.text)

Keeping the headers in one dictionary makes it less likely that one code path silently sends a different identity. A timeout prevents a stalled origin from holding a worker indefinitely. raise_for_status() ensures that an error page is not accidentally processed as data.

Python urllib

Python’s urllib adds a default User-Agent automatically when you do not provide one. Attach your own value to a Request when you need an explicit identity:

from urllib.request import Request, urlopen

request = Request(
    "https://example.org/data",
    headers={
        "User-Agent": "catalog-crawler/1.0 (+https://example.com/crawler-info)",
        "From": "[email protected]",
    },
)

with urlopen(request, timeout=20) as response:
    body = response.read()
print(body)

Set the header with cURL and Node.js

cURL

Use -H for each header. --fail makes HTTP errors fail the command instead of treating an error document as a successful download, while --max-time limits how long the request can wait:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl --fail --max-time 20 
  -H 'User-Agent: catalog-crawler/1.0 (+https://example.com/crawler-info)' 
  -H 'From: [email protected]' 
  https://example.org/data 
  -o data.html

Node.js fetch

Pass the headers in the request options. Check the response status before consuming the body:

const res = await fetch('https://example.org/data', {
  headers: {
    'User-Agent': 'catalog-crawler/1.0 (+https://example.com/crawler-info)',
    'From': '[email protected]'
  }
});

if (!res.ok) {
  throw new Error(`HTTP ${res.status}`);
}

const body = await res.text();
console.log(body);

Browser automation frameworks may manage User-Agent and related client hints for you. If you override them, keep the resulting identity truthful to the program actually making the request. Do not combine a browser-looking header with a non-browser client merely to get around a control.

Why changing User-Agent does not reliably fix a 403

A server can use many signals and policies besides the User-Agent. If a request is denied, changing one header is not a general bypass. Investigate the actual cause:

Symptom Likely issue Responsible next step
403 after switching between browser strings The site’s access policy, authentication, or bot controls reject the client. Read the site’s policy, authenticate if you are authorized, contact the operator, or stop. Do not keep rotating identities.
Requests slow down or begin failing at higher volume Excessive request rate or missing crawl pacing. Reduce concurrency, honor published crawl-delay guidance, add monitoring, and make your contact information available.
HTML is an empty shell while a browser shows data The page depends on JavaScript or another browser-only execution step. Determine whether an authorized API or a permitted rendering workflow exists; a User-Agent alone does not execute JavaScript.
401 or a redirect to login The resource requires authentication. Use credentials only with authorization and follow the service’s terms. A different User-Agent is not authentication.
Accidental processing of an error page The code did not check the HTTP result. Use status checks such as Requests’ raise_for_status() or Node’s res.ok check before parsing.

Also avoid relying on User-Agent sniffing for your own application logic. MDN describes parsing User-Agent strings to identify browser or device type as unreliable and recommends avoiding it unless it is genuinely necessary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run a crawler that operators can recognize

Keep identity and policy handling separate

Your User-Agent identifies the program. Your crawler still needs separate controls for URL discovery, robots rules, authentication, rate limits, retries, logging, and data retention. Changing the header is not a scaling strategy.

Use predictable pacing and observability

At higher volume, record the URL, timestamp, response status, latency, and whether a retry occurred. Set limits before increasing concurrency. If an operator contacts you, a stable product token and a valid From address make it possible to identify and adjust the responsible job.

Keep requests reproducible

Use one documented header value across workers and environments. A sudden mixture of versions makes debugging and policy matching harder. If you release a new crawler version, update the product version deliberately rather than embedding random identifiers.

Respect privacy and legal boundaries

Robots.txt is one part of the decision. Check terms of service, copyright restrictions, authentication rules, and applicable law before collecting or redistributing content. Do not use a misleading User-Agent to conceal activity that a site has prohibited.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup: ScreenshotNeo

If your task is to capture a rendered page rather than download raw HTML, ScreenshotNeo provides a website screenshot API and MCP server. It can accept custom headers, cookies, a user agent, and authorization when a permitted capture requires them, without making you install and operate a browser.

The one-call cURL request below returns a WebP screenshot. See the ScreenshotNeo API documentation for the available parameters:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.org -o shot.webp

Equivalent Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.org"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Equivalent Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.org' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Before capture, it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the result with X-Page-Verdict and X-Billed.
  • An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan.

Sign up for ScreenshotNeo to use the free 1,000-screenshot allowance without adding a card.

FAQ

Does a User-Agent identify a person?

No. It identifies the client program and helps a server recognize the software making the request. It is not proof of a human operator or authorization to access a resource.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does the robots.txt token have to equal the complete header?

No. RFC 9309 describes matching a crawler product token as a substring of the User-Agent header to the applicable robots.txt group. A header such as catalog-crawler/1.0 (+https://example.com/crawler-info) can therefore match a User-agent: catalog-crawler group.

When is a browser User-Agent legitimate?

Only when the requesting implementation is genuinely that browser or a browser automation workflow that accurately represents itself. Copying a browser string onto a different client to evade controls is misleading and does not solve policy, authentication, rate, or JavaScript requirements.

Frequently Asked Questions

Does a User-Agent identify a person?

No. It identifies the client program and is not proof of a human operator or authorization.

Does the robots.txt token have to equal the complete header?

No. The crawler product token can match a substring of the User-Agent value.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When is a browser User-Agent legitimate?

When the requesting implementation is genuinely that browser or an accurately represented browser-automation workflow—not when the string is copied to evade controls.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.