Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no magic header bundle that guarantees a scraper will be allowed through. Use headers to state the representation, language, credentials and client identity your authorized request actually needs; then diagnose redirects, cookies, rate limits and server controls before changing anything. A truthful User-Agent can identify your client, but it is not proof of identity or a bypass for bot protection.

This guide shows a maintainable workflow for authorized scraping in 2026, including robots.txt, Cloudflare-specific behavior, browser versus server runtimes, credential safety and a practical troubleshooting sequence.

What headers can—and cannot—do

HTTP headers influence content negotiation, session state, caching and how an application interprets a request. They do not grant permission. A site can require authentication, validate requests, enforce WAF rules, challenge automation or deny your network independently of the headers you send.

User-Agent is identification, not authentication

Set User-Agent to a stable, honest description of your application, ideally including a contact URL or email when the site’s policy requests one. Do not copy a current desktop browser value and treat it as a bypass. Cloudflare’s Browser Run documentation (updated June 16, 2026) states: “The User-Agent header is not a reliable way to identify Browser Run requests.” In that product, the value is configurable for most methods and can be sent by any HTTP client. Cloudflare documents stronger service identification through non-configurable headers and Web Bot Auth signatures; those mechanisms are specific to supported Cloudflare products, not a universal rule for every website.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Accept and Accept-Language describe the response you can use

Send an Accept value matching the formats your parser really handles, such as text/html,application/xhtml+xml. Use Accept-Language only for languages your workflow actually prefers. Cloudflare Workers guidance describes normalizing these headers for cache variation. That is cache correctness guidance, not evidence that either header prevents blocking.

Accept-Encoding belongs to your HTTP library

Let your client negotiate compression and decompress responses consistently. Cloudflare documents that, for traffic it proxies, the origin sees incoming Accept-Encoding set to br, gzip. This is provider-specific behavior; an origin outside that architecture may see something different.

Cookies represent state

Use a cookie jar when an authorized workflow requires login, consent or a session. Never hard-code another person’s cookies or reuse session cookies across unrelated jobs. Browser JavaScript cannot directly set the Cookie request header; the browser manages cookies. Cloudflare Workers treat cookies as ordinary headers, so code written for one runtime cannot be copied blindly into the other.

Referer, Origin and Sec-Fetch-* are not universal keys

Include these only when the documented application flow requires them and make their values truthful. There is no established universal set of browser-generated headers that unlocks access. Do not invent provider headers such as CF-*, X-Forwarded-* or client-IP fields to impersonate a proxy path; Cloudflare adds or transforms such fields within its own edge-to-origin architecture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check permission before changing a header

Start with the target owner’s terms, API documentation, export options, feeds and crawl policy. robots.txt is a voluntary signal: compliant crawlers can read disallowed paths and a published Crawl-delay: 2 can request a two-second interval, but neither creates access rights nor technically enforces them. Owners that need enforcement must use server-side controls such as authentication, request validation or WAF rules.

Cloudflare’s managed Browser Rendering /crawl endpoint is an example of a compliant option for an authorized site owner. Cloudflare says it discovers URLs from sitemaps and links, supports depth, page-limit and path-scope controls, incremental crawling, HTML/Markdown/structured JSON output, and honors robots.txt including crawl-delay. It also explicitly self-identifies as a bot and cannot bypass Cloudflare bot detection or captchas. Treat it as managed, policy-respecting crawling—not a way around a denial.

A diagnostic workflow for authorized requests

  1. Confirm the permitted interface. Prefer a documented API, export or feed. If the owner denies automation, stop or request permission rather than cycling through spoofed headers.
  2. Reproduce the real request. Use the same URL, method, authentication state and representation as the authorized browser or API flow. Record status, redirect chain, content type and a safe sample of the response body.
  3. Inspect runtime behavior. Verify redirect policy, cookie-jar persistence, compression/decompression and whether your environment permits the header you are trying to set.
  4. Add only documented requirements. Keep the User-Agent truthful. Add language, referer, origin or custom authorization headers only when the application documentation or observed authorized flow calls for them.
  5. Control rate and scope. Honor published crawl delays, use exponential backoff for transient failures, cache unchanged pages and limit concurrency to what the owner permits.
  6. Stop on a continuing denial. A 403, challenge or captcha is a control signal, not proof that one more browser-looking header is the answer. Seek an approved API, ask the owner or discontinue the request.

Runnable header examples

cURL with an honest identity and negotiated formats

curl --compressed 
  -A "ExampleResearchBot/1.0 (+https://example.com/bot-info)" 
  -H "Accept: text/html,application/xhtml+xml" 
  -H "Accept-Language: en" 
  -L --max-redirs 5 
  "https://example.com/articles" -o page.html

--compressed asks cURL to negotiate compression and decompress it. Use -L only when following redirects is safe for the request; do not combine automatic cross-host redirects with credentials unless you have reviewed where they can go.

Python requests with a session and bounded redirects

import requests

url = "https://example.com/articles"
headers = {
    "User-Agent": "ExampleResearchBot/1.0 (+https://example.com/bot-info)",
    "Accept": "text/html,application/xhtml+xml",
    "Accept-Language": "en",
}

with requests.Session() as session:
    session.headers.update(headers)
    response = session.get(url, timeout=(10, 60), allow_redirects=False)
    print(response.status_code, response.headers.get("content-type"))
    if response.is_redirect:
        location = response.headers.get("location")
        print("Redirect requires review:", location)
    response.raise_for_status()
    response.encoding = response.apparent_encoding or response.encoding
    html = response.text

A session stores cookies received during this authorized workflow. If you intentionally follow a redirect, validate the destination and avoid forwarding Authorization or Cookie to another hostname.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Node.js fetch with explicit redirect handling

const url = 'https://example.com/articles';
const res = await fetch(url, {
  redirect: 'manual',
  headers: {
    'User-Agent': 'ExampleResearchBot/1.0 (+https://example.com/bot-info)',
    'Accept': 'text/html,application/xhtml+xml',
    'Accept-Language': 'en'
  }
});

console.log(res.status, res.headers.get('content-type'));
if (res.status >= 300 && res.status < 400) {
  console.log('Review redirect:', res.headers.get('location'));
}
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();

Node’s built-in fetch does not provide a browser cookie jar automatically. Add a maintained cookie-jar library only when the target’s authorized flow requires it, and keep credentials scoped to the intended host.

Redirects, credentials and proxy paths

Cloudflare warns that a Worker fetch() configured to follow redirects can forward sensitive headers such as Cookie and Authorization to the redirect destination, even across hostnames. The safe default for credentialed collection is to inspect redirects, allow only expected hosts and re-create a request without secrets when a host changes.

Cloudflare also documents that it passes request headers to an origin while removing invalid names and applying provider-specific transformations. For example, CF-Connecting-IP is a Cloudflare-to-origin client-IP value, not a header your scraper should manufacture. Treat proxy-added headers as observations about that provider’s architecture, not portable HTTP requirements.

Static HTTP versus a managed browser

Situation Best first option Reason
Documented API or server-rendered HTML Direct HTTP client Lower complexity, explicit headers and predictable parsing.
Authenticated workflow with cookies HTTP session or approved API token Preserves state without pretending to be a browser.
JavaScript-rendered content required by the owner’s workflow Authorized browser automation or managed rendering Executes page code when static HTML is insufficient.
Cloudflare-protected site denying automation Owner-approved API or permission request Changing headers is not a legitimate bypass.

Choose on permission, content requirements, runtime behavior, transparency and operational controls—not on which option supposedly “beats” a vendor’s defenses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability and cost controls

  • Reuse connections: keep a session or HTTP agent alive where the client supports it.
  • Set timeouts: separate connect and read timeouts; never leave requests unbounded.
  • Back off: retry only transient network and 5xx failures, with exponential delays and a maximum attempt count. Do not blindly retry 401, 403 or captcha responses.
  • Cache responsibly: honor freshness headers where practical and use conditional requests such as If-None-Match or If-Modified-Since when supported.
  • Limit concurrency: apply per-host queues and honor the target’s published pacing guidance.
  • Log safely: record status, timing, redirect destinations, content type and request ID; redact authorization headers, cookies and personal data.
  • Validate content: a 200 response can still be a challenge page, login form or empty shell. Check expected content type and markers before parsing.

Common failures and fixes

403 Forbidden

Confirm permission, authentication and request scope. Compare the response with the documented API flow, reduce rate and inspect the body for a policy or challenge message. Do not assume a browser User-Agent will resolve it.

401 Unauthorized

Check token format, expiry, audience and hostname. Keep credentials out of logs and do not send them through an unreviewed redirect.

429 Too Many Requests

Honor Retry-After when present, reduce concurrency, add backoff and cache results. Ask the owner for a quota or API key if the workload is legitimate.

Unexpected language or format

Verify Accept, Accept-Language, URL parameters and cache variation. A proxy or CDN may normalize these values, so inspect the response headers and body rather than guessing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Missing login state

Use a persistent cookie jar in server-side code, complete the approved login flow and verify that cookies are scoped to the correct domain and path. Browser scripts cannot set Cookie directly.

Compressed or unreadable body

Let the HTTP library handle decompression, or disable compression while debugging. Check the response’s Content-Encoding and avoid manually decompressing data your library already decoded.

Redirect to an unexpected host

Switch to manual redirect handling, allow-list expected hosts and strip credentials before any cross-host request. Treat an unexpected destination as a configuration or security failure.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your actual need is a clean visual capture rather than parsed page data, ScreenshotNeo returns a PNG, JPEG, WebP or PDF from one request. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example request (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

You can also use Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Or Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every plan includes the features: full-page and element capture, device presets or custom viewports, retina scale, dark mode, custom CSS/JavaScript, waits, selector hiding, request blocking, cookies and headers, geolocation and timezone, transparent backgrounds, resizing, caching with your TTL, signed links, async webhooks, bulk capture for up to 100 URLs per call, usage API and OpenAPI specification. The parameter names used by other screenshot APIs also work.

Plan Included shots Price
Free 1,000/month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Yearly billing gives two months free. Start with 1,000 free screenshots a month—no card required.

Frequently Asked Questions

Should I rotate User-Agent strings?

Not as a default strategy. Use one truthful identity for your application and follow the target’s documented policy; rotation can make auditing and permission checks harder.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does robots.txt tell me whether scraping is legal?

No. It communicates crawler preferences and may include a delay, but it is not an access-control mechanism or legal authorization. Review the site’s terms and obtain permission where required.

When should I use browser automation instead of HTTP requests?

Use a browser only when the authorized workflow genuinely requires JavaScript rendering or browser-managed state. For APIs and server-rendered pages, direct HTTP is usually simpler and easier to control.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.