There is no magic header bundle that guarantees a scraper will be allowed through. Use headers to state the representation, language, credentials and client identity your authorized request actually needs; then diagnose redirects, cookies, rate limits and server controls before changing anything. A truthful User-Agent can identify your client, but it is not proof of identity or a bypass for bot protection.
This guide shows a maintainable workflow for authorized scraping in 2026, including robots.txt, Cloudflare-specific behavior, browser versus server runtimes, credential safety and a practical troubleshooting sequence.
What headers can—and cannot—do
HTTP headers influence content negotiation, session state, caching and how an application interprets a request. They do not grant permission. A site can require authentication, validate requests, enforce WAF rules, challenge automation or deny your network independently of the headers you send.
User-Agent is identification, not authentication
Set User-Agent to a stable, honest description of your application, ideally including a contact URL or email when the site’s policy requests one. Do not copy a current desktop browser value and treat it as a bypass. Cloudflare’s Browser Run documentation (updated June 16, 2026) states: “The User-Agent header is not a reliable way to identify Browser Run requests.” In that product, the value is configurable for most methods and can be sent by any HTTP client. Cloudflare documents stronger service identification through non-configurable headers and Web Bot Auth signatures; those mechanisms are specific to supported Cloudflare products, not a universal rule for every website.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Accept and Accept-Language describe the response you can use
Send an Accept value matching the formats your parser really handles, such as text/html,application/xhtml+xml. Use Accept-Language only for languages your workflow actually prefers. Cloudflare Workers guidance describes normalizing these headers for cache variation. That is cache correctness guidance, not evidence that either header prevents blocking.
Accept-Encoding belongs to your HTTP library
Let your client negotiate compression and decompress responses consistently. Cloudflare documents that, for traffic it proxies, the origin sees incoming Accept-Encoding set to br, gzip. This is provider-specific behavior; an origin outside that architecture may see something different.
Cookies represent state
Use a cookie jar when an authorized workflow requires login, consent or a session. Never hard-code another person’s cookies or reuse session cookies across unrelated jobs. Browser JavaScript cannot directly set the Cookie request header; the browser manages cookies. Cloudflare Workers treat cookies as ordinary headers, so code written for one runtime cannot be copied blindly into the other.
Referer, Origin and Sec-Fetch-* are not universal keys
Include these only when the documented application flow requires them and make their values truthful. There is no established universal set of browser-generated headers that unlocks access. Do not invent provider headers such as CF-*, X-Forwarded-* or client-IP fields to impersonate a proxy path; Cloudflare adds or transforms such fields within its own edge-to-origin architecture.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Check permission before changing a header
Start with the target owner’s terms, API documentation, export options, feeds and crawl policy. robots.txt is a voluntary signal: compliant crawlers can read disallowed paths and a published Crawl-delay: 2 can request a two-second interval, but neither creates access rights nor technically enforces them. Owners that need enforcement must use server-side controls such as authentication, request validation or WAF rules.
Cloudflare’s managed Browser Rendering /crawl endpoint is an example of a compliant option for an authorized site owner. Cloudflare says it discovers URLs from sitemaps and links, supports depth, page-limit and path-scope controls, incremental crawling, HTML/Markdown/structured JSON output, and honors robots.txt including crawl-delay. It also explicitly self-identifies as a bot and cannot bypass Cloudflare bot detection or captchas. Treat it as managed, policy-respecting crawling—not a way around a denial.
A diagnostic workflow for authorized requests
- Confirm the permitted interface. Prefer a documented API, export or feed. If the owner denies automation, stop or request permission rather than cycling through spoofed headers.
- Reproduce the real request. Use the same URL, method, authentication state and representation as the authorized browser or API flow. Record status, redirect chain, content type and a safe sample of the response body.
- Inspect runtime behavior. Verify redirect policy, cookie-jar persistence, compression/decompression and whether your environment permits the header you are trying to set.
- Add only documented requirements. Keep the User-Agent truthful. Add language, referer, origin or custom authorization headers only when the application documentation or observed authorized flow calls for them.
- Control rate and scope. Honor published crawl delays, use exponential backoff for transient failures, cache unchanged pages and limit concurrency to what the owner permits.
- Stop on a continuing denial. A 403, challenge or captcha is a control signal, not proof that one more browser-looking header is the answer. Seek an approved API, ask the owner or discontinue the request.
Runnable header examples
cURL with an honest identity and negotiated formats
curl --compressed
-A "ExampleResearchBot/1.0 (+https://example.com/bot-info)"
-H "Accept: text/html,application/xhtml+xml"
-H "Accept-Language: en"
-L --max-redirs 5
"https://example.com/articles" -o page.html
--compressed asks cURL to negotiate compression and decompress it. Use -L only when following redirects is safe for the request; do not combine automatic cross-host redirects with credentials unless you have reviewed where they can go.
Python requests with a session and bounded redirects
import requests
url = "https://example.com/articles"
headers = {
"User-Agent": "ExampleResearchBot/1.0 (+https://example.com/bot-info)",
"Accept": "text/html,application/xhtml+xml",
"Accept-Language": "en",
}
with requests.Session() as session:
session.headers.update(headers)
response = session.get(url, timeout=(10, 60), allow_redirects=False)
print(response.status_code, response.headers.get("content-type"))
if response.is_redirect:
location = response.headers.get("location")
print("Redirect requires review:", location)
response.raise_for_status()
response.encoding = response.apparent_encoding or response.encoding
html = response.text
A session stores cookies received during this authorized workflow. If you intentionally follow a redirect, validate the destination and avoid forwarding Authorization or Cookie to another hostname.
Node.js fetch with explicit redirect handling
const url = 'https://example.com/articles';
const res = await fetch(url, {
redirect: 'manual',
headers: {
'User-Agent': 'ExampleResearchBot/1.0 (+https://example.com/bot-info)',
'Accept': 'text/html,application/xhtml+xml',
'Accept-Language': 'en'
}
});
console.log(res.status, res.headers.get('content-type'));
if (res.status >= 300 && res.status < 400) {
console.log('Review redirect:', res.headers.get('location'));
}
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
Node’s built-in fetch does not provide a browser cookie jar automatically. Add a maintained cookie-jar library only when the target’s authorized flow requires it, and keep credentials scoped to the intended host.
Redirects, credentials and proxy paths
Cloudflare warns that a Worker fetch() configured to follow redirects can forward sensitive headers such as Cookie and Authorization to the redirect destination, even across hostnames. The safe default for credentialed collection is to inspect redirects, allow only expected hosts and re-create a request without secrets when a host changes.
Rank #3
Cloudflare also documents that it passes request headers to an origin while removing invalid names and applying provider-specific transformations. For example, CF-Connecting-IP is a Cloudflare-to-origin client-IP value, not a header your scraper should manufacture. Treat proxy-added headers as observations about that provider’s architecture, not portable HTTP requirements.
Static HTTP versus a managed browser
| Situation | Best first option | Reason |
|---|---|---|
| Documented API or server-rendered HTML | Direct HTTP client | Lower complexity, explicit headers and predictable parsing. |
| Authenticated workflow with cookies | HTTP session or approved API token | Preserves state without pretending to be a browser. |
| JavaScript-rendered content required by the owner’s workflow | Authorized browser automation or managed rendering | Executes page code when static HTML is insufficient. |
| Cloudflare-protected site denying automation | Owner-approved API or permission request | Changing headers is not a legitimate bypass. |
Choose on permission, content requirements, runtime behavior, transparency and operational controls—not on which option supposedly “beats” a vendor’s defenses.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsPerformance, reliability and cost controls
- Reuse connections: keep a session or HTTP agent alive where the client supports it.
- Set timeouts: separate connect and read timeouts; never leave requests unbounded.
- Back off: retry only transient network and 5xx failures, with exponential delays and a maximum attempt count. Do not blindly retry 401, 403 or captcha responses.
- Cache responsibly: honor freshness headers where practical and use conditional requests such as
If-None-MatchorIf-Modified-Sincewhen supported. - Limit concurrency: apply per-host queues and honor the target’s published pacing guidance.
- Log safely: record status, timing, redirect destinations, content type and request ID; redact authorization headers, cookies and personal data.
- Validate content: a 200 response can still be a challenge page, login form or empty shell. Check expected content type and markers before parsing.
Common failures and fixes
403 Forbidden
Confirm permission, authentication and request scope. Compare the response with the documented API flow, reduce rate and inspect the body for a policy or challenge message. Do not assume a browser User-Agent will resolve it.
401 Unauthorized
Check token format, expiry, audience and hostname. Keep credentials out of logs and do not send them through an unreviewed redirect.
429 Too Many Requests
Honor Retry-After when present, reduce concurrency, add backoff and cache results. Ask the owner for a quota or API key if the workload is legitimate.
Unexpected language or format
Verify Accept, Accept-Language, URL parameters and cache variation. A proxy or CDN may normalize these values, so inspect the response headers and body rather than guessing.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Missing login state
Use a persistent cookie jar in server-side code, complete the approved login flow and verify that cookies are scoped to the correct domain and path. Browser scripts cannot set Cookie directly.
Compressed or unreadable body
Let the HTTP library handle decompression, or disable compression while debugging. Check the response’s Content-Encoding and avoid manually decompressing data your library already decoded.
Redirect to an unexpected host
Switch to manual redirect handling, allow-list expected hosts and strip credentials before any cross-host request. Treat an unexpected destination as a configuration or security failure.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your actual need is a clean visual capture rather than parsed page data, ScreenshotNeo returns a PNG, JPEG, WebP or PDF from one request. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchExample request (see the ScreenshotNeo API documentation):
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
You can also use Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Or Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every plan includes the features: full-page and element capture, device presets or custom viewports, retina scale, dark mode, custom CSS/JavaScript, waits, selector hiding, request blocking, cookies and headers, geolocation and timezone, transparent backgrounds, resizing, caching with your TTL, signed links, async webhooks, bulk capture for up to 100 URLs per call, usage API and OpenAPI specification. The parameter names used by other screenshot APIs also work.
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000/month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing gives two months free. Start with 1,000 free screenshots a month—no card required.
Frequently Asked Questions
Should I rotate User-Agent strings?
Not as a default strategy. Use one truthful identity for your application and follow the target’s documented policy; rotation can make auditing and permission checks harder.
Does robots.txt tell me whether scraping is legal?
No. It communicates crawler preferences and may include a delay, but it is not an access-control mechanism or legal authorization. Review the site’s terms and obtain permission where required.
When should I use browser automation instead of HTTP requests?
Use a browser only when the authorized workflow genuinely requires JavaScript rendering or browser-managed state. For APIs and server-rendered pages, direct HTTP is usually simpler and easier to control.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

