Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Fetching a web page programmatically means making an HTTP request, checking the response, and reading its body. For static HTML, a server-side client such as Python’s built-in urllib.request can retrieve the document directly. In browser JavaScript, the promise-based fetch() API does the same when the target permits cross-origin access. Neither approach executes the page’s JavaScript or reproduces its layout; dynamic pages require a documented data endpoint or permitted browser automation.
The basic fetch workflow
A reliable fetch has three distinct stages:
- Build and send the request. Usually this is an HTTP
GETfor a URL. - Classify the response. Check the status code, content type, redirects, authentication challenges, rate limits, and transport errors.
- Read and decode the body. Apply the declared character encoding and impose a size limit before parsing HTML, JSON, or another format.
HTTP GET requests ask for a representation of a resource. They have no request body and are defined as safe, idempotent, and cacheable. Use POST or another method only when the destination API requires it or the operation changes server state.
Fetch HTML with Python’s standard library
Python 3 includes urllib.request, so this example needs no third-party package. It sends an identifiable user agent, applies a finite timeout, checks the status and content type, and separates HTTP errors from URL or network failures.
from urllib.request import Request, urlopen
from urllib.error import HTTPError, URLError
url = "https://example.org/"
request = Request(url, headers={"User-Agent": "my-fetcher/1.0"})
try:
with urlopen(request, timeout=10) as response:
status = response.status
content_type = response.headers.get("Content-Type", "")
if status < 200 or status >= 300:
raise RuntimeError(f"HTTP status {status}")
if "text/html" not in content_type.lower():
raise RuntimeError(f"Unexpected content type: {content_type}")
html_bytes = response.read()
html = html_bytes.decode(response.headers.get_content_charset() or "utf-8", errors="replace")
print(html)
except HTTPError as exc:
print(f"HTTP error: {exc.code} {exc.reason}")
except URLError as exc:
print(f"Network or URL error: {exc.reason}")
except TimeoutError:
print("The request timed out")
When no data argument is supplied, Request performs a GET. The Python documentation’s minimal pattern is a context-managed urlopen call followed by response.read(); the additional checks above make that pattern safer for production. urllib.request uses HTTP/1.1 and sends Connection: close, so a high-volume service may benefit from a client that supports connection pooling.
#1 Best Overall
Limit the amount you read
Do not let an untrusted URL consume unlimited memory. Read in chunks and stop at an application-specific maximum, such as 10 MB:
MAX_BYTES = 10 * 1024 * 1024
chunks = []
total = 0
while True:
chunk = response.read(min(64 * 1024, MAX_BYTES - total + 1))
if not chunk:
break
chunks.append(chunk)
total += len(chunk)
if total > MAX_BYTES:
raise RuntimeError("Response exceeds the configured size limit")
html = b"".join(chunks).decode(
response.headers.get_content_charset() or "utf-8", errors="replace"
)
URLs, headers, cookies, and authentication
Validate and normalize the URL before making a request. Restrict schemes to those your application is designed to handle, normally https and, where explicitly needed, http. Add headers with Request; send cookies only when you have a legitimate reason and follow the site’s authentication requirements. Use a truthful user agent rather than impersonating a browser to bypass controls.
Fetch a page with browser JavaScript
The Fetch API returns a promise for a Response. A rejected promise generally indicates a network or permission failure—not an HTTP 404 or 504—so check ok or status yourself. Body readers such as text() and json() are asynchronous.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsRank #2
async function fetchPage(url) {
const response = await fetch(url, { method: "GET" });
if (!response.ok) {
throw new Error(`HTTP ${response.status}`);
}
const contentType = response.headers.get("content-type") || "";
if (!contentType.toLowerCase().includes("text/html")) {
throw new Error(`Unexpected content type: ${contentType}`);
}
return await response.text();
}
fetchPage("https://example.org/")
.then(html => console.log(html))
.catch(error => console.error(error));
Read JSON instead of HTML
async function fetchJson(url) {
const response = await fetch(url, {
headers: { "Accept": "application/json" }
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const type = response.headers.get("content-type") || "";
if (!type.includes("application/json")) {
throw new Error(`Unexpected content type: ${type}`);
}
return response.json();
}
Why browser fetch fails across domains: CORS
Browser scripts are constrained by the same-origin policy. A cross-origin Fetch request can be read only when the destination sends an appropriate Access-Control-Allow-Origin response header (and, for some requests, the other required CORS headers). The browser may first send an OPTIONS preflight when the method or headers are not a simple request.
mode: "no-cors" is not a way to read another site’s HTML. It normally returns an opaque response whose headers and body are unavailable to JavaScript. If the server does not grant CORS access, use one of these permitted designs:
- Fetch from your own server, where browser same-origin restrictions do not apply, subject to the destination’s policies.
- Expose a same-origin backend endpoint that retrieves and validates the target URL.
- Use the destination’s documented cross-origin API.
Do not build an open proxy. Validate schemes and destinations, restrict private-network access, cap response sizes, authenticate your own endpoint, and apply rate limits.
Static HTML is not a rendered web page
An HTTP client receives the server’s response bytes. It does not execute JavaScript, wait for client-side rendering, recreate browser storage, click controls, or calculate layout. A successful 200 therefore does not prove that the content a visitor sees is present in the downloaded HTML.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Prefer a documented data endpoint
For a JavaScript application, inspect its documented API or server-rendered data endpoint and fetch that representation directly when permitted. This is usually faster and more stable than reproducing UI interactions.
Use browser automation when rendering is required
If the only permitted way to obtain the content is to execute scripts, use a browser automation tool that supports the site’s authentication, consent, and interaction requirements. Wait for a specific selector, a known application state, or network idle rather than relying on an arbitrary sleep. Respect robots.txt guidance, rate limits, authentication rules, and terms of service; those conditions are site-specific.
Production controls that prevent fragile fetchers
- Timeouts: Set finite connect and read limits. Cancel work that exceeds them.
- Status handling: Classify redirects, 401/403 authentication failures, 404s, 429 rate limits, and 5xx server errors separately.
- Retries: Retry only transient failures, with exponential backoff and a cap. Do not blindly retry validation errors or authentication failures.
- Encoding: Inspect
Content-Typeand its charset before decoding. Keep raw bytes when you need exact fidelity. - Size: Enforce a maximum response size and, for compressed responses, account for the expanded size.
- Connections: Reuse connections with a pooling client when making many requests; the standard library’s HTTP/1.1 behavior closes connections by default.
- Identity: Send a truthful, identifiable user agent and identify your application where appropriate.
- Security: Validate URLs, protect credentials, avoid server-side request forgery, and never treat downloaded HTML as trusted markup without sanitizing it.
- Observability: Record URL, elapsed time, status, response size, retry count, and a categorized error—without logging cookies or authorization tokens.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Fetch rejects with “Failed to fetch” in a browser | CORS denial, DNS failure, TLS failure, or a blocked request | Inspect the browser console and network panel. If CORS is missing, move the request to an authorized server or use the site’s API. |
Python raises HTTPError |
The server returned an HTTP error status | Use the status code to decide whether to authenticate, correct the URL, back off, or report a permanent failure. |
| Response is HTML instead of JSON | A login page, error page, redirect, or wrong endpoint | Check final URL, status, and Content-Type before parsing. |
| HTML contains no visible article text | Content is inserted by client-side JavaScript | Find a documented data endpoint or use permitted browser automation. |
| Request hangs | No timeout, slow server, stalled TLS, or a never-ending stream | Set connect/read timeouts, cap bytes, cancel the operation, and retry transient failures with backoff. |
| 429 Too Many Requests | Rate limit exceeded | Honor Retry-After when supplied, reduce concurrency, cache results, and review the site’s limits. |
| TLS or certificate error | Invalid certificate, hostname mismatch, or local trust-store problem | Fix the certificate or trust configuration. Do not disable verification in production. |
Choosing the right approach
| Requirement | Best starting point | Important constraint |
|---|---|---|
| Static HTML from a server | Python urllib.request or another server HTTP client |
Handle status, encoding, limits, and timeouts. |
| Request initiated by your web application | Browser fetch() |
Cross-origin reads require CORS permission. |
| Cross-origin data without CORS | Authorized backend fetch or documented API | Secure the proxy against abuse and SSRF. |
| Content created after JavaScript runs | Documented data endpoint or permitted browser automation | An HTTP 200 alone does not mean the rendered page was captured. |
| Many repeated requests | Client with connection pooling, caching, and bounded concurrency | Respect rate limits and use backoff. |
Or skip the browser setup
If your goal is a clean image or PDF of a rendered page rather than raw HTML, ScreenshotNeo makes one request to its website screenshot API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the complete parameter reference in the ScreenshotNeo documentation. The same endpoint also supports full-page captures with lazy images loaded, CSS-selector element shots, dark mode, device presets and custom viewports, retina scale, PDF paper and page options, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, headers, cookies, user agents and authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs are accepted to ease migration.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
For Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
FAQ
Does fetching a page download its images and CSS?
A normal HTTP fetch returns only the requested response. You must discover and request linked resources separately; a browser renderer loads them as part of page execution.
Best Value
Can I use Fetch API from Node.js?
Node.js provides a server-side fetch in current releases, but its availability and defaults depend on the Node version. Apply the same status, timeout, size, and content-type checks shown above.
Is scraping every public page allowed?
Public visibility does not remove contractual, copyright, authentication, robots, or rate-limit constraints. Check the site’s published rules and obtain permission where required.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

