Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Fetching a web page programmatically means making an HTTP request, checking the response, and reading its body. For static HTML, a server-side client such as Python’s built-in urllib.request can retrieve the document directly. In browser JavaScript, the promise-based fetch() API does the same when the target permits cross-origin access. Neither approach executes the page’s JavaScript or reproduces its layout; dynamic pages require a documented data endpoint or permitted browser automation.

The basic fetch workflow

A reliable fetch has three distinct stages:

  1. Build and send the request. Usually this is an HTTP GET for a URL.
  2. Classify the response. Check the status code, content type, redirects, authentication challenges, rate limits, and transport errors.
  3. Read and decode the body. Apply the declared character encoding and impose a size limit before parsing HTML, JSON, or another format.

HTTP GET requests ask for a representation of a resource. They have no request body and are defined as safe, idempotent, and cacheable. Use POST or another method only when the destination API requires it or the operation changes server state.

Fetch HTML with Python’s standard library

Python 3 includes urllib.request, so this example needs no third-party package. It sends an identifiable user agent, applies a finite timeout, checks the status and content type, and separates HTTP errors from URL or network failures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.request import Request, urlopen
from urllib.error import HTTPError, URLError

url = "https://example.org/"
request = Request(url, headers={"User-Agent": "my-fetcher/1.0"})

try:
    with urlopen(request, timeout=10) as response:
        status = response.status
        content_type = response.headers.get("Content-Type", "")
        if status < 200 or status >= 300:
            raise RuntimeError(f"HTTP status {status}")
        if "text/html" not in content_type.lower():
            raise RuntimeError(f"Unexpected content type: {content_type}")
        html_bytes = response.read()
        html = html_bytes.decode(response.headers.get_content_charset() or "utf-8", errors="replace")
        print(html)
except HTTPError as exc:
    print(f"HTTP error: {exc.code} {exc.reason}")
except URLError as exc:
    print(f"Network or URL error: {exc.reason}")
except TimeoutError:
    print("The request timed out")

When no data argument is supplied, Request performs a GET. The Python documentation’s minimal pattern is a context-managed urlopen call followed by response.read(); the additional checks above make that pattern safer for production. urllib.request uses HTTP/1.1 and sends Connection: close, so a high-volume service may benefit from a client that supports connection pooling.

Limit the amount you read

Do not let an untrusted URL consume unlimited memory. Read in chunks and stop at an application-specific maximum, such as 10 MB:

MAX_BYTES = 10 * 1024 * 1024
chunks = []
total = 0
while True:
    chunk = response.read(min(64 * 1024, MAX_BYTES - total + 1))
    if not chunk:
        break
    chunks.append(chunk)
    total += len(chunk)
    if total > MAX_BYTES:
        raise RuntimeError("Response exceeds the configured size limit")
html = b"".join(chunks).decode(
    response.headers.get_content_charset() or "utf-8", errors="replace"
)

URLs, headers, cookies, and authentication

Validate and normalize the URL before making a request. Restrict schemes to those your application is designed to handle, normally https and, where explicitly needed, http. Add headers with Request; send cookies only when you have a legitimate reason and follow the site’s authentication requirements. Use a truthful user agent rather than impersonating a browser to bypass controls.

Fetch a page with browser JavaScript

The Fetch API returns a promise for a Response. A rejected promise generally indicates a network or permission failure—not an HTTP 404 or 504—so check ok or status yourself. Body readers such as text() and json() are asynchronous.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
async function fetchPage(url) {
  const response = await fetch(url, { method: "GET" });
  if (!response.ok) {
    throw new Error(`HTTP ${response.status}`);
  }
  const contentType = response.headers.get("content-type") || "";
  if (!contentType.toLowerCase().includes("text/html")) {
    throw new Error(`Unexpected content type: ${contentType}`);
  }
  return await response.text();
}

fetchPage("https://example.org/")
  .then(html => console.log(html))
  .catch(error => console.error(error));

Read JSON instead of HTML

async function fetchJson(url) {
  const response = await fetch(url, {
    headers: { "Accept": "application/json" }
  });
  if (!response.ok) throw new Error(`HTTP ${response.status}`);
  const type = response.headers.get("content-type") || "";
  if (!type.includes("application/json")) {
    throw new Error(`Unexpected content type: ${type}`);
  }
  return response.json();
}

Why browser fetch fails across domains: CORS

Browser scripts are constrained by the same-origin policy. A cross-origin Fetch request can be read only when the destination sends an appropriate Access-Control-Allow-Origin response header (and, for some requests, the other required CORS headers). The browser may first send an OPTIONS preflight when the method or headers are not a simple request.

mode: "no-cors" is not a way to read another site’s HTML. It normally returns an opaque response whose headers and body are unavailable to JavaScript. If the server does not grant CORS access, use one of these permitted designs:

  • Fetch from your own server, where browser same-origin restrictions do not apply, subject to the destination’s policies.
  • Expose a same-origin backend endpoint that retrieves and validates the target URL.
  • Use the destination’s documented cross-origin API.

Do not build an open proxy. Validate schemes and destinations, restrict private-network access, cap response sizes, authenticate your own endpoint, and apply rate limits.

Static HTML is not a rendered web page

An HTTP client receives the server’s response bytes. It does not execute JavaScript, wait for client-side rendering, recreate browser storage, click controls, or calculate layout. A successful 200 therefore does not prove that the content a visitor sees is present in the downloaded HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prefer a documented data endpoint

For a JavaScript application, inspect its documented API or server-rendered data endpoint and fetch that representation directly when permitted. This is usually faster and more stable than reproducing UI interactions.

Use browser automation when rendering is required

If the only permitted way to obtain the content is to execute scripts, use a browser automation tool that supports the site’s authentication, consent, and interaction requirements. Wait for a specific selector, a known application state, or network idle rather than relying on an arbitrary sleep. Respect robots.txt guidance, rate limits, authentication rules, and terms of service; those conditions are site-specific.

Production controls that prevent fragile fetchers

  • Timeouts: Set finite connect and read limits. Cancel work that exceeds them.
  • Status handling: Classify redirects, 401/403 authentication failures, 404s, 429 rate limits, and 5xx server errors separately.
  • Retries: Retry only transient failures, with exponential backoff and a cap. Do not blindly retry validation errors or authentication failures.
  • Encoding: Inspect Content-Type and its charset before decoding. Keep raw bytes when you need exact fidelity.
  • Size: Enforce a maximum response size and, for compressed responses, account for the expanded size.
  • Connections: Reuse connections with a pooling client when making many requests; the standard library’s HTTP/1.1 behavior closes connections by default.
  • Identity: Send a truthful, identifiable user agent and identify your application where appropriate.
  • Security: Validate URLs, protect credentials, avoid server-side request forgery, and never treat downloaded HTML as trusted markup without sanitizing it.
  • Observability: Record URL, elapsed time, status, response size, retry count, and a categorized error—without logging cookies or authorization tokens.

Common failures and fixes

Symptom Likely cause Fix
Fetch rejects with “Failed to fetch” in a browser CORS denial, DNS failure, TLS failure, or a blocked request Inspect the browser console and network panel. If CORS is missing, move the request to an authorized server or use the site’s API.
Python raises HTTPError The server returned an HTTP error status Use the status code to decide whether to authenticate, correct the URL, back off, or report a permanent failure.
Response is HTML instead of JSON A login page, error page, redirect, or wrong endpoint Check final URL, status, and Content-Type before parsing.
HTML contains no visible article text Content is inserted by client-side JavaScript Find a documented data endpoint or use permitted browser automation.
Request hangs No timeout, slow server, stalled TLS, or a never-ending stream Set connect/read timeouts, cap bytes, cancel the operation, and retry transient failures with backoff.
429 Too Many Requests Rate limit exceeded Honor Retry-After when supplied, reduce concurrency, cache results, and review the site’s limits.
TLS or certificate error Invalid certificate, hostname mismatch, or local trust-store problem Fix the certificate or trust configuration. Do not disable verification in production.

Choosing the right approach

Requirement Best starting point Important constraint
Static HTML from a server Python urllib.request or another server HTTP client Handle status, encoding, limits, and timeouts.
Request initiated by your web application Browser fetch() Cross-origin reads require CORS permission.
Cross-origin data without CORS Authorized backend fetch or documented API Secure the proxy against abuse and SSRF.
Content created after JavaScript runs Documented data endpoint or permitted browser automation An HTTP 200 alone does not mean the rendered page was captured.
Many repeated requests Client with connection pooling, caching, and bounded concurrency Respect rate limits and use backoff.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF of a rendered page rather than raw HTML, ScreenshotNeo makes one request to its website screenshot API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the complete parameter reference in the ScreenshotNeo documentation. The same endpoint also supports full-page captures with lazy images loaded, CSS-selector element shots, dark mode, device presets and custom viewports, retina scale, PDF paper and page options, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, headers, cookies, user agents and authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs are accepted to ease migration.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

For Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

FAQ

Does fetching a page download its images and CSS?

A normal HTTP fetch returns only the requested response. You must discover and request linked resources separately; a browser renderer loads them as part of page execution.

Can I use Fetch API from Node.js?

Node.js provides a server-side fetch in current releases, but its availability and defaults depend on the Node version. Apply the same status, timeout, size, and content-type checks shown above.

Is scraping every public page allowed?

Public visibility does not remove contractual, copyright, authentication, robots, or rate-limit constraints. Check the site’s published rules and obtain permission where required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.