Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use ordinary HTTP requests for pages whose returned HTML already contains the content and links you need. Use a browser renderer when JavaScript adds that content or navigation after the initial response. A practical crawler can inspect both versions, extract real links from <a href> elements, resolve them against the final URL, filter and deduplicate them, then add in-scope URLs to a controlled queue.

Crawling, rendering and indexing are separate operations. A crawler finding or rendering a URL does not prove that a search engine will index it. The example below builds a small, polite Python crawler for inspection—not a search-engine replica.

What crawling a JavaScript site involves

A JavaScript website can return a complete HTML document, a partially populated page, or an application shell that relies on scripts to insert its main content and links. Fetching the response and executing the page are different operations:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • HTTP fetch: retrieve the response and parse its HTML without running page JavaScript. This is simpler and usually uses fewer resources, and it works when the content and links are already in the response.
  • Browser rendering: open the page in a browser engine, let its scripts run, then inspect the rendered DOM. Use this when the response lacks material content or links that appear after execution.

Google Search Central describes its own crawl, render and index stages, including a headless Chromium rendering stage. That is useful context, but it is not a guarantee that your crawler—or another search engine—will behave the same way. Google’s JavaScript SEO documentation also says rendering may take longer than a few seconds, resources must be available, and pages can wait in a rendering queue. An independent crawler should choose its own limits and readiness rules.

Set scope and limits before crawling

Start with one or more seed URLs, then define what the crawler is allowed to visit. These are engineering controls, not requirements attributed here to Google.

  • Scope: allowed hostnames or URL prefixes. Decide whether subdomains, query strings and alternate schemes belong.
  • Depth and volume: maximum link depth and maximum number of pages. A depth of zero can mean seeds only; each newly followed link adds one level.
  • Resource limits: request and browser timeouts, maximum response size, and a delay or concurrency limit appropriate to the site.
  • Policy: identify yourself appropriately, review the site’s terms and robots.txt, and do not bypass access controls. Robots.txt is a crawl convention, not authorization to access restricted content.
  • Output: record each URL, final URL after redirects, response status, whether rendering was used, and any failure reason.

A Python crawler with HTTP-first rendering fallback

This example uses requests and Beautiful Soup for the initial response, then Playwright when that response appears to need JavaScript. It follows only links with an actual href in an anchor element and stays on the seed hostname. The readiness check is intentionally modest: it waits for the page load event, then inspects the DOM. Real sites may need a site-specific selector or other signal.

Install dependencies

Use Python 3.9 or later. Install Playwright’s Python package and the browser binary it needs; keep the installed browser aligned with the Playwright package version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

python -m pip install requests beautifulsoup4 playwright
python -m playwright install chromium

Save and run the crawler

Save as crawl_js.py. Change SEEDS and, where appropriate, MAX_PAGES, MAX_DEPTH and READY_SELECTOR. The sample uses one hostname as its scope and does not try to normalize away meaningful query parameters.

from collections import deque
from urllib.parse import urldefrag, urljoin, urlparse

import requests
from bs4 import BeautifulSoup
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

SEEDS = ["https://example.com/"]
MAX_PAGES = 100
MAX_DEPTH = 2
HTTP_TIMEOUT = 20
BROWSER_TIMEOUT_MS = 30000
READY_SELECTOR = None # Example: "main article"

session = requests.Session()
session.headers.update({"User-Agent": "ExampleResearchCrawler/1.0 (contact: [email protected])"})

def clean_url(url):
# Remove fragments: they are not sent as part of an HTTP request.
return urldefrag(url)[0]

def in_scope(url, allowed_hosts):
parts = urlparse(url)
return parts.scheme in ("http", "https") and parts.hostname in allowed_hosts

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

def links_in(html, base_url, allowed_hosts):
soup = BeautifulSoup(html, "html.parser")
found = set()
for anchor in soup.find_all("a", href=True):
href = anchor.get("href", "").strip()
if not href:
continue
absolute = clean_url(urljoin(base_url, href))
if in_scope(absolute, allowed_hosts):
found.add(absolute)
return found

def should_render(html):
# Heuristic only; tailor this to the site and the content you need.
soup = BeautifulSoup(html, "html.parser")
text = " ".join(soup.stripped_strings)
app_shell = bool(soup.select_one("#root, #app, [data-reactroot]"))
return app_shell and len(text) < 300

def browser_html(browser, url):
page = browser.new_page()
try:
page.goto(url, wait_until="load", timeout=BROWSER_TIMEOUT_MS)
if READY_SELECTOR:
page.locator(READY_SELECTOR).first.wait_for(timeout=BROWSER_TIMEOUT_MS)
return page.content(), page.url
finally:
page.close()

def main():
allowed_hosts = {urlparse(seed).hostname for seed in SEEDS}
queue = deque((clean_url(seed), 0) for seed in SEEDS)
seen = set()

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
try:
while queue and len(seen) < MAX_PAGES:
url, depth = queue.popleft()
if url in seen or not in_scope(url, allowed_hosts):
continue
seen.add(url)
try:
response = session.get(url, timeout=HTTP_TIMEOUT, allow_redirects=True)
final_url = clean_url(response.url)
print(f"HTTP {response.status_code} {url} -> {final_url}")
if not response.ok:
continue
html = response.text
if should_render(html):
try:
html, rendered_url = browser_html(browser, final_url)
final_url = clean_url(rendered_url)
print(f"RENDERED {final_url}")
except PlaywrightTimeoutError:
print(f"TIMEOUT rendering {final_url}")
continue
if depth >= MAX_DEPTH:
continue
for link in sorted(links_in(html, final_url, allowed_hosts)):
if link not in seen:
queue.append((link, depth + 1))
except requests.RequestException as exc:
print(f"REQUEST ERROR {url}: {exc}")
finally:
browser.close()

if __name__ == "__main__":
main()

Run it with python crawl_js.py. The should_render heuristic is a starting point, not a general detector: an empty-looking page can be intentional, and a populated response can still have links added later. For dependable coverage, compare initial and rendered DOMs on representative pages or render pages according to a deliberate policy. Add a delay between requests and a durable queue and log if the crawl is more than a small local job.

How link extraction and queueing work

Extract actual, resolvable links

Google’s link guidance recommends crawlable anchors with a resolvable href. JavaScript can create or insert such anchors; an event handler without a real href, a styled non-anchor element, or an anchor whose destination is not a usable URL is not an equivalent navigation target. Google also recommends History API URLs rather than using hash fragments as separate content routes. The crawler above deliberately ignores fragments because they are not sent in HTTP requests; applications that use client-side fragment routing need a separate, explicit policy.

Resolve against redirects, then filter

Relative links should be resolved against the page’s final URL after redirects, not blindly against the seed URL. The sample does this with urljoin, removes fragments, and limits links to the seed hostname. Expand the scope only if the job requires it. URL normalization is risky if done too aggressively: query parameters, case, trailing slashes and path encodings can affect what a server returns. Deduplicate conservatively and preserve distinctions that may represent different pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect both versions when useful

Google says it can discover links before and after rendering, and that links in the initial response may be discovered faster. A crawler can use the same practical distinction: parse response HTML first, then parse rendered DOM when needed, merging and deduplicating discovered URLs. Rendering every page may improve visibility into script-generated links but consumes more time and resources than parsing the response alone.

Choose a rendering strategy

Approach Use it when Trade-off
HTTP request and HTML parse Initial response contains the target text and links. Less execution overhead; misses content inserted only after JavaScript runs.
Browser rendering Scripts are needed to expose the content or navigation. More resource-intensive and sensitive to readiness conditions, browser binaries and page behavior.
Render selectively You need broad coverage but only some pages are client-rendered. Requires a useful detection rule or site-specific knowledge; a weak heuristic can misclassify pages.

Playwright documents browser launch and page navigation for Chromium, Firefox and WebKit. Choose and document an engine, install its browser binary, and test the target site with the readiness logic you intend to use. A successful Playwright load is evidence about that browser session, not proof that a search engine rendered or indexed the same URL.

When you control the site, prefer renderable architecture

Google currently recommends server-side rendering, static rendering or hydration over dynamic rendering as a long-term approach. Dynamic rendering—serving a different pre-rendered version to bots—was described by Google as a workaround, not a long-term solution, and adds complexity and resource requirements. If you maintain the site, make important content and navigation available in a robust form, and keep crawler-facing content substantially similar to what users receive.

Google-specific considerations

These points describe Google’s documented behavior, not universal rules for every crawler:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Google’s JavaScript SEO guide says blocked pages or JavaScript resources are not rendered by Google Search. Check robots.txt and resource access when diagnosing a Google rendering issue.
  • Google documents a rendering queue and says rendering can take longer than a few seconds; do not treat an immediate absence from search as proof that a page cannot render.
  • Indexing directives such as noindex can affect rendering. Do not assume that a client-side change can always repair an initial response that tells Google not to index the page.
  • Discovering a URL, rendering it and indexing it are distinct stages. A crawler’s output cannot establish Google’s indexing decision.

Troubleshooting common failures

Symptom Likely cause What to check or change
HTTP response has a shell but little useful text Content is populated by client-side JavaScript. Render a representative page and inspect its DOM; replace the generic heuristic with a site-specific readiness selector where possible.
Browser times out waiting for load Page resources continue loading, a request hangs, or the chosen wait event does not fit the site. Use a bounded timeout and an appropriate readiness signal, such as a content selector. Do not wait indefinitely for network idle on sites with persistent connections.
Rendered page is blank or incomplete Script errors, blocked resources, consent gates, bot checks or an application error may prevent expected content from appearing. Record browser console errors and failed requests, inspect the final URL and status, and verify manually in the chosen engine. Do not try to bypass access controls.
Links visible to a person are not queued They may be event-only controls, lack a usable href, be outside the allowed scope, or be added only after an interaction. Inspect the actual rendered anchor markup and destination. Add interaction handling only when the crawl’s legitimate purpose requires it and it is safe and permitted.
Many duplicate URLs appear Tracking parameters, alternate paths or redirect variants may identify equivalent content. Log canonical hints and redirect destinations; define site-specific normalization rules only after confirming which URL variants are equivalent.
Request succeeds but browser capture fails Browser binary/package mismatch, crash, or page-specific failure. Reinstall the Playwright browser binary for the installed package version, capture the exception by URL, and retry only under a bounded retry policy.
Google does not show a discovered page in results Discovery alone does not guarantee rendering or indexing; directives, blocked resources and Google’s scheduling may be relevant. Inspect the page’s directives and accessible resources, then use Google’s own site-owner tools and documentation for the specific URL.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability and cost

HTTP fetching plus parsing is generally cheaper in execution resources than opening a browser, but the documentation cited here gives no benchmark or universal speed ratio. Reuse an HTTP session and browser process, bound concurrency, cap page counts and response sizes, and avoid rendering when the initial response already serves the task. A crawler should distinguish network errors, non-success statuses, browser failures, timeouts, empty rendered content and pages with no links rather than treating all of them as the same result.

For repeatable jobs, persist the queue and visited set, retain status/final-URL/error records, set a per-domain request rate, and make retries limited and observable. A browser can expose links hidden behind JavaScript, but it cannot guarantee those links are accessible to every user agent or search engine.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a link-following crawler: it captures a supplied URL, so your crawler still needs to discover and queue URLs. It can be useful when your immediate task is to render and inspect an individual page without managing a local browser. Its pre-capture cleanup accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. See ScreenshotNeo and the API documentation.

For example, this cURL request captures one page as WebP; it does not extract or follow links:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same one-request pattern in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Or in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does rendering a page mean it has been indexed?

No. Rendering and indexing are separate operations, and neither a local browser result nor a discovered link establishes a search engine’s indexing decision.

Can JavaScript create links that a crawler can follow?

Yes, when it creates ordinary anchors with usable href values. Event-only controls and fake links are not dependable substitutes for that markup.

Which browser engine should I use?

Choose an engine supported by your automation library, test it against the target pages, and document the engine and readiness condition. Results from one engine do not guarantee identical behavior elsewhere.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.