Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use ordinary HTTP requests for pages whose returned HTML already contains the content and links you need. Use a browser renderer when JavaScript adds that content or navigation after the initial response. A practical crawler can inspect both versions, extract real links from <a href> elements, resolve them against the final URL, filter and deduplicate them, then add in-scope URLs to a controlled queue.
Crawling, rendering and indexing are separate operations. A crawler finding or rendering a URL does not prove that a search engine will index it. The example below builds a small, polite Python crawler for inspection—not a search-engine replica.
What crawling a JavaScript site involves
A JavaScript website can return a complete HTML document, a partially populated page, or an application shell that relies on scripts to insert its main content and links. Fetching the response and executing the page are different operations:
Free tools Windows power users keep installed
One-click scans. No signup required.
- HTTP fetch: retrieve the response and parse its HTML without running page JavaScript. This is simpler and usually uses fewer resources, and it works when the content and links are already in the response.
- Browser rendering: open the page in a browser engine, let its scripts run, then inspect the rendered DOM. Use this when the response lacks material content or links that appear after execution.
Google Search Central describes its own crawl, render and index stages, including a headless Chromium rendering stage. That is useful context, but it is not a guarantee that your crawler—or another search engine—will behave the same way. Google’s JavaScript SEO documentation also says rendering may take longer than a few seconds, resources must be available, and pages can wait in a rendering queue. An independent crawler should choose its own limits and readiness rules.
#1 Best Overall
Set scope and limits before crawling
Start with one or more seed URLs, then define what the crawler is allowed to visit. These are engineering controls, not requirements attributed here to Google.
- Scope: allowed hostnames or URL prefixes. Decide whether subdomains, query strings and alternate schemes belong.
- Depth and volume: maximum link depth and maximum number of pages. A depth of zero can mean seeds only; each newly followed link adds one level.
- Resource limits: request and browser timeouts, maximum response size, and a delay or concurrency limit appropriate to the site.
- Policy: identify yourself appropriately, review the site’s terms and robots.txt, and do not bypass access controls. Robots.txt is a crawl convention, not authorization to access restricted content.
- Output: record each URL, final URL after redirects, response status, whether rendering was used, and any failure reason.
A Python crawler with HTTP-first rendering fallback
This example uses requests and Beautiful Soup for the initial response, then Playwright when that response appears to need JavaScript. It follows only links with an actual href in an anchor element and stays on the seed hostname. The readiness check is intentionally modest: it waits for the page load event, then inspects the DOM. Real sites may need a site-specific selector or other signal.
Install dependencies
Use Python 3.9 or later. Install Playwright’s Python package and the browser binary it needs; keep the installed browser aligned with the Playwright package version.
python -m pip install requests beautifulsoup4 playwright
python -m playwright install chromium
Save and run the crawler
Save as crawl_js.py. Change SEEDS and, where appropriate, MAX_PAGES, MAX_DEPTH and READY_SELECTOR. The sample uses one hostname as its scope and does not try to normalize away meaningful query parameters.
from collections import deque
from urllib.parse import urldefrag, urljoin, urlparse
import requests
from bs4 import BeautifulSoup
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallSEEDS = ["https://example.com/"]
MAX_PAGES = 100
MAX_DEPTH = 2
HTTP_TIMEOUT = 20
BROWSER_TIMEOUT_MS = 30000
READY_SELECTOR = None # Example: "main article"
session = requests.Session()
session.headers.update({"User-Agent": "ExampleResearchCrawler/1.0 (contact: [email protected])"})
def clean_url(url):
# Remove fragments: they are not sent as part of an HTTP request.
return urldefrag(url)[0]
def in_scope(url, allowed_hosts):
parts = urlparse(url)
return parts.scheme in ("http", "https") and parts.hostname in allowed_hosts
def links_in(html, base_url, allowed_hosts):
soup = BeautifulSoup(html, "html.parser")
found = set()
for anchor in soup.find_all("a", href=True):
href = anchor.get("href", "").strip()
if not href:
continue
absolute = clean_url(urljoin(base_url, href))
if in_scope(absolute, allowed_hosts):
found.add(absolute)
return found
Rank #3
def should_render(html):
# Heuristic only; tailor this to the site and the content you need.
soup = BeautifulSoup(html, "html.parser")
text = " ".join(soup.stripped_strings)
app_shell = bool(soup.select_one("#root, #app, [data-reactroot]"))
return app_shell and len(text) < 300
def browser_html(browser, url):
page = browser.new_page()
try:
page.goto(url, wait_until="load", timeout=BROWSER_TIMEOUT_MS)
if READY_SELECTOR:
page.locator(READY_SELECTOR).first.wait_for(timeout=BROWSER_TIMEOUT_MS)
return page.content(), page.url
finally:
page.close()
def main():
allowed_hosts = {urlparse(seed).hostname for seed in SEEDS}
queue = deque((clean_url(seed), 0) for seed in SEEDS)
seen = set()
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
try:
while queue and len(seen) < MAX_PAGES:
url, depth = queue.popleft()
if url in seen or not in_scope(url, allowed_hosts):
continue
seen.add(url)
try:
response = session.get(url, timeout=HTTP_TIMEOUT, allow_redirects=True)
final_url = clean_url(response.url)
print(f"HTTP {response.status_code} {url} -> {final_url}")
if not response.ok:
continue
html = response.text
if should_render(html):
try:
html, rendered_url = browser_html(browser, final_url)
final_url = clean_url(rendered_url)
print(f"RENDERED {final_url}")
except PlaywrightTimeoutError:
print(f"TIMEOUT rendering {final_url}")
continue
if depth >= MAX_DEPTH:
continue
for link in sorted(links_in(html, final_url, allowed_hosts)):
if link not in seen:
queue.append((link, depth + 1))
except requests.RequestException as exc:
print(f"REQUEST ERROR {url}: {exc}")
finally:
browser.close()
if __name__ == "__main__":
main()
Run it with python crawl_js.py. The should_render heuristic is a starting point, not a general detector: an empty-looking page can be intentional, and a populated response can still have links added later. For dependable coverage, compare initial and rendered DOMs on representative pages or render pages according to a deliberate policy. Add a delay between requests and a durable queue and log if the crawl is more than a small local job.
How link extraction and queueing work
Extract actual, resolvable links
Google’s link guidance recommends crawlable anchors with a resolvable href. JavaScript can create or insert such anchors; an event handler without a real href, a styled non-anchor element, or an anchor whose destination is not a usable URL is not an equivalent navigation target. Google also recommends History API URLs rather than using hash fragments as separate content routes. The crawler above deliberately ignores fragments because they are not sent in HTTP requests; applications that use client-side fragment routing need a separate, explicit policy.
Resolve against redirects, then filter
Relative links should be resolved against the page’s final URL after redirects, not blindly against the seed URL. The sample does this with urljoin, removes fragments, and limits links to the seed hostname. Expand the scope only if the job requires it. URL normalization is risky if done too aggressively: query parameters, case, trailing slashes and path encodings can affect what a server returns. Deduplicate conservatively and preserve distinctions that may represent different pages.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Inspect both versions when useful
Google says it can discover links before and after rendering, and that links in the initial response may be discovered faster. A crawler can use the same practical distinction: parse response HTML first, then parse rendered DOM when needed, merging and deduplicating discovered URLs. Rendering every page may improve visibility into script-generated links but consumes more time and resources than parsing the response alone.
Choose a rendering strategy
| Approach | Use it when | Trade-off |
|---|---|---|
| HTTP request and HTML parse | Initial response contains the target text and links. | Less execution overhead; misses content inserted only after JavaScript runs. |
| Browser rendering | Scripts are needed to expose the content or navigation. | More resource-intensive and sensitive to readiness conditions, browser binaries and page behavior. |
| Render selectively | You need broad coverage but only some pages are client-rendered. | Requires a useful detection rule or site-specific knowledge; a weak heuristic can misclassify pages. |
Playwright documents browser launch and page navigation for Chromium, Firefox and WebKit. Choose and document an engine, install its browser binary, and test the target site with the readiness logic you intend to use. A successful Playwright load is evidence about that browser session, not proof that a search engine rendered or indexed the same URL.
When you control the site, prefer renderable architecture
Google currently recommends server-side rendering, static rendering or hydration over dynamic rendering as a long-term approach. Dynamic rendering—serving a different pre-rendered version to bots—was described by Google as a workaround, not a long-term solution, and adds complexity and resource requirements. If you maintain the site, make important content and navigation available in a robust form, and keep crawler-facing content substantially similar to what users receive.
Google-specific considerations
These points describe Google’s documented behavior, not universal rules for every crawler:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Google’s JavaScript SEO guide says blocked pages or JavaScript resources are not rendered by Google Search. Check robots.txt and resource access when diagnosing a Google rendering issue.
- Google documents a rendering queue and says rendering can take longer than a few seconds; do not treat an immediate absence from search as proof that a page cannot render.
- Indexing directives such as
noindexcan affect rendering. Do not assume that a client-side change can always repair an initial response that tells Google not to index the page. - Discovering a URL, rendering it and indexing it are distinct stages. A crawler’s output cannot establish Google’s indexing decision.
Troubleshooting common failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| HTTP response has a shell but little useful text | Content is populated by client-side JavaScript. | Render a representative page and inspect its DOM; replace the generic heuristic with a site-specific readiness selector where possible. |
| Browser times out waiting for load | Page resources continue loading, a request hangs, or the chosen wait event does not fit the site. | Use a bounded timeout and an appropriate readiness signal, such as a content selector. Do not wait indefinitely for network idle on sites with persistent connections. |
| Rendered page is blank or incomplete | Script errors, blocked resources, consent gates, bot checks or an application error may prevent expected content from appearing. | Record browser console errors and failed requests, inspect the final URL and status, and verify manually in the chosen engine. Do not try to bypass access controls. |
| Links visible to a person are not queued | They may be event-only controls, lack a usable href, be outside the allowed scope, or be added only after an interaction. |
Inspect the actual rendered anchor markup and destination. Add interaction handling only when the crawl’s legitimate purpose requires it and it is safe and permitted. |
| Many duplicate URLs appear | Tracking parameters, alternate paths or redirect variants may identify equivalent content. | Log canonical hints and redirect destinations; define site-specific normalization rules only after confirming which URL variants are equivalent. |
| Request succeeds but browser capture fails | Browser binary/package mismatch, crash, or page-specific failure. | Reinstall the Playwright browser binary for the installed package version, capture the exception by URL, and retry only under a bounded retry policy. |
| Google does not show a discovered page in results | Discovery alone does not guarantee rendering or indexing; directives, blocked resources and Google’s scheduling may be relevant. | Inspect the page’s directives and accessible resources, then use Google’s own site-owner tools and documentation for the specific URL. |
Performance, reliability and cost
HTTP fetching plus parsing is generally cheaper in execution resources than opening a browser, but the documentation cited here gives no benchmark or universal speed ratio. Reuse an HTTP session and browser process, bound concurrency, cap page counts and response sizes, and avoid rendering when the initial response already serves the task. A crawler should distinguish network errors, non-success statuses, browser failures, timeouts, empty rendered content and pages with no links rather than treating all of them as the same result.
Best Value
For repeatable jobs, persist the queue and visited set, retain status/final-URL/error records, set a per-domain request rate, and make retries limited and observable. A browser can expose links hidden behind JavaScript, but it cannot guarantee those links are accessible to every user agent or search engine.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a link-following crawler: it captures a supplied URL, so your crawler still needs to discover and queue URLs. It can be useful when your immediate task is to render and inspect an individual page without managing a local browser. Its pre-capture cleanup accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. See ScreenshotNeo and the API documentation.
For example, this cURL request captures one page as WebP; it does not extract or follow links:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same one-request pattern in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Or in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.
Recommended Free Tools
Frequently Asked Questions
Does rendering a page mean it has been indexed?
No. Rendering and indexing are separate operations, and neither a local browser result nor a discovered link establishes a search engine’s indexing decision.
Can JavaScript create links that a crawler can follow?
Yes, when it creates ordinary anchors with usable href values. Event-only controls and fake links are not dependable substitutes for that markup.
Which browser engine should I use?
Choose an engine supported by your automation library, test it against the target pages, and document the engine and readiness condition. Results from one engine do not guarantee identical behavior elsewhere.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

