For a TypeScript scraper, start with fetch and an HTML parser such as Cheerio when the data is already in the server’s HTML. Use Playwright when the page needs JavaScript, browser interaction, or browser state to reveal the data. In either case, wait for the specific content you need—not merely the browser’s load event—then validate and save the extracted records with enough logging to diagnose failures.
Choose the right approach
The simplest scraper is usually the most reliable one that can collect the required fields. A browser is not automatically better: it adds startup time, resource use, and more ways for a run to fail. But parsing the initial HTML will not work if the page only inserts its content after JavaScript runs.
| Situation | Use | Why |
|---|---|---|
| Server-rendered HTML and a small collection | Node.js fetch or Axios plus Cheerio |
Request the document and parse its HTML directly, without launching a browser. |
| Content appears after JavaScript, interaction, or browser navigation | Playwright | It runs a browser and provides navigation, locator, and page-event APIs. |
| You need to diagnose redirects or network failures | Playwright request lifecycle events | Events expose request, response, completion, and failure behavior. |
| You are managing many URLs, retries, queues, or proxies | Crawlee or an equivalent crawler framework | Framework-level orchestration is a better fit than hand-building a crawl manager. |
Make the choice by inspecting whether the fields you need are present in the returned HTML. If they are, use a direct request first. If they are not, determine whether JavaScript, a click, a browser session, or a later network response supplies them; then use Playwright and wait for that page-specific signal.
Plan the data and access before coding
- Define the output schema. List each field, its expected type, and whether it may be absent. This gives your selectors and validation code a concrete target.
- Check the site’s terms, available API, and access rules. Review the target’s terms and
/robots.txt, and use permitted access patterns with a conservative request rate. A robots file is normally at the site root and communicates crawler access rules; it is not universal legal permission to collect every public page. The MDN robots.txt guide explains its role, while RFC 9309 states that the rules must be accessible in a file named/robots.txtat the service’s top-level path. - Start with a direct request. If the response contains the fields, keep the scraper lightweight and parse that response.
- Escalate only when necessary. Use a browser when rendered content or an interaction is required. Treat the target site, its terms, the collection purpose, and the relevant geography as factors that may change the access decision.
Robots rules are an access signal, not a search-removal mechanism. Google explains that robots.txt tells search crawlers which URLs they may access; a blocked URL can still be discovered or indexed. Site owners seeking search exclusion need a different mechanism, such as noindex, authentication, or removal. See Google Search Central’s robots.txt guide. Scrapers should also respect applicable terms, privacy obligations, copyright rules, authentication boundaries, and rate limits.
#1 Best Overall
Scrape server-rendered HTML with TypeScript
For a static page, request HTML and parse it without a browser. This example uses Node’s built-in fetch and Cheerio. It checks the HTTP response before extracting records, handles missing fields explicitly, and emits JSON. Install the parser with npm install cheerio; run the TypeScript file using your project’s TypeScript runtime or compile it for Node.js.
import * as cheerio from "cheerio";
type Product = {
name: string;
priceText: string | null;
sourceUrl: string;
retrievedAt: string;
};
async function scrapeProducts(pageUrl: string): Promise<Product[]> {
const response = await fetch(pageUrl, {
headers: { "user-agent": "ExampleResearchBot/1.0" },
signal: AbortSignal.timeout(20_000),
});
if (!response.ok) {
throw new Error(`HTTP ${response.status} for ${pageUrl}`);
}
const contentType = response.headers.get("content-type") ?? "";
if (!contentType.includes("text/html")) {
throw new Error(`Expected HTML, received ${contentType || "unknown type"}`);
}
const html = await response.text();
const $ = cheerio.load(html);
const retrievedAt = new Date().toISOString();
const products: Product[] = [];
$(".product-card").each((_index, element) => {
const card = $(element);
const name = card.find(".product-title").text().trim();
if (!name) return;
const priceText = card.find(".price").first().text().trim() || null;
products.push({ name, priceText, sourceUrl: pageUrl, retrievedAt });
});
return products;
}
const url = process.argv[2];
if (!url) throw new Error("Usage: scraper <page-url>");
scrapeProducts(url)
.then((products) => process.stdout.write(`${JSON.stringify(products, null, 2)}n`))
.catch((error: unknown) => {
console.error(error instanceof Error ? error.message : error);
process.exitCode = 1;
});
Replace .product-card, .product-title, and .price with selectors verified against the target page. Keep selectors narrow: broad selectors can silently match navigation labels, recommendations, or unrelated content. The code deliberately records a source URL and retrieval time so downstream users can trace a record to its input.
When this approach is not enough
Cheerio parses the HTML you give it; it does not run page JavaScript or reproduce a user session. If the initial response lacks the data, check the page’s behavior in a browser and move to Playwright rather than trying to fix missing content with increasingly fragile text matching.
Scrape JavaScript-rendered pages with Playwright
Install Playwright and its browser as part of project setup, then use a locator that represents the content you need. The example waits for a product card to become visible, extracts text from the matching elements, and checks the final HTTP status. It does not assume that a navigation event means the page’s data is ready.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
import { chromium } from "playwright";
type Product = {
name: string;
priceText: string | null;
sourceUrl: string;
retrievedAt: string;
};
async function scrapeRenderedProducts(pageUrl: string): Promise<Product[]> {
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
const response = await page.goto(pageUrl, {
waitUntil: "domcontentloaded",
timeout: 30_000,
});
if (!response) throw new Error(`No main-document response for ${pageUrl}`);
if (!response.ok()) {
throw new Error(`HTTP ${response.status()} for ${pageUrl}`);
}
// Use a condition tied to the data this scraper needs.
await page.locator(".product-card").first().waitFor({
state: "visible",
timeout: 15_000,
});
const retrievedAt = new Date().toISOString();
const products = await page.locator(".product-card").evaluateAll(
(cards): Product[] =>
cards.flatMap((card) => {
const name = card.querySelector(".product-title")?.textContent?.trim();
if (!name) return [];
const priceText =
card.querySelector(".price")?.textContent?.trim() || null;
return [{ name, priceText, sourceUrl: location.href, retrievedAt }];
}),
);
return products;
} finally {
await browser.close();
}
}
const url = process.argv[2];
if (!url) throw new Error("Usage: scraper <page-url>");
scrapeRenderedProducts(url)
.then((products) => process.stdout.write(`${JSON.stringify(products, null, 2)}n`))
.catch((error: unknown) => {
console.error(error instanceof Error ? error.message : error);
process.exitCode = 1;
});
Install with npm install playwright and install the browser required by the project with npx playwright install chromium. Check Playwright’s Page API, navigation guide, and network documentation for the current API details. The callback annotation in the example makes the extracted result type explicit; update the selectors and type to match the actual target.
Wait for the data, not just the page
A page can continue fetching and rendering information after the browser fires load. Playwright’s navigation guide distinguishes navigation commitment, domcontentloaded, and load, and notes that modern pages may continue network activity afterward. Choose the least costly signal that proves the field is ready:
- Known element: wait for the relevant locator to become visible or attached, as in the example.
- Known response: if a specific request supplies the data, wait for that response and then inspect the page or response body.
- Short fixed delay: use only when the site offers no stable signal; fixed delays are brittle because the same page can load at different speeds.
- Network idle: use cautiously. Analytics, polling, and other persistent requests can make network activity a poor proxy for data readiness.
Do not treat “the page loaded” as proof that the selector matched, the data is complete, or the content is valid. A wait timeout should be an explicit failed or incomplete scrape, not a reason to persist an empty record as if it were a successful result.
Diagnose requests, redirects, and failed loads
While developing a browser scraper, subscribe to Playwright’s request, response, requestfinished, and requestfailed events. These help distinguish navigation from failed resources and reveal where a page’s data request goes wrong. A request can complete at the HTTP layer with a 404 or 503 response, so a requestfinished event alone does not mean the response succeeded. Check status codes in your logic.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
page.on("request", (request) => {
console.log("request", request.method(), request.url());
});
page.on("response", (response) => {
console.log("response", response.status(), response.url());
});
page.on("requestfinished", (request) => {
console.log("finished", request.url());
});
page.on("requestfailed", (request) => {
console.error("failed", request.url(), request.failure()?.errorText);
});
For redirects, inspect a request’s redirectedFrom() and redirectedTo() chain. This can expose a target that sends the browser to a login, consent, or error page instead of the expected content. Playwright also supports custom selector engines and documents content-script isolation as a safer choice where page JavaScript might tamper with the page context. These are advanced options, not defaults for a first scraper; begin with built-in locators.
Make a scraper reliable in production
A script that works once is not necessarily a dependable data pipeline. Keep discovery, retrieval, extraction, validation, and persistence as separate steps so that a selector change does not silently corrupt stored output.
- Validate every record. Require essential fields, normalize whitespace and formats, and represent missing optional fields deliberately.
- Detect schema drift. Track empty results, unexpected field counts, and selector failures. A site redesign should be visible rather than quietly producing incomplete data.
- Bound retries and concurrency. Retry transient failures with backoff, stop after a defined limit, and avoid sending unnecessary load to the target.
- Checkpoint and deduplicate. Persist progress for long runs and define a stable key for records so a retry does not create duplicates.
- Cache where permitted. Reuse immutable responses when access rules allow it, and retain retrieval timestamps so freshness is clear.
- Log useful diagnostics. Record request URL, status, retry count, and parser error. Avoid collecting personal information that is not needed for the task.
- Keep provenance. Store source URL, retrieval time, parser version, and selector version alongside output records.
For sustained or larger crawls, consider a crawler framework such as Crawlee or an equivalent that provides queues, retries, and proxy controls. Evaluate its current package and commercial details independently; the framework’s orchestration does not remove the need to follow each target’s access rules.
Troubleshoot common failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| HTTP 404, 403, or 503 | The URL is unavailable, access is refused, or the server is temporarily failing. | Log the status and final URL; check the target’s terms and access rules. Retry only plausibly transient failures, with a limit and backoff. |
| HTML request succeeds but fields are empty | Selectors do not match, or content is inserted by JavaScript. | Inspect the returned HTML and a representative page variant. If the data only appears after rendering or interaction, use Playwright. |
| Playwright navigation succeeds but the locator times out | The page has not exposed the target data, the selector changed, or the scraper landed on a different page. | Inspect the final URL, status, and page content; wait for the actual data condition rather than adding an arbitrary long delay. |
| A request reports failure or unexpected redirect | A network request failed, or the site redirected to another state. | Use request lifecycle logs and inspect the redirect chain to identify the failing URL and response. |
| Scraper returns too few or malformed records | Page variants or a site redesign changed the markup, or optional fields were treated as mandatory. | Validate records, test representative variants, and separate required from optional fields. |
| Runs become slow or unreliable at scale | Too much concurrency, unbounded retries, or crawl management built into a single script. | Bound concurrency, use backoff and checkpoints, and consider a crawler framework for queues and retry management. |
Or skip the browser setup
If the job is to capture a visual record of a page rather than extract structured fields, ScreenshotNeo offers a screenshot API and MCP server. It is not a replacement for a TypeScript parser when you need data fields, but it can return a page image or PDF without building a browser capture flow yourself. The API accepts one GET request; see the ScreenshotNeo API documentation.
Recommended Free Tools
Quick Recap
curl -G "https://api.screenshotneo.com/v1/shot"
-d access_key=YOUR_API_KEY
--data-urlencode url=https://stripe.com
-o shot.webp
Before capture, ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




