Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Playwright when a page’s data depends on browser rendering, interaction, or session state; use a documented API or direct HTTP request when that already provides the data you need. A practical scraper opens the page, waits for a meaningful signal, extracts stable fields, handles pagination and failures explicitly, and closes its browser resources. This guide uses Node.js and JavaScript, with runnable examples for inspecting page traffic and for simpler HTTP requests.

Choose the simplest suitable access method

Playwright is a browser automation library that can also be used for web scraping. It can load pages as a browser does, interact with controls, and expose network activity. That does not mean every scraper should launch a browser for every URL.

  • Use a documented API if it provides the fields you need and your use is authorized. An API is often a more direct interface than extracting data from rendered markup.
  • Use an HTTP request when the required content is available in a direct response and does not depend on browser execution or interaction.
  • Use browser automation when JavaScript rendering, a user-facing interaction, or browser session state is necessary to reach the content.

Playwright’s APIRequestContext can make HTTP requests. Its network APIs can also show requests made by a page, including fetch and XHR. Observing a request can help explain how a page obtains data, but it is not permission to bypass authentication or access controls. Check the site’s rules and use an authorized interface.

Compare options by whether the source offers an authorized API, whether rendering or interaction is genuinely needed, and what rate limits, data-handling duties, and terms apply. Direct requests often involve less browser machinery, but actual speed and completeness depend on the site and workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Playwright and its browser binaries

For a Node.js project, install Playwright and the browser binary it will launch. Each Playwright version expects compatible browser binaries; install or update them when changing Playwright versions. The official browser installation guide includes operating-system dependencies and platform-specific details.

  1. Create a project and install Playwright:

    npm init -y
    npm install playwright
    npx playwright install chromium

  2. Save the scraper below as scrape.js, then run it with:

    node scrape.js https://example.com

  3. If launching the browser reports that an executable is missing, run npx playwright install chromium again for the installed Playwright version. On Linux or other systems with missing libraries, use the installation guide’s instructions for your operating system.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The example uses Chromium explicitly so the browser choice is clear. It is not a claim that the code was tested against a particular Playwright release or operating system.

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Navigate, wait for a page signal, and extract fields

Save this as scrape.js. It opens the supplied URL, waits for a page-specific locator, extracts article headings and links, checks the final response status, and closes the context and browser even if navigation or extraction fails.

const { chromium } = require('playwright');

async function scrape(url) {
  const browser = await chromium.launch();
  const context = await browser.newContext();
  try {
    const page = await context.newPage();
    const response = await page.goto(url, {
      waitUntil: 'domcontentloaded',
      timeout: 30_000,
    });

    if (response && !response.ok()) {
      throw new Error(`Navigation returned HTTP ${response.status()}`);
    }

    // Replace this with a stable signal that identifies the data you need.
    const heading = page.getByRole('heading', { level: 1 });
    await heading.waitFor({ state: 'visible', timeout: 15_000 });

    const result = await page.evaluate(() => ({
      title: document.querySelector('h1')?.textContent?.trim() ?? null,
      links: Array.from(document.querySelectorAll('a[href]')).map((a) => ({
        text: a.textContent?.trim() ?? '',
        href: a.href,
      })),
    }));

    if (!result.title) {
      throw new Error('Expected an h1, but no heading text was found');
    }
    return result;
  } finally {
    await context.close();
    await browser.close();
  }
}

const url = process.argv[2];
if (!url) {
  console.error('Usage: node scrape.js <url>');
  process.exitCode = 2;
} else {
  scrape(url)
    .then((data) => console.log(JSON.stringify(data, null, 2)))
    .catch((error) => {
      console.error(`Scrape failed: ${error.message}`);
      process.exitCode = 1;
    });
}

The example’s h1 is only a placeholder signal. Replace it with a locator that identifies the actual result area—for example, a named product heading or a result-list container. If a page can legitimately have no results, treat that state separately rather than treating every missing item as a successful scrape.

Prefer locators tied to meaning

Playwright calls locators its central mechanism for auto-waiting and retrying. Prefer a role, label, visible text, or an explicit stable contract exposed by the page where it fits. A selector such as article h2 may be reasonable if it reflects the page’s stable content structure; a long CSS or XPath path that depends on incidental wrapper elements is more likely to break when the layout changes. See the locator guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use explicit waits for a meaningful condition, such as a result list becoming visible. Avoid using an arbitrary fixed sleep as the default: a delay can be too short on a slow response and waste time when a page is already ready. No single wait condition is right for every site; identify the signal that corresponds to the data your job requires.

Keep navigation and extraction outcomes distinct

page.goto() returning a response does not by itself mean the desired data loaded. The example checks an HTTP error status and then checks for an expected page element. For a production job, record these as distinct outcomes: navigation error, non-success HTTP status, missing expected content, valid empty state, and successful extraction.

Handle pagination and session state

Follow the site’s actual pagination contract

For page-number navigation, identify the real next-page control and stop when it is absent, disabled, or the page signals an end. For cursor-based data, follow the returned cursor and stop when the documented end condition appears. Keep a visited-page or cursor check where repeated links could otherwise create a loop. Do not assume every site uses a query parameter named page.

Set a deliberate maximum page or item limit for each run, and retain enough information to resume or diagnose a partial run. The appropriate limit and request rate depend on the service’s rules and your permission; there is no universal safe rate established here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Isolate authenticated sessions

Only automate accounts and content you are authorized to access. A Playwright browser context isolates cookies and other storage, which makes separate contexts useful for separate sessions or jobs. When using browser.newContext(), close the context before closing the browser, as in the example. See the browser context guide and Browser API.

A context is not a substitute for authorization, and you should handle credentials and any collected personal data according to the applicable rules. Keep session reuse deliberate: reusing a context can preserve state, while separate contexts start isolated sessions.

Inspect page network traffic when useful

If rendered content is missing or changes after an interaction, observing the page’s requests can help identify when the relevant response arrives. This example logs fetch and XHR responses with their status and URL; use it while investigating a page you are permitted to access.

Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const context = await browser.newContext();
  try {
    const page = await context.newPage();
    page.on('response', async (response) => {
      const request = response.request();
      if (request.resourceType() === 'xhr' || request.resourceType() === 'fetch') {
        console.log(response.status(), response.url());
      }
    });
    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
    await page.getByRole('heading', { level: 1 }).waitFor({ state: 'visible' });
  } finally {
    await context.close();
    await browser.close();
  }
})().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

Playwright can wait for responses and intercept requests through its network APIs. Service workers can make requests invisible to built-in page or context routing; for interception use cases, Playwright’s documentation recommends blocking service workers. Treat this as a debugging or authorized automation capability, not a way to evade a site’s controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use HTTP directly when browser rendering is unnecessary

Playwright’s APIRequestContext is useful when you need an HTTP request without rendering a page. This example checks the status explicitly and reads the response body as text:

const { request } = require('playwright');

(async () => {
  const client = await request.newContext();
  try {
    const response = await client.get('https://example.com/data');
    if (!response.ok()) {
      throw new Error(`HTTP ${response.status()} for ${response.url()}`);
    }
    console.log(await response.text());
  } finally {
    await client.dispose();
  }
})().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

Replace the URL only with an authorized endpoint. A response that completes can still carry an HTTP error: Playwright’s request documentation notes that statuses such as 404 or 503 are responses, not necessarily thrown request failures. Check status and validate the returned content before treating it as usable data. See APIRequestContext.

Respect crawler rules and site permissions

RFC 9309, the IETF’s Robots Exclusion Protocol (September 2022), describes rules crawlers are requested to honor. It explicitly says: “These rules are not a form of access authorization.” Read a site’s robots.txt as crawler guidance, not as a grant of permission. Check the service’s current terms and policies separately, obtain permission when required, and account for applicable law, privacy, copyright, authentication, and rate limits. The standard is available at RFC 9309.

Troubleshoot common failures

Symptom Likely cause What to do
Browser executable missing The browser binary for the installed Playwright version is not installed or does not match it. Run npx playwright install chromium after installing or updating Playwright; consult the browser guide for platform dependencies.
Navigation times out The page is slow, the selected navigation event is unsuitable, or the service is unavailable. Check that the URL is reachable and authorized, use a suitable navigation condition, then wait for the specific content locator with a separate timeout. Do not mask repeated failures with an unlimited timeout.
HTTP 404 or 503, but no thrown navigation error The server completed an HTTP response with an error status. Inspect response.status() and handle the status as a failed or retryable outcome according to the site’s policy.
Locator timeout or missing fields The selector may be brittle, the content has not appeared, or the page has an empty/error state. Inspect the rendered page, choose a stable semantic locator, and distinguish an expected empty result from missing or changed markup.
Content appears in the browser but not in extracted data The extraction may run before the relevant content is rendered, or the visible data may be loaded after an interaction. Wait for the content-specific signal or perform the required authorized interaction. Inspect network responses if that will clarify the page’s loading behavior.
Requests are not visible to routing handlers A service worker may be handling them. For an interception workflow, review Playwright’s service-worker guidance and whether blocking service workers is appropriate for the page.
Later jobs inherit unexpected cookies or state A browser context was reused. Use a separate context when isolation is needed, and close it when that session is finished.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost decisions

Launching and controlling a browser involves more machinery than issuing a direct HTTP request, so avoid browser rendering when a suitable authorized API or direct response already solves the problem. That is an engineering trade-off, not a guaranteed performance ratio: site behavior, workload, browser configuration, and network conditions matter.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For reliability, use bounded timeouts, explicit content checks, status handling, and cleanup in finally. Keep the outcome of each URL observable so a partial batch does not silently look complete. Avoid retrying indefinitely; decide what failures are transient and set retry limits that respect the service’s rules. Store only the data you need and protect session credentials.

Browser scraping has operational costs in compute, maintenance, and the work required to keep selectors aligned with a changing site. There is no verified universal benchmark or success rate to apply to a particular site. Measure your own authorized workload and compare it with a direct API or request where available.

Or skip the browser setup

If your goal is a screenshot rather than structured data, ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. It is not a replacement for extracting records from a page; it is an alternative when the output you need is a capture.

cURL example, adapted to the target URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Before capture, it can accept cookie or consent banners as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies its page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

Frequently Asked Questions

Can Playwright scrape a site that renders content with JavaScript?

Yes, when the content is available to an authorized browser session, Playwright can load the page and wait for a meaningful rendered-content signal before extraction.

Is robots.txt permission to scrape a website?

No. RFC 9309 says robots rules are not access authorization. Check the site’s terms and applicable permissions separately.

Should I use Playwright or a screenshot API?

Use Playwright for browser interaction and data extraction; use a screenshot API when the desired output is an image or PDF capture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.