Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Run the browser on a remote service and connect to it with a compatible automation client such as Playwright or Puppeteer. That lets JavaScript-heavy pages render and supports navigation and interaction without installing a browser on the machine running your scraper. For a single page that needs no interaction, first check whether a stateless extraction API can return the data directly; it may be simpler than controlling a browser.
This guide walks through the decision, a Playwright connection pattern, and the operational details that matter in a recurring scraping pipeline. Use these techniques only after checking the target site’s rules, your rights to collect the data, and applicable legal requirements.
Decide whether you need a browser
A headless browser is useful when the information appears only after JavaScript runs, or when reaching it requires actions such as clicking, scrolling, or navigating through a workflow. Your automation controls a real browser engine without displaying its window. In a cloud setup, that browser runs on a remote host, while your code sends commands and receives results.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Do not add browser automation by default. If you need one page’s content and a stateless scraping or content-extraction endpoint already returns the fields you need, that endpoint may be a better fit. Browserless documents both REST interfaces for scraping/content extraction and browser sessions for controlling Playwright or Puppeteer; they are distinct interfaces for different jobs (Browserless overview, Browserless BaaS).
#1 Best Overall
- Use a stateless extraction API when the job is a straightforward request for page content and you do not need to operate the browser.
- Use a remote browser session when you need JavaScript rendering, navigation, interaction, or to reuse an existing Playwright/Puppeteer workflow.
- Use a broader scraping platform when the job also needs managed scheduling, storage, or monitoring. Apify documents these as platform features, but that does not establish comparative performance or price (Apify platform documentation).
Choose where the browser runs
| Option | Browser control | Who operates the browser infrastructure | Good fit | Questions to resolve |
|---|---|---|---|---|
| Managed remote browser | Full session control through the provider’s supported client and endpoint | The provider operates the hosted service | Moving an existing Playwright or Puppeteer script off a local machine | Supported protocol, language/client, browser engines and versions, session behavior, concurrency, data handling, and provider terms |
| Self-hosted browser service | Full session control, subject to the service and deployment configuration | Your team deploys and operates it | Teams that need to manage the deployment environment themselves | Deployment, browser updates, resource limits, security, monitoring, and maintenance responsibility |
| Stateless scraping API | Usually a request for a result rather than a controllable browser session | The API provider operates its service | Simple extraction that does not require your code to interact with a page | Whether its output covers the required fields and whether it supports the workflow’s constraints |
Browserless documents both its managed cloud service and Docker self-hosting. A managed endpoint removes the need for you to deploy the browser service yourself; self-hosting transfers deployment and ongoing operations to your team (Browserless overview, Browserless API documentation). There is no universal winner on price, speed, reliability, or privacy: those depend on your workload, configuration, and provider terms, and the cited documentation does not establish comparative benchmarks.
Connect Playwright to a remote browser
The connection URL and protocol are provider-specific. Browserless BaaS v2 documents both Chrome DevTools Protocol (CDP) routes and Playwright-native routes. Choose the client that matches the route: the provider warns that pairing the wrong protocol and client fails. Its BaaS v2 documentation also says Selenium/WebDriver is not supported there, so do not assume a Playwright/Puppeteer endpoint will accept Selenium (Browserless BaaS quickstart).
The following is a client-side pattern for a provider that supplies a Playwright-compatible WebSocket endpoint. Replace the environment variable with the exact endpoint URL and authentication format from that provider’s current setup guide. This example uses Playwright’s native connect method; it is not a universal Browserless URL or a substitute for checking the selected route.
Free tools Windows power users keep installed
One-click scans. No signup required.
- Install Playwright and the project dependencies in your client environment.
- Set the provider’s connection URL as a secret environment variable rather than placing credentials in source control.
- Connect, create a page, navigate, wait for the required content, and close the remote browser in a
finallyblock.
npm install playwright
// scrape.mjs
import { chromium } from 'playwright';
const endpoint = process.env.BROWSER_WS_ENDPOINT;
if (!endpoint) throw new Error('Set BROWSER_WS_ENDPOINT to your provider endpoint');
const browser = await chromium.connect(endpoint);
try {
const page = await browser.newPage();
await page.goto('https://example.com', {
waitUntil: 'domcontentloaded',
timeout: 30_000,
});
// Prefer a page-specific readiness condition when you know the target selector.
await page.locator('h1').waitFor({ state: 'visible', timeout: 15_000 });
const result = {
title: await page.title(),
heading: await page.locator('h1').innerText(),
};
console.log(JSON.stringify(result));
} finally {
await browser.close();
}
Some providers instead document a CDP connection method or require an API token as a query parameter or request option. Follow that provider’s current instructions exactly; authentication placement is not standardized. Browserless documents token use for its requests in its overview (Browserless overview). Keep tokens out of logs, error reports, committed configuration files, and URLs exposed to end users.
Use the right readiness condition
domcontentloaded means the initial document has been parsed, not that a client-rendered application has finished loading its data. Waiting for a relevant selector is usually more precise for a known page. A fixed delay can help with a known transient behavior, but it can also waste time or finish too early. Network-idle conditions can be unsuitable for pages that keep requests open in the background. Choose the condition that matches the page and handle timeout as an expected failure, not proof that the page has no data.
Keep browser engines and versions aligned
Playwright supports Chromium, Firefox, and WebKit, and its installation process obtains browser builds associated with the Playwright version. Its documentation recommends keeping Playwright and its browsers updated (Playwright browser documentation). A local browser binary is not automatically available on a remote service: confirm which engine and versions the provider supports, and whether its endpoint uses a compatible Playwright build.
When updating, treat the client library and remote browser as a compatibility pair. Test the connection and key page interactions in a staging job before rolling an update into a recurring scraper. Playwright also documents a headless-shell installation option; whether that option is relevant depends on how you run browsers, and should not be assumed to describe a hosted provider’s environment (Playwright browser documentation).
Add a proxy only for a concrete routing need
Routing can be necessary because the browser must reach a network through a particular egress path or because of an approved network topology. Playwright’s browser API documents HTTP and SOCKS proxy settings; platform providers such as Apify also document proxy functionality (Playwright Browser API, Apify proxy documentation).
A proxy is a network-routing mechanism, not permission to collect data and not a guarantee that a page will load. It does not resolve questions about a site’s terms, access controls, privacy obligations, or applicable law. Verify authorization and target-site rules first, and use only routing arrangements you are entitled to use.
Turn a one-off script into a data pipeline
For recurring work, browser automation is only one stage. Treat the workflow as a pipeline with explicit inputs, outputs, and failure handling rather than assuming that a successful navigation produced usable data.
- Schedule deliberately: run at a cadence appropriate to the data and the target site’s rules. Avoid uncontrolled retry loops.
- Limit concurrency: cap simultaneous sessions based on provider limits, your workload, and the target system’s acceptable request rate.
- Validate extracted data: check required fields, types, freshness, and obvious empty or malformed results before storing them.
- Store results and provenance: record when a run occurred and enough source context to diagnose changes without retaining unnecessary personal or sensitive data.
- Alert on meaningful failures: distinguish a connection failure, navigation timeout, missing selector, and invalid output so operators can act on the cause.
- Plan session cleanup: close pages and browser connections even on exceptions; follow the provider’s session and concurrency guidance.
Apify documents cloud Actors, storage, schedules, monitoring, and proxies as platform capabilities. Browserless documents browser sessions and multi-page crawl jobs. These are descriptions of product surfaces, not independent evidence about speed, reliability, or comparative suitability (Apify platform documentation, Browserless API documentation).
Security, permission, and data handling
Before collecting anything, review the target site’s terms and access rules, the rights attached to the data, and the laws that apply to your use and location. This guide cannot determine whether a specific scraping activity is authorized. Do not treat browser features, proxies, or a provider’s ability to reach a page as evidence of permission.
Best Value
Minimize the data you collect and retain, protect credentials as secrets, restrict access to stored results, and check the browser provider’s data-handling terms against your requirements. If pages expose personal, confidential, or regulated information, assess whether your collection and processing are appropriate before running the job.
Troubleshoot common connection and scraping failures
| Symptom | Likely cause | What to check |
|---|---|---|
| Connection fails immediately | Malformed endpoint, invalid credentials, unavailable session, or client/protocol mismatch | Copy the provider’s current connection URL and auth configuration; verify that the route is intended for your Playwright or Puppeteer client. |
| Browser connects, then navigation times out | Slow or unreachable page, restrictive network path, or a wait condition that never completes | Check reachability and provider session status; use a page-specific selector where suitable; set a bounded timeout and report the failure distinctly. |
| Page loads but extracted fields are empty | Content is delayed, rendered in a different element, or changed on the target page | Wait for the actual content selector, inspect the returned HTML or page state where permitted, and validate selectors against current page structure. |
| Works locally but fails remotely | Remote engine/version, network access, fonts/resources, or provider-specific settings differ | Confirm the provider’s supported browser and version, compare the selected engine, and review its network and session requirements. |
| Unsupported automation client | The service route does not support that client or protocol | Use a documented compatible route and client. For Browserless BaaS v2, its documentation says Selenium/WebDriver is unsupported. |
| Unexpected access block or challenge | The site is restricting access or requires authorization | Stop and review the site’s rules and access requirements; do not interpret a proxy or automation retry as authorization. |
Or skip the browser setup
If your task is to capture a visual screenshot or PDF rather than extract structured data through browser interaction, ScreenshotNeo offers a one-request screenshot API and an MCP server for AI agents. It is not a replacement for a general-purpose scraping workflow or for a Playwright session you need to control. Its stated differentiators are that it accepts cookie/consent banners and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; failed loads, bot checks/CAPTCHAs, blank pages, and cache hits are not billed, with a verdict and billing headers in each response; and AI clients can use its MCP tools, including take_screenshot, get_page_info, and capture_pdf.
For a one-call capture, create an API key and run this cURL example (replace the target URL as needed):
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemscurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python and Node.js examples, along with the other request options, are in the ScreenshotNeo documentation. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo free.
Frequently Asked Questions
Can I use Selenium with a Browserless BaaS v2 endpoint?
Browserless’s BaaS v2 documentation says Selenium/WebDriver is not supported there. Choose a client and route that the provider documents as compatible.
Does running a browser in the cloud make a scrape legal?
No. Hosting and routing do not establish permission. Check the target site’s rules, data rights, and applicable legal requirements for your specific use.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

