Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Selenium Grid lets your WebDriver client run browser sessions remotely, on one machine or across a pool of machines. Your client code still visits pages and interacts with them; Grid routes each new session to a compatible browser slot. Start with Standalone mode at http://localhost:4444, then add Nodes when you need more machines, browser configurations, or concurrent sessions.
What Selenium Grid does—and what it does not do
Grid is the remote browser execution layer for Selenium. A client sends WebDriver commands to Grid, which places the requested browser session on an available machine. The client remains responsible for navigation, waits, element selection, extraction, and saving the results.
That distinction matters for scraping: Grid is neither a scraping framework nor a source of data. It does not decide which pages to collect or turn page content into structured records. Nor does using a browser or distributing sessions grant permission to access a site.
How Grid 4 routes a browser session
A client asks for a session with capabilities, such as a browser type. Grid matches those requirements to an available slot. A slot is a place where a browser session can run; its capabilities limit which requests it can accept.
#1 Best Overall
- Router: receives client requests.
- New Session Queue: holds new session requests while they await a match.
- Distributor: selects a compatible Node slot for a new session.
- Nodes: run the browser sessions.
- Session Map: tracks session IDs and the Nodes running them.
- Event Bus: carries asynchronous messages among Grid components.
In a larger deployment, the client still addresses Grid’s entry point. It does not need to select a particular Node for each request; the Distributor handles session placement.
Start with Standalone mode
The Selenium getting-started guide lists Java 11 or later, browser software, browser drivers, and the Selenium Server JAR as prerequisites. Selenium Manager can configure drivers when enabled. These details can change with Selenium releases, so match the guide and commands to the version you install.
- Install a supported Java runtime and the browser you intend to use.
- Download the Selenium Server JAR for your chosen release from Selenium’s official Grid getting-started guide.
- Start a single-process Grid:
java -jar selenium-server-<version>.jar standalone. Replace<version>with the downloaded JAR’s version. - Point the WebDriver client to
http://localhost:4444. The default address also serves Grid’s browser UI and status endpoint. - Run a small, authorized collection task and confirm that the requested browser appears in Grid before increasing concurrency.
Standalone mode puts all Grid components in one process on one machine. It is useful for local development and debugging, quick suites, and straightforward CI work. It is the simplest way to learn the remote-session pattern before operating multiple machines.
Connect a client with RemoteWebDriver
Java example
This Java example uses Chrome options and the default local Grid endpoint. Ensure the Selenium Java client dependency in your project matches the API for the release you installed.
import java.net.URL;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.remote.RemoteWebDriver;
public class GridScrape {
public static void main(String[] args) throws Exception {
URL grid = new URL("http://localhost:4444");
ChromeOptions options = new ChromeOptions();
WebDriver driver = new RemoteWebDriver(grid, options);
try {
driver.get("https://example.com/");
String title = driver.getTitle();
System.out.println(title);
// Use WebDriver locators and explicit waits for the fields
// your authorized collection needs.
System.out.println(driver.findElement(By.tagName("h1")).getText());
} finally {
driver.quit();
}
}
}
The important pieces are the Grid URL and browser options passed to RemoteWebDriver. Use a different options class for another browser, and request only capabilities your Nodes can satisfy. Always close the session with quit(), including when extraction fails.
Other Selenium client languages
The syntax varies by language, but the pattern is the same: create browser-specific options, construct the language binding’s remote WebDriver using the Grid URL and those options, perform normal WebDriver operations, and quit the session. Do not copy Java constructor syntax into Python, JavaScript, or another binding; use that binding’s RemoteWebDriver API for the Selenium version in use.
Choose a deployment mode
| Mode | Machine layout | When it fits | Trade-offs |
|---|---|---|---|
| Standalone | All components in one process on one machine; that machine supplies the available operating system and browsers. | Learning, local debugging, quick suites, and uncomplicated CI. | Lowest operational overhead, but capacity and failure isolation are limited to the single machine. Use when the required concurrent sessions fit its resources. |
| Hub and Node | A central entry point connects to Nodes, which may run on different machines, operating systems, and browser versions. | Adding or reducing browser capacity, or covering multiple environments. | More machines and configuration to operate; separate Nodes let capacity change without taking down the whole Grid. |
| Distributed | Grid components run separately, ideally on different machines, with ports and internal communication configured. | Operators who need control over component placement and scaling. | Most operational complexity of these patterns; component separation offers placement control but also creates more infrastructure and communication to manage. |
Decide based on the number of machines, required operating systems and browser versions, expected concurrent sessions, operational effort, and how much isolation from a machine or component failure you need. Selenium’s Grid guidance describes the modes and factors involved.
Recommended Free Tools
Run browser work in parallel without overwhelming the Grid
Parallelism means running multiple WebDriver sessions at once, not sending multiple commands through one shared driver object. For a small initial test, create one independent RemoteWebDriver session per worker and ensure the Grid has enough compatible slots. In a Hub-and-Node setup, the Distributor matches each new session request to a compatible available slot.
More concurrent sessions do not automatically mean proportionally more completed pages. Browser startup, page scripts, network delays, page size, and the target site’s own limits all affect elapsed time. Set a deliberate concurrency ceiling, use sensible waits, and avoid repeated requests that are unnecessary for the task.
Estimate capacity, then measure your workload
Grid capacity depends on the number of Nodes, sessions per Node, processors, browser mix, and available machine resources. Selenium’s current getting-started sizing discussion uses around 1 GB of RAM per browser session as a rough reference, not a guarantee; actual requirements vary, and the example defaults may not fit a given environment.
Rank #3
- Begin with fewer concurrent sessions than you think the hardware can support.
- Measure memory, processor use, session creation time, page completion time, and failures with the actual pages and browser versions you need.
- Increase concurrency gradually, watching for resource pressure and rising timeouts.
- Consider smaller Nodes if isolating sessions or limiting the impact of one machine’s failure is important.
There is no universal throughput figure for a scraping workload. Benchmark your own authorized pages and environment; no fixed speedup follows simply from adding Grid Nodes.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesScrape responsibly and keep Grid private
RFC 9309 describes robots.txt as crawler guidance and says its rules are not access authorization. Honor applicable crawler rules, but do not treat the file’s presence—or absence—as proof of legal permission. Robots.txt does not override site terms, access controls, or other obligations. Do not use browser automation to bypass restrictions.
Protect the Grid endpoint as infrastructure. Selenium warns that an externally exposed Grid can give third parties access to internal web applications and files, or allow them to run custom binaries. Restrict access with appropriate firewall rules and allow only trusted clients to reach it. Do not make the default endpoint publicly reachable just to simplify client connectivity.
Or skip the browser setup
If your task is to capture a page as an image or PDF rather than interact with it as a browser-driven scraper, ScreenshotNeo offers a screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
- Cookie and consent banners are accepted, and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed before capture; each step can be turned off.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers say which page verdict and billing result applied.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Troubleshooting common Grid problems
Client cannot connect to localhost:4444
Confirm the Server process is running, the client and Grid are on the same machine when using localhost, and the URL and port match. If the client runs elsewhere, localhost points to the client machine, not the Grid host; use the reachable Grid address and restrict access at the network boundary.
Session request waits or fails to match
The requested browser or capability may not exist on any available Node, or all compatible slots may be occupied. Check the Grid UI or status endpoint at the default Grid address, verify the Node’s registered browser capabilities, and lower concurrency or add compatible capacity.
Browser or driver fails to start
Check that the browser is installed on the machine where the Node runs and that its driver can be configured for the installed browser. Selenium Manager can configure drivers when enabled; otherwise, verify driver availability and compatibility for the specific browser and Selenium release.
Pages time out or extraction returns missing content
A page may load asynchronously, take longer than the client’s timeout, or render content only after interaction. Use a wait for the required element or condition instead of relying only on a fixed short delay, and inspect the page in the same browser configuration. Raising concurrency can worsen resource-related timeouts, so check Node CPU and memory before increasing limits.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Sessions remain open after errors
Put driver.quit() in a cleanup path such as Java’s finally block. Unclosed sessions continue to occupy slots and can make later requests wait even when the scraping code has stopped producing useful work.
Best Value
FAQ
Can I begin on one computer?
Yes. Standalone mode runs Grid in one process on one machine and uses http://localhost:4444 by default.
What changes when I add Nodes?
Nodes run browser sessions. The Distributor assigns each new session request to a Node slot whose capabilities match the request.
Does robots.txt give permission to scrape?
No. RFC 9309 says robots.txt rules are not access authorization; other permissions, terms, and access controls still matter.
Is it safe to expose Grid publicly?
Selenium warns against external exposure because Grid can provide access to internal resources and capabilities to execute custom binaries. Keep it behind restricted network access.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

