October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
browser automation

Selenium WebDriver: A Practical Guide to Browser Automation

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Selenium WebDriver lets a program control a real browser: open pages, find elements, enter text, click controls and verify results. To get started, install a Selenium language binding, have a supported browser available, then create a driver session, wait for the page state your task needs, and end the session with quit(). In current Selenium releases, Selenium Manager can often locate and manage the browser driver automatically, so downloading a driver by hand is not always necessary.

What Selenium WebDriver does

WebDriver is a language-neutral interface for controlling browsers, and the W3C describes it as a Recommendation. Your code uses a Selenium binding for its language; that binding sends commands to a browser-specific driver, which controls the browser. A session can run on the computer executing the script or on a remote machine. Selenium’s WebDriver documentation explains the architecture and supported ways to drive browsers.

This makes Selenium useful for browser-based testing and repetitive interactions where you need to exercise the actual browser UI. It is not a shortcut around application behavior: the browser still loads the page, scripts still run, and your automation must account for changing content and browser differences.

What you need before writing a script

  • A language binding: install Selenium for Python, Java, JavaScript, C#, Ruby or another supported language using that language’s normal package manager.
  • A browser: install or select the browser you intend to automate. Choose it based on the browsers and operating systems relevant to your users or test environment.
  • A browser driver: the implementation that connects Selenium to the browser. Selenium Manager may provide it automatically in current releases; otherwise, provide a compatible driver through PATH or a service configuration.

Exact installation commands depend on the language and package manager; consult Selenium’s getting-started guide for the binding you use. Selenium Manager is shipped with Selenium releases beginning with 4.6 and is invoked by bindings as a fallback when a driver has not been supplied. Its documentation says browser management for Chrome, Firefox and Edge is available from Selenium 4.11.0. These are version-specific details, so check the current documentation and your platform before relying on automatic management: Selenium Manager.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do you still need to download ChromeDriver?

Not necessarily. With a current Selenium release, try creating a Chrome session without setting a driver path; Selenium Manager can detect the browser, resolve a corresponding driver, download it and cache it. If your environment cannot use that flow, download a compatible driver and put it on PATH, or explicitly configure its location in the browser’s service object. The same general choices apply to other browsers, subject to their support and your Selenium version.

Your first Selenium script

A minimal automation follows a repeatable sequence: start a session, navigate, locate a control, interact, inspect the outcome, and quit. The example below uses Python and the public Selenium test page. Install the binding with python -m pip install selenium, then save and run the script:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

# Selenium Manager can supply the driver when one is not configured.
driver = webdriver.Chrome()
try:
    driver.get("https://www.selenium.dev/selenium/web/web-form.html")
    wait = WebDriverWait(driver, 10)

    field = wait.until(EC.visibility_of_element_located((By.NAME, "my-text")))
    field.send_keys("Selenium WebDriver")
    driver.find_element(By.CSS_SELECTOR, "button").click()

    result = wait.until(EC.visibility_of_element_located((By.ID, "message")))
    assert result.text == "Received!"
finally:
    driver.quit()

The ten-second wait sets an upper bound; it does not force the script to pause for the full duration if the condition becomes true earlier. The example uses selectors shown on Selenium’s form page, so if you adapt it to another site, inspect that site’s DOM and choose selectors that identify the intended controls. Selenium’s first-script guide shows corresponding workflows in multiple bindings.

Locate elements deliberately

Use a locator that reflects how the page identifies an element: for example, an ID, a name, CSS selector or accessible text strategy where appropriate. A selector that matches several elements can make an action ambiguous or target the wrong control. Prefer stable identifiers provided for testing when available, and verify the resulting page state rather than assuming a click succeeded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Always end the session

Put cleanup in a finally block or equivalent teardown hook so a failed assertion does not leave browsers running. quit() ends the WebDriver session and closes its associated windows. Closing an individual window is not the same as ending the session; Selenium’s session guidance recommends quitting when the automation is finished: WebDriver drivers and sessions.

Wait for the application, not just the browser

A navigation call waits according to the browser’s page-load strategy, but a completed document load does not prove that a JavaScript application has finished rendering or that a particular control is ready. Dynamic pages can change after the load event. Acting too early is a common source of race conditions and flaky tests.

Use an explicit wait for the condition that matters at that point in the test: an element becoming present, visible or clickable, or a business-relevant state change. Selenium’s wait documentation describes explicit waits and the conditions they can express: Waiting Strategies.

Choose the condition to match the next action

  • Present: the element has entered the DOM, even if it is not yet visible.
  • Visible: the element is displayed and can be inspected or filled.
  • Clickable: the element is in a state suitable for a click, rather than merely existing in the DOM.
  • Application state: a result, status or other meaningful change confirms the action completed.

A fixed sleep can help diagnose whether timing is involved, but it should not be the routine synchronization strategy: it either wastes time on fast runs or remains too short on slow ones. Once a temporary delay confirms a timing issue, replace it with a wait for the relevant condition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Browser options and page-load behavior

Browser options configure session capabilities and behavior. One decision is the page-load strategy. Selenium documents three values:

Strategy Navigation waits for Practical consequence
normal The load event Waits for the document load event, but not necessarily for application-specific rendering.
eager DOMContentLoaded Returns earlier than normal; your test must wait for the UI state it needs.
none The initial page download Returns without waiting for the usual page-load event; explicit synchronization becomes especially important.

These strategies affect when navigation returns, not whether a target element is ready. If you select a faster strategy, pair it with waits for the controls or application state your script uses. See Selenium’s driver options documentation for options and configuration examples.

Run locally, remotely or across a grid

Local browser session

A local session starts the required driver service on the machine running the script and controls a browser there. This is usually the simplest way to begin, debug selectors and reproduce a failure.

Remote session and Selenium Grid

A remote session sends commands to a browser running elsewhere. It requires a remote WebDriver endpoint and browser options describing the requested session. Selenium Grid is the Selenium project’s route for distributing execution across machines and environments; it is useful when tests need to scale beyond the local machine. The exact deployment and capabilities depend on the Grid and browser configuration you operate. Start with Selenium’s WebDriver documentation for local and remote execution concepts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a browser that matches the job

Test the browsers that matter for the application and its users. Consider operating-system coverage, browser-specific behavior, driver availability and whether execution must be remote. Selenium publishes browser-specific guidance for Chrome, Edge, Firefox, Internet Explorer and Safari. Its driver-installation documentation lists Chrome/Chromium, Firefox and Edge across Windows, macOS and Linux; Internet Explorer for Windows; and Safari on macOS High Sierra or later. Opera’s driver no longer works with current Selenium functionality and is officially unsupported. Browser and Selenium compatibility can change, so confirm the current support details before building a matrix: Selenium browser documentation and driver-location guidance.

WebDriver BiDi: when browser events matter

Classic WebDriver is primarily a command-and-response interface. WebDriver BiDi adds a bidirectional WebSocket connection so scripts can receive and react to browser events, including network requests, console messages and JavaScript errors. That can help when a test needs more than page interactions, but functionality depends on support in the target browser and implementation. Check Selenium’s current BiDi documentation for the specific browser and features you plan to use: WebDriver BiDi.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Start by identifying whether the failure is setup-related, synchronization-related or specific to a browser/driver. Selenium’s troubleshooting guidance notes that some apparent Selenium problems originate in the underlying driver, and recommends examining synchronization and trying another browser to help isolate the cause: Troubleshooting.

Symptom Likely cause What to try
Driver executable not found The driver is not on PATH, not configured in a service, or automatic management cannot resolve it in this environment. Check the Selenium version and platform support; try Selenium Manager, add a compatible driver to PATH, or pass its location through the browser’s service configuration.
Element not found The selector does not match the current DOM, the page has not rendered that element, or the page structure changed. Inspect the live DOM and locator; wait for presence or visibility when content is dynamic.
Element is present but action fails The element may be hidden, covered, disabled or not yet interactable. Wait for visibility or clickability, then verify the application state after the action.
Intermittent failures A timing race or browser/driver-specific behavior may be involved. Temporarily add a delay only to diagnose timing, replace it with a condition-based wait, capture useful logs, and compare behavior in another browser.
Session remains after the test exits Cleanup did not run after an error or only a window was closed. Put quit() in a finally block or test teardown.

If Selenium Manager does not work on a constrained platform or architecture, do not assume automatic driver resolution is supported there. Use the documented driver-location options and verify the exact browser, driver and Selenium release combination.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability and cost considerations

For reliability, wait on meaningful conditions and keep each session’s setup and cleanup explicit. For performance, page-load strategies can let navigation return sooner, but they do not make application rendering complete; measure the behavior of your own test flow and retain the waits needed to avoid races. Remote execution and Grid can distribute work, but introduce a remote endpoint and environment configuration that must be maintained.

Selenium is software: the core setup described here is a language binding, browser and browser-driver software. The cited Selenium documentation does not establish a required paid license or a specific commercial execution service, so do not treat any particular vendor or hosted platform as necessary to use WebDriver.

Or skip the browser setup

If your goal is to capture a webpage rather than interact with it as a test user, Selenium may be more setup than you need. ScreenshotNeo is a website screenshot API and MCP server: send a GET request with a URL to receive a PNG, JPEG, WebP or PDF. Its capture flow accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets before the shot; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. AI agents can use its MCP server tools, including take_screenshot, get_page_info and capture_pdf.

Here is a one-call cURL example; replace the target URL with the page you want to capture. See the ScreenshotNeo documentation for API options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.

Frequently asked questions

Can Selenium automate a site without a graphical desktop?

The cited material establishes browser automation and remote execution, but does not specify a universal headless-mode setup for every browser and platform. Check the browser-specific options for your target environment.

Should I use WebDriver or BiDi?

Use WebDriver for ordinary browser commands and interactions. Consider BiDi when receiving browser events such as network or console activity is part of the task, after checking browser and implementation support.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.