Use Selenium to capture the browser, then crop the decoded PNG with OpenCV: image[y1:y2, x1:x2]. The first index is the vertical (row) range and the second is the horizontal (column) range. For a single DOM element, Selenium can save that element directly and you can skip OpenCV.
Choose the right capture method
| Need | Use | Reason |
|---|---|---|
| One element’s rendered box | element.screenshot("element.png") |
Selenium calculates the element capture for you and returns a success boolean. |
| An arbitrary rectangle | Full screenshot, OpenCV decode, then slice | You control exact pixel bounds and can crop several regions from one capture. |
| Several output formats | cv2.imwrite() with the desired extension |
OpenCV selects the encoder from the filename extension. |
The examples below target Selenium’s current Python API documentation (4.49.0) and OpenCV image-operation guidance labeled OpenCV 5.0, with the slicing approach compatible with OpenCV 3.0 and later. Match the code to the versions installed in your environment.
Install Python dependencies and start a browser
Install Selenium, OpenCV’s Python bindings, and NumPy in the environment that will run the capture:
python -m pip install selenium opencv-python numpy
You also need a browser and a compatible Selenium driver. Selenium 4 normally obtains or manages the driver for supported browsers, but locked-down build agents may require you to provision the driver yourself. The browser must be able to reach the target page, including any authentication or network resources it needs.
#1 Best Overall
from selenium import webdriver
driver = webdriver.Chrome()
driver.get("https://example.com")
In production, put the driver lifecycle in try/finally so a failed capture does not leave browser processes running.
Save an arbitrary rectangular area with OpenCV
This complete example captures PNG bytes in memory, decodes them, validates the rectangle, crops it, and writes a PNG file.
import cv2
import numpy as np
from selenium import webdriver
URL = "https://example.com"
OUTPUT = "partial.png"
# Bounds are screenshot pixels: left, top, right, bottom.
x1, y1, x2, y2 = 100, 80, 500, 300
driver = webdriver.Chrome()
try:
driver.get(URL)
# Selenium returns the current window as PNG bytes.
png_bytes = driver.get_screenshot_as_png()
image = cv2.imdecode(
np.frombuffer(png_bytes, dtype=np.uint8),
cv2.IMREAD_COLOR,
)
if image is None:
raise RuntimeError("Could not decode Selenium screenshot")
height, width = image.shape[:2]
if not (0 <= x1 < x2 <= width and 0 <= y1 < y2 <= height):
raise ValueError(
f"Crop bounds are outside screenshot dimensions {width}x{height}"
)
# NumPy/OpenCV uses row (y), then column (x).
crop = image[y1:y2, x1:x2]
if crop.size == 0:
raise ValueError("Crop is empty")
if not cv2.imwrite(OUTPUT, crop):
raise OSError(f"Could not write {OUTPUT}")
finally:
driver.quit()
driver.get_screenshot_as_png() captures the current browser window as PNG bytes. cv2.imdecode() converts those bytes into an image array. The bounds check is important: a reversed or out-of-range slice can produce an empty result instead of the area you intended.
Understand the coordinate order
OpenCV’s Python examples use row-first, column-second indexing: img[10:110, 10:110]. Therefore a rectangle is written as image[y1:y2, x1:x2], never image[x1:x2, y1:y2]. The upper bound follows normal Python slicing and is exclusive, so the output width is x2 - x1 and its height is y2 - y1.
Coordinates are pixels in the decoded screenshot. CSS coordinates from browser scripts are not guaranteed to map one-to-one to those pixels: viewport size, browser scaling, device pixel ratio, and capture behavior can change the relationship. Print width and height, inspect the resulting image, and calibrate bounds for the browser configuration used by your job.
Rank #2
Write JPEG or WebP instead
Change the extension and output path:
if not cv2.imwrite("partial.jpg", crop):
raise OSError("Could not write partial.jpg")
if not cv2.imwrite("partial.webp", crop):
raise OSError("Could not write partial.webp")
OpenCV chooses the file format from the extension. The common screenshot path uses an 8-bit, three-channel BGR image produced by IMREAD_COLOR; format-specific channel and depth constraints still apply.
Capture one element without manual cropping
If the requested area is exactly one rendered WebElement, locate it and let Selenium save its PNG:
from selenium import webdriver
from selenium.webdriver.common.by import By
driver = webdriver.Chrome()
try:
driver.get("https://example.com")
element = driver.find_element(By.CSS_SELECTOR, ".target")
if not element.screenshot("element.png"):
raise OSError("Could not save element.png")
finally:
driver.quit()
element.screenshot_as_png is the equivalent bytes interface if you want to decode or post-process the element with OpenCV:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11png_bytes = element.screenshot_as_png
image = cv2.imdecode(
np.frombuffer(png_bytes, dtype=np.uint8),
cv2.IMREAD_COLOR,
)
if image is None:
raise RuntimeError("Could not decode element screenshot")
This method is preferable for a single element because it avoids estimating its rectangle. Use the full-window method when the target is a freeform region, crosses element boundaries, or requires multiple crops.
Make captures deterministic
Wait for the content you need
A screenshot records the page state at the instant Selenium captures it. Navigate, then wait for a reliable condition such as a target element becoming present and visible before taking the shot. If the page uses lazy loading, scroll the relevant content into view and wait for its image or text to appear. A fixed sleep can work for a small script but is less reliable than a condition tied to the page state.
Rank #3
Control viewport and scaling
Set a known window size before capture when repeatable coordinates matter. Keep the same browser version, operating-system display scale, headless settings, and device-pixel-ratio configuration between runs. Record the decoded image dimensions and fail fast if they differ from the dimensions your crop coordinates expect.
Handle full-page versus viewport shots
get_screenshot_as_png() captures the current window. If you need content below the viewport, use Selenium’s page-size or scrolling strategy first, then verify the resulting image dimensions. Do not assume that a browser’s CSS layout height is the same as the PNG height.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Troubleshooting
The output file is missing or empty
- Check the return value from
element.screenshot()ordriver.save_screenshot();Falseindicates a file-writing failure. - For OpenCV, check the boolean returned by
cv2.imwrite()and verify that the destination directory exists and is writable. - Use an extension with a format encoder such as
.png,.jpg, or.webp.
cv2.imdecode() returns None
The byte buffer is not a valid image. Confirm Selenium returned PNG bytes, pass a NumPy uint8 buffer, and do not accidentally decode an HTML error response or an empty byte string.
The crop is empty or the wrong area
- Print
image.shape[:2]and compare it with your assumed width and height. - Remember the order is
[y1:y2, x1:x2]. - Remember that
x2andy2are exclusive. - Check device scale and browser zoom; recalibrate screenshot pixels against a visible landmark.
The element cannot be found
Use the correct CSS selector, wait for the element to be present, and account for content inside an iframe by switching into that frame before locating the element. An element screenshot captures the element’s rendered box, not an arbitrary rectangle around neighboring content.
The page is blank, blocked, or still loading
Inspect the browser manually or collect page HTML and console logs in your test harness. Authentication redirects, bot checks, network failures, and JavaScript errors can all produce a technically valid screenshot that is not the intended page. Add an explicit readiness condition rather than immediately retrying the same capture.
Rank #4
Performance, reliability, and cost considerations
- Capture once and derive several crops from the same decoded array when you need multiple regions; repeated browser screenshots cost more time than NumPy slicing.
- Keep the image in memory until all crops are written, but release large arrays after processing long batches.
- Use PNG for lossless text and UI evidence; JPEG can be smaller but introduces compression artifacts.
- Always close the driver in
finally. For batch jobs, isolate failures per URL so one navigation does not discard completed files. - Store the browser, Selenium, OpenCV, and operating-system versions with artifacts when pixel-level reproducibility matters.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One request returns a PNG, JPEG, WebP, or PDF, and its cleanup steps accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture. You can turn each cleanup step off when needed. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the result in X-Page-Verdict and X-Billed headers.
Free tools Windows power users keep installed
One-click scans. No signup required.
For a direct replacement for the Selenium setup, see the ScreenshotNeo API documentation and call the endpoint:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers element selection by CSS selector, full-page capture with lazy images loaded, custom CSS and JavaScript, click and wait controls, request blocking, headers and cookies, timezone and geolocation, image resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage reporting, and an OpenAPI specification. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account to get started.
FAQ
Can I crop directly from save_screenshot()?
Yes. Save the PNG to disk, read it with cv2.imread(), then apply the same image[y1:y2, x1:x2] slice. In-memory PNG bytes avoid an unnecessary intermediate file.
Does an element screenshot include content outside the element?
No. It represents the WebElement’s rendered box. Use a full-window capture when the desired rectangle is not exactly one element.
Best Value
Why are my CSS coordinates off by a few pixels?
CSS pixels and screenshot pixels can differ because of device scale, browser zoom, viewport settings, and capture implementation. Measure the actual PNG dimensions and calibrate in the same environment used for automation.
Frequently Asked Questions
Can I crop directly from save_screenshot()?
Yes. Save the PNG, read it with cv2.imread(), and slice image[y1:y2, x1:x2]; capturing PNG bytes first avoids an intermediate file.
Does an element screenshot include content outside the element?
No. It represents the WebElement’s rendered box. Use a full-window capture for a freeform rectangle.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Why are my CSS coordinates off by a few pixels?
Device scale, browser zoom, viewport settings, and capture implementation can make CSS and screenshot pixels differ. Measure and calibrate against the actual PNG.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




