Use a browser automation tool such as Playwright to load the page, inspect its rendered image elements, scroll to reveal lazy-loaded content, and save each selected image resource. The method below downloads the browser-selected image from each visible or discovered <img>; it does not guarantee every image on a site or every asset embedded in CSS, canvas, frames, or custom galleries.
What “all images from a URL” means
A URL identifies a page, not a complete inventory of every asset a site may serve. A practical browser-automation workflow collects image resources exposed by the rendered page under a defined viewport and interaction pattern. That scope may include images inserted after JavaScript runs or after scrolling, but may not include assets that require clicking a gallery, signing in, expanding a section, or triggering another interaction.
There are two common collection goals:
- Save the image the browser chose: collect each image element’s
currentSrc. This reflects the selected source, including a responsivesrcsetchoice. - Collect every declared responsive alternative: parse
srcsetand relevant<picture><source>markup. This can yield several candidates for one displayed image and is a different, broader task.
For a straightforward download of what the page presents, start with currentSrc. MDN describes it as the URL of the image selected by the browser to load: HTMLImageElement.currentSrc.
How the workflow works
- Launch a browser and navigate to the target URL.
- Wait for meaningful page content, then inspect rendered
<img>elements and record their selected URLs and context. - Scroll in increments, pause briefly, and inspect again so lazy-loaded images and dynamically added elements have a chance to appear.
- Deduplicate the URLs, fetch each resource, and save it under a collision-safe filename.
- Record successes and failures, and report the page, time, and extent of scrolling or interaction used.
Do not treat the page’s load event as proof that every lazy image has arrived. MDN notes that lazy-loaded resources can still be pending after that event: HTML image element and loading behavior. Similarly, networkidle is not a universal signal of completion for pages that keep polling or loading content dynamically. Prefer content-aware waits and repeated inspection.
#1 Best Overall
Download browser-selected images with Playwright Python
Install Playwright and its Chromium browser once:
python -m pip install playwright
python -m playwright install chromium
Save this as download_page_images.py. It scrolls the document in steps, repeatedly gathers each image’s currentSrc, waits for each image element to finish loading, then downloads unique image URLs using the browser context’s request client. It writes files into a folder and produces a JSON report mapping each URL to its result.
import asyncio
import hashlib
import json
import os
import re
from pathlib import Path
from urllib.parse import unquote, urlparse
from playwright.async_api import async_playwright
PAGE_URL = "https://example.com"
OUTPUT_DIR = Path("downloaded_images")
SCROLL_PAUSE_MS = 700
def safe_filename(url: str, index: int) -> str:
"""Keep a readable extension when available; hash the URL to avoid collisions."""
parsed = urlparse(url)
name = unquote(os.path.basename(parsed.path))
name = re.sub(r"[^A-Za-z0-9._-]+", "_", name).strip("._")
if not name:
name = "image"
stem, ext = os.path.splitext(name)
if not ext or len(ext) > 8:
ext = ".img"
digest = hashlib.sha256(url.encode("utf-8")).hexdigest()[:12]
return f"{index:04d}_{stem[:70]}_{digest}{ext}"
async def collect_images(page):
return await page.locator("img").evaluate_all("""imgs => imgs.map(img => ({
alt: img.alt || "",
src: img.getAttribute("src") || "",
srcset: img.getAttribute("srcset") || "",
currentSrc: img.currentSrc || "",
complete: img.complete,
naturalWidth: img.naturalWidth,
naturalHeight: img.naturalHeight
}))""")
async def main():
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
report = {"page_url": PAGE_URL, "images": []}
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(accept_downloads=True)
page = await context.new_page()
response = await page.goto(PAGE_URL, wait_until="domcontentloaded", timeout=60000)
report["navigation_status"] = response.status if response else None
# Give app-rendered content a chance to appear, without assuming the whole
# site becomes idle. Scroll and re-query because the DOM may change.
await page.wait_for_timeout(1000)
previous_height = -1
for _ in range(30):
height = await page.locator("body").evaluate("el => el.scrollHeight")
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
await page.wait_for_timeout(SCROLL_PAUSE_MS)
if height == previous_height:
break
previous_height = height
await page.evaluate("window.scrollTo(0, 0)")
# The locator is re-evaluated after scrolling. currentSrc is the selected
# resource; empty or data URLs are not downloaded by this example.
elements = await collect_images(page)
candidates = {}
for item in elements:
url = item["currentSrc"].strip()
if url and not url.startswith("data:"):
candidates.setdefault(url, item)
report["unique_selected_urls"] = len(candidates)
for index, (url, item) in enumerate(candidates.items(), start=1):
entry = {"url": url, "alt": item["alt"], "src": item["src"],
"srcset": item["srcset"], "natural_width": item["naturalWidth"],
"natural_height": item["naturalHeight"]}
try:
# Wait for matching rendered image elements; completion alone is
# not success, so also check naturalWidth before fetching.
await page.locator("img").evaluate_all("""(imgs, target) => {
for (const img of imgs) {
if (img.currentSrc === target) img.loading = 'eager';
}
}""", url)
await page.wait_for_function("""target => [...document.images].some(
img => img.currentSrc === target && img.complete
)""", arg=url, timeout=15000)
state = await page.locator("img").evaluate_all("""(imgs, target) => {
const img = imgs.find(x => x.currentSrc === target);
return img ? {complete: img.complete, naturalWidth: img.naturalWidth} : null;
}""", url)
if not state or state["naturalWidth"] == 0:
raise RuntimeError("Image element completed without a usable naturalWidth")
result = await context.request.get(url, timeout=30000)
entry["http_status"] = result.status
if not result.ok:
raise RuntimeError(f"HTTP {result.status}")
content_type = (result.headers.get("content-type") or "").lower()
if not content_type.startswith("image/"):
raise RuntimeError(f"Response is not reported as an image: {content_type or 'unknown content type'}")
path = OUTPUT_DIR / safe_filename(url, index)
path.write_bytes(await result.body())
entry.update({"saved_as": str(path), "content_type": content_type, "outcome": "saved"})
except Exception as exc:
entry.update({"outcome": "failed", "error": str(exc)})
report["images"].append(entry)
await browser.close()
report["saved"] = sum(x["outcome"] == "saved" for x in report["images"])
report["failed"] = sum(x["outcome"] == "failed" for x in report["images"])
Path("image_download_report.json").write_text(json.dumps(report, indent=2), encoding="utf-8")
print(f"Saved {report['saved']}; failed {report['failed']}; see image_download_report.json")
if __name__ == "__main__":
asyncio.run(main())
Replace PAGE_URL with the page you are authorized to access. The script filters out inline data: images, keeps one URL per selected resource, and uses a URL hash in filenames so different resources with identical basenames do not overwrite each other. It checks both an image element’s completion and nonzero natural width before fetching; the selected URL alone does not prove a successful load.
What the code does not collect
- Alternative
srcsetentries not selected for this browser viewport or device configuration. - CSS background images, images drawn into canvas, assets inside frames, or resources exposed only after clicking or another page-specific interaction.
- Every image from every route on the website. This visits one supplied page URL.
Current source or every responsive candidate?
HTML can provide multiple sources for an image. A <picture> element may offer format or viewport alternatives, while srcset can let the browser choose a resource based on display conditions. The browser-selected currentSrc is the right candidate when the goal is to save what this browser rendered. To archive the markup’s declared choices, inspect the image’s srcset and its parent <picture> sources as well. See MDN’s references for image markup and the picture element.
Parsing srcset correctly matters: commas can separate candidates, while URL and descriptor syntax determine which resource each candidate identifies. Do not treat a filename extension as definitive proof of an image format; servers may use extensionless paths or return a different content type. The example records the response content type and rejects responses not labeled as images.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #2
Direct image fetching versus browser attachment downloads
An ordinary <img> is a page resource, not necessarily a user-triggered file download. Playwright’s download event is intended for attachment downloads initiated by the page; its Download object can save such a file explicitly. Downloads associated with a browser context are temporary and removed when that context closes unless saved elsewhere. See Playwright’s Python download documentation.
For ordinary image elements, collect the selected resource URL and fetch its bytes as the example does. Use a download event when the site itself has a “Download” control that triggers an attachment. Playwright documents browser page and locator APIs for navigating and inspecting rendered content: pages and locators.
Handling other page behaviors
Lazy loading and infinite scroll
The example scrolls to the bottom in increments and repeats until the page height stops changing or its bounded loop ends. This is a practical heuristic, not proof that all content has loaded: some sites load more only near particular elements, after longer pauses, or after a button click. Adapt the interaction to the page, re-query the DOM after each action, and record where the crawl stopped.
Authentication and access controls
Pages behind a login may require a browser context with an authorized session. Do not bypass a site’s access controls, bot checks, or restrictions. A URL that loads for a logged-out browser may not expose the images available to an authenticated user, and a download request can fail if it lacks the session state or headers required by the site.
Rank #3
Frames, galleries, and non-image elements
If expected images are absent, inspect frames separately and interact with gallery controls as a user would. A CSS background is not represented by an <img>; canvas pixels are not ordinary image URLs. Those cases need page-specific inspection or a different capture strategy and should be called out in the collection report.
Or skip the browser setup
If you need a rendered-page screenshot rather than a folder of individual source image files, ScreenshotNeo provides a website screenshot API and MCP server. Its one-call API captures an image or PDF; it does not replace an image-resource downloader when you need each original image file.
For a screenshot of a page, this cURL request saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating page verdict and billing. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Each feature is available on every plan.
Free tools Windows power users keep installed
One-click scans. No signup required.
Sign up for 1,000 free screenshots a month with no card.
Rank #4
Performance, reliability, and cost
Browser automation has setup and runtime costs: it launches a browser, renders the page, scrolls it, and makes a request per unique selected URL. For a long page or a large image set, downloads can dominate runtime. Keep the crawl bounded, avoid unneeded repeated requests, and consider a small concurrency limit if adding parallel downloads. Unlimited parallelism can burden the target site or trigger throttling.
For repeatable runs, store the input page URL, timestamp, viewport, browser choice, number of unique URLs, scroll/interaction method, outcomes, and failure messages. The script’s report provides a starting point. A result count should be described as “unique selected image URLs collected under these conditions,” not “all images on the site.”
Transient network errors may merit a limited retry with a short delay; persistent HTTP errors should remain failures rather than being silently skipped. Preserve the source URL in the report so a saved file can be traced back. Do not assume every response will be a conventional JPEG or PNG, or that one extension accurately represents its encoding.
Recommended Free Tools
Troubleshooting
- No image URLs found: Wait for the page’s content-bearing element, scroll, and query again. Check whether the relevant content is in a frame, background style, canvas, or interaction-gated gallery.
- Images appear in the browser but download fails: Inspect the HTTP status and content type. The server may require session state, deny direct retrieval, or have changed the resource URL. Use only authorized access, and keep failures in the report.
currentSrcis empty or an image has zero natural width: The resource may not yet be selected or loaded, may have failed, or may be a placeholder. Wait for the page-specific state and re-check; do not count discovery as a successful save.- Only one responsive size is saved: That is expected when using
currentSrc. Inspectsrcsetand<picture>sources if the goal is to collect alternatives. - Some lower-page images are missing: Increase the bounded scroll steps or pause, and inspect after each step. Infinite-scroll pages may require repeated scrolling or a “load more” action.
- Files overwrite one another: Use collision-resistant names rather than the basename alone. The example includes a short hash of the source URL.
- Navigation times out: Distinguish navigation failure from a page that rendered useful content but kept background requests open. Use a realistic navigation timeout and wait for a relevant element rather than relying on universal network idleness.
Permissions and responsible use
Being able to view or save an image does not establish permission to reuse it. Check the image’s license and the site’s terms for your intended use. The applicable legal position depends on jurisdiction and context; for consequential reuse questions, seek authoritative legal guidance. Keep requests proportionate and respect the site’s access controls.
Best Value
Frequently Asked Questions
Does this download images from every page on a website?
No. It visits one supplied page URL. Crawling multiple routes requires a separate, authorized URL-discovery workflow.
Does Playwright’s download event save every image displayed on a page?
No. That event handles file downloads initiated by the page; ordinary image elements are fetched as page resources.
Can I save every responsive image variant?
Yes, but the example saves only each browser-selected currentSrc. Collect and parse srcset and relevant picture sources to gather declared alternatives.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

