Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For a static page you’re permitted to access, fetch its HTML, parse image references, turn relative paths into absolute URLs, then download each image. The script below uses Requests and Beautiful Soup, streams files to disk, avoids duplicate URLs and filename overwrites, and reports failures. It cannot guarantee every image a browser displays: JavaScript-rendered content, CSS backgrounds, lazy-loaded images, authentication and host restrictions may need separate handling.

What “all images” means in this method

This method collects image URLs exposed in the HTML response for one webpage. It starts with ordinary <img src="..."> references and, in the example below, also checks common lazy-loading attributes and srcset. It does not crawl every page on a site or automatically reproduce everything a browser eventually renders.

Pages may reveal images only after JavaScript runs, load them as CSS backgrounds, or require a logged-in session. Some sites use custom delivery behavior. A parser inspecting the server-returned HTML cannot infer image URLs that are absent from that HTML. For those cases, use a documented API or authorized export if available, or investigate a rendering approach appropriate to the site’s terms and access rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install the Python packages

The example uses Requests for HTTP requests and streaming, and Beautiful Soup for parsing HTML. Install both in the Python environment where you’ll run the script:

python -m pip install requests beautifulsoup4

Beautiful Soup supports multiple parsers. The script explicitly selects Python’s built-in html.parser, so it does not require an additional parser package.

Download image references from one page

Save this as download_images.py. Replace PAGE_URL with the permitted webpage you want to inspect. The example is illustrative; it has not been presented as a tested download for any particular website.

from pathlib import Path
from urllib.parse import urljoin, urlsplit, unquote
import re

import requests
from bs4 import BeautifulSoup

PAGE_URL = "https://example.com/page"
OUTPUT_DIR = Path("downloaded_images")
TIMEOUT = (10, 45)  # connect timeout, read timeout, in seconds
CHUNK_SIZE = 64 * 1024

# Common attributes used for ordinary and lazy-loaded image URLs.
IMAGE_ATTRIBUTES = ("src", "data-src", "data-original", "data-lazy-src")


def srcset_urls(value):
    """Return URL candidates from a typical comma-separated srcset value."""
    if not value:
        return []
    urls = []
    for candidate in value.split(","):
        fields = candidate.strip().split()
        if fields and fields[0]:
            urls.append(fields[0])
    return urls


def safe_filename(url, index):
    """Use the URL path's final component, with a unique fallback if needed."""
    path_name = unquote(urlsplit(url).path.rsplit("/", 1)[-1]).strip()
    # Remove characters that are unsafe or awkward in common filenames.
    name = re.sub(r"[^A-Za-z0-9._-]+", "_", path_name).strip("._")
    if not name:
        name = f"image_{index}"
    return name[:180]


def unique_path(directory, filename):
    """Avoid replacing an existing file, including same-name URL collisions."""
    candidate = directory / filename
    stem, suffix = candidate.stem, candidate.suffix
    counter = 2
    while candidate.exists():
        candidate = directory / f"{stem}_{counter}{suffix}"
        counter += 1
    return candidate


def main():
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

    try:
        response = requests.get(PAGE_URL, timeout=TIMEOUT)
        response.raise_for_status()
    except requests.RequestException as exc:
        raise SystemExit(f"Could not fetch page: {exc}") from exc

    soup = BeautifulSoup(response.content, "html.parser")
    image_urls = []
    seen = set()

    for img in soup.find_all("img"):
        candidates = []
        for attribute in IMAGE_ATTRIBUTES:
            value = img.get(attribute)
            if value:
                candidates.append(value.strip())
        candidates.extend(srcset_urls(img.get("srcset")))
        candidates.extend(srcset_urls(img.get("data-srcset")))

        for candidate in candidates:
            if not candidate or candidate.startswith(("data:", "blob:")):
                continue
            absolute_url = urljoin(response.url, candidate)
            if urlsplit(absolute_url).scheme not in ("http", "https"):
                continue
            if absolute_url not in seen:
                seen.add(absolute_url)
                image_urls.append(absolute_url)

    print(f"Found {len(image_urls)} distinct image URL(s) in parsed img attributes.")
    if not image_urls:
        print("No downloadable image references were found in the checked attributes.")
        return

    saved = 0
    failed = 0

    for index, image_url in enumerate(image_urls, start=1):
        destination = unique_path(OUTPUT_DIR, safe_filename(image_url, index))
        temporary = destination.with_name(destination.name + ".part")
        try:
            with requests.get(image_url, stream=True, timeout=TIMEOUT) as image_response:
                image_response.raise_for_status()
                content_type = image_response.headers.get("Content-Type", "")
                if content_type.lower().startswith("text/"):
                    raise ValueError(f"server returned text content ({content_type})")
                with temporary.open("wb") as output:
                    for chunk in image_response.iter_content(chunk_size=CHUNK_SIZE):
                        if chunk:
                            output.write(chunk)
            temporary.replace(destination)
            saved += 1
            print(f"Saved {destination} [{content_type or 'content type not stated'}]")
        except (requests.RequestException, OSError, ValueError) as exc:
            failed += 1
            temporary.unlink(missing_ok=True)
            print(f"Failed {image_url}: {exc}")

    print(f"Finished: {saved} saved; {failed} failed; output directory: {OUTPUT_DIR}")


if __name__ == "__main__":
    main()

How the script works

Fetch the page before parsing

requests.get retrieves the page, and raise_for_status() turns HTTP error responses into reported failures rather than treating them as page content. The timeout tuple sets separate connection and response-read limits; adjust them for a slow site only when you have a reason to do so. The script uses response.url as the URL base after redirects, which is important when resolving relative image references from the final page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find and normalize references

Beautiful Soup’s find_all("img") searches image elements in the parsed HTML. The script checks src and several commonly used lazy-loading attributes, then parses candidates from srcset and data-srcset. It uses urljoin rather than string concatenation, so root-relative paths such as /images/a.jpg, path-relative values such as ../a.jpg, and scheme-relative URLs can be resolved against the page URL.

Deduplication happens after URL resolution. Thus, two references resolving to the same absolute URL are requested once. Data and blob URLs are skipped because they are not ordinary HTTP(S) resources the script can fetch in the same way.

Stream each response and handle filenames

Each image response is streamed in chunks using Requests’ iter_content workflow instead of loading the entire response body into memory. A temporary .part file is renamed only after the response completes, helping avoid leaving a partial file under its final name. Existing filenames are preserved by adding a numeric suffix for collisions.

The extension comes from the URL path, not from verification of the image’s bytes. A URL ending in .jpg may return an HTML error page, and a URL without an extension may still return an image. The script rejects a plainly textual content type, but that is not a full image-format validator. If file integrity matters, inspect the response type and validate downloaded bytes with an image library appropriate to your workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Coverage limits and page-specific cases

Responsive images

srcset can list multiple candidates intended for different viewport sizes or pixel densities. The example downloads each candidate it sees; that may mean downloading several versions of what appears visually as one image. Its simple comma split suits common values but is not a complete parser for every unusual or malformed srcset. If you need the single candidate a browser would select, selection depends on the page’s sizes, viewport and device pixel ratio.

Lazy loading and JavaScript

The checked data-* names are common conventions, not a universal standard. A site may use another attribute or populate src only after scrolling or running JavaScript. Inspect the page’s HTML and network behavior to identify its actual convention. If image references are created only in a rendered page, an HTML-only parse will miss them; do not treat this script as a way around access controls.

CSS backgrounds and inline data

Images referenced by CSS, including background images, are not found by searching <img> tags. A page may also embed content as data URLs, which this example deliberately skips. CSS parsing and data extraction require additional handling tailored to how the page represents those assets.

Authentication, redirects and blocked requests

A page or image host may require a legitimate session, cookies or other authorized credentials. The example makes independent requests and does not transfer a browser’s logged-in state. A site may redirect, rate-limit or reject automated requests; a timeout or custom header does not guarantee access and should not be used to bypass restrictions. If the image host requires authentication, use an authorized method supported by the site and protect any credentials you add.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Standard library or Requests?

Beautiful Soup is the parser in either approach; it is not a downloader. Python’s urllib.request can retrieve URLs without adding an HTTP-client dependency. Requests offers a convenient interface and a straightforward streamed-save pattern using iter_content. Choose based on your project’s dependencies and preferred error-handling interface.

Choice Dependency and workflow Useful consideration
urllib.request Python standard library; offers URL retrieval functions such as urlretrieve. Python documents ContentTooShortError for retrievals shorter than a reported Content-Length. Handle incomplete downloads and errors explicitly.
Requests Third-party package; supports a convenient request interface and streamed response iteration. For saved downloads, use stream=True and iterate Response.iter_content rather than keeping a large response in memory.
Beautiful Soup Separate HTML parsing package; searches parsed markup. It locates references present in parsed HTML; it does not render JavaScript or download image files by itself.

Responsible use, performance and recovery

  • Check permission and site terms. Downloading a file does not grant the right to republish it. Confirm your intended use is allowed, especially for commercial reuse.
  • Understand robots.txt correctly. Google Search Central describes robots.txt as telling search engine crawlers which URLs they can access and as a way to manage crawler traffic, including media files. It is not a security mechanism and does not grant copyright permission or decide whether your use is lawful.
  • Keep requests measured. A page with many references creates many separate downloads. Avoid running repeated, high-volume requests against a host; respect applicable site guidance and stop if requests are rejected or rate-limited.
  • Expect partial results. The script reports individual failures and continues to later URLs. Review the final saved/failed counts and the reported URL-specific errors rather than assuming every reference succeeded.
  • Retry deliberately. For a transient timeout, you can rerun after considering a longer timeout or an appropriate delay. The script avoids overwriting existing files, so reruns may create suffixed duplicates; remove or organize prior output if you want a clean run.

Or skip the browser setup

If your goal is a visual record of the page rather than downloading its original image files, ScreenshotNeo can return a screenshot or PDF through one GET request. It is not a replacement for extracting the page’s image assets. Its clean-shot options accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets before capture, and each step can be turned off. Bot checks, blank pages, timeouts and failed loads are not billed; the response identifies the page verdict and billing status. An MCP server provides screenshot tools for Claude, Cursor and other MCP clients. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo and its API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/page -o shot.webp

Sign up for 1,000 free screenshots a month, with no card required.

Sources and documentation

The relevant primary documentation is the Beautiful Soup documentation for parsing HTML, the Requests Quickstart for requests and streamed downloads, the Python 3.14.7 urllib.request documentation for URL retrieval and incomplete-transfer behavior, and Google Search Central’s robots.txt introduction for crawler access and the limits of robots.txt. These sources explain their respective components; the combined safeguards in the illustrative script are not a claim of universal compatibility.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does this download images from every page on a website?

No. It fetches the single URL assigned to PAGE_URL. A site-wide crawl is a different task with additional scope, rate and permission considerations.

Why might the saved file have the wrong extension?

The example derives its name from the URL path, which need not describe the returned bytes. Check the response content type and validate the file if format certainty matters.

Can I reuse images I downloaded?

Not automatically. Download access does not establish permission to republish or use an image commercially; check the applicable rights and terms.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.