Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a bulk image downloader as four separate steps: fetch a page, find the image URLs it exposes, retrieve each image as bytes, and save each file under a safe local name. The Python example below uses Requests and Beautiful Soup, handles per-image errors, streams large files, avoids accidental overwrites, and limits the batch. It is a starting point—not a universal scraper: the selectors and navigation logic must match the site you are allowed to access.

How a bulk image downloader works

A downloader is easier to maintain when page discovery and file transfer are separate. The page parser can change when a site changes its markup without requiring a rewrite of the code that streams and saves image data.

  1. Fetch: request the page that lists or displays images.
  2. Discover: parse its HTML and select image elements or links, then resolve relative URLs.
  3. Retrieve: request each image URL and check whether the response succeeded.
  4. Save: stream the response bytes to disk with a safe, unique filename and record failures.

This method applies when the target makes image URLs available in its HTML. Some pages populate content with JavaScript, require authentication, or use a site-specific endpoint. A selector copied from one site is not a general-purpose way to discover images on every site.

Install Python dependencies

The example uses Requests for HTTP requests and Beautiful Soup for HTML parsing. Install both in the Python environment you will use to run the script:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.
python -m pip install requests beautifulsoup4

Save the code below as bulk_image_downloader.py. It accepts a page URL, a CSS selector, an output folder, and a maximum number of images. The selector is deliberately configurable because it must be chosen for the target site’s actual HTML.

Runnable Python downloader

from __future__ import annotations

import argparse
import mimetypes
import re
import time
from pathlib import Path
from urllib.parse import unquote, urljoin, urlsplit

import requests
from bs4 import BeautifulSoup

CHUNK_SIZE = 64 * 1024
USER_AGENT = "BulkImageDownloader/1.0 (contact: [email protected])"


def safe_filename(image_url: str, content_type: str | None, index: int) -> str:
    """Create a simple filename from the URL, with an extension fallback."""
    raw_name = unquote(Path(urlsplit(image_url).path).name)
    name = re.sub(r"[^A-Za-z0-9._-]+", "_", raw_name).strip("._")
    if not name:
        name = f"image_{index:04d}"

    # If the URL has no usable suffix, use the response media type where possible.
    if not Path(name).suffix:
        media_type = (content_type or "").split(";", 1)[0].strip().lower()
        extension = mimetypes.guess_extension(media_type) or ".img"
        name += extension
    return name


def unique_path(folder: Path, filename: str) -> Path:
    """Avoid overwriting an existing file by adding a numeric suffix."""
    candidate = folder / filename
    stem, suffix = candidate.stem, candidate.suffix
    counter = 1
    while candidate.exists():
        candidate = folder / f"{stem}_{counter}{suffix}"
        counter += 1
    return candidate


def discover_image_urls(session: requests.Session, page_url: str, selector: str) -> list[str]:
    response = session.get(page_url, timeout=(10, 30))
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")
    found: list[str] = []

    for element in soup.select(selector):
        # For img tags, prefer the ordinary src attribute. Sites may use a
        # different lazy-loading attribute; adapt this to the page if needed.
        candidate = element.get("src") or element.get("data-src")
        # An anchor selector can also be used to collect links to image files.
        if not candidate and element.name == "a":
            candidate = element.get("href")
        if candidate:
            absolute_url = urljoin(response.url, candidate.strip())
            if absolute_url not in found:
                found.append(absolute_url)
    return found


def download_one(
    session: requests.Session,
    image_url: str,
    output_dir: Path,
    index: int,
) -> Path:
    with session.get(image_url, stream=True, timeout=(10, 60)) as response:
        response.raise_for_status()
        filename = safe_filename(image_url, response.headers.get("Content-Type"), index)
        destination = unique_path(output_dir, filename)
        # Write to a temporary file first. A failed or interrupted transfer
        # will not leave a partial file that looks complete.
        temporary = destination.with_name(destination.name + ".part")
        try:
            with temporary.open("wb") as output:
                for chunk in response.iter_content(chunk_size=CHUNK_SIZE):
                    if chunk:
                        output.write(chunk)
            temporary.replace(destination)
        except Exception:
            temporary.unlink(missing_ok=True)
            raise
    return destination


def main() -> int:
    parser = argparse.ArgumentParser(description="Download images exposed by a page's HTML.")
    parser.add_argument("page_url", help="Page that contains the images")
    parser.add_argument("--selector", default="img", help="CSS selector, for example 'article img'")
    parser.add_argument("--output", default="downloaded_images", help="Output folder")
    parser.add_argument("--limit", type=int, default=10, help="Maximum image URLs to attempt")
    parser.add_argument("--delay", type=float, default=1.0, help="Seconds between image requests")
    args = parser.parse_args()

    if args.limit < 1:
        parser.error("--limit must be at least 1")
    if args.delay < 0:
        parser.error("--delay cannot be negative")

    output_dir = Path(args.output)
    output_dir.mkdir(parents=True, exist_ok=True)

    headers = {"User-Agent": USER_AGENT}
    with requests.Session() as session:
        session.headers.update(headers)
        try:
            image_urls = discover_image_urls(session, args.page_url, args.selector)
        except requests.RequestException as exc:
            print(f"Could not fetch or parse the listing page: {exc}")
            return 1

        selected = image_urls[:args.limit]
        print(f"Found {len(image_urls)} image URL(s); attempting {len(selected)}.")
        successes = 0
        failures = 0
        for index, image_url in enumerate(selected, start=1):
            try:
                saved = download_one(session, image_url, output_dir, index)
                successes += 1
                print(f"OK   {image_url} -> {saved}")
            except requests.RequestException as exc:
                failures += 1
                print(f"FAIL {image_url}: {exc}")
            except OSError as exc:
                failures += 1
                print(f"FAIL {image_url}: file error: {exc}")

            if index < len(selected) and args.delay:
                time.sleep(args.delay)

    print(f"Finished: {successes} saved, {failures} failed.")
    return 0 if failures == 0 else 2


if __name__ == "__main__":
    raise SystemExit(main())

For example, if the target page has ordinary image tags inside an article, run:

python bulk_image_downloader.py "https://example.com/gallery" --selector "article img" --output images --limit 10 --delay 1

Replace the example URL and selector with values appropriate to the target. The script requests at most 10 image URLs by default and pauses one second between image requests. Those are conservative starting defaults, not a universal rule or a statement of any site’s limits. The official Automate the Boring Stuff with Python, 3rd Edition, uses a 10-download default and a one-second pause in its XKCD example for that tutorial’s context.

Rank #2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
  • Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Adapt discovery to the target page

Choose a selector based on the page’s markup

Start by inspecting the page HTML and identify the elements that point to the images you intend to retrieve. The default img selector finds image elements, while a narrower selector such as article img can avoid logos, avatars, and unrelated page artwork. If the image is linked rather than embedded, use a selector for the relevant anchors and the script will use their href values when no image source is present.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Resolve relative and lazy-loaded URLs

The script uses urljoin() so a value such as /media/photo.jpg becomes an absolute URL relative to the page’s final address. It checks src and data-src, two common locations, but sites may store lazy-loaded images in another attribute or in a srcset. Change the discovery function to read the field the target actually uses; do not assume every page uses the same convention.

When HTML is not enough

If a page’s images are inserted only after JavaScript runs, a basic HTTP request may not contain the image elements. First check whether the site documents a data endpoint or offers another permitted way to retrieve the image URLs. If it does not, browser rendering may be necessary. The cited tutorial demonstrates a known HTML layout; it does not establish how an unspecified site structures its pages.

Rank #3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
  • Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Why the implementation streams and checks responses

  • Binary output: image responses are written as bytes, not decoded as text, so the file contents are preserved.
  • Streaming: stream=True and iter_content() write chunks incrementally instead of holding an entire image in memory. Requests documents streaming downloads, sessions, connection pooling, timeouts, and response handling in its documentation.
  • Finite timeouts: the connect/read timeout tuples prevent an individual request from waiting indefinitely. A timeout is not a total job deadline; a large batch may take longer than any one request.
  • Status validation: raise_for_status() treats unsuccessful HTTP responses as failures instead of saving an error page under an image filename.
  • Temporary files: data is first written with a .part suffix, then renamed after a successful transfer. This reduces the chance that an interrupted download is mistaken for a complete image.
  • Collision protection: if a URL basename already exists, a numeric suffix is added instead of overwriting the earlier file. This is a practical collision safeguard, not a guarantee that two differently named files contain different images.

Requests or Python’s urllib

Requests is used above because its API provides sessions, streaming, timeouts, and response handling in one compact workflow. Python’s standard library also offers urllib.request, including URL opening, request headers, handlers, and file-like response objects; its HOWTO demonstrates copying a response stream to a temporary file, and the API documentation describes its interfaces. Choose Requests when its higher-level session and response API suits the project; choose urllib when keeping dependencies to the standard library matters. The cited documentation does not establish a performance winner between them.

Rate, storage, and operational safeguards

Begin with a small batch

Use --limit to validate the selector and output before scaling up. Increase the cap only after confirming that the target site allows the intended access and that your script is selecting the right images. Keep the delay configurable and avoid parallel requests until you understand the site’s documented limits and the effect on its service.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep a useful record of failures

The example reports each failed URL and continues to the next image, so one broken link does not abort the batch. For recurring jobs, write those results to a CSV or log file with the URL, error, and timestamp; then retry only transient failures under a bounded retry policy. Do not retry indefinitely: repeated failures may indicate a permanent denial, removed image, or incorrect URL.

Rank #4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
  • Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Plan for disk and filename edge cases

Ensure the output folder has enough free space, particularly for full-resolution originals. URL basenames can be absent, misleading, or shared across images, so the script creates a fallback name and avoids overwriting existing files. If preserving original filenames is important, store a manifest mapping each source URL to its saved path. For untrusted inputs or a larger application, also enforce an output-root boundary and validate downloaded content rather than trusting a filename or response header alone.

Check permission and site rules

Before running a bulk job, consult the target site’s terms, technical documentation, and applicable access restrictions. This generic workflow cannot establish whether a particular site permits automated retrieval, what rate it allows, whether authentication is needed, or whether you have rights to save or reuse the images. Those answers depend on the target and context.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

  • The script finds zero images: inspect the returned page HTML and test a selector that matches its structure. The page may use another lazy-loading attribute, require JavaScript, or return different markup to automated requests.
  • Images return 403 or 401: the server is denying the request or requires authorization. Check the site’s documented access method and permitted headers or authentication; do not try to bypass access controls.
  • Saved files are tiny or are not images: inspect the HTTP status and response content. A server may have returned an error or redirect page; keep raise_for_status() and verify the actual content before treating it as an image.
  • Some requests time out: connectivity, server response time, or a large file may be responsible. The script reports the item and continues; adjust finite timeout values only when appropriate, and avoid an unbounded wait.
  • Duplicate-looking images are saved under different names: distinct URLs may serve identical content. The example prevents filename overwrites but does not perform content hashing or deduplication; add a hash-based manifest if deduplication is a requirement.
  • Existing files remain after a crash: completed files are retained, while an interrupted transfer should leave at most a .part file that the exception handler removes. If a process is forcibly terminated, inspect and remove stale partial files before another run.

Or skip the browser setup

If the task is to capture a rendered page rather than download original image assets individually, ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. Its clean-shot options accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, this cURL call captures the page at Stripe and writes a WebP file:

Best Value
Sale
UnionSine 500GB Ultra Slim Portable External Hard Drive HDD-USB 3.0
  • [Upgraded Version] - This external hard drive features a mirrored logo stripe combined with a striped anti-slip design, and the rounded corners of the casing make it easier to grip. The stripes also have a heat dissipation function, ensuring stable and fast data transfer.
  • 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
  • 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
  • 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
  • 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month with no card.

Further reading

The official online Chapter 13, “Web Scraping,” of Al Sweigart’s Automate the Boring Stuff with Python, 3rd Edition, walks through an XKCD image-download example and an Image Site Downloader practice project. Its approach is a useful reference for the page-fetch, parse, download, and save sequence, while its selector and navigation logic remain specific to that example.

Quick Recap

SaleBestseller No. 1
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$129.99
Bestseller No. 2
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
Seagate Portable 5TB External Hard Drive HDD – USB 3.0 for PC, Mac, PS4, & Xbox - 1-Year Rescue Service (STGX5000400), Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$229.99
Bestseller No. 3
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.80
Bestseller No. 4
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
Seagate Portable 4TB External Hard Drive HDD – USB 3.0, 1-Year Rescue
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$189.99

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.