Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To extract images from an HTML file, parse every <img> and <picture> element, collect src, srcset, and <source> references, resolve relative URLs, decode data: images, and download or copy the resulting bytes. A static parser can inventory everything present in the saved markup. If JavaScript inserts images later, first capture the post-render DOM or network requests with a browser, then run the same extraction logic.

This guide shows a repeatable Python workflow for local files and remote pages, explains responsive-image selection, handles inline Base64 data, and includes fixes for malformed HTML, duplicate names, failed downloads, and JavaScript-rendered content.

Decide what “extract” means

There are three different outcomes people call image extraction:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • URL inventory: produce a list of image references without downloading them.
  • Asset download: fetch remote resources or copy local files and preserve their original bytes.
  • Conversion: decode or transform assets into another format. Conversion is a separate step and can change quality or metadata.

The procedure below downloads external images and writes inline data-URI images to disk. It does not claim that you have permission to reuse the images. Licensing, site terms, and local law still apply to each source.

How image references appear in HTML

Standard img elements

An img normally has a src URL. It may also have srcset, which lists alternative versions for different viewport widths or pixel densities. Keep src as the fallback; a downloader that reads only one attribute can miss the higher-resolution alternatives.

picture art direction

picture groups one or more source elements, often with media or type conditions, followed by a fallback img. Collect every source[srcset] candidate and the fallback img. A static extractor records possibilities; it does not know which candidate a particular browser would select without evaluating the media and type conditions.

Inline data URIs

A src beginning with data: already contains the image bytes. Decode Base64 payloads locally instead of sending an HTTP request. Non-Base64 data payloads are percent-encoded text and need URL decoding before writing bytes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prepare the Python environment

The example uses Beautiful Soup for parsing, Requests for HTTP, and only standard-library modules for URL handling and decoding.

python -m pip install beautifulsoup4 requests

Beautiful Soup offers three useful parser choices: Python’s built-in html.parser (no extra parser dependency), lxml (usually faster when installed), and html5lib (browser-like error recovery). Invalid markup can produce different trees with different parsers, so switch parsers when a malformed document yields surprising results.

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Complete extractor for a saved HTML file or known page URL

Save this as extract_images.py. Set html_path to your file. If the file came from a known web page, replace base_url with that page’s real URL so relative references resolve correctly.

from pathlib import Path
from urllib.parse import urljoin, urlparse, unquote
from base64 import b64decode
import mimetypes
import re
import hashlib
import requests
from bs4 import BeautifulSoup

html_path = Path("page.html")
base_url = "https://example.com/articles/page.html"
out_dir = Path("extracted-images")
out_dir.mkdir(exist_ok=True)

html = html_path.read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")

refs = []
for img in soup.find_all("img"):
    src = img.get("src")
    if src:
        refs.append(src.strip())
    srcset = img.get("srcset")
    if srcset:
        refs.extend(part.strip().split()[0] for part in srcset.split(",") if part.strip())
for source in soup.select("picture source[srcset]"):
    srcset = source.get("srcset", "")
    refs.extend(part.strip().split()[0] for part in srcset.split(",") if part.strip())

# Preserve order while removing exact duplicate references.
refs = list(dict.fromkeys(refs))

allowed_remote = {"http", "https"}
def safe_name(index, suffix, source):
    digest = hashlib.sha256(source.encode("utf-8")).hexdigest()[:10]
    return out_dir / f"image-{index:04d}-{digest}{suffix}"

def suffix_for(content_type=None, source=""):
    if content_type:
        value = content_type.split(";", 1)[0].strip().lower()
        ext = mimetypes.guess_extension(value)
        if ext:
            return ext
    path_suffix = Path(urlparse(source).path).suffix
    return path_suffix if re.fullmatch(r"\.[A-Za-z0-9]{1,8}", path_suffix or "") else ".bin"

for index, ref in enumerate(refs, 1):
    if ref.startswith("data:"):
        header, payload = ref.split(",", 1)
        media_type = header.split(";", 1)[0].split(":", 1)[1]
        data = b64decode(payload) if ";base64" in header.lower() else unquote(payload).encode("utf-8")
        target = safe_name(index, suffix_for(media_type, ""), ref)
        target.write_bytes(data)
        print(f"saved {target} (inline data)")
        continue

    absolute = urljoin(base_url, ref)
    scheme = urlparse(absolute).scheme.lower()
    if scheme not in allowed_remote:
        print(f"skipped unsupported scheme: {absolute}")
        continue
    response = requests.get(absolute, timeout=30)
    response.raise_for_status()
    content_type = response.headers.get("Content-Type", "")
    # Do not trust a misleading filename extension; prefer the response type.
    target = safe_name(index, suffix_for(content_type, absolute), absolute)
    target.write_bytes(response.content)
    print(f"saved {target} ({content_type or 'unknown type'})")

The script deduplicates identical references, creates collision-resistant names, rejects non-HTTP schemes, uses a timeout, and prefers the server’s Content-Type when choosing an extension. It still preserves the downloaded bytes exactly; it does not transcode them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Resolve local paths versus web URLs

Self-contained local archives

If an HTML archive and its images are stored together, a relative reference such as images/logo.png should be resolved against html_path.parent, then copied from the filesystem. Do not issue an HTTP request for a path that is meant to be local. Add a local branch before the network branch:

from urllib.parse import unquote

candidate = (html_path.parent / unquote(ref)).resolve()
if candidate.is_file():
    target.write_bytes(candidate.read_bytes())

For security, ensure the resolved path remains inside the archive directory before copying; otherwise a crafted ../ reference could read an unrelated local file.

Remote pages

Use the page’s final, canonical URL as base_url, not the URL of your script or working directory. urljoin then handles root-relative paths such as /media/a.webp, directory-relative paths, and absolute URLs. A page downloaded after a redirect may require the response’s final URL as the base.

Parse responsive candidates correctly

Each comma-separated srcset item consists of a URL followed by an optional descriptor such as 480w or 2x. The extraction code stores the first token—the URL—and intentionally keeps every candidate. If you need the exact image a browser would display, you must additionally evaluate viewport width, device pixel ratio, the sizes attribute, and any picture media/type conditions. Without that context, selecting the first or largest candidate is only a policy choice, not a browser-equivalent result.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle Base64 and other data URLs

Base64 data URLs have a header such as data:image/png;base64, followed by encoded bytes. The script writes those bytes directly. A data URL without ;base64 contains URL-encoded data; the script decodes it as UTF-8 bytes. For very large inline assets, stream or validate the declared media type before writing, and impose a size limit if the HTML is untrusted.

Why a static parser misses JavaScript images

Python’s html.parser exposes attributes in the markup but returns the contents of script and style without parsing them as HTML. An image created only after JavaScript runs therefore will not appear in the original file’s img list.

Render first, then extract

  1. Open the page in a browser automation tool.
  2. Wait for the application’s content or network activity to finish.
  3. Save the post-render DOM (the document after scripts have modified it), or record image requests from the browser’s network log.
  4. Run the same src, srcset, picture, and data-URI extraction against that rendered HTML.

Rendering is also necessary when images are represented only by CSS background properties or when a click, scroll, consent action, or lazy-load threshold causes the URL to be inserted. A browser can expose the computed style or network request; a static HTML parser cannot infer it.

Parser and download decisions

Situation Recommended approach Trade-off
Simple, well-formed file html.parser No parser dependency; less browser-like recovery
Large batch where speed matters lxml Fast, but requires an additional dependency
Malformed markup requiring browser-style recovery html5lib More forgiving, generally slower
Images present only after scripts run Render with a browser, then parse More setup and resource use, but sees the post-render page
Need original bytes Download/copy without conversion Output format and metadata remain source-dependent

Troubleshooting common failures

Nothing is found

Inspect the file for img, picture, and source elements. If the page is an application shell, render it first. Also check whether the “images” are CSS backgrounds, SVG markup, or URLs embedded in script data rather than HTML attributes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

Relative URLs return 404

Set base_url to the actual page URL, including its directory path. A base of https://example.com resolves a relative reference differently from https://example.com/articles/page.html.

Downloads are HTML error pages

Some servers return a login page, bot challenge, or error document with a successful HTTP status. Inspect Content-Type and the first bytes before saving. Supply required cookies, authorization, or a user agent only when you are allowed to access the resource.

Duplicate files overwrite one another

Different URLs can end with the same filename. The sample uses a short SHA-256 digest of each source in the filename, preventing accidental overwrites while retaining a traceable mapping.

Malformed HTML gives inconsistent results

Try lxml or html5lib and compare the parsed tree. Beautiful Soup notes that parser choice changes how invalid markup is repaired.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Character or Base64 errors occur

Read HTML using the document’s declared encoding when known. For data URLs, split only at the first comma, detect ;base64 case-insensitively, and reject malformed payloads rather than silently writing corrupt files.

The process is slow or times out

Use a finite request timeout, process URLs in batches, and avoid downloading the same reference twice. For large collections, stream responses to disk and limit concurrency to what the origin permits. Caching can reduce repeated requests, but honor cache headers and site policies.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It can render a URL and return a PNG, JPEG, WebP, or PDF, which is useful when your goal is a visual capture rather than downloading each original image asset. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

Use the documented parameters and options in the ScreenshotNeo documentation. A one-call capture with cURL is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const buffer = Buffer.from(await res.arrayBuffer());

Relevant capture controls include full-page screenshots with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, click or wait conditions, ad/tracker/request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, a chosen cache TTL, signed public image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, PDF output, and usage reporting. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; all features are available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account to start.

Safety, permissions, and reproducibility

  • Extracting bytes does not grant redistribution rights; verify the image license and the source site’s terms.
  • Do not send credentials, private URLs, or personal data to a third-party renderer unless your organization permits it.
  • Record the original reference, resolved URL, response status, content type, and retrieval time alongside each file so the extraction can be audited.
  • Respect robots directives, rate limits, authentication boundaries, and applicable law when downloading remote resources.

Frequently Asked Questions

Can I extract images embedded as SVG markup rather than an image URL?

Yes, but that is a different path: locate the inline <svg> element, serialize it, and save it as an SVG file or rasterize it with a separate image-conversion tool. The src/srcset workflow handles image references, not inline SVG serialization.

How can I tell whether two different URLs contain the same image?

Compare the downloaded bytes, for example with SHA-256 hashes. URL equality alone cannot prove content equality, because different URLs may return identical bytes and one URL may change over time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I choose the largest srcset candidate?

Only if that is your stated policy. The browser’s choice depends on viewport, device pixel ratio, sizes, and picture conditions; a static extractor cannot reproduce that choice without those inputs.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.