Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To extract images from an HTML file, parse every <img> and <picture> element, collect src, srcset, and <source> references, resolve relative URLs, decode data: images, and download or copy the resulting bytes. A static parser can inventory everything present in the saved markup. If JavaScript inserts images later, first capture the post-render DOM or network requests with a browser, then run the same extraction logic.
This guide shows a repeatable Python workflow for local files and remote pages, explains responsive-image selection, handles inline Base64 data, and includes fixes for malformed HTML, duplicate names, failed downloads, and JavaScript-rendered content.
Decide what “extract” means
There are three different outcomes people call image extraction:
- URL inventory: produce a list of image references without downloading them.
- Asset download: fetch remote resources or copy local files and preserve their original bytes.
- Conversion: decode or transform assets into another format. Conversion is a separate step and can change quality or metadata.
The procedure below downloads external images and writes inline data-URI images to disk. It does not claim that you have permission to reuse the images. Licensing, site terms, and local law still apply to each source.
#1 Best Overall
How image references appear in HTML
Standard img elements
An img normally has a src URL. It may also have srcset, which lists alternative versions for different viewport widths or pixel densities. Keep src as the fallback; a downloader that reads only one attribute can miss the higher-resolution alternatives.
picture art direction
picture groups one or more source elements, often with media or type conditions, followed by a fallback img. Collect every source[srcset] candidate and the fallback img. A static extractor records possibilities; it does not know which candidate a particular browser would select without evaluating the media and type conditions.
Inline data URIs
A src beginning with data: already contains the image bytes. Decode Base64 payloads locally instead of sending an HTTP request. Non-Base64 data payloads are percent-encoded text and need URL decoding before writing bytes.
Recommended Free Tools
Prepare the Python environment
The example uses Beautiful Soup for parsing, Requests for HTTP, and only standard-library modules for URL handling and decoding.
python -m pip install beautifulsoup4 requests
Beautiful Soup offers three useful parser choices: Python’s built-in html.parser (no extra parser dependency), lxml (usually faster when installed), and html5lib (browser-like error recovery). Invalid markup can produce different trees with different parsers, so switch parsers when a malformed document yields surprising results.
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Complete extractor for a saved HTML file or known page URL
Save this as extract_images.py. Set html_path to your file. If the file came from a known web page, replace base_url with that page’s real URL so relative references resolve correctly.
from pathlib import Path
from urllib.parse import urljoin, urlparse, unquote
from base64 import b64decode
import mimetypes
import re
import hashlib
import requests
from bs4 import BeautifulSoup
html_path = Path("page.html")
base_url = "https://example.com/articles/page.html"
out_dir = Path("extracted-images")
out_dir.mkdir(exist_ok=True)
html = html_path.read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
refs = []
for img in soup.find_all("img"):
src = img.get("src")
if src:
refs.append(src.strip())
srcset = img.get("srcset")
if srcset:
refs.extend(part.strip().split()[0] for part in srcset.split(",") if part.strip())
for source in soup.select("picture source[srcset]"):
srcset = source.get("srcset", "")
refs.extend(part.strip().split()[0] for part in srcset.split(",") if part.strip())
# Preserve order while removing exact duplicate references.
refs = list(dict.fromkeys(refs))
allowed_remote = {"http", "https"}
def safe_name(index, suffix, source):
digest = hashlib.sha256(source.encode("utf-8")).hexdigest()[:10]
return out_dir / f"image-{index:04d}-{digest}{suffix}"
def suffix_for(content_type=None, source=""):
if content_type:
value = content_type.split(";", 1)[0].strip().lower()
ext = mimetypes.guess_extension(value)
if ext:
return ext
path_suffix = Path(urlparse(source).path).suffix
return path_suffix if re.fullmatch(r"\.[A-Za-z0-9]{1,8}", path_suffix or "") else ".bin"
for index, ref in enumerate(refs, 1):
if ref.startswith("data:"):
header, payload = ref.split(",", 1)
media_type = header.split(";", 1)[0].split(":", 1)[1]
data = b64decode(payload) if ";base64" in header.lower() else unquote(payload).encode("utf-8")
target = safe_name(index, suffix_for(media_type, ""), ref)
target.write_bytes(data)
print(f"saved {target} (inline data)")
continue
absolute = urljoin(base_url, ref)
scheme = urlparse(absolute).scheme.lower()
if scheme not in allowed_remote:
print(f"skipped unsupported scheme: {absolute}")
continue
response = requests.get(absolute, timeout=30)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
# Do not trust a misleading filename extension; prefer the response type.
target = safe_name(index, suffix_for(content_type, absolute), absolute)
target.write_bytes(response.content)
print(f"saved {target} ({content_type or 'unknown type'})")
The script deduplicates identical references, creates collision-resistant names, rejects non-HTTP schemes, uses a timeout, and prefers the server’s Content-Type when choosing an extension. It still preserves the downloaded bytes exactly; it does not transcode them.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Resolve local paths versus web URLs
Self-contained local archives
If an HTML archive and its images are stored together, a relative reference such as images/logo.png should be resolved against html_path.parent, then copied from the filesystem. Do not issue an HTTP request for a path that is meant to be local. Add a local branch before the network branch:
from urllib.parse import unquote
candidate = (html_path.parent / unquote(ref)).resolve()
if candidate.is_file():
target.write_bytes(candidate.read_bytes())
For security, ensure the resolved path remains inside the archive directory before copying; otherwise a crafted ../ reference could read an unrelated local file.
Remote pages
Use the page’s final, canonical URL as base_url, not the URL of your script or working directory. urljoin then handles root-relative paths such as /media/a.webp, directory-relative paths, and absolute URLs. A page downloaded after a redirect may require the response’s final URL as the base.
Rank #3
Parse responsive candidates correctly
Each comma-separated srcset item consists of a URL followed by an optional descriptor such as 480w or 2x. The extraction code stores the first token—the URL—and intentionally keeps every candidate. If you need the exact image a browser would display, you must additionally evaluate viewport width, device pixel ratio, the sizes attribute, and any picture media/type conditions. Without that context, selecting the first or largest candidate is only a policy choice, not a browser-equivalent result.
Free tools Windows power users keep installed
One-click scans. No signup required.
Handle Base64 and other data URLs
Base64 data URLs have a header such as data:image/png;base64, followed by encoded bytes. The script writes those bytes directly. A data URL without ;base64 contains URL-encoded data; the script decodes it as UTF-8 bytes. For very large inline assets, stream or validate the declared media type before writing, and impose a size limit if the HTML is untrusted.
Why a static parser misses JavaScript images
Python’s html.parser exposes attributes in the markup but returns the contents of script and style without parsing them as HTML. An image created only after JavaScript runs therefore will not appear in the original file’s img list.
Render first, then extract
- Open the page in a browser automation tool.
- Wait for the application’s content or network activity to finish.
- Save the post-render DOM (the document after scripts have modified it), or record image requests from the browser’s network log.
- Run the same
src,srcset,picture, and data-URI extraction against that rendered HTML.
Rendering is also necessary when images are represented only by CSS background properties or when a click, scroll, consent action, or lazy-load threshold causes the URL to be inserted. A browser can expose the computed style or network request; a static HTML parser cannot infer it.
Parser and download decisions
| Situation | Recommended approach | Trade-off |
|---|---|---|
| Simple, well-formed file | html.parser |
No parser dependency; less browser-like recovery |
| Large batch where speed matters | lxml |
Fast, but requires an additional dependency |
| Malformed markup requiring browser-style recovery | html5lib |
More forgiving, generally slower |
| Images present only after scripts run | Render with a browser, then parse | More setup and resource use, but sees the post-render page |
| Need original bytes | Download/copy without conversion | Output format and metadata remain source-dependent |
Troubleshooting common failures
Nothing is found
Inspect the file for img, picture, and source elements. If the page is an application shell, render it first. Also check whether the “images” are CSS backgrounds, SVG markup, or URLs embedded in script data rather than HTML attributes.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallRank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Relative URLs return 404
Set base_url to the actual page URL, including its directory path. A base of https://example.com resolves a relative reference differently from https://example.com/articles/page.html.
Downloads are HTML error pages
Some servers return a login page, bot challenge, or error document with a successful HTTP status. Inspect Content-Type and the first bytes before saving. Supply required cookies, authorization, or a user agent only when you are allowed to access the resource.
Duplicate files overwrite one another
Different URLs can end with the same filename. The sample uses a short SHA-256 digest of each source in the filename, preventing accidental overwrites while retaining a traceable mapping.
Malformed HTML gives inconsistent results
Try lxml or html5lib and compare the parsed tree. Beautiful Soup notes that parser choice changes how invalid markup is repaired.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsCharacter or Base64 errors occur
Read HTML using the document’s declared encoding when known. For data URLs, split only at the first comma, detect ;base64 case-insensitively, and reject malformed payloads rather than silently writing corrupt files.
Best Value
The process is slow or times out
Use a finite request timeout, process URLs in batches, and avoid downloading the same reference twice. For large collections, stream responses to disk and limit concurrency to what the origin permits. Caching can reduce repeated requests, but honor cache headers and site policies.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It can render a URL and return a PNG, JPEG, WebP, or PDF, which is useful when your goal is a visual capture rather than downloading each original image asset. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
Use the documented parameters and options in the ScreenshotNeo documentation. A one-call capture with cURL is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const buffer = Buffer.from(await res.arrayBuffer());
Relevant capture controls include full-page screenshots with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, click or wait conditions, ad/tracker/request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, a chosen cache TTL, signed public image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, PDF output, and usage reporting. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; all features are available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account to start.
Safety, permissions, and reproducibility
- Extracting bytes does not grant redistribution rights; verify the image license and the source site’s terms.
- Do not send credentials, private URLs, or personal data to a third-party renderer unless your organization permits it.
- Record the original reference, resolved URL, response status, content type, and retrieval time alongside each file so the extraction can be audited.
- Respect robots directives, rate limits, authentication boundaries, and applicable law when downloading remote resources.
Frequently Asked Questions
Can I extract images embedded as SVG markup rather than an image URL?
Yes, but that is a different path: locate the inline <svg> element, serialize it, and save it as an SVG file or rasterize it with a separate image-conversion tool. The src/srcset workflow handles image references, not inline SVG serialization.
How can I tell whether two different URLs contain the same image?
Compare the downloaded bytes, for example with SHA-256 hashes. URL equality alone cannot prove content equality, because different URLs may return identical bytes and one URL may change over time.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Should I choose the largest srcset candidate?
Only if that is your stated policy. The browser’s choice depends on viewport, device pixel ratio, sizes, and picture conditions; a static extractor cannot reproduce that choice without those inputs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

