To scrape images from a web page, request the page HTML, parse its <img> elements, resolve each image reference to an absolute URL, then download and save the bytes with a validated extension. The basic approach works for server-rendered pages. JavaScript galleries, consent overlays and access controls require a permitted rendered-page or official API workflow instead.
What you need before downloading anything
Use Python 3, a network connection and a destination directory with enough space for the files. The third-party example below uses requests and beautifulsoup4:
python -m pip install requests beautifulsoup4
You can use only the standard library (urllib.request, urllib.parse, urllib.robotparser) when installing packages is undesirable. Requests is generally easier to read and gives convenient status, header and timeout handling.
Check permission and site rules first
Read the target site’s terms and robots.txt, honor published crawl delays and rate limits, and stop if automated retrieval is disallowed. Do not bypass a login, CAPTCHA, bot check or other explicit access restriction. Collecting an image for private analysis is legally different from republishing it; copyright, licenses and attribution requirements still apply.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
A robust one-page image scraper
This script handles relative URLs, lazy-loading attributes, duplicates, redirects, non-image responses and deterministic names. It also limits response size and pauses between downloads so a small job does not become an uncontrolled crawler.
from pathlib import Path
from urllib.parse import urljoin, urlparse
import hashlib
import mimetypes
import time
import requests
from bs4 import BeautifulSoup
PAGE_URL = "https://example.com/gallery"
OUT_DIR = Path("images")
MAX_IMAGE_BYTES = 20 * 1024 * 1024
REQUEST_TIMEOUT = (10, 30) # connect, read seconds
USER_AGENT = "image-research-bot/1.0 (+https://example.com/contact)"
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
page = session.get(PAGE_URL, timeout=REQUEST_TIMEOUT, allow_redirects=True)
page.raise_for_status()
soup = BeautifulSoup(page.content, "html.parser")
# Keep the final page URL as the base when redirects changed the address.
base_url = page.url
candidates = []
for tag in soup.select("img"):
# Sites commonly put the real file in data-src/data-lazy-src and a
# thumbnail or placeholder in src.
raw = (tag.get("data-src") or tag.get("data-lazy-src") or
tag.get("data-original") or tag.get("src"))
if raw:
candidates.append(raw)
# Also inspect srcset, choosing the largest declared candidate.
for tag in soup.select("img[srcset], source[srcset]"):
srcset = tag.get("srcset", "")
entries = []
for item in srcset.split(","):
parts = item.strip().split()
if parts:
width = 0
if len(parts) > 1 and parts[1].endswith("w"):
try:
width = int(parts[1][:-1])
except ValueError:
pass
entries.append((width, parts[0]))
if entries:
candidates.append(max(entries)[1])
OUT_DIR.mkdir(parents=True, exist_ok=True)
seen = set()
saved = 0
for raw in candidates:
image_url = urljoin(base_url, raw)
parsed = urlparse(image_url)
if parsed.scheme not in {"http", "https"} or image_url in seen:
continue
seen.add(image_url)
try:
with session.get(image_url, timeout=REQUEST_TIMEOUT, stream=True,
allow_redirects=True) as response:
response.raise_for_status()
content_type = response.headers.get("content-type", "").split(";", 1)[0].lower()
if not content_type.startswith("image/"):
print("Skipping non-image:", image_url, content_type)
continue
length = response.headers.get("content-length")
if length and int(length) > MAX_IMAGE_BYTES:
print("Skipping oversized image:", image_url)
continue
extension = mimetypes.guess_extension(content_type) or ".bin"
# Hash the URL so repeated runs produce stable, collision-resistant names.
stem = hashlib.sha256(image_url.encode("utf-8")).hexdigest()[:16]
destination = OUT_DIR / f"{stem}{extension}"
total = 0
with destination.open("wb") as output:
for chunk in response.iter_content(chunk_size=64 * 1024):
if not chunk:
continue
total += len(chunk)
if total > MAX_IMAGE_BYTES:
raise ValueError("image exceeded size limit")
output.write(chunk)
saved += 1
print(f"Saved {destination} ({total:,} bytes)")
except (requests.RequestException, ValueError) as error:
print(f"Failed {image_url}: {error}")
time.sleep(0.25)
print(f"Saved {saved} unique images")
Run it after replacing PAGE_URL. The script reads the response bytes rather than assuming the page is UTF-8, and it writes image bytes in binary mode. A server can return an HTTP success status for an HTML error page, which is why the content-type check matters.
How image URLs are represented in HTML
src and relative paths
An image may be written as /media/photo.jpg, ../assets/photo.webp or a complete HTTPS URL. urljoin(page_url, raw) applies the document’s base URL correctly, including pages reached through redirects.
Lazy loading attributes
For performance, a page may put a tiny placeholder in src and the actual file in data-src, data-lazy-src, data-original or a site-specific attribute. Inspect the site’s markup and add the relevant attributes; there is no universal lazy-loading name.
Free tools Windows power users keep installed
One-click scans. No signup required.
srcset and the full-resolution choice
srcset can list several widths, such as a 400-pixel thumbnail and a 1600-pixel file. The example selects the candidate with the largest declared w descriptor. That is a useful approximation, not a guarantee of the original: some sites omit descriptors, use pixel-density (2x) values, or generate signed URLs that expire.
Background images and linked originals
CSS background-image values and image URLs inside JSON are not <img> tags. If the site’s markup exposes them in inline styles or a documented data endpoint, parse those sources explicitly. A thumbnail may also be wrapped in an <a> element pointing to the original; follow that link only when the site’s structure and terms make the relationship clear.
Rank #2
Saving files safely and correctly
Never derive a filename directly from an untrusted URL path. Query strings can contain characters that are invalid on some operating systems, and two different URLs can map to the same basename. A counter (image_0001) is simple; a short hash of the canonical URL is stable across runs. Keep a CSV or database containing the source URL, final URL, timestamp, status, content type and local filename if you need provenance.
Use the HTTP Content-Type header to select an initial extension, but treat it as a hint. Servers occasionally mislabel files, and formats such as SVG can contain active markup. For higher assurance, inspect the file signature and decode it with an image library such as Pillow before presenting or transforming it. Reject unexpected formats and enforce a byte limit while streaming, as the example does.
When a plain parser finds no images
The page is JavaScript-rendered
Requests and urllib receive the server’s initial HTML; they do not execute the browser’s JavaScript. If the response contains an empty gallery shell, inspect network requests in a browser’s developer tools and use an authorized JSON endpoint or export. Where permitted, use a browser-rendering tool that waits for the gallery to load, then capture the rendered DOM. Do not use rendering to evade anti-bot controls.
The image is behind authentication
Public HTML does not grant access to a private image. Use an account and API that you are authorized to use, pass the required cookies or bearer token securely, and never hard-code credentials in a script or commit them to source control.
Consent banners, popups and chat widgets obscure the page
Those overlays usually do not remove the underlying URL from HTML, but they can block a browser-rendered capture. Accept or configure consent according to the site’s policy, or use an authorized capture service that handles overlays before rendering.
Scaling from one page to a reusable crawler
A production collector needs more than a loop. Add the following deliberately:
Recommended Free Tools
- URL queue and scope: normalize URLs, restrict hosts and paths, and enforce a maximum page count.
- Deduplication: track both page URLs and image URLs; consider fragment removal and clearly documented query normalization.
- Retries: retry transient 429 and 5xx responses with exponential backoff and respect
Retry-After. Do not repeatedly retry 4xx permission failures. - Rate limiting: use a per-host delay, a small connection pool and a descriptive user agent.
- Caching: store response metadata and conditional request headers so reruns do not download unchanged files.
- Observability: log redirect destinations, status, content type, byte count and failure reason; retain enough metadata to reproduce a decision.
- Shutdown safety: write to a temporary file and rename it only after a complete download, so an interrupted transfer is not mistaken for a valid image.
For a large collection, persist the queue and metadata rather than keeping them only in memory. Separate discovery, download and validation stages so a failed file can be retried without crawling every page again.
Standard-library alternative with urllib
urllib.request.urlopen() returns a response whose bytes can be read and saved, while urllib.parse.urljoin() resolves links. This small example is suitable for a controlled, one-off page:
from pathlib import Path
from urllib.parse import urljoin
from urllib.request import Request, urlopen
from bs4 import BeautifulSoup
page_url = "https://example.com/gallery"
request = Request(page_url, headers={"User-Agent": "image-research-bot/1.0"})
with urlopen(request, timeout=20) as response:
html = response.read()
final_url = response.geturl()
soup = BeautifulSoup(html, "html.parser")
out = Path("images"); out.mkdir(exist_ok=True)
for number, tag in enumerate(soup.select("img"), 1):
raw = tag.get("data-src") or tag.get("src")
if not raw:
continue
image_url = urljoin(final_url, raw)
image_request = Request(image_url, headers={"User-Agent": "image-research-bot/1.0"})
with urlopen(image_request, timeout=20) as image_response:
content_type = image_response.headers.get_content_type()
if not content_type.startswith("image/"):
continue
data = image_response.read()
(out / f"image_{number:04d}.bin").write_bytes(data)
For real jobs, add the same size, extension, retry, validation and policy checks as the Requests version. urllib.robotparser can read a site’s robots.txt rules before you enqueue URLs.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server when you need a rendered result rather than writing and maintaining browser automation. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; failed loads, bot checks or CAPTCHAs and blank pages are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For a screenshot of a rendered page, make one GET request (replace the URL with the page you are authorized to access):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request in Python is:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
See the ScreenshotNeo documentation for all 63 options: full-page and selector captures, 12 device presets or custom viewports, dark mode, retina scale, lazy-image loading, PDF paper and page-range controls, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs, usage data and the OpenAPI specification. Existing parameter names used by other screenshot APIs also work for easier migration.
The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing provides two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to begin.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
403, 429 or repeated timeouts
A 403 usually indicates permission or policy enforcement; a 429 means you are sending requests too quickly. Stop, check the site’s rules, slow down and use an official API or export. Increase a timeout only when the site is legitimately slow; it does not fix blocked access.
Every downloaded file is HTML
Print the final URL, status and Content-Type. You may have been redirected to a login page, an error document or a bot challenge. Do not save it with an image extension; resolve authorization or choose a permitted source.
Only thumbnails are saved
Inspect srcset, lazy attributes, anchor links and network requests for the original. The largest srcset candidate is not necessarily the source file, and some originals require a signed request.
Duplicate or corrupted files
Deduplicate canonical URLs, stream to a temporary path, enforce a size limit and validate the file signature or decode it before renaming. Preserve the response metadata so you can identify a server that changed content under one URL.
Beautiful Soup returns zero <img> tags
Log the first part of the response and its final URL. You may have fetched a consent or login page, a JavaScript shell, or a non-HTML response. If the images appear only after script execution, use the site’s authorized endpoint or a permitted rendered-page workflow.
FAQ
Can I scrape any image I can see in a browser?
No. Visibility does not grant permission to automate collection or redistribution. Follow the site’s terms, robots guidance, authentication boundary and the image’s license.
Best Value
Should I choose Requests or urllib?
Choose urllib when avoiding dependencies is important; choose Requests for a clearer session, timeout and streaming interface. Both still receive only the HTML the server sends.
Is a URL ending in .jpg always a JPEG?
No. Query parameters, redirects and incorrect server headers are common. Validate the response and, when correctness matters, inspect the file bytes with an image decoder.
Frequently Asked Questions
Can I scrape any image I can see in a browser?
No. Visibility does not grant permission to automate collection or redistribution. Follow the site’s terms, robots guidance, authentication boundary and the image’s license.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Should I choose Requests or urllib?
Choose urllib when avoiding dependencies is important; choose Requests for a clearer session, timeout and streaming interface. Both still receive only the HTML the server sends.
Is a URL ending in .jpg always a JPEG?
No. Query parameters, redirects and incorrect server headers are common. Validate the response and, when correctness matters, inspect the file bytes with an image decoder.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




