Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteFor a static page, image scraping is a short pipeline: request the HTML, parse its <img> elements, read candidate URL attributes, resolve relative paths against the page URL, filter out irrelevant assets, and download only what the site permits you to collect. The Python example below implements that workflow with timeouts, duplicate removal, safe filenames, HTTP-status checks and a modest delay.
Choose the supported data source first
Before writing a scraper, check whether the site offers an API, feed, export or other documented interface. A supported interface is usually more stable than parsing presentation HTML and may provide image metadata, licensing information and pagination directly.
If no suitable interface exists, inspect the target host’s /robots.txt, terms and access instructions. The Robots Exclusion Protocol (RFC 9309) describes rules that crawlers are requested to honor; its introduction also says, “These rules are not a form of access authorization.” Google describes robots.txt as a way to tell search-engine crawlers which URLs they can access, not as a security mechanism. An allow rule is not a copyright license, and a disallow rule is not the only legal consideration.
- Confirm that the pages and data are public and do not expose personal or confidential information.
- Keep request rates low, add pauses for large collections and avoid parallel bursts that could burden the host.
- Record the source URL and, where available, the site’s license or attribution requirements.
Static HTML versus a rendered page
| Approach | Best fit | Trade-off |
|---|---|---|
| HTTP fetch plus an HTML parser such as Beautiful Soup | The delivered HTML already contains image URLs | Lightweight and direct, but it does not run client-side JavaScript |
| Browser-rendered extraction | Images appear only after scripts run or after scrolling/interacting | Can observe rendered state, but browser automation needs its own current, site-specific setup and controls |
View the page source or inspect the response body before choosing. If the image markup is absent from the initial response, a parser alone cannot discover it. Deferred loading, client-side rendering and interaction can require a browser-based workflow; do not assume that a successful browser view means the same HTML was sent by the server.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11#1 Best Overall
Install the Python dependencies
The example uses the widely used requests HTTP client and Beautiful Soup’s bs4 parser.
python -m pip install requests beautifulsoup4
Save the script as scrape_images.py. It is intended for pages whose image URLs are present in HTML. Review and adapt the filtering rule for each site rather than downloading every asset indiscriminately.
Complete Python scraper for image URLs and files
from __future__ import annotations
import hashlib
import mimetypes
import re
import time
from pathlib import Path
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
PAGE_URL = "https://example.com/gallery"
OUTPUT_DIR = Path("images")
REQUEST_DELAY_SECONDS = 0.5
TIMEOUT_SECONDS = 30
# Keep this narrow for the target site. An empty tuple accepts every img URL.
ALLOWED_HOSTS = {urlparse(PAGE_URL).netloc}
def safe_name(image_url: str, content_type: str | None) -> str:
"""Create a collision-resistant filename without trusting the URL path."""
path_name = Path(urlparse(image_url).path).name
path_name = re.sub(r"[^A-Za-z0-9._-]", "_", path_name)
stem = Path(path_name).stem or "image"
suffix = Path(path_name).suffix.lower()
if suffix not in {".jpg", ".jpeg", ".png", ".gif", ".webp", ".svg", ".avif"}:
guessed = mimetypes.guess_extension((content_type or "").split(";", 1)[0])
suffix = guessed or ".bin"
digest = hashlib.sha256(image_url.encode("utf-8")).hexdigest()[:12]
return f"{stem}-{digest}{suffix}"
def scrape_images(page_url: str) -> list[tuple[str, Path]]:
session = requests.Session()
session.headers.update({
"User-Agent": "image-collector/1.0 (contact: [email protected])"
})
page_response = session.get(page_url, timeout=TIMEOUT_SECONDS)
page_response.raise_for_status()
soup = BeautifulSoup(page_response.text, "html.parser")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
seen: set[str] = set()
saved: list[tuple[str, Path]] = []
for tag in soup.find_all("img"):
raw_url = tag.get("src")
if not raw_url:
continue
image_url = urljoin(page_url, raw_url.strip())
parsed = urlparse(image_url)
if parsed.scheme not in {"http", "https"}:
continue
if ALLOWED_HOSTS and parsed.netloc not in ALLOWED_HOSTS:
continue
if image_url in seen:
continue
seen.add(image_url)
# Replace this with page-specific checks, such as a CSS class or path.
alt = (tag.get("alt") or "").lower()
classes = " ".join(tag.get("class", [])).lower()
if "logo" in alt or "logo" in classes or "avatar" in classes:
continue
try:
response = session.get(image_url, timeout=TIMEOUT_SECONDS, stream=True)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if not content_type.startswith("image/"):
response.close()
continue
filename = safe_name(image_url, content_type)
destination = OUTPUT_DIR / filename
with destination.open("wb") as output:
for chunk in response.iter_content(chunk_size=64 * 1024):
if chunk:
output.write(chunk)
response.close()
saved.append((image_url, destination))
print(f"saved {image_url} -> {destination}")
except requests.RequestException as error:
print(f"skipped {image_url}: {error}")
time.sleep(REQUEST_DELAY_SECONDS)
return saved
if __name__ == "__main__":
results = scrape_images(PAGE_URL)
print(f"Downloaded {len(results)} image(s)")
Run it with:
python scrape_images.py
The script first fetches the page, then finds img tags and reads src. urljoin turns values such as ../images/photo.jpg into an absolute URL. It rejects non-HTTP schemes, cross-host URLs (unless you change ALLOWED_HOSTS), duplicates and responses whose content type is not an image. The filename combines a cleaned path component with a hash of the URL, so two different URLs are less likely to overwrite one another.
Extract URLs without downloading files
If you only need a manifest, remove the download block and write the resolved URLs to a text or CSV file. This is useful for review, deduplication and rights checks before collecting binary data.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #2
import csv
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
page_url = "https://example.com/gallery"
html = requests.get(page_url, timeout=30).text
soup = BeautifulSoup(html, "html.parser")
urls = []
for image in soup.find_all("img"):
value = image.get("src")
if value:
urls.append(urljoin(page_url, value.strip()))
with open("image_urls.csv", "w", newline="", encoding="utf-8") as file:
writer = csv.writer(file)
writer.writerow(["image_url"])
writer.writerows(sorted(set(urls)))
Review this list for logos, tracking pixels, placeholders, unrelated page art and externally hosted assets. A page can contain many img elements that are not the photographs you want.
Relative URLs, redirects and host safety
HTML may contain absolute URLs, root-relative paths such as /media/a.jpg, or document-relative paths such as ../media/a.jpg. Python’s urllib.parse.urljoin resolves these against the page URL. An important edge case is that an absolute second argument replaces the base host; therefore validate the final parsed hostname before requesting it. Decide explicitly whether a CDN or image host is allowed instead of silently following every destination.
HTTP redirects can also move a request to another host. For collections with strict boundaries, inspect response.url after the request and reject destinations outside your approved host set. Do not send credentials or private cookies to an unexpected host.
Filtering modern and irrelevant image markup
Separate content images from page furniture
Use surrounding context: a gallery container, a known CSS class, an alt pattern, or a URL path used by the site’s media system. Exclude logos, icons, spacer images and avatars where they are not part of the collection. Test the selector on a small sample before scaling up.
Free tools Windows power users keep installed
One-click scans. No signup required.
When src is absent
Some pages put a deferred URL in another attribute or create the element after JavaScript runs. The static workflow can only extract what the retrieved HTML contains. Responsive markup may expose multiple candidate sizes; choose a policy for which variant to keep and verify the site’s current HTML documentation before relying on a particular attribute. The evidence here does not establish a universal recipe for every responsive-image or JavaScript framework.
When a browser is required
If “view source” has no image URL but the rendered page displays images, use a browser-rendered process appropriate to the site, or prefer an official API. Account for consent dialogs, lazy loading, scrolling, authentication and rate limits. Do not present a browser result as proof that the static scraper succeeded.
Reliability and performance practices
- Bound every request. Use connect/read timeouts and catch HTTP and network exceptions.
- Check status and media type. A URL ending in
.jpgcan return an error page; inspect the status andContent-Type. - Stream large files. Iterating over response chunks avoids loading an entire image into memory.
- Deduplicate. Keep a set of resolved URLs; some pages repeat the same image in several components.
- Pause. A delay between image requests reduces load. For a large job, checkpoint the manifest so an interruption does not restart everything.
- Log decisions. Record skipped URLs and reasons (non-image response, timeout, duplicate or policy filter) for later review.
- Do not invent throughput expectations. Actual speed depends on the target server, image sizes, network and your delay settings.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Zero images found | Images are injected by JavaScript, are in a different element, or the response is not the page you expected | Save and inspect the returned HTML, verify the URL and look for an API or rendered workflow |
| Many tiny or irrelevant files | Global img selection includes logos, icons and placeholders |
Filter by container, class, path, dimensions or metadata specific to the page |
404 or 403 |
Stale URL, access policy, required headers or a protected resource | Check the site’s documented access method; do not try to bypass controls |
| Downloaded files are HTML | Error page or redirect returned under an image-looking URL | Check status, final URL and Content-Type before saving |
| Wrong host contacted | An extracted absolute URL or redirect replaced the base host | Validate parsed and final hostnames against an explicit allow-list |
| Requests time out | Slow server, large file or transient network problem | Use bounded retries with backoff, lower concurrency and preserve a resume manifest |
Rights, robots.txt and privacy
Downloading a publicly visible file is not the same as having permission to republish it. The U.S. Copyright Office states that “The original authorship appearing on a website may be protected by copyright,” which can include photographs. Its fair-use guidance says the result depends on all the circumstances; there is no universal number of images, words or percentage that automatically qualifies.
For a project that republishes, trains a model, sells a collection or otherwise uses images beyond personal analysis, identify the copyright owner and license, obtain permission where needed and preserve required attribution. These are U.S. sources; the outcome can differ by jurisdiction, license, purpose and facts. Robots.txt expresses crawler preferences and traffic management, not ownership or a blanket right to use the files.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Or skip the browser setup
When your goal is a clean visual capture rather than the original image files, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL and can return PNG, JPEG, WebP or PDF. The service accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for parameters and response handling. Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper settings, custom CSS and JavaScript, clicks, selector or network-idle waits, request/resource blocking, headers, cookies, user-agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data and an OpenAPI specification.
ScreenshotNeo also has an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. Every plan includes every feature: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I scrape images from any public page?
No. Public visibility does not settle crawler preferences, copyright, licensing, privacy or contractual restrictions. Check the site’s rules and your intended use before collecting or republishing files.
Why does Beautiful Soup miss images I can see?
The initial HTTP response may not contain them. Client-side rendering, deferred loading or interaction can add image elements later; inspect the response and choose an appropriately documented rendered or API-based method.
Best Value
Should I save the original filename?
Not by default. URL paths can collide, contain unsafe characters or be misleading. A sanitized name plus a URL-derived identifier, as in the example, is safer.
Is a robots.txt allow rule permission to reuse an image?
No. Robots.txt concerns crawler access guidance and traffic management. Copyright, license terms, privacy and other applicable rules remain separate questions.
Frequently Asked Questions
Can I scrape images from any public page?
No. Public visibility does not settle crawler preferences, copyright, licensing, privacy or contractual restrictions. Check the site’s rules and your intended use before collecting or republishing files.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Why does Beautiful Soup miss images I can see?
The initial HTTP response may not contain them. Client-side rendering, deferred loading or interaction can add image elements later; inspect the response and choose an appropriately documented rendered or API-based method.
Should I save the original filename?
Not by default. URL paths can collide, contain unsafe characters or be misleading. A sanitized name plus a URL-derived identifier is safer.
Is a robots.txt allow rule permission to reuse an image?
No. Robots.txt concerns crawler access guidance and traffic management. Copyright, license terms, privacy and other applicable rules remain separate questions.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools




