To find every discoverable image, combine five passes: parse each page for img and picture references, expand every srcset candidate, inspect lazy-load attributes and CSS url(...) values, render JavaScript when needed, and parse image sitemaps. Normalize each URL against its source page, remove fragments for deduplication, and retain provenance so you know where every result came from.
The coverage problem: “all images” has several meanings
A static HTML parser can find images present in the response returned by a server. It cannot see an image that JavaScript creates later, a CSS background hidden in a downloaded stylesheet, or an asset listed only in a sitemap. Treat image discovery as layered collection rather than a single selector.
| Pass | What it finds | Trade-off |
|---|---|---|
| HTML parsing | img[src], fallback images, srcset, and picture source candidates |
Fast and reproducible, but limited to the initial response |
| Lazy-load inspection | Site-specific attributes such as data-src and data-srcset |
Useful without a browser, but attribute names are not standardized |
| CSS inspection | Inline and stylesheet background-image:url(...) assets |
Requires downloading stylesheets and parsing CSS |
| Rendered browser | Images inserted or selected after JavaScript, including network requests | More time, memory, and operational complexity |
| Sitemap discovery | Image URLs absent from page HTML, including CDN-hosted assets | Completeness depends on the site’s publishing practice |
“All” should therefore mean all URLs your chosen passes can discover, not a guarantee that private, blocked, or unreferenced files will be revealed.
Start with a static Python extractor
Install the two dependencies:
python -m pip install requests beautifulsoup4
The script below handles ordinary images, responsive candidates, common lazy-load attributes, inline CSS, linked stylesheets, URL normalization, and provenance. It prints one JSON record per unique URL.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches#1 Best Overall
- Compatible with Nintendo Switch 2’s new GameChat mode
- Auto-Light Balance: RightLight boosts brightness by up to 50%, reducing shadows so you look your best—compared to previous-generation Logitech webcams (1)
- Privacy with a Slide: The integrated webcam cover makes it easy to get total, reliable privacy when you're not on a video call
- Built-In Mic: The built-in microphone lets others hear you clearly during video calls
- Easy Plug-And-Play: The Brio 101 works with most video calling platforms, including Microsoft Teams, Zoom and Google Meet—no hassle; it just works
from collections import deque
from urllib.parse import urljoin, urldefrag, urlparse
import json
import re
import requests
from bs4 import BeautifulSoup
URL = 'https://example.com/'
TIMEOUT = 20
CSS_URL = re.compile(r'url\(\s*["\']?([^"\')]+)', re.I)
session = requests.Session()
session.headers['User-Agent'] = 'image-discovery/1.0 (contact: [email protected])'
def normalize(raw, base):
raw = raw.strip()
if not raw or raw.startswith(('data:', 'blob:', 'javascript:')):
return None
absolute = urljoin(base, raw)
absolute, _fragment = urldefrag(absolute)
parsed = urlparse(absolute)
if parsed.scheme not in ('http', 'https'):
return None
return absolute
def add(record, raw, base, attribute):
value = normalize(raw, base)
if value:
record.setdefault(value, set()).add(attribute)
def srcset_values(value):
# Handles normal URL + descriptor pairs such as 320w and 2x.
for candidate in value.split(','):
parts = candidate.strip().split()
if parts:
yield parts[0]
def css_values(text):
for match in CSS_URL.finditer(text):
yield match.group(1).strip()
def extract_page(page_url):
response = session.get(page_url, timeout=TIMEOUT)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
found = {}
for tag in soup.select('img, source'):
for name in ('src', 'data-src', 'data-original', 'data-lazy-src'):
if tag.get(name):
add(found, tag[name], page_url, name)
if tag.get('srcset'):
for candidate in srcset_values(tag['srcset']):
add(found, candidate, page_url, 'srcset')
if tag.get('data-srcset'):
for candidate in srcset_values(tag['data-srcset']):
add(found, candidate, page_url, 'data-srcset')
for tag in soup.select('[style]'):
for raw in css_values(tag.get('style', '')):
add(found, raw, page_url, 'inline-css')
for link in soup.select('link[rel="stylesheet"][href]'):
css_url = normalize(link['href'], page_url)
if not css_url:
continue
css = session.get(css_url, timeout=TIMEOUT)
css.raise_for_status()
for raw in css_values(css.text):
add(found, raw, css_url, 'stylesheet-css')
return [{'url': image, 'source_page': page_url,
'attributes': sorted(attributes)}
for image, attributes in sorted(found.items())]
if __name__ == '__main__':
print(json.dumps(extract_page(URL), indent=2))
The img fallback matters even inside a picture element; responsive markup can contain several valid candidates. Preserve query strings because they may select a size, format, or transformation. Removing only the fragment prevents the same resource from being counted repeatedly when links differ only by #anchor.
Turn the extractor into a small crawl
For multiple pages, maintain a queue of same-site links, a visited set, and a request limit. Add links selected with a[href], normalize them, and enqueue only the hostnames you explicitly allow. Store each record’s source page and attribute instead of returning only a bare URL. That audit trail lets you distinguish an HTML reference from a CSS or sitemap discovery.
Responsive images: collect every candidate
Do not stop at img[src]. A srcset may list several widths (320w, 768w) or pixel-density variants (1x, 2x). Under picture, each source[srcset] can represent a different format or media condition, while the nested img[src] remains the fallback. Collect all candidates; selecting the one a browser would display requires evaluating media conditions and device characteristics.
The simple comma split in the example works for conventional URL-and-descriptor lists. If a site uses unusual values, use a standards-aware srcset parser and keep the original descriptor with each record so downstream code can choose an appropriate rendition.
Recommended Free Tools
Find images hidden by lazy loading
Many sites put the real URL in attributes such as data-src, data-srcset, data-original, or data-lazy-src and leave a placeholder in src. These names are site-specific, so inspect the markup and add selectors for the site you are analyzing. Also look for JSON state embedded in script tags; a framework may keep image URLs there before inserting them into the DOM.
Rank #2
- Compatible with Nintendo Switch 2’s new GameChat mode
- Crisp HD 720p/30 fps video calls with diagonal 55° field of view and auto light correction. Compatible with popular platforms including Skype and Zoom.
- The built-in noise-reducing mic makes sure your voice comes across clearly up to 1.5 meters away, even if you’re in busy surroundings.
- C270’s RightLight 2 feature adjusts to lighting conditions, producing brighter, contrasted images to help you look good in all your conference calls.
- The adjustable universal clip lets you attach the camera securely to your screen or laptop, or fold the clip and set the webcam on a shelf. You’re always ready for your next video call.
If the initial response contains no useful URL, render the page. A headless browser can scroll to trigger viewport-based lazy loading, wait for network activity to settle, and then inspect the post-render DOM. Capture network requests as well: an image requested by JavaScript may never appear as a normal HTML attribute.
Rendered pass with Playwright
Install Playwright and its browser once:
python -m pip install playwright
python -m playwright install chromium
This example records image responses and final DOM attributes. It intentionally limits scrolling and waiting so you can set resource bounds for a larger crawl.
from urllib.parse import urldefrag
from playwright.sync_api import sync_playwright
page_url = 'https://example.com/'
seen = set()
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.on('response', lambda response: seen.add(urldefrag(response.url)[0])
if response.request.resource_type == 'image' else None)
page.goto(page_url, wait_until='networkidle', timeout=60000)
for _ in range(4):
page.evaluate('window.scrollBy(0, window.innerHeight)')
page.wait_for_timeout(500)
for value in page.locator('img, source').evaluate_all(
"els => els.flatMap(e => [e.src, e.srcset, e.dataset.src, e.dataset.srcset])"):
if value:
for item in value.split(','):
seen.add(urldefrag(item.strip().split()[0])[0])
browser.close()
for image in sorted(seen):
print(image)
Use rendering only where static parsing is insufficient. Browser sessions consume substantially more CPU, memory, and bandwidth than HTTP requests, and JavaScript can trigger third-party requests you did not intend to crawl.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Extract CSS background images
Decorative images commonly live in background-image declarations rather than image elements. Scan every inline style and downloaded stylesheet for url(...). Resolve stylesheet-relative paths against the stylesheet URL, not the HTML page URL. Ignore data: URLs unless your goal is to inventory embedded bytes; they are not separately fetchable image URLs.
A regular expression is adequate for ordinary CSS, but production crawlers should use a CSS parser to handle escaped characters, comments, nested rules, and quoted parentheses correctly. Keep the stylesheet URL as provenance.
Rank #3
- 【Full HD 1080P Webcam】Powered by a 1080p FHD two-MP CMOS, the NexiGo N60 Webcam produces exceptionally sharp and clear videos at resolutions up to 1920 x 1080 with 30fps. The 3.6mm glass lens provides a crisp image at fixed distances and is optimized between 19.6 inches to 13 feet, making it ideal for almost any indoor use.
- 【Wide Compatibility】Works with USB 2.0/3.0, no additional drivers required. Ready to use in approximately one minute or less on any compatible device. Compatible with Mac OS X 10.7 and higher / Windows 7, 8, 10 & 11 / Android 4.0 or higher / Linux 2.6.24 / Chrome OS 29.0.1547 / Ubuntu Version 10.04 or above. Not compatible with XBOX/PS4/PS5.
- 【Built-in Noise-Cancelling Microphone】The built-in noise-canceling microphone reduces ambient noise to enhance the sound quality of your video. Great for Zoom / Facetime / Video Calling / OBS / Twitch / Facebook / YouTube / Conferencing / Gaming / Streaming / Recording / Online School.
- 【USB Webcam with Privacy Protection Cover】The privacy cover blocks the lens when the webcam is not in use. It's perfect to help provide security and peace of mind to anyone, from individuals to large companies. 【Note:】Please contact our support for firmware update if you have noticed any audio delays.
- 【Wide Compatibility】Works with USB 2.0/3.0, no additional drivers required. Ready to use in approximately one minute or less on any compatible device. Compatible with Mac OS X 10.7 and higher / Windows 7, 10 & 11, Pro / Android 4.0 or higher / Linux 2.6.24 / Chrome OS 29.0.1547 / Ubuntu Version 10.04 or above. Not compatible with XBOX/PS4/PS5.
Use sitemaps to discover what pages omit
Check /sitemap.xml, sitemap indexes, and any sitemap URL declared in robots.txt. Image sitemaps can contain an image:image element with an image:loc URL, including an image hosted on a separate CDN domain. Parse every child sitemap recursively, enforce a maximum count, and deduplicate against page-derived results.
A sitemap is a discovery source, not proof of completeness. Sites may omit old assets, publish stale entries, or list only selected images. Record the sitemap URL and XML location for each result so later verification is possible.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Respect robots.txt, terms, and request budgets
Fetch and honor the applicable robots.txt rules before crawling. Robots.txt is a crawler directive, not an access-control boundary: a disallowed URL may still be indexed when linked elsewhere, and a permitted URL is not automatically authorized for every use. Follow the site’s terms, identify your client, rate-limit requests, and stop on repeated failures.
For a large crawl, use conditional requests (ETag and Last-Modified), connection pooling, exponential backoff for transient 5xx responses, and a bounded queue. Cache downloaded HTML and CSS, but do not silently reuse a response after its freshness assumptions expire.
Choose an approach by coverage and cost
- Need a quick inventory of a few pages: static HTML plus inline and linked CSS.
- Need responsive variants: collect every
srcsetand preserve descriptors. - Need lazy or JavaScript-created assets: render with a headless browser and observe image network requests.
- Need site-wide discovery: combine a controlled crawl with sitemap indexes and image extensions.
- Need reproducibility: save the page URL, attribute, timestamp, HTTP status, and normalized URL for every record.
Separate discovery from downloading. First produce a deduplicated manifest; then apply your own allowlist, file-type checks, size limits, and licensing review before fetching binaries. An image URL is not automatically permission to copy or republish the image.
Rank #4
- 1080P Webcam with Cover for Video Calls - EMEET computer webcam provides design and Optimization for professional video streaming. Realistic 1920 x 1080p video, 5-layer anti-glare lens, providing smooth video. C960 computer camera delivers 1920x1080 video with fixed focus (11.8–118.1 inches), so as to provide a clearer image. C960 USB webcam has a cover and can be removed automatically to meet your needs for privacy. For optimal image performance, use the webcam in a well-lit environment.
- Built-in 2 Omnidirectional Mics - EMEET webcam with microphone for desktop features 2 built-in omnidirectional microphones, picking up your voice to create clear audio for communication. When installing the webcam, select EMEET C960 as the default microphone input device in your computer and video applications and select C960 as the default device in Zoom/Teams and ensure microphone permissions are enabled for proper use. Please note that C960 does not include built-in speakers.
- Automatic Light Adjustment - Automatic exposure adjustment is applied in EMEET HD webcam 1080p so that the streaming webcam can deliver stable image performance. EMEET C960 camera for computer also features color adjustment and exposure optimization to help you look your best. For optimal video quality, it is recommended to use the webcam in normal or well-lit environments and select suitable video settings in your application. Proper lighting helps achieve a clearer and more balanced image.
- Plug-and-Play & Upgraded USB Connectivity - New C960 webcam features both USB Type-A & A-to-C adapter connections for wider compatibility. For stable performance, connect the webcam directly to the computer's main USB port and ensure the device is recognized correctly. If a hub or docking station is used, please ensure it provides sufficient power and stable data transmission, as limited ports may affect performance. 90° wide-angle lens captures more participants without frequent adjustments.
- High Compatibility & Multi Application - C960 webcam for laptop is compatible with Windows 10/11, macOS 10.14+, and Android TV 7.0+. Not supported: Windows Hello, TVs, tablets, or game consoles. It works with Zoom, Teams, Facetime, Google Meet, YouTube and more. Please select C960 webcam as the default camera and microphone device in your application and ensure camera/microphone permissions are enabled, especially on macOS. (Tips: Incompatible with Windows Hello)
Common failures and fixes
Only a few images are returned
Cause: the page uses picture, srcset, or lazy attributes that your selector ignores. Fix: collect source[srcset], all responsive candidates, and the site’s observed data-* attributes.
Free tools Windows power users keep installed
One-click scans. No signup required.
CSS images are missing
Cause: you scanned HTML but not linked stylesheets, or resolved paths against the wrong base. Fix: download each stylesheet, parse url(...), and resolve relative to that stylesheet’s URL.
The HTML has no image URLs
Cause: JavaScript obtains data after load or an API response supplies the gallery. Fix: render, inspect network requests, and identify the underlying JSON endpoint; apply the endpoint’s access and rate limits.
Relative URLs become invalid
Cause: string concatenation ignores the document or stylesheet base. Fix: use urljoin, then remove only fragments with urldefrag.
Requests fail or return a challenge page
Cause: rate limits, authentication, geo restrictions, or bot checks. Fix: slow down, supply authorization only when you have permission, honor robots and terms, and do not attempt to bypass a CAPTCHA or access control.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
- Compatible with Nintendo Switch 2’s new GameChat mode
- HD lighting adjustment and autofocus: The Logitech webcam automatically fine-tunes the lighting, producing bright, razor-sharp images even in low-light settings. This makes it a great webcam for streaming and an ideal web camera for laptop use
- Advanced capture software: Easily create and share video content with this Logitech camera that is suitable for use as a desktop computer camera or a monitor webcam
- Stereo audio with dual mics: Capture natural sound during calls and recorded videos with this 1080p webcam, great as a video conference camera or a computer webcam
- Full HD 1080p video calling and recording at 30 fps. You'll make a strong impression with this PC webcam that features crisp, clearly detailed, and vibrantly colored video
Duplicate URLs remain
Cause: the same asset appears in HTML, CSS, and a sitemap, or differs only by a fragment. Fix: deduplicate normalized URLs while retaining every source record and attribute.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when you need a rendered result rather than writing and operating a browser crawler. A single request can capture a page after it accepts the cookie or consent banner and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
Use the documented endpoint and parameters shown in the ScreenshotNeo documentation:
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
For Python:
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://example.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
For Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo also supports full-page capture, element selection, dark mode, device presets and custom viewports, retina scale, custom CSS and JavaScript, click and wait actions, request and resource blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 shots per month without a card. Paid plans start at $5 for 3,000 shots; higher plans are Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000). Yearly billing gives two months free, and every feature is available on every plan. Sign up for the free ScreenshotNeo plan to try it without a card.
Frequently Asked Questions
Should I save the original HTML and CSS?
Yes. Store the response bodies or content hashes alongside your manifest when reproducibility matters; pages and stylesheets can change while you are processing the results.
How do I tell whether two different URLs serve the same file?
Fetch permitted assets and compare bytes or a cryptographic hash after redirects. URL normalization alone cannot prove that two CDN variants are identical.
Can an image sitemap replace crawling?
No. It can reveal assets omitted from page markup, but its coverage depends on how consistently the site maintains its sitemap.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

