October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
HTML

Image Extractor from HTML: Find Image URLs, srcset Candidates, and Picture Sources

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To extract image references from HTML, collect each <img> element’s src and srcset, then inspect any enclosing <picture> element for alternative <source> entries. That produces a useful inventory of image URLs declared in markup—but it does not necessarily tell you which responsive candidate a browser selected, find CSS background images, or identify which images matter to the page’s main content.

The right extraction method depends on what “all images” means for your task: markup references, browser-selected resources, CSS images, or content-relevant images. The examples below extract markup references and preserve the information needed to interpret them.

What counts as an image in HTML?

An image inventory can mean several different things. Decide which one you need before writing an extractor, because a simple list of URLs cannot answer every version of the question.

  • Markup references: URLs explicitly written in image-related HTML attributes, such as img[src], img[srcset], and source[srcset] inside a picture.
  • The browser’s selected image: the candidate a browser uses after evaluating responsive-image conditions. A markup inventory can preserve the candidates, but does not by itself resolve the browser’s choice.
  • Images loaded from CSS: background images can be declared in stylesheets or inline styles, outside an img-only pass.
  • Content-relevant images: the images that illustrate an article or form part of its main content, rather than logos, icons, advertisements, or other page elements. Finding these is a filtering task, not merely URL collection.

The scripts here target the first scope: image references in the HTML you provide. They retain responsive candidates and descriptors rather than claiming that every URL is currently displayed. They do not comprehensively inspect CSS, execute page JavaScript, or determine which images are editorially relevant.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which HTML elements and attributes should you extract?

Start with img[src]

The src attribute is the basic image reference on an img element. Record its raw value, along with the element’s other relevant attributes, rather than treating it as the only possible resource. An image element can also have a srcset, which offers alternative candidates.

Keep every srcset candidate

A srcset value contains one or more candidate URLs, optionally followed by a descriptor. Width descriptors look like 640w; pixel-density descriptors look like 2x. Keep both the URL and its descriptor in your output. With width descriptors, the sizes attribute helps the browser evaluate which candidate is appropriate for the rendered image size. The presence of a URL in srcset does not mean that it is the one currently displayed.

For example, a markup inventory should preserve the candidates in this element:

<img src="/images/card-small.jpg"
     srcset="/images/card-640.jpg 640w, /images/card-1280.jpg 1280w"
     sizes="(max-width: 700px) 100vw, 700px"
     alt="A sample card">

Here, the extractor can list the fallback src and both srcset candidates. It should not assert which candidate a browser uses without accounting for the browser environment and the image’s responsive conditions.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect picture and its source elements

A picture can contain one or more source elements before its fallback img. A browser can evaluate a source’s media, type, and srcset when choosing an image resource. To inventory the alternatives, collect the srcset values from those sources as well as the nested img attributes. Record the source conditions with each candidate so you can see why an option might apply.

<picture>
  <source media="(min-width: 900px)"
          srcset="/images/landscape-wide.jpg 1200w"
          type="image/jpeg">
  <source media="(max-width: 899px)"
          srcset="/images/landscape-small.webp"
          type="image/webp">
  <img src="/images/landscape-fallback.jpg"
       alt="A landscape">
</picture>

The source list is an inventory, not a browser simulation. Which resource is selected depends on the applicable conditions and the browser’s environment.

Set a separate policy for CSS images

An extractor that parses only img and source elements can miss CSS background images. The same distinction applies whether CSS is inline or in a stylesheet: those references are not covered by the markup-only scripts below. If your requirement includes CSS, state that explicitly and handle stylesheets or rendered styles as a separate part of the job. There is no single comprehensive CSS-discovery technique established here, so do not label an img-only result “every image on the page.”

Extract markup image references with Python

This runnable Python example uses only the standard library. It reads an HTML file, parses image elements and picture sources, and prints JSON with each raw attribute plus parsed srcset candidates. It deliberately preserves URLs as written: it does not resolve relative paths against a page URL, download files, select a responsive candidate, or parse CSS.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from html.parser import HTMLParser
import json
import sys


def parse_srcset(value):
    """Return (URL, descriptor) pairs from a conventional srcset value."""
    if not value:
        return []

    candidates = []
    # In conventional srcset syntax, commas separate candidates. This simple
    # parser is suitable for typical HTML attributes; unusual URL content
    # containing commas may require a parser tailored to that source.
    for item in value.split(","):
        parts = item.strip().split()
        if not parts:
            continue
        url = parts[0]
        descriptor = " ".join(parts[1:]) or None
        candidates.append({"url": url, "descriptor": descriptor})
    return candidates


class ImageExtractor(HTMLParser):
    def __init__(self):
        super().__init__(convert_charrefs=True)
        self.images = []
        self.picture_sources = []
        self.in_picture = False

    def handle_starttag(self, tag, attrs):
        attrs = dict(attrs)
        if tag == "picture":
            self.in_picture = True
        elif tag == "source" and self.in_picture:
            srcset = attrs.get("srcset", "")
            self.picture_sources.append({
                "srcset": srcset or None,
                "candidates": parse_srcset(srcset),
                "media": attrs.get("media"),
                "type": attrs.get("type"),
                "sizes": attrs.get("sizes"),
            })
        elif tag == "img":
            srcset = attrs.get("srcset", "")
            self.images.append({
                "src": attrs.get("src"),
                "srcset": srcset or None,
                "candidates": parse_srcset(srcset),
                "sizes": attrs.get("sizes"),
                "alt": attrs.get("alt"),
                "inside_picture": self.in_picture,
            })

    def handle_endtag(self, tag):
        if tag == "picture":
            self.in_picture = False


if len(sys.argv) != 2:
    raise SystemExit("Usage: python extract_images.py page.html")

with open(sys.argv[1], encoding="utf-8") as html_file:
    parser = ImageExtractor()
    parser.feed(html_file.read())

print(json.dumps({
    "images": parser.images,
    "picture_sources": parser.picture_sources,
}, ensure_ascii=False, indent=2))

Save it as extract_images.py, then run python extract_images.py page.html. The output separates fallback img records from sources found inside picture elements. Each image record says whether its img was inside a picture element; each source record keeps media, type, and sizes alongside the candidates.

The example uses a straightforward comma-separated split for srcset. That is adequate for typical values, but URLs containing unusual comma usage can make a simplistic parser ambiguous. If you process untrusted or unusually formatted HTML, validate the parser against your inputs rather than assuming this compact implementation covers every edge case.

Use JavaScript when the HTML is already in Node.js

If a program already has an HTML string and a DOM implementation, the extraction logic is similar: query image elements, retain raw attributes, and inspect picture sources. This example expects a browser-like document; Node.js does not provide one by default. Supply an HTML parser or a browser DOM as appropriate for your application. The code does not fetch a page or execute its scripts.

function parseSrcset(value = "") {
  return value
    .split(",")
    .map((item) => item.trim())
    .filter(Boolean)
    .map((item) => {
      const [url, ...descriptorParts] = item.split(/s+/);
      return {
        url,
        descriptor: descriptorParts.join(" ") || null,
      };
    });
}

function extractImageReferences(document) {
  const images = [...document.querySelectorAll("img")].map((img) => {
    const srcset = img.getAttribute("srcset");
    return {
      src: img.getAttribute("src"),
      srcset,
      candidates: parseSrcset(srcset || ""),
      sizes: img.getAttribute("sizes"),
      alt: img.getAttribute("alt"),
      insidePicture: img.parentElement?.tagName === "PICTURE",
    };
  });

  const pictureSources = [...document.querySelectorAll("picture source")].map((source) => {
    const srcset = source.getAttribute("srcset");
    return {
      srcset,
      candidates: parseSrcset(srcset || ""),
      media: source.getAttribute("media"),
      type: source.getAttribute("type"),
      sizes: source.getAttribute("sizes"),
    };
  });

  return { images, pictureSources };
}

The Node.js version of insidePicture checks for a direct parent, which matches the usual structure. If your input allows wrappers or you need to account for nonstandard markup, use img.closest("picture") instead. As with the Python example, this output is a declaration inventory, not a guarantee of browser selection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Normalize URLs only when you know the page URL

HTML commonly uses relative references such as /images/photo.jpg or ../assets/photo.jpg. The extractors preserve those strings instead of guessing a base URL. To turn them into absolute URLs, you need the URL of the document they came from and must resolve each reference against that base. A fragment, a data URL, or a site-specific path can require additional handling, so keep the raw value as well as any normalized value in your output.

Do not collapse distinct candidates just because they look similar. A small fallback, a high-density version, and a wide-screen source can represent different files for different conditions. Keeping the original attributes makes later auditing or reprocessing possible.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When a markup extractor is not enough

To find the browser’s current choice

You need to evaluate the relevant responsive conditions in the browser environment you care about. Viewport, device characteristics, and supported image formats can affect which candidate is appropriate. A static inventory is useful for discovery, but it is not a report of the browser’s selected file.

To include background images

Extend the scope to CSS rather than assuming the HTML parser found them. A page can refer to background images through styles, which an img/picture collector will not see. Clearly distinguish any CSS pass from the markup output, and do not claim exhaustive CSS coverage unless your implementation actually handles the stylesheets and conditions relevant to your pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To identify the article’s meaningful images

First collect references, then apply a separate relevance rule. A URL collector has no reliable basis for knowing whether an image is the lead illustration, a logo, or a decorative asset. Research on relevant-image extraction has treated browser-rendering information as one way to distinguish content from page boilerplate; that is a different problem from parsing image URLs.

For JavaScript-created or protected content

The examples operate on HTML they receive. They do not establish how a particular site handles JavaScript-created images, authentication-gated pages, canvas output, blob URLs, or site protections. Those cases depend on the page and implementation. Test against the actual pages and access conditions in your workflow, and label your output according to what the method observed.

Troubleshooting common extraction problems

  • No image appears in the result: check that the source HTML really contains an img element. The page may represent the visual through CSS or create it dynamically; this markup-only method will not discover those automatically.
  • You see a fallback URL but not the expected responsive file: inspect srcset and any enclosing picture sources. The fallback src is not a complete inventory of responsive alternatives.
  • The extracted candidate is not the file visible in your browser: the output lists declared choices, not the browser’s selected resource. Record the relevant source conditions and evaluate selection in the browser environment you are targeting.
  • A relative URL does not open by itself: preserve it as a raw reference and resolve it against the original document URL before using it elsewhere.
  • Your candidate list looks malformed: check how your parser splits srcset, especially if the input contains unusual URL formatting. Compare the parsed records with the original attribute instead of discarding that raw value.
  • The list contains logos, icons, or decorative images: URL extraction reports references, not editorial relevance. Add a separate filtering stage appropriate to your page type.
  • A screenshot shows an image that the markup inventory lacks: the two methods observe different things. A screenshot displays a rendered page; the scripts above inventory declared HTML references. Neither output should be described as the other.

Or skip the browser setup

If you need a rendered visual check alongside your markup inventory, ScreenshotNeo can capture a page as an image or PDF through one GET request. A screenshot is useful for checking what a page looks like; it is not a substitute for extracting and parsing image URLs. The request below captures a page image. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; those cleanup steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response indicates the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.

Frequently Asked Questions

Does extracting an image URL download the image file?

No. The examples collect attribute values from HTML; downloading and storing the referenced files is a separate step.

Can I use the extracted list to identify the main image on an article page?

Not by itself. The list reports declared references, while deciding which one is content-relevant requires a separate filtering method.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.