Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To get a website’s favicon from its URL, fetch the page, inspect its <link> elements for icon relations, resolve each href against the page URL, download viable candidates, and apply a documented selection rule. If no declaration is usable, try the conventional /favicon.ico path. Keep Apple touch icons separate when you need an iOS home-screen image, and do not confuse successful extraction with Google Search displaying a favicon.

The extraction pipeline

A reliable webpage-to-icon API should make five decisions explicit:

  1. Fetch the final document. Follow redirects and retain the final URL, because relative icon paths must be resolved against the document that supplied them.
  2. Parse icon declarations. Inspect every <link> element whose rel value includes icon, the historical shortcut icon, apple-touch-icon, or apple-touch-icon-precomposed. Google documents these relations and permits relative or absolute href values, including icons hosted on a CDN (Google Search Central).
  3. Resolve and retrieve candidates. Convert relative references with URL resolution, then request each candidate with appropriate timeouts, redirect limits and content checks.
  4. Select or return results. Use media, type and sizes to choose the best candidate, as described by MDN’s rel documentation. If callers need control, return all viable candidates and their metadata instead of silently discarding them.
  5. Apply a fallback. When no usable declaration exists, try the site root’s /favicon.ico. Browsers and applications commonly use this convention, but it is not guaranteed (MDN; Website Icon standard).

Return provenance in your API response: source relation, resolved URL, HTTP status, MIME type, byte length, declared sizes, and whether the value came from the fallback. That makes a “missing icon” diagnosis possible instead of returning an unexplained 404.

Parsing link elements without missing candidates

Normalize the rel attribute

HTML permits multiple space-separated relation tokens and different capitalization. Lowercase and split the rel value before testing it. Treat shortcut plus icon as the historical combination, while accepting ordinary icon as the primary relation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Resolve href exactly

Use a standards-compliant URL resolver, not string concatenation. For example, a page at https://example.com/docs/page with href="../icons/site.png" resolves to https://example.com/icons/site.png. Preserve query strings, fragments (removed for an HTTP request), ports and non-default schemes according to your client’s policy. Reject unsupported schemes such as javascript: and avoid following references to local files.

Keep Apple touch icons distinct

Apple touch icons represent the image used when an iOS user saves a Web Clip. MDN notes that iOS does not use the ordinary rel="icon" link type for that purpose (MDN). Offer a field such as kind: "favicon" or kind: "apple-touch-icon" rather than treating the largest touch icon as the browser favicon automatically.

Choosing among multiple declarations

A page can publish several sizes, formats and media conditions. A practical deterministic policy is:

  1. Discard candidates whose media condition is false for the requested environment; if you do not evaluate media queries, state that limitation and retain the candidates.
  2. Prefer a supported, specific MIME type over an omitted or generic type. Verify the response’s Content-Type and, where safe, the file signature.
  3. Prefer an icon whose declared sizes is at least the requested display size, choosing the smallest sufficient image; otherwise choose the largest available candidate.
  4. Use relation priority: ordinary icon first, historical shortcut icon next, and Apple touch icons only when requested.
  5. Break ties consistently, for example by document order, and expose the rule in your API documentation.

sizes="any" commonly denotes a scalable format such as SVG. If your raster pipeline cannot safely render SVG, report it as unsupported rather than returning an empty image. If a selected URL returns an error or an invalid body, continue to the next candidate; MDN describes moving to another resource when the selected one is unsuitable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fallback to /favicon.ico

Only attempt the root fallback after scanning declarations (or when the HTML cannot be fetched). Construct it from the origin, not from the current directory: new URL('/favicon.ico', pageUrl). A successful response still needs validation. Servers sometimes return an HTML error page with status 200, redirect the path to a login page, or provide an unsupported media type. Check status, a reasonable size limit, and image decoding before marking the fallback usable. The conventional path is common behavior, not a guarantee that a file exists or that every client uses it.

Reference implementation in Python

The following script returns every discovered candidate, then tries the root fallback when no declared icon can be downloaded. It uses only standard URL resolution plus HTML parsing and leaves final image decoding to your imaging library.

Rank #2
Basics Fashion Design 03: Construction
  • Used Book in Good Condition
import sys
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup

RELATIONS = {"icon", "shortcut icon", "apple-touch-icon", "apple-touch-icon-precomposed"}

def relation_kind(rel_tokens):
    tokens = set(rel_tokens)
    if "apple-touch-icon-precomposed" in tokens:
        return "apple-touch-icon"
    if "apple-touch-icon" in tokens:
        return "apple-touch-icon"
    if "icon" in tokens:
        return "favicon"
    if {"shortcut", "icon"}.issubset(tokens):
        return "favicon"
    return None

def extract(url):
    r = requests.get(url, timeout=20, allow_redirects=True,
                     headers={"User-Agent": "favicon-extractor/1.0"})
    r.raise_for_status()
    final_url = r.url
    soup = BeautifulSoup(r.text, "html.parser")
    candidates = []
    for tag in soup.find_all("link"):
        raw_rel = tag.get("rel", [])
        tokens = [str(x).lower() for x in raw_rel] if isinstance(raw_rel, list) else str(raw_rel).lower().split()
        kind = relation_kind(tokens)
        href = tag.get("href")
        if not kind or not href:
            continue
        resolved = urljoin(final_url, href.split("#", 1)[0])
        if urlparse(resolved).scheme not in ("http", "https"):
            continue
        candidates.append({"kind": kind, "url": resolved,
                           "type": tag.get("type"), "sizes": tag.get("sizes"),
                           "media": tag.get("media")})
    for candidate in candidates:
        try:
            icon = requests.get(candidate["url"], timeout=20, allow_redirects=True)
            if icon.ok and icon.content:
                candidate["status"] = icon.status_code
                candidate["content_type"] = icon.headers.get("content-type")
                candidate["bytes"] = len(icon.content)
        except requests.RequestException as exc:
            candidate["error"] = str(exc)
    if not any(c.get("bytes") for c in candidates):
        fallback = urljoin(final_url, "/favicon.ico")
        icon = requests.get(fallback, timeout=20, allow_redirects=True)
        if icon.ok and icon.content:
            candidates.append({"kind": "favicon", "url": icon.url,
                               "fallback": True, "status": icon.status_code,
                               "content_type": icon.headers.get("content-type"),
                               "bytes": len(icon.content)})
    return {"page_url": final_url, "candidates": candidates}

if __name__ == "__main__":
    print(extract(sys.argv[1]))

Install dependencies with pip install requests beautifulsoup4. In production, add HTML-size limits, decompression limits, an image decoder, SSRF protection, and a candidate-selection function tailored to the caller’s requested size and format.

HTTP API design and security

Suggested response shape

Return page_url, selected, candidates, and fallback_used. For each candidate include the original relation, resolved URL, media, type, sizes, final URL after redirects, status, MIME type and validation error. A caller asking for an Apple touch icon should be able to filter by kind.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Protect the fetcher

  • Allow only HTTP and HTTPS and block loopback, link-local, private and cloud-metadata IP ranges after DNS resolution.
  • Limit redirects, response bytes, decompressed bytes and total per-page time.
  • Do not forward arbitrary caller credentials to third-party icon hosts. If custom headers are supported, redact them from logs.
  • Cache by normalized URL with an explicit TTL, and include cache status in the response.
  • Use a browser only when JavaScript is required to produce the HTML; most static icon declarations are available in the initial response.

Why Google Search favicon rules are different

Google Search Central describes eligibility for a favicon shown beside a search result, not a universal extraction specification. Google says its crawler must access both the home page and icon; the icon must be square and at least 8 by 8 pixels, with larger than 48 by 48 pixels recommended. Supported formats listed there are BMP, GIF, ICO, PNG, JPEG, PPM and TIFF, and a stable URL is recommended (Google Search Central). Even if those conditions are met, Google states: “A favicon isn’t guaranteed to appear in Google Search results, even if all guidelines are met.” Your extractor should therefore report retrieval and validation independently from Search appearance.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

No link tag and no image

The page may genuinely publish no icon, may block your request, or may generate markup only after JavaScript runs. Check the final response body, try /favicon.ico, and, if permitted, use a rendering browser as a second-stage fetcher.

Relative URL downloads the wrong file

You probably resolved against the requested URL instead of the final response URL, or concatenated strings. Resolve with a URL library against the final document URL.

200 response is not an image

Inspect the MIME type and decode the bytes. Login pages, bot challenges and custom error documents frequently return HTML with status 200.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Only one of several icons is returned

Document your ranking across media, type and sizes. If consumers need different resolutions or formats, expose all validated candidates and let them select.

SVG is present but your pipeline fails

Support SVG rendering explicitly or mark the candidate unsupported. Do not silently substitute an unrelated PNG.

Or skip the browser setup

If your broader workflow already needs clean page captures, ScreenshotNeo can fetch a URL through its screenshot API; it is not a favicon parser, so use the extraction algorithm above when you need icon declarations and metadata. ScreenshotNeo removes cookie banners, newsletter popups and chat widgets before capture, bills only clean shots (bot checks, blank pages, timeouts, failed loads and cache hits are not billed), provides an MCP server for AI agents, and includes 1,000 screenshots each month free with no card; paid plans start at $5 for 3,000 shots.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for parameters. Sign up free for 1,000 screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Should an extractor return the favicon URL or image bytes?

Return both when practical: the resolved, final URL is useful for caching and provenance, while bytes (or a validated proxy URL) spare callers another request and preserve the exact resource you inspected.

Can a favicon be discovered from robots.txt or a sitemap?

Those files do not define the page’s icon relation. Start with HTML link elements and the conventional root fallback.

Is an Apple touch icon a better favicon because it is larger?

Not automatically. It serves an iOS Web Clip use case; classify it separately and select it only when the caller requests that kind.

The Bottom Line

Inspect declared icon links first, resolve them against the final document URL, select with media/type/sizes, preserve metadata, and use /favicon.ico only as a conventional fallback. Retrieval success is separate from Google Search eligibility.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.