To extract a website logo automatically, fetch the site’s homepage, collect declared logo and icon URLs, inspect structured data and its web manifest, then use a browser-rendered pass when the mark is only present in CSS or JavaScript. Save every candidate with its source and metadata: a favicon, app icon, share image, and primary logo are not interchangeable, and an image URL alone does not grant permission to reuse the artwork.
What counts as a website logo?
A website may expose several images that appear brand-related: a primary wordmark or symbol in the header, a square favicon, an app icon, a social-sharing banner, or a partner badge. Automated extraction should identify candidates, not pretend every image is the official primary logo. Keep source labels and confidence information so a person or downstream system can choose the right asset.
For one-off work and controlled sites, a static HTTP fetch plus an HTML parser is usually the simplest starting point. Add structured data and manifest parsing for broader coverage. Use a browser-rendered pass if client-side code, inline SVG, or CSS backgrounds hide the asset from the raw HTML.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Logo Design. Global Brands | $23.30 | Buy on Amazon |
| 2 |
|
Principles of Logo Design: A Practical Guide to Creating Effective Signs, Symbols, and Icons | $22.30 | Buy on Amazon |
| 3 |
|
Logo, revised edition | $22.04 | Buy on Amazon |
| 4 |
|
Logo Design (Bibliotheca Universalis) (Multilingual Edition) | $13.99 | Buy on Amazon |
Use a layered extraction pipeline
- Fetch the canonical homepage. Follow redirects, record the final URL and origin, and note when you retrieved it. Respect robots rules, access controls, and the website’s terms before crawling. Resolve relative asset paths against the page that declared them.
- Collect declared icons. Parse
link[rel]declarations foricon,shortcut icon,apple-touch-icon, andapple-touch-icon-precomposed. Icon URLs can be relative or absolute; retain each declaration’s attributes rather than discarding size and type details. Google documents these rel values and URL handling in its favicon guidance. - Read organization structured data. Inspect JSON-LD, microdata, or RDFa for
Organization.logo. It may be a URL or anImageObject. Google recommends that organization information appear on the homepage or a page describing the organization, and that the logo image be crawlable and indexable; its current guidance sets a 112×112-pixel minimum. See Google’s Organization documentation. - Inspect the web app manifest. If the page links a manifest, parse its
iconsarray. Preserve each icon’s size, purpose, MIME type, and density metadata; a manifest icon may be designed for an app launcher rather than a header. - Collect social metadata separately. Record
og:image,twitter:image, and equivalent share-image declarations as social candidates. They can be wide promotional banners rather than logo files. - Render the page if needed. A browser pass can reveal inline SVGs, CSS
background-imageassets, images inserted by JavaScript, and metadata created after client-side rendering. Firecrawl’s documented Website Logo Extractor combines browser rendering with schema.org data, icon links, manifest icons, Open Graph, and Twitter images: Firecrawl’s workflow description. - Validate and rank. Check response status, content type, actual image dimensions, transparency, aspect ratio, and whether the file is a logo rather than a generic icon or banner. Prefer an explicit organization logo, then a prominent header mark, then a high-resolution icon, while retaining the other candidates for review.
- Preserve provenance. Store the declared URL, final URL after redirects, retrieval timestamp, MIME type, dimensions, content hash, source field or page location, and any available license or terms information. Convert or resize only after retaining the original asset and its provenance.
Extract candidates with a static HTTP request
A static parser is efficient for small batches and sites whose relevant declarations are present in the returned HTML. The example below uses Python and Beautiful Soup to retrieve a homepage, resolve icon and social-image URLs, and extract basic organization-logo candidates from JSON-LD. It does not execute JavaScript or discover CSS backgrounds; use the browser fallback described later when those matter.
Free tools Windows power users keep installed
One-click scans. No signup required.
Install the dependencies with python -m pip install requests beautifulsoup4, then save the following as extract_logo_candidates.py:
#1 Best Overall
import json
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
def find_organization_logos(value):
"""Return logo values found on Organization-like JSON-LD objects."""
found = []
if isinstance(value, list):
for item in value:
found.extend(find_organization_logos(item))
elif isinstance(value, dict):
kind = value.get("@type", [])
kinds = kind if isinstance(kind, list) else [kind]
if any(str(item).endswith("Organization") for item in kinds):
logo = value.get("logo")
if logo:
found.append(logo)
for key, item in value.items():
if key not in ("logo",):
found.extend(find_organization_logos(item))
return found
def logo_url(value, base_url):
if isinstance(value, str):
return urljoin(base_url, value)
if isinstance(value, dict):
url = value.get("url") or value.get("contentUrl")
if url:
return urljoin(base_url, url)
return None
def extract_candidates(homepage):
response = requests.get(
homepage,
headers={"User-Agent": "LogoCandidateExtractor/1.0"},
timeout=20,
allow_redirects=True,
)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if "html" not in content_type.lower():
raise ValueError(f"Expected HTML, received {content_type!r}")
final_url = response.url
soup = BeautifulSoup(response.text, "html.parser")
candidates = []
icon_rels = {
"icon", "shortcut icon", "apple-touch-icon",
"apple-touch-icon-precomposed",
}
for tag in soup.find_all("link", href=True):
rels = tag.get("rel", [])
rel_text = " ".join(rels).lower() if isinstance(rels, list) else str(rels).lower()
if rel_text in icon_rels or any(rel in rel_text.split() for rel in icon_rels):
candidates.append({
"kind": "declared-icon",
"url": urljoin(final_url, tag["href"]),
"rel": rel_text,
"sizes": tag.get("sizes"),
"type": tag.get("type"),
})
for tag in soup.find_all("meta"):
key = (tag.get("property") or tag.get("name") or "").lower()
if key in {"og:image", "twitter:image"} and tag.get("content"):
candidates.append({
"kind": "social-share-candidate",
"url": urljoin(final_url, tag["content"]),
"source": key,
})
for script in soup.find_all("script", type="application/ld+json"):
if not script.string and not script.get_text(strip=True):
continue
try:
data = json.loads(script.string or script.get_text())
except json.JSONDecodeError:
continue
for logo in find_organization_logos(data):
url = logo_url(logo, final_url)
if url:
candidates.append({"kind": "organization-logo", "url": url})
return {
"requested_url": homepage,
"final_url": final_url,
"final_origin": f"{urlparse(final_url).scheme}://{urlparse(final_url).netloc}",
"retrieved_at_utc": response.headers.get("Date"),
"candidates": candidates,
}
if __name__ == "__main__":
import sys
if len(sys.argv) != 2:
raise SystemExit("Usage: python extract_logo_candidates.py https://example.com")
print(json.dumps(extract_candidates(sys.argv[1]), indent=2))
Run it as python extract_logo_candidates.py https://example.com. The output is a candidate inventory, not proof that any file is the current primary logo. The server’s Date header is optional and may not be the exact retrieval time; production systems should record their own UTC timestamp. This compact example handles common JSON-LD shapes but does not fully interpret every possible JSON-LD graph, microdata, or RDFa pattern.
Fetch and validate each candidate asset
Once a candidate URL is found, fetch it separately. Do not trust the file extension: servers can redirect, return an HTML error page with status 200, or serve a different image type than the URL suggests. Record status and the response’s content type, then inspect the bytes with an image library to confirm dimensions and format. Reject unexpectedly large responses and impose timeouts and redirect limits in production.
- Dimensions: distinguish small square favicons from large transparent wordmarks and social banners. An icon can be valid as a favicon and unsuitable for a document header.
- Transparency and aspect ratio: preserve transparency when present; a white background can make a dark logo appear blank on a light preview.
- Content identity: compare candidates visually or with a human review when the ranking is ambiguous. A declared share image may be an advertisement, not a logo.
- Deduplication: hash downloaded bytes to detect identical assets exposed through multiple metadata fields, but retain all source declarations.
- Safe fetching: if processing arbitrary domains, guard against server-side request forgery. Restrict private, loopback, and link-local destinations, re-check redirect targets, and set response-size limits.
When to add browser rendering
Use a rendered browser when the homepage’s raw source lacks the visible mark, the site is a JavaScript-heavy single-page app, or you need to confirm which candidate is actually prominent in the header. A browser can inspect the rendered DOM for images and inline SVG, examine computed styles for CSS backgrounds, and wait for client-side content. It costs more CPU and time than an HTTP fetch, and may encounter bot checks or access restrictions; do not attempt to bypass protections.
A practical two-pass design keeps work proportional to need: run the static pass first, then render only when it finds no plausible candidate or when confidence is below your threshold. Keep the browser’s page URL, viewport, and capture timestamp with the discovered element or asset URL. A screenshot helps a reviewer understand context, but it does not replace downloading and validating the original image.
Rank #2
Choose an approach for your scale and site mix
| Approach | Best for | Strengths | Limitations |
|---|---|---|---|
| Static HTTP fetch and HTML parser | Small batches and controlled sites | Low overhead, deterministic, easy to cache | Misses client-rendered and CSS-only assets |
| Static parser plus JSON-LD, manifest, and social metadata | General-purpose crawler | Broad coverage without running a full browser | Metadata may be stale, missing, or semantically ambiguous |
| Headless browser | JavaScript-heavy sites and visual confirmation | Sees rendered DOM, CSS backgrounds, and dynamically inserted assets | More CPU, latency, anti-bot friction, and operational cost |
| Hosted brand API | Large-scale enrichment and normalization | Can provide a consistent schema, delivery, and less crawler maintenance | Evaluate pricing, quotas, freshness, coverage, terms, and vendor dependence |
Compare options on source coverage, fidelity to the primary logo, JavaScript and CSS handling, output formats and dimensions, throughput, rate limits, freshness, and rights to reuse. Do not assume a universal extraction success rate: there is no established benchmark figure that applies across arbitrary websites.
Hosted logo and browser tools
ScreenshotNeo is a website screenshot API and MCP server for developers. It can help with the rendered inspection stage when you need a screenshot of a page, but a screenshot is not the same as extracting the source logo asset: use page markup, browser inspection, or a brand API to obtain the original candidate URL and file.
Brandfetch documents a Brand API for logos, colors, fonts, and company details, with data primarily from first-party websites and managed social profiles; its current product documentation describes coverage of 50 million brands. See Brandfetch’s developer documentation and its products page for the product options. Firecrawl’s Website Logo Extractor is a browser-rendered, no-code-oriented alternative, with documented output spanning organization logos, icon links, manifest icons, and social images.
Recommended Free Tools
Do not build a new integration around the old public Clearbit Logo API: Clearbit’s support documentation says it was sunset on December 1, 2025, and that it no longer sells new Logo API subscriptions. It notes that some customers may access logos through the Enrichment API. See Clearbit’s sunset notice.
Rank #3
Or skip the browser setup
To inspect a rendered page, ScreenshotNeo can return a screenshot in one GET request. Its clean-shot options accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the result identified in X-Page-Verdict and X-Billed headers. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents. It is not a substitute for extracting the original image URL when you need the logo file itself.
One-call cURL example (see the ScreenshotNeo API documentation for options):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card required.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Handle favicon and logo expectations correctly
Favicons are a fallback identifier, not a guarantee of a suitable primary logo. Google says a favicon must be square and at least 8×8 pixels, recommends larger than 48×48 pixels, and supports BMP, GIF, ICO, PNG, JPEG, PPM, and TIFF. Google also notes that a favicon is not guaranteed to appear in Search even when guidelines are met. See the favicon documentation. That is guidance for Google Search’s favicon handling; it does not mean every listed file is appropriate for reuse as a full-size logo.
Structured-data requirements have a different purpose: Google’s Organization documentation specifies a minimum 112×112-pixel logo image and says it must be crawlable and indexable. The markup designates an image representing the organization and may help Google use it in results; it is not a general logo-download endpoint. See the current Organization guidance and Google’s earlier explanation of organization markup.
Troubleshoot common extraction failures
- No candidates found: confirm the requested URL redirects to the intended canonical homepage and that the response is HTML. Inspect the rendered page; the logo may be injected by JavaScript or applied as a CSS background.
- Only tiny icons appear: treat them as fallback candidates. Search structured data, inspect the header’s rendered image and CSS, and review the manifest. Do not upscale a tiny favicon and label it a primary logo.
- JSON-LD parsing fails: malformed or non-JSON scripts can be skipped, as in the example. Log the parse failure, inspect other JSON-LD blocks, and add microdata/RDFa parsing if your target sites rely on those formats.
- Candidate URL returns an error or HTML: follow redirects, check status and content type, and verify that the site permits access. A valid-looking URL does not ensure the asset remains available.
- The selected image is a banner or badge: preserve its declaration source and rank it lower than explicit organization markup or a prominent header asset. Use a review threshold and expose alternatives instead of silently choosing.
- Browser rendering is blocked or slow: respect access controls and avoid evasion. Apply bounded timeouts, render only after the static pass needs it, and record the failure distinctly from “no logo found.”
- The result looks blank after conversion: check for transparency, background color, format support, and image dimensions before flattening or resizing; retain the original bytes for reprocessing.
Separate finding a logo from permission to reuse it
Extraction establishes where an image was published, not who owns it or whether your intended use is allowed. Before republishing or using an asset commercially, review the site’s terms and any license or permissions that apply. Keep the original URL, retrieval date, and available rights notes with the asset, and seek permission when the intended use is unclear.
Frequently Asked Questions
Can a favicon be used as the website’s main logo?
Sometimes, but it may be monochrome, outdated, or too small for a larger placement. Treat it as a fallback and check the header or organization-logo metadata for a better candidate.
Does an Organization logo in structured data guarantee that Google will show it?
No. The markup identifies the organization’s representative image and may help Google use it; it does not guarantee a particular Search display.
Quick Recap
Is a social preview image usually the logo?
Not necessarily. Open Graph and Twitter images often serve as wide social banners, so keep them labeled as share-image candidates.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

