A reliable custom link checker is a small crawler-and-probe pipeline, not a single HTTP request. Start with a seed URL, fetch pages within a defined scope, resolve and normalize every link, probe each URL with a HEAD request plus a GET fallback, preserve redirects and exact failures, and emit a report that tells you where each problem originated.
What the checker must do
A useful checker answers more than “valid” or “broken.” For every discovered reference, retain the source page, the original spelling, normalized URL, status code, redirect chain, final destination, content type, elapsed time, and a distinct error class when no HTTP response exists.
- Reachable: 2xx response.
- Redirected: 3xx response or a request whose history contains redirects.
- Client response: 4xx, including authentication and permission responses.
- Server response: 5xx.
- Non-HTTP failure: DNS failure, refused connection, TLS error, timeout, unsupported scheme, or parser failure.
A successful response does not prove that a page contains the intended content, that a JavaScript-rendered link works, or that an authenticated visitor can access it. Treat those as separate validation problems.
Choose scope and safety limits first
Accept these settings rather than allowing an unrestricted crawl:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
- Seed URL and maximum pages and links.
- Allowed schemes:
httpandhttpsonly. - Same-origin mode (recommended for a site audit) or an explicit host allow-list.
- Concurrency limit, per-host delay, timeout, and maximum redirect hops.
- A descriptive user-agent, such as
FreedomLinkChecker/1.0 (+https://example.com/contact). - TLS verification enabled by default; do not disable it to hide certificate errors.
Fetch each origin’s /robots.txt and honor rules for your user-agent. The W3C link checker documents this behavior and uses a recognizable W3C-checklink user-agent rule. Robots.txt is an access-policy signal, not a license to ignore rate limits.
Resolve and normalize links correctly
HTML commonly contains relative references such as /docs, ../pricing, and guide#install. Resolve them against the page URL with Python’s urljoin, remove fragments with urldefrag, then compare scheme and hostname case-insensitively. Keep the unmodified spelling for the report.
Always apply scheme, host, and scope checks after joining. An absolute URL supplied in an attribute can otherwise escape a same-origin crawl. Fragment identifiers identify a document location and should not cause the same resource to be fetched repeatedly.
A runnable Python checker
Install Requests with python -m pip install requests. Save the following as link_checker.py and run python link_checker.py https://example.com. It performs a bounded breadth-first crawl, uses a Requests session, tries HEAD first, falls back to GET when HEAD is unsupported, follows redirects, and prints JSON.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsRank #2
import sys, json, time
from collections import deque
from html.parser import HTMLParser
from urllib.parse import urljoin, urldefrag, urlsplit
import requests
class LinkParser(HTMLParser):
def __init__(self):
super().__init__()
self.links = []
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
key = "href" if tag in {"a", "area", "link"} else "src"
if tag in {"a", "area", "link", "img", "script", "iframe", "source"} and attrs.get(key):
self.links.append((tag, attrs[key]))
def normalize(base, raw):
absolute = urljoin(base, raw)
absolute, _ = urldefrag(absolute)
parts = urlsplit(absolute)
if parts.scheme not in {"http", "https"} or not parts.hostname:
return None
return absolute
def in_scope(url, seed_host):
return urlsplit(url).hostname.lower() == seed_host.lower()
def probe(session, url, timeout=10):
started = time.monotonic()
try:
response = session.head(url, allow_redirects=True, timeout=timeout)
# Some servers reject HEAD or return an unusable result.
if response.status_code in {405, 501} or response.status_code == 200 and not response.headers.get("content-type"):
response = session.get(url, allow_redirects=True, timeout=timeout, stream=True)
return {
"status": response.status_code,
"content_type": response.headers.get("content-type"),
"redirects": [{"status": r.status_code, "url": r.url, "location": r.headers.get("location")} for r in response.history],
"final_url": response.url,
"elapsed_ms": round((time.monotonic() - started) * 1000),
}
except requests.RequestException as exc:
return {"error": type(exc).__name__, "detail": str(exc), "elapsed_ms": round((time.monotonic() - started) * 1000)}
def main(seed):
seed = normalize(seed, seed)
if not seed: raise SystemExit("Seed must be an http or https URL")
seed_host = urlsplit(seed).hostname
session = requests.Session()
session.headers["User-Agent"] = "FreedomLinkChecker/1.0"
queue, queued, visited, report = deque([seed]), {seed}, set(), []
max_pages, max_links, timeout = 100, 1000, 10
while queue and len(visited) < max_pages and len(report) < max_links:
page = queue.popleft()
if page in visited: continue
visited.add(page)
page_result = probe(session, page, timeout)
page_result.update({"source_page": page, "discovered_url": page, "normalized_url": page})
report.append(page_result)
if page_result.get("error") or page_result.get("status", 0) >= 400:
continue
try:
response = session.get(page, allow_redirects=True, timeout=timeout)
if "html" not in response.headers.get("content-type", "").lower():
continue
parser = LinkParser(); parser.feed(response.text)
except (requests.RequestException, ValueError) as exc:
report.append({"source_page": page, "error": type(exc).__name__, "detail": str(exc)})
continue
for tag, raw in parser.links:
target = normalize(response.url, raw)
if not target or not in_scope(target, seed_host):
continue
result = probe(session, target, timeout)
result.update({"source_page": page, "discovered_url": raw, "normalized_url": target})
report.append(result)
if target not in queued and len(queue) + len(visited) < max_pages:
queue.append(target); queued.add(target)
if len(report) >= max_links: break
print(json.dumps(report, indent=2))
if __name__ == "__main__":
if len(sys.argv) != 2: raise SystemExit("Usage: python link_checker.py https://example.com")
main(sys.argv[1])
This example deliberately keeps limits visible. Before using it on a large or hostile site, add robots parsing, per-host politeness delays, a bounded worker queue, maximum redirect hops, transient-error backoff, and a persistent cache.
HEAD versus GET
MDN defines HEAD as requesting the metadata that a GET would return, without the response body. It can reduce bandwidth, so it is a sensible first probe. However, servers and proxies sometimes block or mishandle HEAD. Retry with GET when the server returns 405 or 501, supplies an unusable result, or when you must validate body-dependent resources. Use stream=True when only headers are needed and close responses promptly.
| Strategy | Strength | Trade-off |
|---|---|---|
| HEAD first, GET fallback | Lower normal bandwidth while retaining compatibility | Two requests for servers that mishandle HEAD |
| GET first | Works with servers that do not implement HEAD | More bytes, time, and load |
Redirects and final destinations
Redirect responses start with 3 and include a Location header. Preserve every response in response.history and the final URL. A 301 or 308 is generally permanent; 302, 303, and 307 have different temporary and method semantics. Report the chain rather than marking a redirected link as simply broken, because a redirect can reveal an obsolete URL, a loop, or an unexpected external destination.
Redirect checks worth adding
- Stop after a configured hop count.
- Flag loops and excessive chains.
- Apply scope rules to every destination, not only the first URL.
- Retain the original source and spelling so editors can fix the actual page.
Robots, rate limits, and concurrency
A production crawler should fetch and cache robots.txt per origin, evaluate the rules for its user-agent, and skip disallowed URLs. Use a queue with a visited set, bounded workers, and a per-host delay. Retry only transient failures (for example, connection resets or selected 5xx responses), with exponential backoff and a maximum attempt count. Do not retry malformed URLs, unsupported schemes, authentication failures, or a deterministic 404.
Prevent accidental abuse
- Reject non-HTTP schemes before any request.
- Limit pages, extracted links, response size, redirect hops, and total run time.
- Do not let a user-supplied redirect escape an allowed host set.
- Keep certificate verification on unless a controlled test environment explicitly requires otherwise.
Make the report actionable
Store JSON or CSV fields for source_page, discovered_url, normalized_url, status, error_class, redirect_chain, final_url, content_type, and elapsed_ms. Add a suggested action such as “fix relative-path typo,” “review redirect target,” “retry later,” or “request credentials.” Group failures by source page, and distinguish an external outage from a local-content mistake.
Interpret common outcomes
| Outcome | Likely action |
|---|---|
| 404 or 410 on an internal URL | Correct, remove, or replace the source link. |
| 401 or 403 | Decide whether authentication is expected; do not call it universally broken. |
| Repeated timeout | Check host health, timeout policy, and whether the resource is intentionally slow. |
| 5xx from an external host | Record the outage and retry later before editing your content. |
| Unsupported scheme | Skip it or add an explicit handler; never pass it to an HTTP client blindly. |
Single-page checker or full crawler?
A single-page checker is easier to run in a pre-publish hook and has predictable load. A crawler finds site-wide failures but needs queueing, scope, robots handling, caching, and politeness controls. Start with one page for editor feedback; schedule bounded crawls for broader audits.
Troubleshooting
“HEAD returns 405 or 501”
Use the GET fallback shown above. Some CDNs also return misleading HEAD headers, so configure specific hosts or content types to use GET directly.
“Relative links become the wrong host”
Resolve with urljoin(response.url, raw), remove fragments, then enforce the host allow-list. Never validate scope on the raw attribute.
“The checker reports a timeout for a page that loads in a browser”
Compare DNS, TLS, proxy, user-agent, and authentication conditions. Increase the timeout only within a global run limit; a browser may also execute JavaScript that a Requests client cannot.
“Everything is 403”
Identify your crawler with a descriptive user-agent, honor robots rules, slow down, and verify whether the site requires authentication. Do not bypass an access control.
“The same URL appears many times”
Deduplicate after joining and fragment removal. Decide whether query strings are meaningful; if you canonicalize them, preserve the original URL in output.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your workflow also needs clean visual captures of pages, ScreenshotNeo provides a one-call screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
Recommended Free Tools
cURL (see the ScreenshotNeo documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
An MCP server lets Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf. Every plan includes the features; 1,000 screenshots per month are free without a card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Best Value
FAQ
Can a link checker validate JavaScript-only links?
Not with Requests alone. Use a browser automation layer for links created after script execution, while retaining the same normalization, scope, and reporting rules.
Should fragments be checked?
Use fragment removal for network deduplication. If validating in-page anchors, run a separate pass that checks whether each fragment target exists in the fetched document.
Is a 3xx response an error?
It is a redirect outcome, not automatically a failure. Flag unexpected, looping, cross-origin, or excessively long chains according to your policy.
Frequently Asked Questions
Can a link checker validate JavaScript-only links?
Not with Requests alone. Use a browser automation layer for links created after script execution, while retaining the same normalization, scope, and reporting rules.
Should fragments be checked?
Use fragment removal for network deduplication. If validating in-page anchors, run a separate pass that checks whether each fragment target exists in the fetched document.
Is a 3xx response an error?
It is a redirect outcome, not automatically a failure. Flag unexpected, looping, cross-origin, or excessively long chains according to your policy.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

