Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a link extractor on the HTML response for one page, then classify each resolved URL against a domain rule. A useful result includes the destination, anchor text, fragment, and relevant rel values—not just a flat list of strings. The workflow below covers a single page, a controlled crawl, JavaScript-rendered content, duplicate handling, and reliable internal/external classification.

What a link extractor actually does

A link extractor parses a fetched response and selects link-bearing elements that match its configuration. In the common case, that means <a href="..."> and <area href="...">. The extractor resolves relative references against the page URL and can return metadata alongside the destination.

For each result, keep these fields:

  • URL: the resolved destination, normally without a fragment when you are deduplicating destinations.
  • Anchor text: the visible text inside the link.
  • Fragment: the part after #, useful when links target a heading or another in-page location.
  • Relationship: values such as nofollow, sponsored, or ugc from the rel attribute.
  • Classification: internal or external according to a domain policy that you define.

Extracting one response is not the same as crawling a site. A crawler puts extracted URLs into a queue, applies scope and access rules, fetches more pages, and repeats the process. Scrapy’s documented LinkExtractor is designed for this response-level step and can be used from a spider that follows the returned links.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decide your scope before writing code

Page-level extraction versus a crawl

For a link inventory of one page, fetch once and extract once. For a site-wide audit, set a queue, a maximum page count or depth, a per-host policy, and a delay or concurrency limit. Do not let every external URL become a crawl target by accident.

#1 Best Overall
Klein Tools VDV526-200 LAN Scout Jr Cable Tester Ethernet Cable Tester Kit
  • VERSATILE CABLE TESTING: Cable tester for data (RJ45) terminated cables and patch cords, ensuring comprehensive testing capabilities
  • LARGE BACKLIT LCD: Backlit LCD display enables easy reading of pin-to-pin wiremap results, even in low-lit areas
  • COMPREHENSIVE FAULT DETECTION: Test for Open, Short, Miswire, Split-Pair faults, Cross-over, and Shield, providing thorough fault detection
  • INTUITIVE USER INTERFACE: User-friendly interface with three buttons and simple, easy-to-identify test responses, ensuring a smooth testing experience
  • MULTIPLE TONE GENERATOR STYLES: Tone on a single wire, wire pair, or all 8 conductor wires using the multiple style tone generator (solid/warble); requires probe Cat. No. VDV500-123 (sold separately)

What counts as internal?

Choose and document the boundary. The strict rule in the example code is “same hostname as the starting URL.” That treats www.example.com and example.com as different sites, and it treats a subdomain as external. A broader policy can allow a set of approved hostnames, but it should be explicit. Redirect destinations are a separate decision: classify the URL as written, or fetch it and record the final host as an additional field.

Which HTML is being inspected?

A server-response extractor sees only the HTML returned by the request. Links inserted later by JavaScript, links behind a click handler, and content loaded by an API call will not appear unless you render the page in a browser or obtain the underlying data another way. Always state whether your result came from raw HTML or a rendered DOM.

Option 1: a complete Python extractor for one page

Install the two dependencies first:

python -m pip install requests beautifulsoup4

Save the following as extract_links.py. It resolves relative URLs, separates fragments, records rel values, skips non-web schemes, applies a same-host internal rule, and removes duplicate destinations while retaining the first occurrence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import sys
from urllib.parse import urldefrag, urljoin, urlparse

import requests
from bs4 import BeautifulSoup


def extract(page_url: str):
    response = requests.get(
        page_url,
        headers={"User-Agent": "link-audit/1.0"},
        timeout=30,
        allow_redirects=True,
    )
    response.raise_for_status()

    # Use the final response URL as the base for relative links after redirects.
    base_url = response.url
    base_host = (urlparse(base_url).hostname or "").lower()
    soup = BeautifulSoup(response.text, "html.parser")
    results = []
    seen = set()

    for tag in soup.find_all(["a", "area"]):
        raw = tag.get("href")
        if not raw:
            continue
        raw = raw.strip()
        if raw.lower().startswith(("mailto:", "tel:", "javascript:", "data:")):
            continue

        absolute = urljoin(base_url, raw)
        destination, fragment = urldefrag(absolute)
        parsed = urlparse(destination)
        if parsed.scheme not in ("http", "https") or not parsed.hostname:
            continue

        key = destination
        if key in seen:
            continue
        seen.add(key)

        host = parsed.hostname.lower()
        rel = [value.lower() for value in tag.get("rel", [])]
        results.append({
            "url": destination,
            "fragment": fragment or None,
            "text": " ".join(tag.get_text(" ", strip=True).split()),
            "nofollow": "nofollow" in rel,
            "rel": rel,
            "internal": host == base_host,
        })

    return {
        "requested_url": page_url,
        "final_url": base_url,
        "links": results,
    }


if __name__ == "__main__":
    if len(sys.argv) != 2:
        raise SystemExit("Usage: python extract_links.py https://example.com/page")
    print(json.dumps(extract(sys.argv[1]), indent=2, ensure_ascii=False))

Run it with:

python extract_links.py https://example.com/page

The script deduplicates only after removing fragments. Thus /guide#intro and /guide#install become one destination with the first fragment retained. If fragment-level navigation matters, use a key of (destination, fragment) instead. URL query strings remain significant; do not strip tracking parameters unless your audit policy says they are equivalent.

Option 2: the same job in Node.js

This version uses Node’s built-in fetch and the cheerio parser:

npm install cheerio
import * as cheerio from "cheerio";

const input = process.argv[2];
if (!input) throw new Error("Usage: node extract-links.mjs https://example.com/page");

const response = await fetch(input, {
  headers: { "user-agent": "link-audit/1.0" },
  redirect: "follow"
});
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);

const finalUrl = response.url;
const baseHost = new URL(finalUrl).hostname.toLowerCase();
const html = await response.text();
const $ = cheerio.load(html);
const seen = new Set();
const links = [];

$("a[href], area[href]").each((_, element) => {
  const raw = $(element).attr("href").trim();
  if (/^(mailto:|tel:|javascript:|data:)/i.test(raw)) return;
  const resolved = new URL(raw, finalUrl);
  if (!["http:", "https:"].includes(resolved.protocol)) return;
  const fragment = resolved.hash.slice(1) || null;
  resolved.hash = "";
  const url = resolved.href;
  if (seen.has(url)) return;
  seen.add(url);
  const rel = ($(element).attr("rel") || "").toLowerCase().split(/s+/).filter(Boolean);
  links.push({
    url,
    fragment,
    text: $(element).text().replace(/s+/g, " ").trim(),
    nofollow: rel.includes("nofollow"),
    rel,
    internal: resolved.hostname.toLowerCase() === baseHost
  });
});

console.log(JSON.stringify({ requested_url: input, final_url: finalUrl, links }, null, 2));

Both scripts intentionally use raw response HTML. A browser automation workflow is required when completeness depends on client-side rendering.

Rank #2
Klein Tools VDV501-851 Scout Pro 3 Tester Starter Set Cable Tester
  • VERSATILE CABLE TESTING: Cable tester tests voice (RJ11/12), data (RJ45), and video (coax F-connector) terminated cables, providing clear results for comprehensive testing on unenergized Ethernet cables (not designed to test PoE)
  • EXTENDED CABLE LENGTH MEASUREMENT: Measure cable length up to 2000 feet (610 m), allowing for precise cable length determination
  • COMPREHENSIVE FAULT DETECTION: Test for Open, Short, Miswire, or Split-Pair faults, ensuring thorough fault detection and identification
  • BACKLIT LCD DISPLAY: Backlit LCD screen displays cable length, wiremap, cable ID, and test results, ensuring easy readability in various lighting conditions
  • EFFICIENT CABLE TRACING: Trace cables, wire pairs, and individual conductor wires using the multiple style tone generator (requires analog probe Cat. No. VDV500-123, sold separately), simplifying cable tracing tasks

Using Scrapy when you need filters or a crawl

Scrapy’s documented LxmlLinkExtractor returns Link objects and supports controls that become cumbersome in a one-off script:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Tags and attributes to scan (the documented defaults are a, area, and href).
  • Allowed or denied domains, URL regular expressions, and file-extension filters.
  • XPath or CSS regions, and link-text filters.
  • A process_value callback that can transform or discard an attribute before filtering.
  • Duplicate suppression and optional URL canonicalization.

Canonicalization deserves care. It can change the URL presented to the server. When the goal is to follow links robustly, the Scrapy documentation recommends retaining its default canonicalize=False; enable canonicalization only when your reporting definition calls for normalized URLs. Preserve the original href in your output so an audit can explain any transformation.

A typical crawl architecture is: parse the current response, extract links, classify them, keep only allowed crawl targets, and schedule those targets with a visited set. Add depth, page-count, host, and rate limits before running against a production site.

Internal and external URL classification

Policy Internal result When it is useful
Exact hostname Only the starting host Auditing one host without subdomains
Approved-host set Several named hosts Organizations using separate documentation or blog hosts
Registrable-domain policy Subdomains grouped together Brand-level inventories, provided you handle public-suffix boundaries correctly
Final-host policy Based on redirect destination Outbound-link or redirect analysis

Do not silently mix these policies. Record the starting URL, the rule, and whether classification uses the written URL or the final URL after redirects. A link to a different hostname may still be owned by the same organization, while a same-host URL can redirect elsewhere.

What to include in an audit result

A CSV or JSON export is more useful than a column of URLs. Include:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Source page and extraction timestamp.
  • Original href and resolved URL.
  • Final URL and HTTP status if you performed a follow-up request.
  • Anchor text and fragment.
  • Internal/external classification and the rule used.
  • rel tokens, including whether nofollow is present.
  • Occurrence count and source positions when duplicate placement matters.
  • Whether the value came from raw HTML or a rendered DOM.

Keep duplicate occurrences when the question is “where is this link used?” Keep one row per destination when the question is “which destinations exist?” These are different reports.

Rank #3
NOYAFA NF-8508 Network Cable Tester with Optical Power Meter
  • Multifunctional NOYAFA NF-8508 Network Cable Tester: There are nine features to meet your needs. Continuity Testing, Cable Scan, Port Flash, Length Measurement, POE Power Supply Test, QC testing, Optical Power Meter, VFL and NVC function.It is perfectly suited for various engineering cabling projects, network troubleshooting, network equipment maintenance and testing scenarios. Its precise cable scanning and fault localization capabilities help you effortlessly pinpoint the root cause of issues.
  • 7 WAVELENGTHS OPTICAL POWER METER: NF-8508 network cable tester can measure 7 standard wavelengths, 850/1300/1310/1490/1550/1625/1650, power detecting range(dBm): -70 ~ +10. Its power detection range spans from -70 dBm to +10 dBm, supporting FC/SC/ST connectors. It enables precise fiber optic power measurement, helping users efficiently assess fiber signal strength and ensure healthy fiber link operation. It effortlessly detects attenuation issues within fibers, thereby safeguarding fiber network stability.
  • High Efficiency Visual Fault Locator: Easy identification of fiber breakpoints, poor connections, bending or cracking. Excellent for finding the right fiber to splice or quickly finding a break. Emmiting Energy: standard wavelenth: 650nm. Fast flashing, slow flashing, high precison.The built-in self-calibration ensures stable long-term performance, and Class IIIa laser (output<5mW) ensures safe daily operation.
  • PORT FLASHING:The indicator light on the connection port in the NF-8508 device flashes to help accurately locate the cable. Displays port information, including operating speed, duplex mode, and negotiation settings. Port lights flash on the same screen to show the port's operating speed, making it easy to pinpoint lines and ports.
  • PoE Testing and Cable Length Test: PoE testing can check cable mapping polarity and voltage of PoE network switches, withstand 60VDC. Automatically detects and switches between 10M/100M/1000M modes, Includes cable tracking, short circuit test, interruption of circuit test and etc The RJ45 cable tester can quickly measure the length of the cable with a range of 200m. Not only network cables, but also phone lines and BNC cables.

Browser rendering and visual verification

Raw extraction is fast and deterministic, but it can miss links added by scripts. A browser-rendered workflow should wait for the page state you need, then inspect the live DOM. It also introduces cookie banners, newsletter popups, chat widgets, bot checks, timing differences, and higher resource use. If you need to confirm what a human sees, capture a rendered page separately from the link inventory; a screenshot is evidence of appearance, not a replacement for parsing href values.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a link extractor. It is useful when you need a clean visual check of a page before manually investigating a rendering discrepancy. Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—let Claude, Cursor, or another MCP client request captures.

One request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for parameters and response details:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/page -o page.webp

ScreenshotNeo has 1,000 free shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account to try it.

Filtering, normalization, and duplicate strategy

Skip non-page schemes

Decide whether to report or exclude mailto:, tel:, javascript:, data:, and custom application schemes. They are not HTTP destinations and should not be mixed into an HTTP status report.

Preserve versus normalize

Keep the original href for auditability. Store a second resolved form for comparison. Removing fragments is usually sensible for destination-level counts; changing case, decoding characters, sorting query parameters, or removing trailing slashes can alter server behavior. Apply only transformations you can explain.

Duplicates

Repeated navigation links are normal. Count them when measuring placement, but deduplicate when producing a destination inventory. A visited set for a crawl should use the same normalization policy as the report, or the crawler may revisit equivalent-looking URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability, and access limits

  • Fetch responsibly: set timeouts, identify your user agent, honor the site’s access policy, and limit concurrency.
  • Cache responses: save the HTML and response URL so you can rerun parsing without refetching.
  • Separate extraction from validation: first collect links; then perform bounded status checks. A link can be valid even when a later check is blocked or rate-limited.
  • Handle encodings: use the response’s declared encoding where possible and retain replacement characters rather than silently dropping text.
  • Record failures: distinguish DNS errors, timeouts, HTTP errors, redirects, and parser errors.
  • Control crawl scope: deny external hosts by default, cap depth and page count, and avoid infinite calendar, search, or session URLs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

The list is empty

Inspect the saved response. You may have received an interstitial, an error page, or a page whose links are injected by JavaScript. Check the HTTP status, final URL, and response body before changing the parser.

Rank #4
Sale
iMBAPrice - RJ45 Network Cable Tester for Lan Phone RJ45/RJ11/RJ12/CAT5/CAT6/CAT7 UTP Wire Test Tool
  • Automatically runs all tests and checks for continuity, open, shorted and crossed wire pairs. Visible LED status display.
  • Cable state testing (2-wire): Line DC detecting, anode and cathode determination,Ringing signal detecting open, short and cross circuit testing
  • Cable Type: RJ11 Telephone cable and RJ45 LAN cable
  • Connectors: Ethernet Cat 5, Ethernet Cat 5e, Ethernet Cat 6, Ethernet Cat 7, RJ11 6P and RJ45 8P
  • Power Source: DC9V Battery Required (not included)

Relative links point to the wrong host

Resolve against the final response URL, not the command-line input, because redirects can change the document base. Also account for an HTML <base href> element if the target site uses one.

Internal links are labeled external

Print the parsed hostnames and compare them with your stated policy. Decide whether www, language subdomains, documentation hosts, and alternate domains belong in the approved set.

Too many duplicates

Choose whether fragments, query strings, trailing slashes, and URL case are significant. Apply the decision consistently to both the visited set and the output key; retain the raw href for review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The server returns 403 or 429

Do not bypass access controls aggressively. Slow the crawl, reduce concurrency, identify your client, and use an authorized authenticated session if you have one. Treat a blocked fetch as an unverified result rather than evidence that the destination is broken.

Important links appear only after interaction

Use browser automation or inspect the network/API request that supplies the content. A raw-HTML extractor cannot infer links that do not exist in the fetched markup.

A practical workflow

  1. Write down the starting URL, hostname boundary, rendering mode, and whether fragments count separately.
  2. Fetch the page with a timeout and save the response metadata.
  3. Extract configured tags and attributes, resolve URLs, and preserve anchor and relationship metadata.
  4. Filter schemes and destinations according to the policy.
  5. Export both the original href and normalized/resolved values.
  6. For a crawl, enqueue only approved targets and enforce depth, page, host, and rate limits.
  7. Run optional status checks separately, recording redirects and failures without overwriting extraction results.
  8. Use a rendered browser check when JavaScript-generated content or visual state is material.

FAQ

Does a nofollow value prevent extraction?

No. It is metadata on the link. Your extractor can report it, while a crawler may use it as one input to its own following policy.

Best Value
Network Ethernet Cable Tester for LAN RJ45 RJ11 CAT5 CAT5E CAT6 CAT6A CAT7, Ethernet Wire Tester Tool UTP/STP Continuity Test for Telephone Line Finder Home Repair (HT812A)
  • Multi-Function Network Cable Tester: Supports RJ45 (CAT5, CAT5e, CAT6, CAT6A, CAT7) and RJ11 telephone cables. Quickly detects continuity, short circuits, open wires, miswiring, and cable shielding status, ensuring your LAN or phone lines are correctly wired and ready to use.
  • Fast/Slow Mode with LED Indicators: Switch between fast and slow scan speeds to identify wiring issues more precisely. LED lights on both master and remote units show wire order, making it easy to spot errors like open pairs or misaligned pins at a glance.
  • Split-Type Design for Long-Distance Testing: Master and remote units can be detached and used separately, allowing you to test both ends of a long cable run, ideal for wall-mounted ports, long runs, or structured cabling. Perfect for home, office, or professional IT setups.
  • Compact, Lightweight & Durable: Ergonomically designed with sturdy ABS housing, this pocket-sized tester is ideal for on-the-go network engineers, DIYers, and electricians. It’s your go-to toolkit for cable maintenance, upgrades, or new installations.
  • Safe & Easy to Use: Simple one-button operation makes testing quick and hassle-free. LED indicators clearly show wiring status, while the G light instantly identifies shielded (FTP/STP) or unshielded (UTP) cables. Supports safe testing of telephone lines with typical voltages under 48-72V, ideal for both home and professional use.

Should URL fragments be sent to the server?

No. Fragments are interpreted by the client after the HTTP request. Keep them as metadata when the destination section matters, but do not expect a server status check to validate the fragment itself.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can an extractor prove that every visible link is present?

Only within its rendering boundary. A raw-response parser can be complete for the HTML it received, but it cannot prove completeness for links created later by scripts, user interaction, or an authenticated browser session.

Frequently Asked Questions

Does a nofollow value prevent extraction?

No. It is metadata on the link. Your extractor can report it, while a crawler may use it as one input to its own following policy.

Should URL fragments be sent to the server?

No. Fragments are interpreted by the client after the HTTP request. Keep them as metadata when the destination section matters, but do not expect a server status check to validate the fragment itself.

Can an extractor prove that every visible link is present?

Only within its rendering boundary. A raw-response parser can be complete for the HTML it received, but it cannot prove completeness for links created later by scripts, user interaction, or an authenticated browser session.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.