Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

First fetch the page or Markdown file at the URL; then parse its content as Markdown. Python’s urllib.parse can split and resolve URLs, but it does not extract Markdown links or validate a URL against a web standard. For Markdown links—including reference links—and email autolinks, use a CommonMark parser rather than trying to treat the document as one long URL string.

What you are extracting—and from where

A URL identifies a resource; it is not the resource’s Markdown content. To extract links or email addresses, separate the task into two stages: obtain the content, then interpret it according to its format. If the URL points to a Markdown file, parse the returned text as Markdown. If it points to an ordinary website, the response may be HTML rather than Markdown, so a Markdown parser is not the right parser for that response.

Be clear about what “email addresses” means. CommonMark recognizes email autolinks such as <[email protected]>, whose destination is mailto:[email protected]. It does not follow that every address-shaped string in ordinary prose is a Markdown link. Nor does extracting an address prove that the mailbox exists or accepts mail.

The code below extracts Markdown link destinations and email autolinks from a Markdown document fetched over HTTP or HTTPS. It resolves relative link destinations against the document URL, and prints the Markdown links and email addresses separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse the URL before you fetch the document

Python’s urllib.parse provides URL-component parsing, recombination, and relative-reference resolution. Its familiar components include scheme, network location, path, query, and fragment; urlparse also exposes path parameters. The documentation calls netloc a legacy term; RFC 3986 uses “authority.” See the Python 3.15 urllib.parse documentation.

from urllib.parse import urlparse, urljoin

page_url = "https://example.com/guides/start.md?lang=en#intro"
parts = urlparse(page_url)

print(parts.scheme)    # https
print(parts.netloc)    # example.com
print(parts.path)      # /guides/start.md
print(parts.query)     # lang=en
print(parts.fragment)  # intro

absolute_link = urljoin(page_url, "../contact.md")
print(absolute_link)   # https://example.com/contact.md

Parsing is not validation. Python notes that urllib.parse combines historical behavior and aspects of multiple conventions, and cannot be claimed compliant with either RFC 3986 or the WHATWG URL standard. If your application requires a particular standard’s treatment of unusual inputs, check those cases against that requirement rather than treating a successful parse as proof of validity.

Use a Markdown parser for links and email autolinks

CommonMark defines several forms that a text search can miss or misread: inline links, reference links, URI autolinks, and email autolinks. For example, an inline link contains both visible label text and a destination; a reference link gets its destination from a separate definition elsewhere in the document. A parser can connect those pieces according to the Markdown syntax. CommonMark describes autolinks as absolute URIs and email addresses inside angle brackets; its email pattern is non-normative, not a guarantee of deliverability. Read the CommonMark specification for the syntax rules.

This example uses the Python commonmark parser. Install it in the same Python environment in which you run the script:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m pip install commonmark

Save the following as extract_markdown.py, then pass it an HTTP or HTTPS URL that returns Markdown:

import sys
from urllib.parse import urljoin, urlparse
from urllib.request import Request, urlopen

from commonmark import Parser


def fetch_markdown(page_url):
    parts = urlparse(page_url)
    if parts.scheme not in {"http", "https"} or not parts.netloc:
        raise ValueError("Pass a complete HTTP or HTTPS URL")

    request = Request(page_url, headers={"User-Agent": "MarkdownLinkExtractor/1.0"})
    with urlopen(request, timeout=30) as response:
        final_url = response.geturl()
        content = response.read()

    # This example assumes the Markdown response is UTF-8.
    return content.decode("utf-8"), final_url


def extract(markdown_text, base_url):
    document = Parser().parse(markdown_text)
    walker = document.walker()
    links = []
    emails = []

    while True:
        event = walker.nxt()
        if event is None:
            break
        node, entering = event
        if not entering or node.t not in {"link", "image"}:
            continue

        destination = node.destination or ""
        if destination.lower().startswith("mailto:"):
            # Keep the address portion; query parameters are not part of it.
            address = destination[7:].split("?", 1)[0]
            if address:
                emails.append(address)
        elif destination:
            links.append(urljoin(base_url, destination))

    # Preserve document order while removing repeated destinations.
    return list(dict.fromkeys(links)), list(dict.fromkeys(emails))


def main():
    if len(sys.argv) != 2:
        raise SystemExit("Usage: python extract_markdown.py MARKDOWN_URL")

    markdown_text, final_url = fetch_markdown(sys.argv[1])
    links, emails = extract(markdown_text, final_url)

    print("Markdown link and image destinations:")
    for link in links:
        print(link)

    print("Email autolinks:")
    for address in emails:
        print(address)


if __name__ == "__main__":
    main()

Run it like this:

python extract_markdown.py https://example.com/guide.md

The parser produces link nodes for Markdown link syntax, including reference links, and email autolinks have mailto: destinations. The example includes image destinations in the first output list because Markdown images also have destinations; remove "image" from the node-type set if you want links only. The script resolves relative destinations using the final URL after a redirect, so a link such as ../contact is interpreted relative to the document actually returned.

Choose the right output for your use case

Keep destinations, not just visible labels

For a Markdown link such as [Support](../support), the human-readable label is “Support,” while the destination is ../support. The script prints the destination after resolving it against the page URL. If your next step needs both label and destination—for example, building a link inventory—collect the parser node’s child text as well as node.destination. Do not substitute the visible label for the URL.

Decide whether image destinations count

Markdown image syntax has a destination too. The sample deliberately includes images in the destinations list and labels it accordingly. For a report limited to clickable text links, change the accepted node types to {"link"}. Make that choice explicit so the output matches the question your downstream code is answering.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not mistake address-shaped text for an email autolink

The email list contains addresses represented as CommonMark email autolinks, not every string that resembles an address in prose or code. Broadening the search to plain text requires additional rules for what counts as an address, and extraction alone still cannot establish that it is deliverable. Keep syntax extraction and mailbox verification as separate requirements.

Handle URL fragments and queries deliberately

The input URL’s query and fragment identify parts of the requested resource or page location; they are not themselves Markdown link destinations. The script fetches the supplied URL and resolves relative Markdown destinations against the final response URL. If your application needs to preserve a particular fragment or query from a base URL, test that case and confirm it matches the URL behavior your application expects.

Common failures and what to check

  • The URL parses but the fetch fails. Parsing only splits the input into components; it does not establish that the resource is reachable. Check the scheme, hostname, network access, server response, and whether the URL actually serves Markdown.
  • The parser finds no links. Confirm the fetched body is Markdown rather than HTML, plain text, or a redirect/error page. Inspect a small portion of the returned text and verify that the source uses Markdown link syntax.
  • Relative links point somewhere unexpected. Resolve against the final response URL, not necessarily the originally requested address. The code uses response.geturl() for that base. If the document uses a different base convention, apply that convention explicitly.
  • Some email addresses are missing. The example extracts CommonMark email autolinks, not unlinked address-like text. Add a separate, clearly specified text-extraction policy if the requirement includes prose addresses.
  • Characters appear corrupted. The sample assumes UTF-8. If the server returns content in another encoding, decode according to the response’s actual encoding rather than silently treating the output as correct.
  • A URL looks valid but behaves differently in another system. Python cautions that urllib.parse is not a claim of compliance with RFC 3986 or WHATWG URL. Test inputs at the edge of your application’s accepted URL rules.
  • The destination list includes images or duplicates. The sample includes image nodes and removes duplicate output destinations while preserving order. Restrict node types or retain duplicates if your reporting rules require different behavior.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Operational notes for a reliable extractor

The sample uses a 30-second network timeout, rejects non-HTTP(S) input, and assumes a UTF-8 Markdown response. Those are explicit choices in this example, not universal requirements. For a production service that accepts URLs from other users, also define allowed hosts, response-size limits, redirect policy, and content-type handling according to your environment; otherwise a URL that is syntactically acceptable may still be unsuitable to fetch. Add retries only if your application’s failure policy calls for them, and avoid retrying indefinitely.

Deduplication is useful for a compact inventory, but it discards repetition and the locations where each destination occurred. If you need an audit trail, store the source node or document position for each occurrence instead of converting output to a unique list. Likewise, preserve raw destinations separately if you need to distinguish the Markdown spelling from its resolved URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

ScreenshotNeo captures a page visually; it does not return Markdown link nodes or extract email addresses. Use the parser method above when you need structured destinations. If you also need a clean visual capture of the page, this single request returns a screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Frequently asked questions

Does a parsed URL prove that it is safe to request?

No. Component parsing and fetching are separate operations. Applications that accept user-supplied URLs should set their own host, redirect, and resource limits before making requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Will extracting an email address tell me whether it is active?

No. Syntax identifies an address-like destination; it does not confirm that a mailbox exists or can receive messages.

Can I use this method on any webpage?

Only when the response you parse is Markdown. A regular webpage commonly supplies HTML, which requires an HTML-aware extraction approach rather than CommonMark parsing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.