The reliable way to extract URLs is a two-stage pipeline: first locate URL-like text, then trim surrounding punctuation and validate each candidate with a real URL parser. A regular expression is good at finding candidates; it is not a complete validator. This approach handles prose punctuation, relative links, fragments, internationalized domains and security checks without treating every regex match as a usable link.
The URL-extraction pipeline
- Locate candidates. Search for explicit schemes such as
https://,http://and, when required,ftp://. Protocol-relative references beginning with//need a trusted base URL. - Trim context. Remove wrappers such as quotes, angle brackets and sentence punctuation only when they are outside the URL. Balanced parentheses can be legitimate path characters.
- Parse. Use a standards-aware URL API to separate scheme, host, path, query and fragment.
- Apply policy. Allow only schemes, hosts, ports and credential rules that your application actually supports.
- Normalize and deduplicate. Keep the original spelling for display, but compare a carefully normalized form. Do not blindly lowercase paths or decode percent escapes.
This separation matters because URI syntax permits characters that commonly appear as prose delimiters. RFC 3986 describes parsing a URI reference into components and warns that punctuation can be mistaken for part of the URI.
Choose the right extractor for your input
| Input | Preferred method | Why |
|---|---|---|
| Controlled HTML | HTML parser and link nodes | Reads href attributes without guessing where a link starts or ends. |
| Markdown | Markdown parser | Understands inline links, reference links and angle-bracket autolinks. |
| Plain prose, logs or chat | Candidate regex plus URL parser | Finds links embedded in otherwise unstructured text. |
| Known base URL with relative links | Parser plus explicit URL resolution | Turns /docs or ../img.png into absolute URLs only when the base is trusted. |
Python: a production-minded extractor
The following script finds HTTP, HTTPS and FTP candidates, removes common trailing punctuation, parses them with Python’s standard library, removes fragments and rejects malformed or disallowed values.
import re
from urllib.parse import urlsplit, urldefrag
candidate_re = re.compile(r'(?i)\b(?:https?|ftp)://[^\s<>"\']+')
ALLOWED_SCHEMES = {"http", "https", "ftp"}
def trim_trailing(raw):
value = raw
# Remove punctuation usually added by surrounding prose.
while value and value[-1] in ".,;:!?":
value = value[:-1]
# A closing delimiter is removable only when it is unmatched.
pairs = [("(", ")"), ("[", "]"), ("{", "}")]
changed = True
while changed and value:
changed = False
for opening, closing in pairs:
if value.endswith(closing) and value.count(closing) > value.count(opening):
value = value[:-1]
changed = True
return value
def extract_urls(text):
results = []
for raw in candidate_re.findall(text):
cleaned = trim_trailing(raw)
try:
parts = urlsplit(cleaned)
except ValueError:
continue
if parts.scheme not in ALLOWED_SCHEMES or not parts.netloc:
continue
# Accessing hostname can raise for malformed ports or brackets.
try:
host = parts.hostname
port = parts.port
except ValueError:
continue
if not host:
continue
url_without_fragment, _fragment = urldefrag(cleaned)
results.append(url_without_fragment)
return results
text = 'Read https://example.com/docs, then visit <https://example.org/a(b)>.'
print(extract_urls(text))
urlsplit() separates the components without fetching the address. urldefrag() removes a fragment when your application treats /page#section and /page as the same resource. Keep the fragment if it is meaningful to your output.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems#1 Best Overall
- USB-C 2-in-1 storage OTG: The Lexar JumpDrive Dual Drive D40E features USB Type-A and Type-C connectors in a slim, portable form factor for easy device compatibility
- Transfer speeds up to 100MB/s: Based on internal testing, performance may vary depending upon the host device, interface, and usage conditions. 1MB=1,000,000 bytes
- Plug and Play: Widely compatible with USB Type-C smartphones, tablets, laptops, Macs, and traditional Type-A devices, no software installation required. The 360° swivel design allows for easy switching between connectors without the hassle of losing a cap
- Durable & Compact: The Lexar D40E USB memory stick features a metal enclosure, withstands temperatures from 0° to 50° C (32°F to 122°F), and is lightweight at 26g with dimensions of 70.4 x 16.9 x 11.7mm
- Security & Warranty: Securely protects files using an advanced security software solution with 256-bit AES encryption. Backed by a Lexar 3-year limited warranty
Resolving relative references in Python
A value such as /docs/page is not an absolute URL. Resolve it only with a known, trusted base:
from urllib.parse import urljoin
absolute = urljoin("https://example.com/guide/", "/docs/page")
# https://example.com/docs/page
Never invent a base URL from untrusted text. If no trustworthy base exists, retain the value as a relative reference and label it accordingly.
JavaScript: extract and validate with URL
function trimTrailing(raw) {
let value = raw.replace(/[.,;:!?]+$/, "");
let changed = true;
while (changed) {
changed = false;
for (const [open, close] of [["(", ")"], ["[", "]"], ["{", "}"]]) {
if (value.endsWith(close) && (value.match(new RegExp("\\" + close, "g")) || []).length > (value.match(new RegExp("\\" + open, "g")) || []).length) {
value = value.slice(0, -1);
changed = true;
}
}
}
return value;
}
function extractUrls(text, baseUrl) {
const rough = text.match(/\b(?:https?|ftp):\/\/[^\s<>"']+/gi) ?? [];
return rough.flatMap(raw => {
const cleaned = trimTrailing(raw);
try {
const parsed = new URL(cleaned, baseUrl);
if (!["http:", "https:", "ftp:"].includes(parsed.protocol)) return [];
if (!parsed.hostname) return [];
parsed.hash = ""; // Remove this line if fragments must be preserved.
return [parsed.href];
} catch {
return [];
}
});
}
console.log(extractUrls("See https://example.com/a?x=1."));
The URL constructor performs parsing and normalization. In environments that support it, URL.canParse() can provide a quick validity check before constructing the object. A relative string is accepted only when the supplied base URL is valid.
Rank #2
- High-speed USB 3.0 performance of up to 150MB/s(1) [(1) Write to drive up to 15x faster than standard USB 2.0 drives (4MB/s); varies by drive capacity. Up to 150MB/s read speed. USB 3.0 port required. Based on internal testing; performance may be lower depending on host device, usage conditions, and other factors; 1MB=1,000,000 bytes]
- Transfer a full-length movie in less than 30 seconds(2) [(2) Based on 1.2GB MPEG-4 video transfer with USB 3.0 host device. Results may vary based on host device, file attributes and other factors]
- Transfer to drive up to 15 times faster than standard USB 2.0 drives(1)
- Sleek, durable metal casing
- Easy-to-use password protection for your private files(3) [(3)Password protection uses 128-bit AES encryption and is supported by Windows 7, Windows 8, Windows 10, and Mac OS X v10.9 plus; Software download required for Mac, visit the SanDisk SecureAccess support page]
Regex patterns: useful, but deliberately incomplete
A practical locator for common absolute links is:
(?i)b(?:https?|ftp)://[^s<>"']+
It intentionally stops at whitespace, angle brackets and quotes. It will still capture a period at the end of a sentence, a closing parenthesis from prose, or a malformed host. That is why trimming and parser validation must follow it.
Protocol-relative and bare domains
To support //cdn.example.com/file.js, find that form separately and resolve it with a trusted base scheme. Bare text such as example.com is ambiguous: it may be a hostname, a product name or an email fragment. Only support bare domains with a clearly defined product policy, and validate them more strictly than explicit-scheme URLs.
HTML and Markdown require parsers
If the source is HTML, select a[href], area[href] and any other link-bearing elements defined by your application. Read the attribute value, then parse and validate it. Do not run a prose regex over HTML: script contents, attributes and encoded entities create false positives.
Rank #3
- What You Get - 2 pack 64GB genuine USB 2.0 flash drives, 12-month warranty and lifetime friendly customer service
- Great for All Ages and Purposes – the thumb drives are suitable for storing digital data for school, business or daily usage. Apply to data storage of music, photos, movies and other files
- Easy to Use - Plug and play USB memory stick, no need to install any software. Support Windows 7 / 8 / 10 / Vista / XP / Unix / 2000 / ME / NT Linux and Mac OS, compatible with USB 2.0 and 1.1 ports
- Convenient Design - 360°metal swivel cap with matt surface and ring designed zip drive can protect USB connector, avoid to leave your fingerprint and easily attach to your key chain to avoid from losing and for easy carrying
- Brand Yourself - Brand the flash drive with your company's name and provide company's overview, policies, etc. to the newly joined employees or your customers
For Markdown, use a Markdown parser’s link and autolink nodes. A link destination may be surrounded by parentheses or angle brackets, and escaped characters can change its meaning. Parser output also lets you distinguish visible text from the actual destination.
Cleaning punctuation without breaking real URLs
- Sentence punctuation: trim a final comma or period when it is outside the URL.
- Parentheses: preserve balanced parentheses, which are legal in many paths; remove only an unmatched closing delimiter.
- Wrappers: support forms such as
<https://example.com>, quoted URLs and legacyURL:prefixes when your input format uses them. - Line wrapping: treat inserted newlines or spaces as suspicious. Do not silently join wrapped text unless the source format documents that behavior.
- Fragments: remove them only when your application does not need in-page targets.
Validation and security policy
Parsing answers “can this string be interpreted as a URL?” It does not answer “should my service navigate to or fetch it?” Apply an explicit policy before using results.
- Allow only the schemes you need, commonly
httpsand optionallyhttp. Rejectjavascript:,data:and unexpected custom schemes when links will be opened or fetched. - Require a nonempty host for network URLs.
- Set acceptable port rules and reject malformed bracketed IPv6 or invalid ports.
- Treat userinfo such as
https://user:[email protected]as sensitive. Remove it, reject it or require a specific business reason. - Be cautious with unusual IP representations, internationalized domains and redirects. Apply your network’s SSRF protections separately from string validation.
- Do not fetch a candidate merely because a regex matched it.
Normalization and deduplication
Store both the original candidate and a normalized comparison value when auditability matters. URL libraries may normalize case in schemes and hosts, resolve dot segments or encode characters. Do not lowercase paths, decode percent escapes or remove a trailing slash unless the target scheme and your application define those operations as equivalent. Deduplicate only according to those known semantics.
Rank #4
- GOOD VALUE PACKAGE - 1 Pack 32GB Memory Stick USB 2.0 Flash Drives with great cost performance and high quality.
- BIG CAPACITY - The available capacity: 29.10GB-29.8GB, You can save the data of movies, music, photos, designs, programs, manuals, handouts in a high speed.Good performance in digital data storing, transferring and sharing with families, friends, workmates, clients and machines.
- EASY TO USE & PLUG AND WORK - Support windows 7 / 8 / 10 / Vista / XP / 2000 / ME / NT Linux and Mac OS, Compatible with USB2.0 and below.
- TWISTTURN DESIGN & EASY CARRY - The metal clip rotates 360° round the ABS plastic body which with rubber oil skin feeling finish. The capless design can avoid lossing of cap, and providing efficient protection to the USB port.
- WARRANTY & SUPPORT - SIMMAX logo is laser printed on the USB connector surface, our products are of good quality and we promise that any problem about the product within one year since you buy.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| A period or comma appears in every result | Regex match stopped only at whitespace | Trim sentence punctuation, then parse. |
| Valid parenthesized paths are truncated | All closing parentheses were removed | Remove a delimiter only when closing characters outnumber opening characters. |
/docs is rejected |
It is a relative reference | Resolve with a trusted base or keep it relative. |
| HTML extraction finds scripts and styles | Regex was run over markup | Parse HTML and read link attributes. |
| Malformed ports crash parsing | Port or bracket validation is lazy in some APIs | Catch parser exceptions and access host/port inside the guarded block. |
| Dangerous links are accepted | Syntax validation was mistaken for policy validation | Allow-list schemes and enforce host, credential and network rules. |
| Duplicate links remain after normalization | Different textual forms were compared literally | Use a documented normalized key while preserving original text. |
Performance and reliability
For large documents, scan once with a compiled regex or streaming tokenizer, then parse each candidate. Put a sensible maximum input size and candidate count in place so hostile text cannot consume unbounded CPU or memory. Cache parsed results when the same document is processed repeatedly. Network fetching is a separate, slower operation: use timeouts, redirect limits, DNS and private-network protections, and never make extraction depend on a successful fetch unless your product explicitly requires reachability checks.
Or skip the browser setup
If your workflow also needs a clean image or PDF of the page that contains the text, ScreenshotNeo provides a website screenshot API and MCP server. It is not a URL parser; use the extraction pipeline above for text. It can, however, capture the source page before a human or agent reviews it.
One GET request returns an image or PDF. See the ScreenshotNeo API documentation for all options:
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
- 【16GB Flash Drive】USB flash drives with 16GB capacity, meet your needs of daily use on work, school, home and travelling for photos, music, videos, files storage and transfer. IMEASON thumb drives can be used to store different files, easy to data backup.
- 【Metal Swivel Cap Design】USB thumb drive is metal swivel cover provides extra protection for the usb thumbdrive connector, no usb drive cap to lose; keychain design makes it easier to carry without worrying lose it.
- 【Wide Compatibility】USB drive supports Windows 7/8/10/11 / Vista / XP / Unix / 2000 / ME / NT Linux and Mac OS, also Supports USB 2.0 and 1.1 ports. USB Stick support TV, desktop, notebook computer, car, audio and other device. The USB Memory Stick is your great data storage and transfer companion with traveling and working.
- 【Easy to use】usb memory stick is plug and play without any software installation. Just simply plug the Flashdrive into the port of your USB-compatible devices such as computer, laptop to start data storage or transmission.
- 【What You Get】16 GB USB Flash Drive Thumb Drive, The default format of the usb storage flash drive is FAT32.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, cookie and consent banners, newsletter popups and chat widgets are removed. Bot checks, blank pages and failed loads are not billed, and response headers identify the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info and capture_pdf for AI clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently Asked Questions
Should I remove URL fragments?
Only if your application treats in-page anchors as irrelevant. Otherwise preserve the fragment and deduplicate using a policy that distinguishes page identity from navigation targets.
Can one regex validate every possible URL?
No. URL syntax, scheme-specific rules and application security policy are separate concerns; use a locator regex followed by a parser and explicit checks.
How should email addresses be handled?
Treat them as a separate extraction type. A URL regex should not infer a web address from an email domain unless your product explicitly defines that behavior.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Is checking that a URL responds part of extraction?
No. Reachability is a later network operation with its own timeout, redirect and security controls.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

