Load the HTML into Cheerio, select anchors with $('a'), and read each anchor’s href attribute. Use .attr('href') for the literal value in the markup, map the selection to collect every link, and use .prop('href') with a document URL when you need absolute URLs.
The shortest working example
This complete ES module loads two anchors and returns their raw href strings:
import * as cheerio from 'cheerio';
const html = `
<a href="/docs">Docs</a>
<a href="https://example.com/blog">Blog</a>
`;
const $ = cheerio.load(html);
const links = $('a').map((_, element) => $(element).attr('href')).get();
console.log(links);
// [ '/docs', 'https://example.com/blog' ]
The a selector matches every anchor element. attr('href') reads the text of the attribute exactly as it appears in the source, and get() converts Cheerio’s collection into a normal JavaScript array.
Install Cheerio first if it is not already in your project:
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
npm install cheerio
Cheerio’s selector syntax is covered in the official selecting guide, while its manipulation guide documents reading attributes and properties.
Get one link or all links
Read the first matching anchor
const $ = cheerio.load(html);
const firstHref = $('a').attr('href');
console.log(firstHref);
When a selection contains several elements, attr('href') reads the first one. If no anchor matches, or the first matching anchor has no href attribute, the result is undefined.
Collect every href
const hrefs = $('a')
.map((_, element) => $(element).attr('href'))
.get();
This preserves document order. It can include undefined for anchors without an href, so filter those values when your output must contain only actual attributes:
const hrefs = $('a')
.map((_, element) => $(element).attr('href'))
.get()
.filter((href) => typeof href === 'string' && href.length > 0);
Select a narrower group
Use normal CSS selectors when the page contains several kinds of links:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
const navigation = $('nav a').map((_, el) => $(el).attr('href')).get();
const externalCandidates = $('a[href^="http"]').map((_, el) => $(el).attr('href')).get();
const articleLinks = $('article a[href]').map((_, el) => $(el).attr('href')).get();
The [href] attribute selector excludes anchors that do not have an href at all. A prefix selector such as [href^="http"] is only a quick filter; it does not prove that a value is valid or that it points to another site.
Raw href values versus absolute URLs
Cheerio does not rewrite the string returned by attr('href'). For <a href="/docs">, the result is /docs. That is useful when you need to reproduce the source markup, compare templates, or preserve exactly what the publisher supplied.
Rank #2
| Method | Result | Document URL required? | Best use |
|---|---|---|---|
attr('href') |
Literal attribute string | No | Preserve or inspect source markup |
prop('href') |
URL resolved against the document URL | Yes for relative values | Fetch, compare, or export absolute links |
$.extract({ links: [{ selector: 'a', value: 'href' }] }) |
All selected values in a declarative shape | Resolution depends on a document URL | Combine links with other extracted fields |
Resolve a relative href with a base URI
import * as cheerio from 'cheerio';
const $ = cheerio.load('<a href="/docs">Docs</a>', {
baseURI: 'https://example.com/articles/page.html',
});
console.log($('a').prop('href'));
// https://example.com/docs
The base URI tells Cheerio which origin and path to use. Without it, a relative attribute remains relative. Cheerio’s troubleshooting documentation explains this distinction and its manipulation guide documents prop('href').
Use a URL-aware loader
When you load a page directly from a URL with Cheerio’s URL loader, the document URL is available for property resolution. The same principle applies: use attr when you want the original text and prop when you want Cheerio’s URL-aware property.
Recommended Free Tools
Use the extract API for structured output
Cheerio’s extract method is convenient when links are one field in a larger result:
const data = $.extract({
links: [{ selector: 'a', value: 'href' }],
});
console.log(data);
// { links: ['/docs', 'https://example.com/blog'] }
An array descriptor collects all matches. A selector descriptor without the array returns the first match. The value: 'href' descriptor uses Cheerio’s property API, so relative values are resolved only when the loaded document has a URL or a configured base URI.
You can also extract text and links from repeated records:
const cards = $.extract({
products: [{
selector: '.product-card',
value: {
name: '.name',
href: { selector: 'a', value: 'href' },
},
}],
});
For a simple list of anchors, map(...).get() is usually easier to read; use extract when the output has several related fields.
Build a dependable link extractor
Keep the source value and resolved value together
Storing both forms avoids losing information and makes later decisions explicit:
const links = $('a[href]').map((_, element) => {
const anchor = $(element);
const raw = anchor.attr('href');
const absolute = anchor.prop('href');
return {
text: anchor.text().trim(),
raw,
absolute,
};
}).get();
If no base URI was supplied, absolute may still be a relative value. Do not label it absolute unless you have provided a document URL and checked the result.
Resolve values yourself when you need strict control
The standard URL constructor is useful when you want explicit error handling or need to retain special schemes:
const base = 'https://example.com/articles/page.html';
const resolved = $('a[href]').map((_, element) => {
const raw = $(element).attr('href');
try {
return { raw, url: new URL(raw, base).href };
} catch {
return { raw, url: null };
}
}).get();
This lets you identify malformed values instead of allowing one bad attribute to stop the entire extraction. Values such as fragments, mailto:, and javascript: are not ordinary HTTP pages; classify or exclude them according to your application rather than silently treating every string as a fetchable web URL.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRemove duplicates only when that is your requirement
const uniqueHrefs = [...new Set(
$('a[href]')
.map((_, el) => $(el).prop('href'))
.get()
)];
Deduplication can hide meaningful differences if the same destination appears in separate navigation areas or if query strings and fragments matter. Decide whether identity means the literal string, the resolved URL, or a normalized URL before using a Set.
Loading HTML correctly
Cheerio parses the markup you provide. Its normal loader treats input as a complete document and may add missing document structure. If you are parsing a fragment copied from a component rather than a full page, use Cheerio’s fragment mode as described in the official troubleshooting guide.
Rank #4
const fragment = '<a href="/one">One</a><a href="/two">Two</a>';
const $ = cheerio.load(fragment, null, false);
const hrefs = $('a').map((_, el) => $(el).attr('href')).get();
When your input comes from an HTTP client, pass the response body to Cheerio and retain the final response URL as the base URI if redirects or relative links matter. Cheerio itself is the parser; your HTTP client is responsible for downloading the response.
Why a link may be missing
The page creates it with JavaScript
Cheerio is not a web browser. It does not execute page JavaScript, wait for client-side rendering, click controls, or run framework hydration. If an anchor is inserted only after browser code runs, it will not exist in the static HTML passed to Cheerio. The official introduction points to browser automation tools such as Puppeteer or Playwright, and DOM emulation such as jsdom, for cases that require execution.
Before changing selectors, inspect the exact response body. A link visible in a browser’s Elements panel may be absent from the original HTTP response.
The anchor has no href
Some markup uses an element styled as a link without an href. Select with a[href] when those should be excluded, or inspect the element’s other attributes if your application intentionally supports them.
The selector is too narrow
Start with $('a').length, then test progressively specific selectors such as main a or a[href]. A typo in a class name or an unexpected iframe boundary can produce an empty selection even though links exist elsewhere in the response.
Troubleshooting table
| Symptom | Likely cause | Fix |
|---|---|---|
undefined from attr('href') |
No matching anchor, or the first match has no href | Check $('a').length, use a[href], and inspect the input HTML. |
| Only one link is returned | attr() reads the first element in a selection |
Use map(...).get() or an array descriptor in $.extract. |
| A relative path stays relative | attr() returns the literal source value |
Supply baseURI or a URL-aware loader and read prop('href'). |
| Browser shows a link but Cheerio finds none | The link is inserted after JavaScript executes | Obtain server-rendered HTML or use browser automation before parsing. |
Unexpected html/body wrappers |
Document mode adds structure to a fragment | Load the input in fragment mode when it is not a complete document. |
| Extraction stops on one malformed URL | URL construction throws for an invalid value | Wrap new URL() in try/catch and retain the raw value for review. |
Performance and reliability practices
- Parse once and reuse the same
$function for all selectors instead of loading identical HTML repeatedly. - Use a specific selector such as
article a[href]when site navigation and footer links are irrelevant; this reduces downstream filtering. - Keep extraction separate from downloading. Cheerio does not fetch every discovered URL, so queue and rate-limit follow-up requests in your HTTP layer.
- Record the source URL, retrieval time, and raw href when results may need auditing. Relative links cannot be reproduced correctly without their base document.
- Treat untrusted HTML as data. Validate destinations before making requests, and apply your application’s policy for non-HTTP schemes, credentials, localhost addresses, and redirects.
- Do not assume a successful parse means the page was complete. A server response can be an error page, a consent wall, or a partial shell even when Cheerio reports no parsing error.
Or skip the browser setup
If your immediate need is a rendered visual capture rather than programmatic href extraction, ScreenshotNeo provides a single-request website screenshot API. It is separate from Cheerio: it returns a PNG, JPEG, WebP, or PDF, not an array of links. It is useful when you need to verify what a browser-rendered page looks like before deciding how to obtain its HTML.
Best Value
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for all parameters. Equivalent examples:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and whether the request was billed.
- An MCP server exposes
take_screenshot,get_page_info, andcapture_pdfto Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is included on every plan.
When you are ready to try it, sign up for the free ScreenshotNeo plan.
Choosing the right output
- Choose
attr('href')when fidelity to the source HTML matters. - Choose
prop('href')with a known base URI when downstream code needs absolute destinations. - Choose
map(...).get()for a straightforward array, or$.extractwhen links belong to a larger structured record. - Use a browser-capable renderer first when the links are created only by client-side JavaScript; Cheerio can parse that rendered HTML after you obtain it.
Frequently Asked Questions
Does Cheerio request each href it finds?
No. Cheerio parses the HTML string you provide; extracting an href does not download or follow the destination. Make any additional HTTP requests explicitly in your application.
Why does the same relative href resolve differently on two pages?
Relative URLs are interpreted against the document URL. Supply the correct page URL as baseURI (or use a URL-aware loader) for each document before reading prop('href').
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Can Cheerio discover links hidden inside an iframe?
Only if the iframe document’s HTML is separately retrieved and loaded. The parent document contains the iframe element, not the child page’s anchors.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




