October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Cheerio

How to Get Links in Cheerio (Read Every href, Resolve URLs, and Handle Dynamic Pages)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Load the HTML into Cheerio, select anchors with $('a'), and read each anchor’s href attribute. Use .attr('href') for the literal value in the markup, map the selection to collect every link, and use .prop('href') with a document URL when you need absolute URLs.

The shortest working example

This complete ES module loads two anchors and returns their raw href strings:

import * as cheerio from 'cheerio';

const html = `
  <a href="/docs">Docs</a>
  <a href="https://example.com/blog">Blog</a>
`;

const $ = cheerio.load(html);
const links = $('a').map((_, element) => $(element).attr('href')).get();

console.log(links);
// [ '/docs', 'https://example.com/blog' ]

The a selector matches every anchor element. attr('href') reads the text of the attribute exactly as it appears in the source, and get() converts Cheerio’s collection into a normal JavaScript array.

Install Cheerio first if it is not already in your project:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
npm install cheerio

Cheerio’s selector syntax is covered in the official selecting guide, while its manipulation guide documents reading attributes and properties.

Get one link or all links

Read the first matching anchor

const $ = cheerio.load(html);
const firstHref = $('a').attr('href');
console.log(firstHref);

When a selection contains several elements, attr('href') reads the first one. If no anchor matches, or the first matching anchor has no href attribute, the result is undefined.

Collect every href

const hrefs = $('a')
  .map((_, element) => $(element).attr('href'))
  .get();

This preserves document order. It can include undefined for anchors without an href, so filter those values when your output must contain only actual attributes:

const hrefs = $('a')
  .map((_, element) => $(element).attr('href'))
  .get()
  .filter((href) => typeof href === 'string' && href.length > 0);

Select a narrower group

Use normal CSS selectors when the page contains several kinds of links:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const navigation = $('nav a').map((_, el) => $(el).attr('href')).get();
const externalCandidates = $('a[href^="http"]').map((_, el) => $(el).attr('href')).get();
const articleLinks = $('article a[href]').map((_, el) => $(el).attr('href')).get();

The [href] attribute selector excludes anchors that do not have an href at all. A prefix selector such as [href^="http"] is only a quick filter; it does not prove that a value is valid or that it points to another site.

Raw href values versus absolute URLs

Cheerio does not rewrite the string returned by attr('href'). For <a href="/docs">, the result is /docs. That is useful when you need to reproduce the source markup, compare templates, or preserve exactly what the publisher supplied.

Method Result Document URL required? Best use
attr('href') Literal attribute string No Preserve or inspect source markup
prop('href') URL resolved against the document URL Yes for relative values Fetch, compare, or export absolute links
$.extract({ links: [{ selector: 'a', value: 'href' }] }) All selected values in a declarative shape Resolution depends on a document URL Combine links with other extracted fields

Resolve a relative href with a base URI

import * as cheerio from 'cheerio';

const $ = cheerio.load('<a href="/docs">Docs</a>', {
  baseURI: 'https://example.com/articles/page.html',
});

console.log($('a').prop('href'));
// https://example.com/docs

The base URI tells Cheerio which origin and path to use. Without it, a relative attribute remains relative. Cheerio’s troubleshooting documentation explains this distinction and its manipulation guide documents prop('href').

Use a URL-aware loader

When you load a page directly from a URL with Cheerio’s URL loader, the document URL is available for property resolution. The same principle applies: use attr when you want the original text and prop when you want Cheerio’s URL-aware property.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the extract API for structured output

Cheerio’s extract method is convenient when links are one field in a larger result:

const data = $.extract({
  links: [{ selector: 'a', value: 'href' }],
});

console.log(data);
// { links: ['/docs', 'https://example.com/blog'] }

An array descriptor collects all matches. A selector descriptor without the array returns the first match. The value: 'href' descriptor uses Cheerio’s property API, so relative values are resolved only when the loaded document has a URL or a configured base URI.

You can also extract text and links from repeated records:

const cards = $.extract({
  products: [{
    selector: '.product-card',
    value: {
      name: '.name',
      href: { selector: 'a', value: 'href' },
    },
  }],
});

For a simple list of anchors, map(...).get() is usually easier to read; use extract when the output has several related fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a dependable link extractor

Keep the source value and resolved value together

Storing both forms avoids losing information and makes later decisions explicit:

const links = $('a[href]').map((_, element) => {
  const anchor = $(element);
  const raw = anchor.attr('href');
  const absolute = anchor.prop('href');

  return {
    text: anchor.text().trim(),
    raw,
    absolute,
  };
}).get();

If no base URI was supplied, absolute may still be a relative value. Do not label it absolute unless you have provided a document URL and checked the result.

Resolve values yourself when you need strict control

The standard URL constructor is useful when you want explicit error handling or need to retain special schemes:

const base = 'https://example.com/articles/page.html';

const resolved = $('a[href]').map((_, element) => {
  const raw = $(element).attr('href');
  try {
    return { raw, url: new URL(raw, base).href };
  } catch {
    return { raw, url: null };
  }
}).get();

This lets you identify malformed values instead of allowing one bad attribute to stop the entire extraction. Values such as fragments, mailto:, and javascript: are not ordinary HTTP pages; classify or exclude them according to your application rather than silently treating every string as a fetchable web URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Remove duplicates only when that is your requirement

const uniqueHrefs = [...new Set(
  $('a[href]')
    .map((_, el) => $(el).prop('href'))
    .get()
)];

Deduplication can hide meaningful differences if the same destination appears in separate navigation areas or if query strings and fragments matter. Decide whether identity means the literal string, the resolved URL, or a normalized URL before using a Set.

Loading HTML correctly

Cheerio parses the markup you provide. Its normal loader treats input as a complete document and may add missing document structure. If you are parsing a fragment copied from a component rather than a full page, use Cheerio’s fragment mode as described in the official troubleshooting guide.

const fragment = '<a href="/one">One</a><a href="/two">Two</a>';
const $ = cheerio.load(fragment, null, false);
const hrefs = $('a').map((_, el) => $(el).attr('href')).get();

When your input comes from an HTTP client, pass the response body to Cheerio and retain the final response URL as the base URI if redirects or relative links matter. Cheerio itself is the parser; your HTTP client is responsible for downloading the response.

Why a link may be missing

The page creates it with JavaScript

Cheerio is not a web browser. It does not execute page JavaScript, wait for client-side rendering, click controls, or run framework hydration. If an anchor is inserted only after browser code runs, it will not exist in the static HTML passed to Cheerio. The official introduction points to browser automation tools such as Puppeteer or Playwright, and DOM emulation such as jsdom, for cases that require execution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before changing selectors, inspect the exact response body. A link visible in a browser’s Elements panel may be absent from the original HTTP response.

The anchor has no href

Some markup uses an element styled as a link without an href. Select with a[href] when those should be excluded, or inspect the element’s other attributes if your application intentionally supports them.

The selector is too narrow

Start with $('a').length, then test progressively specific selectors such as main a or a[href]. A typo in a class name or an unexpected iframe boundary can produce an empty selection even though links exist elsewhere in the response.

Troubleshooting table

Symptom Likely cause Fix
undefined from attr('href') No matching anchor, or the first match has no href Check $('a').length, use a[href], and inspect the input HTML.
Only one link is returned attr() reads the first element in a selection Use map(...).get() or an array descriptor in $.extract.
A relative path stays relative attr() returns the literal source value Supply baseURI or a URL-aware loader and read prop('href').
Browser shows a link but Cheerio finds none The link is inserted after JavaScript executes Obtain server-rendered HTML or use browser automation before parsing.
Unexpected html/body wrappers Document mode adds structure to a fragment Load the input in fragment mode when it is not a complete document.
Extraction stops on one malformed URL URL construction throws for an invalid value Wrap new URL() in try/catch and retain the raw value for review.

Performance and reliability practices

  • Parse once and reuse the same $ function for all selectors instead of loading identical HTML repeatedly.
  • Use a specific selector such as article a[href] when site navigation and footer links are irrelevant; this reduces downstream filtering.
  • Keep extraction separate from downloading. Cheerio does not fetch every discovered URL, so queue and rate-limit follow-up requests in your HTTP layer.
  • Record the source URL, retrieval time, and raw href when results may need auditing. Relative links cannot be reproduced correctly without their base document.
  • Treat untrusted HTML as data. Validate destinations before making requests, and apply your application’s policy for non-HTTP schemes, credentials, localhost addresses, and redirects.
  • Do not assume a successful parse means the page was complete. A server response can be an error page, a consent wall, or a partial shell even when Cheerio reports no parsing error.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your immediate need is a rendered visual capture rather than programmatic href extraction, ScreenshotNeo provides a single-request website screenshot API. It is separate from Cheerio: it returns a PNG, JPEG, WebP, or PDF, not an array of links. It is useful when you need to verify what a browser-rendered page looks like before deciding how to obtain its HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for all parameters. Equivalent examples:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and whether the request was billed.
  • An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is included on every plan.

When you are ready to try it, sign up for the free ScreenshotNeo plan.

Choosing the right output

  • Choose attr('href') when fidelity to the source HTML matters.
  • Choose prop('href') with a known base URI when downstream code needs absolute destinations.
  • Choose map(...).get() for a straightforward array, or $.extract when links belong to a larger structured record.
  • Use a browser-capable renderer first when the links are created only by client-side JavaScript; Cheerio can parse that rendered HTML after you obtain it.

Frequently Asked Questions

Does Cheerio request each href it finds?

No. Cheerio parses the HTML string you provide; extracting an href does not download or follow the destination. Make any additional HTTP requests explicitly in your application.

Why does the same relative href resolve differently on two pages?

Relative URLs are interpreted against the document URL. Supply the correct page URL as baseURI (or use a URL-aware loader) for each document before reading prop('href').

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can Cheerio discover links hidden inside an iframe?

Only if the iframe document’s HTML is separately retrieved and loaded. The parent document contains the iframe element, not the child page’s anchors.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.