DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
Cheerio

How to Parse HTML in JavaScript: DOMParser, Fetch, Security, and Node.js

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In a browser, parse an HTML string with DOMParser, then query the detached document with normal selectors:

const parser = new DOMParser();
const doc = parser.parseFromString(htmlString, "text/html");
const title = doc.querySelector("title")?.textContent.trim() ?? "";
const links = [...doc.querySelectorAll("a")].map(a => ({
  text: a.textContent.trim(),
  href: a.href
}));

This creates an in-memory Document; it does not fetch a URL and it does not make untrusted markup safe. Fetching, parsing, extracting, sanitizing, and inserting are separate operations.

Parse an HTML string in a browser

DOMParser.parseFromString() accepts a string (or TrustedHTML) and a supported MIME type, then returns a Document. With text/html, the result is a complete detached document with html, head, and body nodes, even if the input is only a fragment.

function parseHtml(htmlString) {
  const parser = new DOMParser();
  return parser.parseFromString(htmlString, "text/html");
}

const html = `<!doctype html>
  <html>
    <head><title>Example</title></head>
    <body><main><h1>Hello</h1></main></body>
  </html>`;

const doc = parseHtml(html);
console.log(doc.querySelector("title")?.textContent); // Example
console.log(doc.querySelector("main h1")?.textContent); // Hello

The detached tree is separate from the visible page. Scripts in the parsed HTML are marked non-executable, and inline event handlers do not run while the document remains detached. The method is broadly available across browsers; MDN documents support since July 2015.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract text, attributes, and links

Use textContent for text and DOM selectors for structure. Use getAttribute() when you need the literal attribute value. The href property can resolve a relative URL against the document base URL, so choose deliberately.

const cards = [...doc.querySelectorAll("article.card")].map(card => ({
  heading: card.querySelector("h2")?.textContent.trim() ?? "",
  url: card.querySelector("a")?.href ?? "",
  summary: card.querySelector("p")?.textContent.trim() ?? ""
}));

const rawHref = doc.querySelector("a")?.getAttribute("href") ?? "";
console.log({ cards, rawHref });

Fetch a page, then parse its response

The parser never downloads a URL. Retrieve the response separately, check the HTTP status, convert the body to text, and pass that text to DOMParser.

async function fetchDocument(url) {
  const response = await fetch(url);
  if (!response.ok) {
    throw new Error(`HTTP ${response.status}`);
  }
  const html = await response.text();
  return new DOMParser().parseFromString(html, "text/html");
}

const doc = await fetchDocument("/page.html");
const mainText = doc.querySelector("main")?.textContent.trim() ?? "";
console.log(mainText);

Origin and CORS constraints

Browser fetch() is subject to normal same-origin and CORS rules. A page can parse HTML already delivered by your application, but it cannot automatically read arbitrary third-party pages unless the server permits that cross-origin request. A server-side fetch or a service that captures the rendered page is appropriate when browser policy blocks the request.

Static HTML versus rendered HTML

response.text() gives you the response body. It does not execute the page’s JavaScript or wait for client-rendered content. If a site builds its article list after hydration, the HTML response may not contain those nodes. In that case, use the site’s data endpoint, run a browser automation workflow, or capture a rendered page rather than assuming DOMParser can see content that was never in the string.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fragments: template and contextual fragments

Use DOMParser when you want a queryable document. For a small fragment intended for insertion, a <template> element or document.createRange().createContextualFragment() preserves fragment context.

const template = document.createElement("template");
template.innerHTML = "<li class="item">One</li>";
const item = template.content.firstElementChild;

const range = document.createRange();
range.selectNode(document.body);
const fragment = range.createContextualFragment("<p>A paragraph</p>");
document.body.append(fragment);

Fragment parsing still requires sanitization when the source is untrusted. The choice between a full detached document and a fragment is about the tree you need, not about safety.

Security: parsing is not sanitization

MDN describes parseFromString() as an injection sink. Detached parsing is inert, but unsafe elements can become active if you later insert them into the live DOM. Treat parsing and sanitization as different steps.

Sanitize before insertion

Use a reviewed sanitizer policy, commonly DOMPurify, and Trusted Types where your application supports them. The policy below illustrates the boundary; configure the sanitizer for your application’s allowed markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const policy = trustedTypes.createPolicy("html", {
  createHTML: input => DOMPurify.sanitize(input)
});

const safeDoc = new DOMParser().parseFromString(
  policy.createHTML(untrustedHtml),
  "text/html"
);

const safeMain = safeDoc.querySelector("main");
if (safeMain) {
  document.querySelector("#preview").replaceChildren(
    ...safeMain.childNodes
  );
}

Do not assume that selecting nodes, reading text, or serializing with outerHTML makes markup safe. Validate URLs and attributes as well as element names, and avoid assigning untrusted strings to innerHTML without a sanitizer.

XML, XHTML, and SVG parsing

Pass an XML MIME type when XML rules are required. Supported choices include text/xml, application/xml, application/xhtml+xml, and image/svg+xml. These modes are not HTML error recovery.

const xmlDoc = new DOMParser().parseFromString(
  xmlString,
  "application/xml"
);

if (xmlDoc.querySelector("parsererror")) {
  throw new Error("Malformed XML");
}

const ids = [...xmlDoc.querySelectorAll("item")]
  .map(item => item.getAttribute("id"));

HTML parsing repairs malformed markup according to browser parsing rules. XML parsing can produce a parsererror node instead. Do not use an XML MIME type merely to obtain stricter HTML validation.

Parse HTML in Node.js with Cheerio

Node.js has no browser DOM by default. Cheerio is a common choice for selector-based extraction and transformation: provide the HTML to load(), then query with a jQuery-like API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import * as cheerio from "cheerio";

const html = `<table>
  <tr><td>A</td><td>1</td></tr>
  <tr><td>B</td><td>2</td></tr>
</table>`;

const $ = cheerio.load(html);
const rows = $("table tr").map((_, row) => ({
  cells: $(row).find("td").map((_, cell) => $(cell).text().trim()).get()
})).get();

console.log(rows);

Cheerio parser behavior

Cheerio defaults to parse5, which treats input as a complete document and can add html, head, and body. That matters when you compare serialized output or expect a fragment to remain a fragment. Configure htmlparser2 when you need more forgiving parsing or performance characteristics such as lower memory use, and test the resulting tree because parser behavior can differ from browser parsing.

Cheerio’s loadBuffer, decodeStream, and fromURL helpers use Node.js APIs. If a URL comes from a user, review URL loading as a server-side security boundary: protect internal services, restrict protocols and hosts, and apply timeouts and response-size limits. Cheerio’s selector and serialization operations do not sanitize output; the calling application remains responsible for safe rendering.

Choosing the right parser

Choice Best fit Main trade-off
Browser DOMParser Existing browser code and detached DOM queries Requires a browser environment; sanitize before live-DOM insertion
template or contextual fragments Small fragments that need insertion context Context affects parsing; untrusted input still needs sanitization
Cheerio load Node.js scraping, transformation, and CSS-selector extraction Dependency and document-wrapping behavior must be understood
Cheerio with htmlparser2 Forgiving or performance-sensitive parsing Tree and serialization behavior can differ from parse5 and browsers

Performance and reliability practices

  • Parse once and reuse the resulting document instead of reparsing the same string for every selector.
  • Use specific selectors and extract only required fields; converting a huge subtree to a string creates avoidable memory pressure.
  • For fetched content, check response.ok, impose an application timeout, and handle network failures before parsing.
  • Limit server-side response size when processing untrusted URLs.
  • Expect missing nodes: optional chaining and nullish defaults prevent a single absent heading from crashing an extraction job.
  • Test malformed HTML and fragments separately from well-formed full documents.
  • Keep parsing synchronous work off latency-sensitive UI paths when documents are very large; consider a worker or server-side pipeline.

Troubleshooting common failures

“DOMParser is not defined”

You are likely running browser code in Node.js, a test runner without DOM globals, or a server process. Use Cheerio in Node.js, provide a DOM implementation in tests, or move the parsing step into a browser context.

The result is empty

Inspect the original response text. The URL may have returned an error page, a login challenge, or a shell whose content is inserted later by JavaScript. Check the HTTP status and verify that your selector matches the returned markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Relative links are wrong

element.href resolves URLs using the document base. If the HTML was detached from its original page or has no base URL, use getAttribute("href") for the raw value or add a deliberate base URL before resolving links.

Malformed XML does not throw

XML parsing commonly reports failure by adding a parsererror element. Check for it explicitly; HTML mode instead performs error recovery.

Inserted HTML executes or creates a vulnerability

Parsing did not sanitize it. Sanitize untrusted input before insertion, use a Trusted Types policy where available, and validate URL-bearing attributes.

Cheerio output has unexpected html/body wrappers

That is normal with the default parse5 document mode. Decide whether you need a complete document or a fragment, then configure the parser and test serialization accordingly.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your real goal is a clean image or PDF of a URL rather than extracting nodes, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all options, including PNG, JPEG, WebP, PDF, full-page lazy-image loading, CSS-selector element capture, device presets, custom JavaScript and CSS, request blocking, cookies, headers, geolocation, signed links, asynchronous webhooks, and bulk capture.

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));

ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.

Frequently Asked Questions

Does DOMParser download external images, stylesheets, or scripts?

No. It parses the string you provide into a detached tree. Fetching the source document and rendering its external resources are separate operations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use DOMParser to scrape any website from client-side JavaScript?

Only when the browser is allowed to read the response, normally through same-origin rules or a server that grants CORS access. Otherwise use a permitted server-side or rendered-page workflow.

When should I choose Cheerio instead of a browser parser?

Choose Cheerio for Node.js extraction and transformation when you already have the HTML string. Use a real browser when the content depends on client-side JavaScript, layout, or other rendering behavior.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.