The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →In a browser, parse an HTML string with DOMParser, then query the detached document with normal selectors:
const parser = new DOMParser();
const doc = parser.parseFromString(htmlString, "text/html");
const title = doc.querySelector("title")?.textContent.trim() ?? "";
const links = [...doc.querySelectorAll("a")].map(a => ({
text: a.textContent.trim(),
href: a.href
}));
This creates an in-memory Document; it does not fetch a URL and it does not make untrusted markup safe. Fetching, parsing, extracting, sanitizing, and inserting are separate operations.
Parse an HTML string in a browser
DOMParser.parseFromString() accepts a string (or TrustedHTML) and a supported MIME type, then returns a Document. With text/html, the result is a complete detached document with html, head, and body nodes, even if the input is only a fragment.
function parseHtml(htmlString) {
const parser = new DOMParser();
return parser.parseFromString(htmlString, "text/html");
}
const html = `<!doctype html>
<html>
<head><title>Example</title></head>
<body><main><h1>Hello</h1></main></body>
</html>`;
const doc = parseHtml(html);
console.log(doc.querySelector("title")?.textContent); // Example
console.log(doc.querySelector("main h1")?.textContent); // Hello
The detached tree is separate from the visible page. Scripts in the parsed HTML are marked non-executable, and inline event handlers do not run while the document remains detached. The method is broadly available across browsers; MDN documents support since July 2015.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
Extract text, attributes, and links
Use textContent for text and DOM selectors for structure. Use getAttribute() when you need the literal attribute value. The href property can resolve a relative URL against the document base URL, so choose deliberately.
const cards = [...doc.querySelectorAll("article.card")].map(card => ({
heading: card.querySelector("h2")?.textContent.trim() ?? "",
url: card.querySelector("a")?.href ?? "",
summary: card.querySelector("p")?.textContent.trim() ?? ""
}));
const rawHref = doc.querySelector("a")?.getAttribute("href") ?? "";
console.log({ cards, rawHref });
Fetch a page, then parse its response
The parser never downloads a URL. Retrieve the response separately, check the HTTP status, convert the body to text, and pass that text to DOMParser.
async function fetchDocument(url) {
const response = await fetch(url);
if (!response.ok) {
throw new Error(`HTTP ${response.status}`);
}
const html = await response.text();
return new DOMParser().parseFromString(html, "text/html");
}
const doc = await fetchDocument("/page.html");
const mainText = doc.querySelector("main")?.textContent.trim() ?? "";
console.log(mainText);
Origin and CORS constraints
Browser fetch() is subject to normal same-origin and CORS rules. A page can parse HTML already delivered by your application, but it cannot automatically read arbitrary third-party pages unless the server permits that cross-origin request. A server-side fetch or a service that captures the rendered page is appropriate when browser policy blocks the request.
Static HTML versus rendered HTML
response.text() gives you the response body. It does not execute the page’s JavaScript or wait for client-rendered content. If a site builds its article list after hydration, the HTML response may not contain those nodes. In that case, use the site’s data endpoint, run a browser automation workflow, or capture a rendered page rather than assuming DOMParser can see content that was never in the string.
Free tools Windows power users keep installed
One-click scans. No signup required.
Fragments: template and contextual fragments
Use DOMParser when you want a queryable document. For a small fragment intended for insertion, a <template> element or document.createRange().createContextualFragment() preserves fragment context.
Rank #2
const template = document.createElement("template");
template.innerHTML = "<li class="item">One</li>";
const item = template.content.firstElementChild;
const range = document.createRange();
range.selectNode(document.body);
const fragment = range.createContextualFragment("<p>A paragraph</p>");
document.body.append(fragment);
Fragment parsing still requires sanitization when the source is untrusted. The choice between a full detached document and a fragment is about the tree you need, not about safety.
Security: parsing is not sanitization
MDN describes parseFromString() as an injection sink. Detached parsing is inert, but unsafe elements can become active if you later insert them into the live DOM. Treat parsing and sanitization as different steps.
Sanitize before insertion
Use a reviewed sanitizer policy, commonly DOMPurify, and Trusted Types where your application supports them. The policy below illustrates the boundary; configure the sanitizer for your application’s allowed markup.
Recommended Free Tools
const policy = trustedTypes.createPolicy("html", {
createHTML: input => DOMPurify.sanitize(input)
});
const safeDoc = new DOMParser().parseFromString(
policy.createHTML(untrustedHtml),
"text/html"
);
const safeMain = safeDoc.querySelector("main");
if (safeMain) {
document.querySelector("#preview").replaceChildren(
...safeMain.childNodes
);
}
Do not assume that selecting nodes, reading text, or serializing with outerHTML makes markup safe. Validate URLs and attributes as well as element names, and avoid assigning untrusted strings to innerHTML without a sanitizer.
XML, XHTML, and SVG parsing
Pass an XML MIME type when XML rules are required. Supported choices include text/xml, application/xml, application/xhtml+xml, and image/svg+xml. These modes are not HTML error recovery.
const xmlDoc = new DOMParser().parseFromString(
xmlString,
"application/xml"
);
if (xmlDoc.querySelector("parsererror")) {
throw new Error("Malformed XML");
}
const ids = [...xmlDoc.querySelectorAll("item")]
.map(item => item.getAttribute("id"));
HTML parsing repairs malformed markup according to browser parsing rules. XML parsing can produce a parsererror node instead. Do not use an XML MIME type merely to obtain stricter HTML validation.
Parse HTML in Node.js with Cheerio
Node.js has no browser DOM by default. Cheerio is a common choice for selector-based extraction and transformation: provide the HTML to load(), then query with a jQuery-like API.
import * as cheerio from "cheerio";
const html = `<table>
<tr><td>A</td><td>1</td></tr>
<tr><td>B</td><td>2</td></tr>
</table>`;
const $ = cheerio.load(html);
const rows = $("table tr").map((_, row) => ({
cells: $(row).find("td").map((_, cell) => $(cell).text().trim()).get()
})).get();
console.log(rows);
Cheerio parser behavior
Cheerio defaults to parse5, which treats input as a complete document and can add html, head, and body. That matters when you compare serialized output or expect a fragment to remain a fragment. Configure htmlparser2 when you need more forgiving parsing or performance characteristics such as lower memory use, and test the resulting tree because parser behavior can differ from browser parsing.
Cheerio’s loadBuffer, decodeStream, and fromURL helpers use Node.js APIs. If a URL comes from a user, review URL loading as a server-side security boundary: protect internal services, restrict protocols and hosts, and apply timeouts and response-size limits. Cheerio’s selector and serialization operations do not sanitize output; the calling application remains responsible for safe rendering.
Choosing the right parser
| Choice | Best fit | Main trade-off |
|---|---|---|
Browser DOMParser |
Existing browser code and detached DOM queries | Requires a browser environment; sanitize before live-DOM insertion |
template or contextual fragments |
Small fragments that need insertion context | Context affects parsing; untrusted input still needs sanitization |
Cheerio load |
Node.js scraping, transformation, and CSS-selector extraction | Dependency and document-wrapping behavior must be understood |
Cheerio with htmlparser2 |
Forgiving or performance-sensitive parsing | Tree and serialization behavior can differ from parse5 and browsers |
Performance and reliability practices
- Parse once and reuse the resulting document instead of reparsing the same string for every selector.
- Use specific selectors and extract only required fields; converting a huge subtree to a string creates avoidable memory pressure.
- For fetched content, check
response.ok, impose an application timeout, and handle network failures before parsing. - Limit server-side response size when processing untrusted URLs.
- Expect missing nodes: optional chaining and nullish defaults prevent a single absent heading from crashing an extraction job.
- Test malformed HTML and fragments separately from well-formed full documents.
- Keep parsing synchronous work off latency-sensitive UI paths when documents are very large; consider a worker or server-side pipeline.
Troubleshooting common failures
“DOMParser is not defined”
You are likely running browser code in Node.js, a test runner without DOM globals, or a server process. Use Cheerio in Node.js, provide a DOM implementation in tests, or move the parsing step into a browser context.
Rank #4
The result is empty
Inspect the original response text. The URL may have returned an error page, a login challenge, or a shell whose content is inserted later by JavaScript. Check the HTTP status and verify that your selector matches the returned markup.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Relative links are wrong
element.href resolves URLs using the document base. If the HTML was detached from its original page or has no base URL, use getAttribute("href") for the raw value or add a deliberate base URL before resolving links.
Malformed XML does not throw
XML parsing commonly reports failure by adding a parsererror element. Check for it explicitly; HTML mode instead performs error recovery.
Inserted HTML executes or creates a vulnerability
Parsing did not sanitize it. Sanitize untrusted input before insertion, use a Trusted Types policy where available, and validate URL-bearing attributes.
Cheerio output has unexpected html/body wrappers
That is normal with the default parse5 document mode. Decide whether you need a complete document or a fragment, then configure the parser and test serialization accordingly.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
Or skip the browser setup
If your real goal is a clean image or PDF of a URL rather than extracting nodes, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for all options, including PNG, JPEG, WebP, PDF, full-page lazy-image loading, CSS-selector element capture, device presets, custom JavaScript and CSS, request blocking, cookies, headers, geolocation, signed links, asynchronous webhooks, and bulk capture.
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.
Frequently Asked Questions
Does DOMParser download external images, stylesheets, or scripts?
No. It parses the string you provide into a detached tree. Fetching the source document and rendering its external resources are separate operations.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsCan I use DOMParser to scrape any website from client-side JavaScript?
Only when the browser is allowed to read the response, normally through same-origin rules or a server that grants CORS access. Otherwise use a permitted server-side or rendered-page workflow.
When should I choose Cheerio instead of a browser parser?
Choose Cheerio for Node.js extraction and transformation when you already have the HTML string. Use a real browser when the content depends on client-side JavaScript, layout, or other rendering behavior.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




