Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Html Agility Pack (HAP) parses HTML that your .NET code has already obtained; it does not fetch pages or run their JavaScript. A typical scraper therefore has two distinct jobs: request a page with an HTTP client, then load the returned HTML into HAP and query its DOM with XPath. This guide shows that flow, how to handle missing or changing content, and when a browser-rendered capture is needed instead.
What Html Agility Pack does—and what it does not
HAP is a .NET library that turns an HTML string or document into a read/write document object model (DOM). You can navigate that structure with XPath; the project also advertises XSLT support. Its maintainers describe its parser as “very tolerant of real world malformed HTML.” That is a design characteristic, not a guarantee that a particular XPath will match a particular page.
HAP is the parsing component, not a web browser. Installing it does not make an HTTP request, execute page JavaScript, bypass a CAPTCHA or other access control, or ensure a scrape succeeds. If a page’s data is inserted only after client-side scripts run, the HTML returned by a basic HTTP request may not contain that data. Look for an official API or data feed first; otherwise, use a separate browser-rendering approach and feed its resulting HTML to a parser if structured extraction is still needed.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Before automating access, check the target site’s documentation and applicable permissions. A successful HTTP response is not itself permission to collect or reuse the content.
#1 Best Overall
Install HAP and create a small .NET scraper
At the time of the reviewed NuGet listing, HtmlAgilityPack was listed at version 1.13.0. The listing included .NET 8.0 and .NET Standard 2.0 as target frameworks, as well as compatibility information for additional frameworks. Package versions and compatibility listings can change, so check NuGet for the current release and the target frameworks in your own project.
1. Create a console project and add the package
For a .NET 8 console project, run:
dotnet new console --framework net8.0 --name HapScraper
cd HapScraper
dotnet add package HtmlAgilityPack --version 1.13.0
The version pin reflects the NuGet listing at the time of the reviewed material; use the current version shown on NuGet if you intentionally want a newer release. The package page also provides a PackageReference installation form for projects where package versions are managed in the project file.
2. Fetch the HTML, then parse it
This complete illustrative console program fetches an HTML response, checks the HTTP status, loads the response text into HAP, and extracts product cards from a sample document structure. Replace the sample URL and XPath with ones that match the page and response you are permitted to process. The markup and selectors below are examples, not selectors verified against a live site.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Rank #2
using HtmlAgilityPack;
using var http = new HttpClient();
http.DefaultRequestHeaders.UserAgent.ParseAdd("ExampleScraper/1.0 (contact: [email protected])");
var url = "https://example.com/catalog";
using var response = await http.GetAsync(url);
response.EnsureSuccessStatusCode();
var html = await response.Content.ReadAsStringAsync();
var document = new HtmlDocument();
document.LoadHtml(html);
var cards = document.DocumentNode.SelectNodes("//article[contains(concat(' ', normalize-space(@class), ' '), ' product-card ')]");
if (cards is null)
{
Console.WriteLine("No product cards matched. Check the response HTML and XPath.");
return;
}
foreach (var card in cards)
{
var nameNode = card.SelectSingleNode(".//*[contains(concat(' ', normalize-space(@class), ' '), ' product-name ')]");
var linkNode = card.SelectSingleNode(".//a[@href]");
var priceNode = card.SelectSingleNode(".//*[contains(concat(' ', normalize-space(@class), ' '), ' price ')]");
var name = CleanText(nameNode?.InnerText);
var price = CleanText(priceNode?.InnerText);
var href = linkNode?.GetAttributeValue("href", "");
var absoluteLink = Uri.TryCreate(url, UriKind.Absolute, out var baseUri)
&& Uri.TryCreate(baseUri, href, out var resolvedUri)
? resolvedUri.ToString()
: "";
if (string.IsNullOrWhiteSpace(name))
{
Console.WriteLine("Skipping a card with no product name.");
continue;
}
Console.WriteLine($"{name} | {price} | {absoluteLink}");
}
static string CleanText(string? value)
{
if (string.IsNullOrWhiteSpace(value)) return "";
var decoded = HtmlEntity.DeEntitize(value);
return string.Join(" ", decoded.Split((char[]?)null, StringSplitOptions.RemoveEmptyEntries));
}
The null-conditional access on nameNode and priceNode makes missing descendants a normal case rather than a null-reference exception. GetAttributeValue supplies an empty-string default if the link has no href. The XPath uses a whitespace-aware class test: matching @class='product-card' alone would fail when an element has multiple class names. Relative XPath beginning with . searches inside the current card, not from the root of the document.
Build extraction around the response you actually receive
Inspect before writing selectors
Do not assume the browser’s visible page and an HTTP response contain identical markup. Save or log a small, appropriately redacted portion of the response, check its content type and status, and search for the text or attributes you intend to extract. If the expected data is absent from the response, XPath changes will not make it appear. Confirm whether the site exposes an API or feed, or whether the content depends on browser-side execution.
Choose resilient XPath expressions
Prefer meaningful attributes and relationships over brittle positional paths such as /html/body/div[2]/div[4]/span. A path based on an element’s class, link destination, or nearby label is usually easier to understand and maintain. Avoid assuming that class names are unique or permanent. For repeated records, select the record containers first and query each container relative to itself, as in the example.
XPath text matching can be sensitive to whitespace and nested elements. Selecting a node and cleaning its InnerText is often more practical than expecting one exact text node. Decode HTML entities and normalize whitespace before comparing or storing text. Keep the original response or a representative fixture during development so selector changes can be checked against actual markup.
Validate types and required fields
Parsed values are strings until your code validates and converts them. Check required fields explicitly, parse numbers and dates with the expected culture and format, and record which records were rejected and why. Do not silently convert a missing price or date into a plausible-looking default. If a value is optional, represent it as optional in your output model; if it is required, fail or skip that record deliberately.
For production work, separate the fetch step from a parsing function that accepts an HTML string. That makes it possible to test extraction against saved fixtures without repeatedly requesting the target site. Add fixtures for missing fields, unexpected whitespace, multiple matching elements, and a changed page structure. HAP’s malformed-markup tolerance can help it build a DOM from imperfect input, but your own checks still need to detect missing or shifted content.
Rank #4
When to choose HAP, CSS selectors, or a browser
| Need | Practical fit | What to consider |
|---|---|---|
| XPath queries against returned HTML | Html Agility Pack | HAP documents XPath and XSLT support and describes tolerant handling of malformed HTML. Check the target framework and validate selectors against the actual response. |
| CSS-selector workflow with HAP | Universal.HtmlAgilityPack | This is a separate package that advertises CSS selector support by converting selectors to XPath. It is an add-on choice, not a built-in HAP capability. |
| HTML5-specification-based parsing and CSS selectors | AngleSharp | Its project description centers on HTML5 and W3C specifications and CSS selectors. Compare it against your parsing and framework requirements rather than assuming one library is universally more accurate. |
| Data rendered only after page scripts run | An API/data feed or a browser-rendering approach | A basic HTTP response parsed by HAP may not include client-side content. HAP itself is not the renderer. |
The choice depends on the input and workflow: whether the needed content is present in returned HTML, whether your team prefers XPath or CSS selectors, how standards-based parsing matters to the project, which frameworks you target, and how much maintenance the site’s changing markup will require. The cited project descriptions do not establish blanket performance rankings or extraction-accuracy percentages.
Common failures and how to diagnose them
- The selector returns null or no matches. First inspect the exact response body and confirm the expected node is present. Check for a redirect, an error page, different markup, a missing class, or content that only appears after JavaScript runs. Then adjust the XPath to the observed structure.
- The request fails or returns an unexpected status. Check the exception and HTTP status before parsing. Network errors, redirects, rate limits, and server errors are fetch-stage problems, not XPath problems. Handle transient failures deliberately, respect the site’s access rules, and avoid retry loops that increase load.
- The parser finds a page, but extracted fields are empty. A matching container does not prove its expected child exists. Use null-safe node access, inspect the child markup, and decide whether to skip the record or emit an explicit missing value.
- Text contains entities, line breaks, or extra spaces. Decode entities with
HtmlEntity.DeEntitize, collapse whitespace, and validate the normalized result. Preserve punctuation and meaningful separators; do not strip characters indiscriminately. - Links are relative. Resolve them against the page URI using
Uri.TryCreate, as the example does. Validate the base URI and handle absent or malformedhrefvalues rather than concatenating strings. - Results change after a site redesign. Treat selectors as dependencies on a remote document structure. Keep parsing tests with representative HTML, log match counts and validation failures, and review changes when the observed structure no longer fits the expected fields.
Performance, reliability, and cost considerations
There is no benchmark or extraction-accuracy figure established here for HAP, so do not choose it based on an assumed speed ranking. For a scraper, total runtime and reliability also depend on network latency, response size, target-site behavior, request frequency, and how much of the DOM your code processes. Measure your own workload with representative responses rather than extrapolating from another site.
Reuse an HttpClient for multiple requests rather than creating one for every URL. Set sensible timeouts and bounded concurrency, and avoid fetching the same page repeatedly when a cached response is appropriate. A timeout or failed response should be handled as a fetch failure; it is not evidence that the HTML parser is at fault. For large jobs, record the URL, status, elapsed time, and parsing outcome without retaining sensitive response data unnecessarily.
Best Value
HAP is a NuGet dependency, and the reviewed material does not establish a separate usage charge. Your operational costs may still include compute, bandwidth, storage, and any infrastructure or third-party services you use to obtain or render pages.
Or skip the browser setup
If your task is to capture what a page looks like rather than extract structured values from raw HTML, ScreenshotNeo is a screenshot API and MCP server. It does not replace HAP when you need DOM values, but it can avoid setting up and operating a browser for screenshot work. Its API accepts a URL and returns an image or PDF; the request below saves a WebP response.
Quick Recap
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, failed loads, and cache hits are not billed. An MCP server lets AI agents use screenshot tools, and 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

