Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: choose AngleSharp for a new project that parses static HTML, HtmlAgilityPack when you need established XPath code, Microsoft.Playwright for JavaScript-heavy or multi-browser sites, Selenium.WebDriver for an existing WebDriver operation, and PuppeteerSharp for Chrome/Chromium DevTools workflows. ScrapySharp and CsQuery are maintenance choices for older applications, not first picks for new development.

The key decision is whether the data exists in the HTTP response or appears only after JavaScript runs. Parsers read returned markup; browser automation loads a page, executes scripts, and exposes the rendered DOM.

At a glance: which C# scraper should you use?

Library Best fit Selectors/API Runs JavaScript? Browser engines 2026 recommendation
AngleSharp New static-HTML projects Standards-oriented DOM, CSS selectors No None Best modern parser
HtmlAgilityPack Established XPath code Node tree, XPath No None Best established XPath parser
Microsoft.Playwright JavaScript-heavy sites and cross-browser coverage Locator and page API Yes Chromium, Firefox, WebKit Best broad browser tool
Selenium.WebDriver Organizations already using WebDriver WebDriver API, Selenium.Support Yes Browser/driver combinations Best ecosystem choice
PuppeteerSharp 25.12.0 Chrome-only automation Puppeteer-style API, DevTools Protocol Yes Chrome/Chromium Best Chrome-focused control
ScrapySharp 3.0.0 Existing legacy applications HtmlAgilityPack plus jQuery-like CSS helpers Limited browser simulation; not a JavaScript engine None Use only after compatibility review
CsQuery 1.3.4 Legacy .NET Framework projects jQuery-style DOM and CSS2/CSS3 selectors No None Keep existing code running; prefer AngleSharp for new work

There is no directly comparable primary benchmark or adoption statistic establishing a universal speed or popularity winner. Parser-only code normally consumes fewer CPU and memory resources than a real browser, while browser automation provides rendering and interaction at a higher operational cost.

1. AngleSharp: best modern static HTML parser

AngleSharp implements a standards-oriented HTML5 DOM and browser-like querySelector and querySelectorAll traversal. Its current target list includes netstandard2.0, net8.0, and net10.0, making it a practical fit for current .NET services.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When it fits

  • The server returns the content you need in the initial HTML.
  • You want CSS selectors and a browser-like DOM rather than XPath.
  • Malformed markup needs HTML5-style parsing instead of brittle string processing.

Runnable C# example

using AngleSharp;
using System.Net.Http;

var url = "https://example.com/products";
using var http = new HttpClient();
var html = await http.GetStringAsync(url);
var context = BrowsingContext.New(Configuration.Default);
var document = await context.OpenAsync(req => req.Content(html).Address(url));
foreach (var card in document.QuerySelectorAll(".product-card"))
{
    var name = card.QuerySelector(".name")?.TextContent.Trim();
    Console.WriteLine(name);
}

AngleSharp does not execute arbitrary page JavaScript. If the initial response contains only an application shell, move to a browser library.

2. HtmlAgilityPack: best established XPath parser

HtmlAgilityPack builds a navigable node tree and is widely paired with HttpClient. It is a strong choice when an existing codebase, team, or data-extraction specification already uses XPath.

Runnable C# example

using HtmlAgilityPack;
using System.Net.Http;

using var client = new HttpClient();
var html = await client.GetStringAsync("https://example.com/products");
var doc = new HtmlDocument();
doc.LoadHtml(html);
foreach (var node in doc.DocumentNode.SelectNodes("//article[contains(@class,'product')]") ?? new HtmlNodeCollection(null))
{
    var title = node.SelectSingleNode(".//h2")?.InnerText.Trim();
    Console.WriteLine(title);
}

The null-handling around SelectNodes matters: a page redesign can legitimately produce no matches. HtmlAgilityPack parses what you downloaded; add Playwright, Selenium, or PuppeteerSharp when the desired nodes are created by scripts.

3. Microsoft.Playwright: best for JavaScript-heavy and multi-browser sites

Playwright for .NET is the official language port of Playwright and automates Chromium, Firefox, and WebKit through one API. It is the broadest default when rendering fidelity, locator auto-waiting, and cross-browser behavior matter.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Setup and example

Add the Microsoft.Playwright package, then install the browser binaries required by your project using the install command documented for that package version. Keep those browser revisions aligned with your deployment image.

using Microsoft.Playwright;

using var playwright = await Playwright.CreateAsync();
await using var browser = await playwright.Chromium.LaunchAsync(new BrowserTypeLaunchOptions
{
    Headless = true
});
var page = await browser.NewPageAsync();
await page.GotoAsync("https://example.com/products", new PageGotoOptions
{
    WaitUntil = WaitUntilState.NetworkIdle
});
var titles = await page.Locator(".product-card .name").AllTextContentsAsync();
foreach (var title in titles)
    Console.WriteLine(title.Trim());

Why choose it

  • One API covers three browser engines.
  • Locators wait for elements and reduce timing races.
  • You can interact with menus, logins, scrolling, and consent dialogs before extraction.

Browser workers need more memory and startup time than an HTTP parser. Reuse a browser process, limit concurrency, and close contexts so long crawls do not leak resources.

4. Selenium.WebDriver: best WebDriver ecosystem

Selenium’s .NET API is supplied through Selenium.WebDriver and commonly Selenium.Support. It remains the pragmatic choice when your organization already manages WebDriver infrastructure, browser drivers, grids, or shared test-team knowledge.

Runnable C# example

using OpenQA.Selenium;
using OpenQA.Selenium.Chrome;

var options = new ChromeOptions();
options.AddArgument("--headless=new");
options.AddArgument("--no-sandbox");
using IWebDriver driver = new ChromeDriver(options);
driver.Navigate().GoToUrl("https://example.com/products");
var cards = driver.FindElements(By.CssSelector(".product-card .name"));
foreach (var card in cards)
    Console.WriteLine(card.Text.Trim());

Selenium is a browser automation stack, not a lightweight HTML parser. Driver and browser version mismatches, grid capacity, and explicit waits are operational concerns. Use explicit waits for a known condition rather than fixed sleeps whenever possible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. PuppeteerSharp: best Chrome/Chromium DevTools control

PuppeteerSharp 25.12.0 is a .NET port of the official Node.js Puppeteer API. It controls headless or headed Chrome/Chromium through the Chrome DevTools Protocol, making it suitable for single-page applications, screenshots, PDFs, and Chrome-specific workflows.

Runnable C# example

using PuppeteerSharp;

await new BrowserFetcher().DownloadAsync();
await using var browser = await Puppeteer.LaunchAsync(new LaunchOptions
{
    Headless = true
});
await using var page = await browser.NewPageAsync();
await page.GoToAsync("https://example.com/products", WaitUntilNavigation.Networkidle0);
var names = await page.EvaluateFunctionAsync<string[]>(@"() =>
    [...document.querySelectorAll('.product-card .name')]
      .map(e => e.textContent.trim())");
foreach (var name in names)
    Console.WriteLine(name);

Choose PuppeteerSharp when Chrome is the target and DevTools-level control is valuable. If Firefox or WebKit coverage is a requirement, Playwright is a better fit.

6. ScrapySharp: a legacy combined helper

ScrapySharp 3.0.0 combines a browser-simulating web client with an HtmlAgilityPack extension that offers jQuery-like CSS selection. NuGet lists its last update as 2018-10-02. That age does not automatically break an existing application, but it increases dependency and runtime-compatibility risk.

Use it when

  • An older application already depends on its APIs.
  • You have tests covering its request, parsing, and encoding behavior.
  • You have verified compatibility with the target .NET runtime and current TLS requirements.

For a new project, use AngleSharp for static documents or a maintained browser automation library for rendered pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. CsQuery: legacy jQuery-style parsing

CsQuery 1.3.4 provides an HTML parser, CSS selector engine, and jQuery-style DOM API for .NET Framework 4 and C#. Its package description includes CSS2 and CSS3 selector support, but the package line is old.

Keep CsQuery when replacing it would destabilize a legacy .NET Framework application. For new code, AngleSharp offers a more current standards-oriented DOM and modern target frameworks.

How to choose: a practical decision framework

Start with the response

  1. Request the URL with an HTTP client.
  2. Inspect the response body, not just the browser’s final screen.
  3. If the required text or links are present, parse with AngleSharp or HtmlAgilityPack.
  4. If the response is an app shell and data appears after scripts run, use Playwright, Selenium, or PuppeteerSharp.

Choose selectors deliberately

  • Use CSS selectors with AngleSharp and browser locators when markup is component-oriented.
  • Use XPath with HtmlAgilityPack when existing extraction rules already rely on it.
  • Prefer stable attributes such as data-testid over generated class names.

Account for operations

Parsers are usually cheaper to run and easier to scale. Browser workers require browser binaries, sandbox configuration, memory limits, navigation timeouts, and concurrency control. Cache responses where permitted, use backoff for transient failures, identify your client honestly, and respect the site’s terms, robots guidance, authentication boundaries, and applicable law.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

Empty results with AngleSharp or HtmlAgilityPack

Cause: the content is injected by JavaScript, hidden behind an interaction, or returned only after an API call. Fix: inspect the raw response; switch to Playwright, Selenium, or PuppeteerSharp, or call the underlying data endpoint when its use is authorized.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Browser timeout

Cause: a slow third-party request, a never-idle analytics connection, or an unreachable host. Fix: set a bounded navigation timeout, wait for a specific selector instead of network idle, and record the URL and stage that failed.

Works locally but fails in a container

Cause: missing browser binaries, OS libraries, sandbox permissions, or fonts. Fix: build those dependencies into the image, run the package’s browser-install step during image creation, and use the container’s supported sandbox configuration rather than disabling security blindly.

Stale or duplicated records

Cause: pagination, infinite scroll, retries, or changing content. Fix: use a stable item key, persist crawl checkpoints, deduplicate before writing, and make retries idempotent.

Driver or browser mismatch in Selenium

Cause: incompatible browser and driver versions. Fix: pin compatible versions in the deployment image and treat upgrades as a tested release change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup: ScreenshotNeo

If your goal is a clean page image or PDF rather than extracting structured fields, ScreenshotNeo is the alternative to try first. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; each response reports the page verdict and billing status in headers.

One GET request returns PNG, JPEG, WebP, or PDF. The API also supports full-page and element capture, device and viewport settings, JavaScript and CSS, waits, request blocking, cookies, headers, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and a usage API. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for all parameters.

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan, and yearly billing provides two months free. Create a free ScreenshotNeo account to start.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Can a parser replace a browser?

Only when the required data is already in the HTTP response. A parser does not execute arbitrary JavaScript or perform clicks.

Which library supports all major browser engines?

Microsoft.Playwright provides one .NET API for Chromium, Firefox, and WebKit.

Should I migrate every old ScrapySharp or CsQuery project?

No. Keep a stable legacy dependency when its tests and runtime are acceptable; migrate when you need current frameworks, security updates, or browser rendering.

Is there a universally fastest C# scraper?

No comparable primary benchmark establishes one. Workload, selector complexity, network latency, browser startup, and concurrency determine real performance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.