October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
.NET

Web Scraping in C#: From Basics to Production-Ready Code in 2026

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I scrape a website with C#? Start by identifying where the data lives. If the server sends the required markup in its initial response, use a reused HttpClient, validate the response, parse the HTML with Html Agility Pack or AngleSharp, normalize and validate fields, then persist results. If JavaScript creates the content or interaction is required, use Playwright for .NET instead of trying to make an HTML parser execute browser code.

A production scraper is more than a loop over URLs: it has bounded concurrency, cancellation, timeouts, status and size checks, observable failures, resilient selectors, DNS-aware connection management, and an access policy that considers robots.txt, terms, authentication and applicable law.

How do I scrape a website with C#?

Use this pipeline for a page whose data is present in the initial HTML:

  1. Build a request with an explicit timeout and cancellation token.
  2. Send it through a reused HttpClient or an IHttpClientFactory-created client.
  3. Check the HTTP status, content type and response size before parsing.
  4. Parse the document and select nodes with XPath or CSS selectors.
  5. Normalize whitespace, parse numbers and dates, and validate required fields.
  6. Persist accepted records and log rejected pages with enough context to diagnose markup changes.

Do not create and dispose an HttpClient for every URL. Microsoft recommends either a long-lived client with PooledConnectionLifetime or clients created by IHttpClientFactory. DNS is resolved when a connection is created; a long-lived connection can therefore retain an old endpoint until its pool connection is replaced.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A small, defensive static-page scraper

using System.Net.Http.Headers;
using HtmlAgilityPack;

public sealed record Product(string Name, decimal? Price);

public sealed class ProductScraper
{
    private readonly HttpClient _http;

    public ProductScraper(HttpClient http) => _http = http;

    public async Task<Product?> FetchAsync(Uri uri, CancellationToken cancellationToken)
    {
        using var request = new HttpRequestMessage(HttpMethod.Get, uri);
        request.Headers.UserAgent.ParseAdd("ExampleResearchBot/1.0 ([email protected])");

        using var response = await _http.SendAsync(
            request, HttpCompletionOption.ResponseHeadersRead, cancellationToken);
        response.EnsureSuccessStatusCode();

        if (response.Content.Headers.ContentLength is > 10_000_000)
            throw new InvalidDataException("Response exceeds the configured size limit.");

        var html = await response.Content.ReadAsStringAsync(cancellationToken);
        var doc = new HtmlDocument();
        doc.LoadHtml(html);

        var nameNode = doc.DocumentNode.SelectSingleNode("//h1[contains(@class,'product-name')]");
        var priceNode = doc.DocumentNode.SelectSingleNode("//span[contains(@class,'price')]");
        var name = Normalize(nameNode?.InnerText);
        if (string.IsNullOrWhiteSpace(name)) return null;

        decimal? price = decimal.TryParse(
            Normalize(priceNode?.InnerText),
            System.Globalization.NumberStyles.Currency,
            System.Globalization.CultureInfo.InvariantCulture,
            out var parsed) ? parsed : null;

        return new Product(name, price);
    }

    private static string Normalize(string? value) =>
        string.Join(' ', (value ?? string.Empty).Split((char[]?)null,
            StringSplitOptions.RemoveEmptyEntries));
}

Register the client once in dependency injection:

services.AddHttpClient<ProductScraper>(client =>
{
    client.Timeout = TimeSpan.FromSeconds(90);
});

For a long-lived design, configure the handler’s PooledConnectionLifetime according to how often the target’s DNS may change. Microsoft’s commonly shown 15-minute value is illustrative, not a universal production setting. A factory-created client is preferable when your application already uses named or typed clients.

Which C# library should I use for web scraping?

Situation Choice Why
Initial response contains the data HttpClient + Html Agility Pack Lightweight fetching with XPath-oriented parsing.
Initial response contains the data and you prefer standards-style DOM APIs HttpClient + AngleSharp CSS selectors and a browser-like DOM model.
Data appears only after JavaScript runs Playwright for .NET Runs Chromium, Firefox or WebKit and supports browser events and interaction.
Legacy networking code Migrate to HttpClient Microsoft documents WebRequest, WebClient and ServicePoint as obsolete beginning with .NET 6.

There is no established benchmark here that makes one parser universally superior. Choose based on selector style, document quirks, team familiarity and how likely the target markup is to change. Whichever parser you select, treat missing nodes as expected input: emit a parse-failure record instead of silently writing empty values.

AngleSharp example

using AngleSharp;

var context = BrowsingContext.New(Configuration.Default);
var document = await context.OpenAsync(req => req.Content(html));
var title = document.QuerySelector("h1.product-name")?.TextContent.Trim();

Can C# scrape JavaScript-rendered pages?

Yes, but only with a JavaScript-capable runtime. An ordinary HTTP parser sees the server response; it does not execute scripts that later construct the DOM. Playwright’s official .NET port automates Chromium, Firefox and WebKit, and exposes page request, response, completion and failure events.

using Microsoft.Playwright;

using var playwright = await Playwright.CreateAsync();
await using var browser = await playwright.Chromium.LaunchAsync(new BrowserTypeLaunchOptions
{
    Headless = true
});
var page = await browser.NewPageAsync();
page.Response += (_, response) =>
    Console.WriteLine($"{response.Status} {response.Url}");

await page.GotoAsync("https://example.com/catalog",
    new PageGotoOptions { WaitUntil = WaitUntilState.NetworkIdle });
await page.WaitForSelectorAsync(".product-card");
var cards = await page.Locator(".product-card").AllTextContentsAsync();

Install the browser binaries required by your deployment and account for their disk, startup and patching overhead. Listen to request failures and responses when debugging. A 404 or 503 response can still be a completed browser request, so inspect status explicitly. Before launching a browser, check whether the page calls a public, permitted data endpoint; direct HTTP may be simpler. Browser automation does not override authentication requirements, site controls or legal restrictions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use HttpClient, HtmlAgilityPack, AngleSharp, or Playwright?

Decide along six axes:

  • Content location: initial HTML favors HTTP plus a parser; post-load DOM favors Playwright.
  • Runtime cost: an HTTP client is lightweight; browsers require binaries and more memory.
  • Interaction: clicks, cookies, dialogs and page events require a browser.
  • Selection ergonomics: Html Agility Pack is commonly XPath-oriented; AngleSharp provides CSS/DOM APIs.
  • Operations: both designs still need cancellation, bounded concurrency, status checks, logging and output validation.
  • Access constraints: terms, robots rules, authentication and applicable law apply to either approach.

Production reliability: lifecycle, limits and failure handling

Cancellation, timeouts and status

Pass a cancellation token from the job or request boundary. Set a finite client timeout, but also cancel individual operations when a job is stopped. Call EnsureSuccessStatusCode only after deciding how to record non-success responses. Preserve status, URL and a short error body in structured logs; never log credentials or sensitive page content by default.

Response size and malformed markup

Enforce a maximum response size before loading a document into memory. Check content type when the target can return files or error pages. Parsers may recover from malformed HTML, but recovery can move nodes; validate required fields and alert on sudden parse-failure spikes.

Concurrency, pacing and retries

Use a bounded worker count and site-appropriate pacing. There is no universal safe concurrency, timeout or retry number: capacity and instructions differ by site. Retry only transient failures and operations that are safe to repeat; do not blindly retry authentication failures, validation errors or a server explicitly refusing access. Honor cancellation while waiting between attempts.

Selector and data quality checks

Prefer stable attributes over deeply positional XPath. Keep selectors in configuration when targets change independently of code. Validate identifiers, currencies, dates and required relationships, and store the source URL and retrieval time with each record. A successful HTTP 200 is not proof that extraction succeeded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots.txt: is it permission to scrape?

No. RFC 9309 (IETF, September 2022) states: “These rules are not a form of access authorization.” Robots Exclusion Protocol instructions are crawler directions, not a grant of legal permission or a replacement for authentication and security controls.

  • A successfully retrieved, parseable /robots.txt is to be followed by crawlers that honor the protocol.
  • The standard says cached files generally should not be used for more than 24 hours unless the file is unreachable.
  • If a server or network error makes the file unreachable, the crawler must assume complete disallow.
  • A 4xx “unavailable” response is treated differently under the protocol; do not reduce every missing-file case to “allowed.”
  • The RFC specifies a 500 kibibyte minimum parsing limit and gives 30 days as an example duration for handling an undefined or unavailable file. These are protocol details, not request-rate advice.

Separately evaluate the site’s terms, authorization, access controls, copyright and privacy obligations in the relevant jurisdiction. Public visibility alone does not establish that a particular use is lawful.

Or skip the browser setup

For screenshot or rendered-page capture, ScreenshotNeo provides a single-request API and an MCP server for Claude, Cursor and other MCP clients. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

Use the ScreenshotNeo API documentation for all options, including full-page or element capture, device and retina settings, PDF output, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, caching, signed links, asynchronous jobs, bulk capture and usage reporting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

An MCP server lets AI agents take screenshots without you packaging browser binaries. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Every request returns 403 or 429

Check authorization, terms and robots instructions, identify your client honestly, slow and bound concurrency, and stop if the site denies access. Do not try to bypass a control.

The parser finds no records but the browser shows them

Inspect the raw response. If records are absent, switch to Playwright or, where permitted, identify the page’s underlying data request.

Intermittent DNS or stale endpoints

Use IHttpClientFactory or configure an appropriate PooledConnectionLifetime; do not recreate a client per request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright hangs during navigation

Set a finite timeout, use a narrower wait condition than network idle when third-party resources never settle, and capture request-failure and response-status logs.

Fields suddenly become empty

Record selector misses as failures, compare a saved sample of the response, and update selectors only after confirming the markup change.

Frequently Asked Questions

Does scraping require a browser in every C# project?

No. A reused HttpClient and an HTML parser are sufficient when the needed data is in the initial response.

Can I treat a 200 response as a successful scrape?

No. HTTP success and extraction success are separate; validate required fields and record parse failures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Are robots.txt rules a legal safe harbor?

No. RFC 9309 describes crawler instructions, not access authorization; assess the site’s terms, controls and applicable law.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.