How do I scrape a website with C#? Start by identifying where the data lives. If the server sends the required markup in its initial response, use a reused HttpClient, validate the response, parse the HTML with Html Agility Pack or AngleSharp, normalize and validate fields, then persist results. If JavaScript creates the content or interaction is required, use Playwright for .NET instead of trying to make an HTML parser execute browser code.
A production scraper is more than a loop over URLs: it has bounded concurrency, cancellation, timeouts, status and size checks, observable failures, resilient selectors, DNS-aware connection management, and an access policy that considers robots.txt, terms, authentication and applicable law.
How do I scrape a website with C#?
Use this pipeline for a page whose data is present in the initial HTML:
- Build a request with an explicit timeout and cancellation token.
- Send it through a reused
HttpClientor anIHttpClientFactory-created client. - Check the HTTP status, content type and response size before parsing.
- Parse the document and select nodes with XPath or CSS selectors.
- Normalize whitespace, parse numbers and dates, and validate required fields.
- Persist accepted records and log rejected pages with enough context to diagnose markup changes.
Do not create and dispose an HttpClient for every URL. Microsoft recommends either a long-lived client with PooledConnectionLifetime or clients created by IHttpClientFactory. DNS is resolved when a connection is created; a long-lived connection can therefore retain an old endpoint until its pool connection is replaced.
Recommended Free Tools
#1 Best Overall
A small, defensive static-page scraper
using System.Net.Http.Headers;
using HtmlAgilityPack;
public sealed record Product(string Name, decimal? Price);
public sealed class ProductScraper
{
private readonly HttpClient _http;
public ProductScraper(HttpClient http) => _http = http;
public async Task<Product?> FetchAsync(Uri uri, CancellationToken cancellationToken)
{
using var request = new HttpRequestMessage(HttpMethod.Get, uri);
request.Headers.UserAgent.ParseAdd("ExampleResearchBot/1.0 ([email protected])");
using var response = await _http.SendAsync(
request, HttpCompletionOption.ResponseHeadersRead, cancellationToken);
response.EnsureSuccessStatusCode();
if (response.Content.Headers.ContentLength is > 10_000_000)
throw new InvalidDataException("Response exceeds the configured size limit.");
var html = await response.Content.ReadAsStringAsync(cancellationToken);
var doc = new HtmlDocument();
doc.LoadHtml(html);
var nameNode = doc.DocumentNode.SelectSingleNode("//h1[contains(@class,'product-name')]");
var priceNode = doc.DocumentNode.SelectSingleNode("//span[contains(@class,'price')]");
var name = Normalize(nameNode?.InnerText);
if (string.IsNullOrWhiteSpace(name)) return null;
decimal? price = decimal.TryParse(
Normalize(priceNode?.InnerText),
System.Globalization.NumberStyles.Currency,
System.Globalization.CultureInfo.InvariantCulture,
out var parsed) ? parsed : null;
return new Product(name, price);
}
private static string Normalize(string? value) =>
string.Join(' ', (value ?? string.Empty).Split((char[]?)null,
StringSplitOptions.RemoveEmptyEntries));
}
Register the client once in dependency injection:
services.AddHttpClient<ProductScraper>(client =>
{
client.Timeout = TimeSpan.FromSeconds(90);
});
For a long-lived design, configure the handler’s PooledConnectionLifetime according to how often the target’s DNS may change. Microsoft’s commonly shown 15-minute value is illustrative, not a universal production setting. A factory-created client is preferable when your application already uses named or typed clients.
Which C# library should I use for web scraping?
| Situation | Choice | Why |
|---|---|---|
| Initial response contains the data | HttpClient + Html Agility Pack |
Lightweight fetching with XPath-oriented parsing. |
| Initial response contains the data and you prefer standards-style DOM APIs | HttpClient + AngleSharp |
CSS selectors and a browser-like DOM model. |
| Data appears only after JavaScript runs | Playwright for .NET | Runs Chromium, Firefox or WebKit and supports browser events and interaction. |
| Legacy networking code | Migrate to HttpClient |
Microsoft documents WebRequest, WebClient and ServicePoint as obsolete beginning with .NET 6. |
There is no established benchmark here that makes one parser universally superior. Choose based on selector style, document quirks, team familiarity and how likely the target markup is to change. Whichever parser you select, treat missing nodes as expected input: emit a parse-failure record instead of silently writing empty values.
AngleSharp example
using AngleSharp;
var context = BrowsingContext.New(Configuration.Default);
var document = await context.OpenAsync(req => req.Content(html));
var title = document.QuerySelector("h1.product-name")?.TextContent.Trim();
Can C# scrape JavaScript-rendered pages?
Yes, but only with a JavaScript-capable runtime. An ordinary HTTP parser sees the server response; it does not execute scripts that later construct the DOM. Playwright’s official .NET port automates Chromium, Firefox and WebKit, and exposes page request, response, completion and failure events.
using Microsoft.Playwright;
using var playwright = await Playwright.CreateAsync();
await using var browser = await playwright.Chromium.LaunchAsync(new BrowserTypeLaunchOptions
{
Headless = true
});
var page = await browser.NewPageAsync();
page.Response += (_, response) =>
Console.WriteLine($"{response.Status} {response.Url}");
await page.GotoAsync("https://example.com/catalog",
new PageGotoOptions { WaitUntil = WaitUntilState.NetworkIdle });
await page.WaitForSelectorAsync(".product-card");
var cards = await page.Locator(".product-card").AllTextContentsAsync();
Install the browser binaries required by your deployment and account for their disk, startup and patching overhead. Listen to request failures and responses when debugging. A 404 or 503 response can still be a completed browser request, so inspect status explicitly. Before launching a browser, check whether the page calls a public, permitted data endpoint; direct HTTP may be simpler. Browser automation does not override authentication requirements, site controls or legal restrictions.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteRank #2
Should I use HttpClient, HtmlAgilityPack, AngleSharp, or Playwright?
Decide along six axes:
- Content location: initial HTML favors HTTP plus a parser; post-load DOM favors Playwright.
- Runtime cost: an HTTP client is lightweight; browsers require binaries and more memory.
- Interaction: clicks, cookies, dialogs and page events require a browser.
- Selection ergonomics: Html Agility Pack is commonly XPath-oriented; AngleSharp provides CSS/DOM APIs.
- Operations: both designs still need cancellation, bounded concurrency, status checks, logging and output validation.
- Access constraints: terms, robots rules, authentication and applicable law apply to either approach.
Production reliability: lifecycle, limits and failure handling
Cancellation, timeouts and status
Pass a cancellation token from the job or request boundary. Set a finite client timeout, but also cancel individual operations when a job is stopped. Call EnsureSuccessStatusCode only after deciding how to record non-success responses. Preserve status, URL and a short error body in structured logs; never log credentials or sensitive page content by default.
Response size and malformed markup
Enforce a maximum response size before loading a document into memory. Check content type when the target can return files or error pages. Parsers may recover from malformed HTML, but recovery can move nodes; validate required fields and alert on sudden parse-failure spikes.
Concurrency, pacing and retries
Use a bounded worker count and site-appropriate pacing. There is no universal safe concurrency, timeout or retry number: capacity and instructions differ by site. Retry only transient failures and operations that are safe to repeat; do not blindly retry authentication failures, validation errors or a server explicitly refusing access. Honor cancellation while waiting between attempts.
Selector and data quality checks
Prefer stable attributes over deeply positional XPath. Keep selectors in configuration when targets change independently of code. Validate identifiers, currencies, dates and required relationships, and store the source URL and retrieval time with each record. A successful HTTP 200 is not proof that extraction succeeded.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Robots.txt: is it permission to scrape?
No. RFC 9309 (IETF, September 2022) states: “These rules are not a form of access authorization.” Robots Exclusion Protocol instructions are crawler directions, not a grant of legal permission or a replacement for authentication and security controls.
- A successfully retrieved, parseable
/robots.txtis to be followed by crawlers that honor the protocol. - The standard says cached files generally should not be used for more than 24 hours unless the file is unreachable.
- If a server or network error makes the file unreachable, the crawler must assume complete disallow.
- A 4xx “unavailable” response is treated differently under the protocol; do not reduce every missing-file case to “allowed.”
- The RFC specifies a 500 kibibyte minimum parsing limit and gives 30 days as an example duration for handling an undefined or unavailable file. These are protocol details, not request-rate advice.
Separately evaluate the site’s terms, authorization, access controls, copyright and privacy obligations in the relevant jurisdiction. Public visibility alone does not establish that a particular use is lawful.
Or skip the browser setup
For screenshot or rendered-page capture, ScreenshotNeo provides a single-request API and an MCP server for Claude, Cursor and other MCP clients. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
Use the ScreenshotNeo API documentation for all options, including full-page or element capture, device and retina settings, PDF output, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, caching, signed links, asynchronous jobs, bulk capture and usage reporting.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
An MCP server lets AI agents take screenshots without you packaging browser binaries. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Rank #4
Troubleshooting common failures
Every request returns 403 or 429
Check authorization, terms and robots instructions, identify your client honestly, slow and bound concurrency, and stop if the site denies access. Do not try to bypass a control.
The parser finds no records but the browser shows them
Inspect the raw response. If records are absent, switch to Playwright or, where permitted, identify the page’s underlying data request.
Intermittent DNS or stale endpoints
Use IHttpClientFactory or configure an appropriate PooledConnectionLifetime; do not recreate a client per request.
Playwright hangs during navigation
Set a finite timeout, use a narrower wait condition than network idle when third-party resources never settle, and capture request-failure and response-status logs.
Best Value
Fields suddenly become empty
Record selector misses as failures, compare a saved sample of the response, and update selectors only after confirming the markup change.
Frequently Asked Questions
Does scraping require a browser in every C# project?
No. A reused HttpClient and an HTML parser are sufficient when the needed data is in the initial response.
Can I treat a 200 response as a successful scrape?
No. HTTP success and extraction success are separate; validate required fields and record parse failures.
Are robots.txt rules a legal safe harbor?
No. RFC 9309 describes crawler instructions, not access authorization; assess the site’s terms, controls and applicable law.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




