October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
.NET

Getting Started with Web Scraping in C#: Fetch HTML, Parse It, and Know When to Use a Browser

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with an ordinary HTTP request, not a browser. In C#, reuse an HttpClient, request the page asynchronously, verify the response, and parse the returned HTML with a DOM library such as AngleSharp. Move to Playwright for .NET only when the useful content is created by browser-side JavaScript or requires interaction. Always check the site’s robots.txt, terms and permissions first: RFC 9309 defines robots rules, but explicitly says they are not access authorization.

The smallest responsible C# scraping workflow

  1. Choose a page you are allowed to access and identify the data you need.
  2. Check robots.txt and the site’s terms. Do not use scraping to bypass authentication, paywalls or other access controls.
  3. Reuse an HttpClient and make an asynchronous request.
  4. Check the HTTP status and inspect the response body.
  5. Parse the markup with AngleSharp or Html Agility Pack.
  6. Extract text or attributes with precise selectors.
  7. Add pacing, cancellation, error handling and a clear stop condition.
  8. If the response does not contain the data because a browser must execute JavaScript, use Playwright for .NET instead.

An HTTP client fetches bytes, a parser turns markup into a queryable document, and browser automation runs a real browser. They solve different problems and are not interchangeable.

Set up a console project

Use a current .NET SDK and create a console application:

dotnet new console -n CSharpScraper
cd CSharpScraper
dotnet add package AngleSharp

AngleSharp exposes a standards-oriented HTML DOM with familiar querySelector and querySelectorAll methods. Html Agility Pack is another established option, especially if its API already fits your project. Check the package’s current target frameworks before pinning a version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fetch a page with a reused HttpClient

Microsoft describes HttpClient as the class that sends HTTP requests and receives HTTP responses from a URI. For ordinary retrieval, use asynchronous APIs and inspect the response before parsing.

using System.Net;

using var handler = new HttpClientHandler
{
    AutomaticDecompression = DecompressionMethods.All
};

using var client = new HttpClient(handler)
{
    Timeout = TimeSpan.FromSeconds(30)
};

client.DefaultRequestHeaders.UserAgent.ParseAdd("CSharpScraper/1.0 (contact: [email protected])");

using var response = await client.GetAsync(
    "https://example.com/",
    HttpCompletionOption.ResponseHeadersRead);

if (!response.IsSuccessStatusCode)
{
    Console.Error.WriteLine($"HTTP {(int)response.StatusCode} {response.ReasonPhrase}");
    return;
}

var html = await response.Content.ReadAsStringAsync();
Console.WriteLine($"Received {html.Length} characters");

ResponseHeadersRead lets your code begin handling the response without buffering it through the full HttpClient pipeline. For a small page, the simpler GetStringAsync is also adequate, but it does not give you the same explicit status-check step.

Why client lifetime matters

Do not construct and dispose an HttpClient for every URL. Repeated creation can prevent effective connection reuse and cause avoidable socket and DNS problems. A long-lived client with an appropriate PooledConnectionLifetime, or an IHttpClientFactory in an ASP.NET Core application, is the normal production choice. Configure cookies deliberately: a shared handler can also share cookie state between requests.

var handler = new SocketsHttpHandler
{
    PooledConnectionLifetime = TimeSpan.FromMinutes(5),
    AutomaticDecompression = DecompressionMethods.All
};

using var client = new HttpClient(handler)
{
    Timeout = TimeSpan.FromSeconds(30)
};

Inspect more than the status code

  • Check the final status code and reason phrase.
  • Log the final request URI after redirects when diagnosing unexpected pages.
  • Inspect Content-Type before treating a response as HTML.
  • Enforce a maximum body size for untrusted targets.
  • Honor cancellation tokens and timeouts.

Parse HTML with AngleSharp

Parsing is separate from downloading. The parser builds a DOM that you can query with CSS selectors; it does not, by itself, execute arbitrary page JavaScript.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
using AngleSharp;
using AngleSharp.Dom;

var config = Configuration.Default;
var context = BrowsingContext.New(config);
var document = await context.OpenAsync(req => req.Content(html));

foreach (var link in document.QuerySelectorAll("a"))
{
    var text = link.TextContent.Trim();
    var href = link.GetAttribute("href");
    if (!string.IsNullOrWhiteSpace(text) && !string.IsNullOrWhiteSpace(href))
        Console.WriteLine($"{text} -> {href}");
}

var headings = document.QuerySelectorAll("h1, h2, h3")
    .Select(e => e.TextContent.Trim())
    .Where(t => t.Length > 0)
    .ToList();

Selectors that survive small layout changes

Prefer semantic elements, stable IDs and meaningful classes over deeply nested selectors such as div:nth-child(3) > div > span. Select a product card, then query fields relative to that card:

foreach (var card in document.QuerySelectorAll("article.product-card"))
{
    var name = card.QuerySelector("h2")?.TextContent.Trim();
    var price = card.QuerySelector("[data-price]")?.GetAttribute("data-price");
    Console.WriteLine($"{name}: {price}");
}

Normalize whitespace and treat missing nodes as normal input, not as an exceptional crash. Store the source URL and retrieval time with each record so that downstream users can trace a value back to its page.

When plain HTTP is enough—and when it is not

Need Starting point What it does
Retrieve a page or endpoint HttpClient Sends HTTP requests and receives responses; you handle status, lifetime and connection behavior.
Query returned HTML AngleSharp or Html Agility Pack Builds a DOM and exposes selector APIs; parsing alone does not run page JavaScript.
Browser-dependent content Playwright for .NET Automates Chromium, Firefox and WebKit when execution, interaction or rendering is required.

Download a page first and inspect the raw HTML. If the desired text, links or data attributes are present, a browser is unnecessary. If the response is only an application shell and the data appears after scripts run, use browser automation.

Use Playwright for JavaScript-rendered pages

Playwright for .NET provides one API over Chromium, Firefox and WebKit. It is heavier than an HTTP request because it starts a browser process, but it can execute scripts, wait for selectors and perform interactions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
dotnet add package Microsoft.Playwright
dotnet build
# Install the browser binaries using the command shown by the package's current documentation
using Microsoft.Playwright;

using var playwright = await Playwright.CreateAsync();
await using var browser = await playwright.Chromium.LaunchAsync(new BrowserTypeLaunchOptions
{
    Headless = true
});

var page = await browser.NewPageAsync();
await page.GotoAsync("https://example.com/app", new PageGotoOptions
{
    WaitUntil = WaitUntilState.NetworkIdle,
    Timeout = 30_000
});

await page.WaitForSelectorAsync("article.product-card");
var cards = await page.Locator("article.product-card").AllTextContentsAsync();
foreach (var card in cards)
    Console.WriteLine(card.Trim());

Use explicit waits for a meaningful selector rather than an arbitrary long sleep. Browser runs need more CPU, memory and operational care, so reserve them for pages that genuinely require execution. Never treat browser automation as permission to defeat a login, CAPTCHA or access control.

Responsible crawling design

Robots.txt and permission

Fetch and read the target site’s robots policy before crawling. RFC 9309 is the IETF Robots Exclusion Protocol specification (published September 2022); its rules are instructions for automated crawlers, not authorization. A permissive file does not override terms, copyright, privacy obligations or technical access controls.

Rate, identity and scope

  • Request only the pages and fields you need.
  • Use a clear, truthful User-Agent where appropriate.
  • Space requests and cap concurrency; no universal rate limit applies to every site.
  • Cache pages you have already retrieved and stop on repeated failures.
  • Protect credentials, cookies and personal data in logs.
  • Provide a cancellation path and a finite URL queue.

Pagination and duplicate data

Follow pagination only while the next link is valid and new. Track canonical URLs or stable IDs to avoid loops. Normalize URLs before deduplication, and persist progress so a restart does not begin from page one.

Production reliability and cost considerations

  • Timeouts: Set both an overall request timeout and cancellation for the job.
  • Retries: Retry transient network failures and selected 5xx responses with bounded exponential backoff; do not aggressively retry 4xx responses.
  • Memory: Stream or reject unexpectedly large bodies. Browser pages consume substantially more resources than HTTP requests.
  • Observability: Record URL, status, elapsed time, parser errors and a content hash, while redacting secrets.
  • Change detection: Keep selectors centralized and alert when expected elements disappear.
  • Concurrency: Limit per-host concurrency and respect the site’s capacity.

There is no published universal benchmark that makes one parser or browser faster for every workload. Measure your own pages, selectors and concurrency settings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

Symptom Likely cause Fix
403 or 429 Permission, bot policy or excessive request rate Stop, verify authorization and robots rules, slow down and contact the site owner where appropriate.
200 response but no data Data is injected by JavaScript or loaded from an API Inspect raw HTML and network behavior; switch to Playwright only if browser execution is required and permitted.
Selector returns zero nodes Markup changed, wrong selector or different response variant Save a redacted sample, inspect it, prefer stable attributes and add a selector test.
Task hangs Slow server, streaming response or browser wait with no condition Set timeouts, pass cancellation tokens and wait for a specific selector or bounded state.
Wrong language or consent page Locale, cookies or headers changed the response Set an appropriate Accept-Language, handle cookies deliberately and verify the final URL and title.
Socket exhaustion New HttpClient per request Reuse a client or use IHttpClientFactory.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server when your goal is a clean visual capture rather than extracting structured fields. Its request accepts a URL and returns PNG, JPEG, WebP or PDF. Before capture it can accept cookie/consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. The MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for authentication, options and response details. The same call in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
await Bun.write('shot.webp', res);

ScreenshotNeo includes full-page and element capture, device and retina settings, dark mode, PDF controls, custom CSS and JavaScript, selector waits, request blocking, headers, cookies, user agents, timezone and geolocation, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture for up to 100 URLs per call, usage data and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, and every feature is available on every plan. Sign up free.

FAQ

Can I scrape a site just because robots.txt allows it?

No. Robots rules are not access authorization. Check permissions, terms, privacy and other applicable obligations.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does AngleSharp execute JavaScript?

No. It parses the HTML you give it. Use a permitted browser automation tool when the content appears only after scripts run.

Should I use Html Agility Pack or AngleSharp?

Both are viable .NET parsers. Choose based on selector and DOM behavior, project compatibility and the API your team can maintain.

Is a screenshot API a replacement for a data scraper?

No. A screenshot is visual output; structured extraction still requires an HTTP/parser or browser workflow and selectors for the data.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.