Start with an ordinary HTTP request, not a browser. In C#, reuse an HttpClient, request the page asynchronously, verify the response, and parse the returned HTML with a DOM library such as AngleSharp. Move to Playwright for .NET only when the useful content is created by browser-side JavaScript or requires interaction. Always check the site’s robots.txt, terms and permissions first: RFC 9309 defines robots rules, but explicitly says they are not access authorization.
The smallest responsible C# scraping workflow
- Choose a page you are allowed to access and identify the data you need.
- Check
robots.txtand the site’s terms. Do not use scraping to bypass authentication, paywalls or other access controls. - Reuse an
HttpClientand make an asynchronous request. - Check the HTTP status and inspect the response body.
- Parse the markup with AngleSharp or Html Agility Pack.
- Extract text or attributes with precise selectors.
- Add pacing, cancellation, error handling and a clear stop condition.
- If the response does not contain the data because a browser must execute JavaScript, use Playwright for .NET instead.
An HTTP client fetches bytes, a parser turns markup into a queryable document, and browser automation runs a real browser. They solve different problems and are not interchangeable.
Set up a console project
Use a current .NET SDK and create a console application:
dotnet new console -n CSharpScraper
cd CSharpScraper
dotnet add package AngleSharp
AngleSharp exposes a standards-oriented HTML DOM with familiar querySelector and querySelectorAll methods. Html Agility Pack is another established option, especially if its API already fits your project. Check the package’s current target frameworks before pinning a version.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Fetch a page with a reused HttpClient
Microsoft describes HttpClient as the class that sends HTTP requests and receives HTTP responses from a URI. For ordinary retrieval, use asynchronous APIs and inspect the response before parsing.
using System.Net;
using var handler = new HttpClientHandler
{
AutomaticDecompression = DecompressionMethods.All
};
using var client = new HttpClient(handler)
{
Timeout = TimeSpan.FromSeconds(30)
};
client.DefaultRequestHeaders.UserAgent.ParseAdd("CSharpScraper/1.0 (contact: [email protected])");
using var response = await client.GetAsync(
"https://example.com/",
HttpCompletionOption.ResponseHeadersRead);
if (!response.IsSuccessStatusCode)
{
Console.Error.WriteLine($"HTTP {(int)response.StatusCode} {response.ReasonPhrase}");
return;
}
var html = await response.Content.ReadAsStringAsync();
Console.WriteLine($"Received {html.Length} characters");
ResponseHeadersRead lets your code begin handling the response without buffering it through the full HttpClient pipeline. For a small page, the simpler GetStringAsync is also adequate, but it does not give you the same explicit status-check step.
Why client lifetime matters
Do not construct and dispose an HttpClient for every URL. Repeated creation can prevent effective connection reuse and cause avoidable socket and DNS problems. A long-lived client with an appropriate PooledConnectionLifetime, or an IHttpClientFactory in an ASP.NET Core application, is the normal production choice. Configure cookies deliberately: a shared handler can also share cookie state between requests.
Rank #2
var handler = new SocketsHttpHandler
{
PooledConnectionLifetime = TimeSpan.FromMinutes(5),
AutomaticDecompression = DecompressionMethods.All
};
using var client = new HttpClient(handler)
{
Timeout = TimeSpan.FromSeconds(30)
};
Inspect more than the status code
- Check the final status code and reason phrase.
- Log the final request URI after redirects when diagnosing unexpected pages.
- Inspect
Content-Typebefore treating a response as HTML. - Enforce a maximum body size for untrusted targets.
- Honor cancellation tokens and timeouts.
Parse HTML with AngleSharp
Parsing is separate from downloading. The parser builds a DOM that you can query with CSS selectors; it does not, by itself, execute arbitrary page JavaScript.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteusing AngleSharp;
using AngleSharp.Dom;
var config = Configuration.Default;
var context = BrowsingContext.New(config);
var document = await context.OpenAsync(req => req.Content(html));
foreach (var link in document.QuerySelectorAll("a"))
{
var text = link.TextContent.Trim();
var href = link.GetAttribute("href");
if (!string.IsNullOrWhiteSpace(text) && !string.IsNullOrWhiteSpace(href))
Console.WriteLine($"{text} -> {href}");
}
var headings = document.QuerySelectorAll("h1, h2, h3")
.Select(e => e.TextContent.Trim())
.Where(t => t.Length > 0)
.ToList();
Selectors that survive small layout changes
Prefer semantic elements, stable IDs and meaningful classes over deeply nested selectors such as div:nth-child(3) > div > span. Select a product card, then query fields relative to that card:
foreach (var card in document.QuerySelectorAll("article.product-card"))
{
var name = card.QuerySelector("h2")?.TextContent.Trim();
var price = card.QuerySelector("[data-price]")?.GetAttribute("data-price");
Console.WriteLine($"{name}: {price}");
}
Normalize whitespace and treat missing nodes as normal input, not as an exceptional crash. Store the source URL and retrieval time with each record so that downstream users can trace a value back to its page.
Rank #3
When plain HTTP is enough—and when it is not
| Need | Starting point | What it does |
|---|---|---|
| Retrieve a page or endpoint | HttpClient |
Sends HTTP requests and receives responses; you handle status, lifetime and connection behavior. |
| Query returned HTML | AngleSharp or Html Agility Pack | Builds a DOM and exposes selector APIs; parsing alone does not run page JavaScript. |
| Browser-dependent content | Playwright for .NET | Automates Chromium, Firefox and WebKit when execution, interaction or rendering is required. |
Download a page first and inspect the raw HTML. If the desired text, links or data attributes are present, a browser is unnecessary. If the response is only an application shell and the data appears after scripts run, use browser automation.
Use Playwright for JavaScript-rendered pages
Playwright for .NET provides one API over Chromium, Firefox and WebKit. It is heavier than an HTTP request because it starts a browser process, but it can execute scripts, wait for selectors and perform interactions.
Recommended Free Tools
dotnet add package Microsoft.Playwright
dotnet build
# Install the browser binaries using the command shown by the package's current documentation
using Microsoft.Playwright;
using var playwright = await Playwright.CreateAsync();
await using var browser = await playwright.Chromium.LaunchAsync(new BrowserTypeLaunchOptions
{
Headless = true
});
var page = await browser.NewPageAsync();
await page.GotoAsync("https://example.com/app", new PageGotoOptions
{
WaitUntil = WaitUntilState.NetworkIdle,
Timeout = 30_000
});
await page.WaitForSelectorAsync("article.product-card");
var cards = await page.Locator("article.product-card").AllTextContentsAsync();
foreach (var card in cards)
Console.WriteLine(card.Trim());
Use explicit waits for a meaningful selector rather than an arbitrary long sleep. Browser runs need more CPU, memory and operational care, so reserve them for pages that genuinely require execution. Never treat browser automation as permission to defeat a login, CAPTCHA or access control.
Responsible crawling design
Robots.txt and permission
Fetch and read the target site’s robots policy before crawling. RFC 9309 is the IETF Robots Exclusion Protocol specification (published September 2022); its rules are instructions for automated crawlers, not authorization. A permissive file does not override terms, copyright, privacy obligations or technical access controls.
Rate, identity and scope
- Request only the pages and fields you need.
- Use a clear, truthful User-Agent where appropriate.
- Space requests and cap concurrency; no universal rate limit applies to every site.
- Cache pages you have already retrieved and stop on repeated failures.
- Protect credentials, cookies and personal data in logs.
- Provide a cancellation path and a finite URL queue.
Pagination and duplicate data
Follow pagination only while the next link is valid and new. Track canonical URLs or stable IDs to avoid loops. Normalize URLs before deduplication, and persist progress so a restart does not begin from page one.
Production reliability and cost considerations
- Timeouts: Set both an overall request timeout and cancellation for the job.
- Retries: Retry transient network failures and selected 5xx responses with bounded exponential backoff; do not aggressively retry 4xx responses.
- Memory: Stream or reject unexpectedly large bodies. Browser pages consume substantially more resources than HTTP requests.
- Observability: Record URL, status, elapsed time, parser errors and a content hash, while redacting secrets.
- Change detection: Keep selectors centralized and alert when expected elements disappear.
- Concurrency: Limit per-host concurrency and respect the site’s capacity.
There is no published universal benchmark that makes one parser or browser faster for every workload. Measure your own pages, selectors and concurrency settings.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 or 429 | Permission, bot policy or excessive request rate | Stop, verify authorization and robots rules, slow down and contact the site owner where appropriate. |
| 200 response but no data | Data is injected by JavaScript or loaded from an API | Inspect raw HTML and network behavior; switch to Playwright only if browser execution is required and permitted. |
| Selector returns zero nodes | Markup changed, wrong selector or different response variant | Save a redacted sample, inspect it, prefer stable attributes and add a selector test. |
| Task hangs | Slow server, streaming response or browser wait with no condition | Set timeouts, pass cancellation tokens and wait for a specific selector or bounded state. |
| Wrong language or consent page | Locale, cookies or headers changed the response | Set an appropriate Accept-Language, handle cookies deliberately and verify the final URL and title. |
| Socket exhaustion | New HttpClient per request |
Reuse a client or use IHttpClientFactory. |
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when your goal is a clean visual capture rather than extracting structured fields. Its request accepts a URL and returns PNG, JPEG, WebP or PDF. Before capture it can accept cookie/consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. The MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for authentication, options and response details. The same call in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
await Bun.write('shot.webp', res);
ScreenshotNeo includes full-page and element capture, device and retina settings, dark mode, PDF controls, custom CSS and JavaScript, selector waits, request blocking, headers, cookies, user agents, timezone and geolocation, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture for up to 100 URLs per call, usage data and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, and every feature is available on every plan. Sign up free.
FAQ
Can I scrape a site just because robots.txt allows it?
No. Robots rules are not access authorization. Check permissions, terms, privacy and other applicable obligations.
Free tools Windows power users keep installed
One-click scans. No signup required.
Does AngleSharp execute JavaScript?
No. It parses the HTML you give it. Use a permitted browser automation tool when the content appears only after scripts run.
Should I use Html Agility Pack or AngleSharp?
Both are viable .NET parsers. Choose based on selector and DOM behavior, project compatibility and the API your team can maintain.
Is a screenshot API a replacement for a data scraper?
No. A screenshot is visual output; structured extraction still requires an HTTP/parser or browser workflow and selectors for the data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




