To capture an HTML table in ASP.NET, download the page with HttpClient, parse the response with a DOM library such as Html Agility Pack, select the intended table, iterate both th and td cells, normalize each cell’s text, and map the rows to objects or an export format. This approach handles nested spans and imperfect markup far more safely than regular expressions.
The capture pipeline
A reliable implementation separates five jobs:
- Fetch: receive the actual HTML response with
HttpClient. - Parse: build a document object model (DOM) with Html Agility Pack (HAP), a free, open-source NuGet package for reading and writing HTML with XPath/XSLT support.
- Select: target a stable table id, class, or narrowly scoped XPath.
- Normalize: read descendant text, decode entities, and trim whitespace.
- Map: convert cells to a DTO,
DataTable, CSV, JSON, or database record.
Do not select the first table blindly. Pages often contain navigation, layout, or nested tables.
Prerequisites and package setup
The example works in an ASP.NET Core application or another modern .NET project. Add HAP from NuGet:
dotnet add package HtmlAgilityPack
Use a supported .NET SDK and configure outbound network access from the server. The target site may still require authentication, cookies, a specific user agent, or permission to automate access; those policies are site-specific.
Recommended Free Tools
#1 Best Overall
Complete C# example
This service downloads a page, finds <table id="results">, includes header and data cells, and returns strongly typed rows.
using System.Net;
using System.Net.Http.Headers;
using HtmlAgilityPack;
public sealed record ResultRow(string Name, string Status, string Amount);
public sealed class HtmlTableReader
{
private readonly HttpClient _http;
public HtmlTableReader(HttpClient http)
{
_http = http;
_http.DefaultRequestHeaders.UserAgent.Add(
new ProductInfoHeaderValue("TableReader", "1.0"));
}
public async Task<IReadOnlyList<ResultRow>> ReadAsync(
Uri pageUri, CancellationToken cancellationToken = default)
{
using var response = await _http.GetAsync(pageUri, cancellationToken);
response.EnsureSuccessStatusCode();
var html = await response.Content.ReadAsStringAsync(cancellationToken);
var doc = new HtmlDocument();
doc.LoadHtml(html);
var table = doc.DocumentNode.SelectSingleNode(
"//table[@id='results']");
if (table is null)
throw new InvalidOperationException("Table #results was not found.");
var output = new List<ResultRow>();
var rows = table.SelectNodes(".//tr") ?? Enumerable.Empty<HtmlNode>();
foreach (var row in rows)
{
var cells = row.SelectNodes("./th|./td");
if (cells is null || cells.Count < 3)
continue;
var values = cells.Select(ReadCell).ToArray();
output.Add(new ResultRow(values[0], values[1], values[2]));
}
return output;
}
private static string ReadCell(HtmlNode cell)
{
var decoded = WebUtility.HtmlDecode(cell.InnerText) ?? string.Empty;
return string.Join(" ", decoded
.Split((char[]?)null, StringSplitOptions.RemoveEmptyEntries));
}
}
Register the client with IHttpClientFactory rather than constructing a new client for every request:
builder.Services.AddHttpClient<HtmlTableReader>(client =>
{
client.Timeout = TimeSpan.FromSeconds(30);
});
Inject HtmlTableReader into a controller, Razor Page, background service, or endpoint and call ReadAsync. Pass a cancellation token from the request so a disconnected client does not leave unnecessary work running.
Selecting the correct table
Stable id
var table = doc.DocumentNode.SelectSingleNode("//table[@id='results']");
An id is usually the least ambiguous selector, but confirm it exists in the downloaded response rather than only in browser developer tools.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Class or attribute
var table = doc.DocumentNode.SelectSingleNode(
"//table[contains(concat(' ', normalize-space(@class), ' '), ' data-grid ')]");
The expanded class expression avoids matching a class that merely contains the same characters as a word.
Rank #2
Narrow XPath scope
var table = doc.DocumentNode.SelectSingleNode(
"//main[@id='content']//table[@data-kind='orders']");
Use .//tr beneath the table. It remains correct when a browser inserts a tbody element or the source nests rows differently. Do not assume rows are direct children.
CSS selectors and alternatives
HAP’s core workflow is XPath. Aspose.HTML for .NET provides CSS selectors, URL or file loading, link extraction, and export-oriented examples, and may suit a project that needs a supported commercial component. AngleSharp is another HTML5-parser option in the .NET ecosystem; verify its current API and licensing for your application. Compare parser tolerance, selector model, loading features, export helpers, maintenance, and licensing before standardizing.
Headers, nested markup, and malformed HTML
Select ./th|./td, not only td. Otherwise a header row disappears and column positions can shift. HAP’s InnerText includes text inside descendant elements such as span, a, and custom inline tags. Decode entities such as & with WebUtility.HtmlDecode, then trim or collapse whitespace.
Rows do not always have the same number of cells. A colspan can make a visual table contain fewer nodes than its apparent columns. Decide whether to skip short rows, pad them, or expand spans according to the target site’s rules. Validate conversions before assigning values:
if (!decimal.TryParse(values[2], NumberStyles.Number,
CultureInfo.InvariantCulture, out var amount))
{
// Log the row and handle the site's currency/locale convention.
}
Keep the raw HTML or row text in diagnostic logs (subject to privacy rules) when a mapping fails. This makes markup changes easier to identify.
Exporting captured rows
JSON
var rows = await reader.ReadAsync(uri, cancellationToken);
var json = JsonSerializer.Serialize(rows,
new JsonSerializerOptions { WriteIndented = true });
await File.WriteAllTextAsync("results.json", json, cancellationToken);
CSV
For CSV, quote every field containing a comma, quote, or line break, and double embedded quotes. Do not concatenate values with commas without escaping. A dedicated CSV writer is safer for production exports.
DataTable
var table = new DataTable();
table.Columns.Add("Name");
table.Columns.Add("Status");
table.Columns.Add("Amount");
foreach (var row in rows)
table.Rows.Add(row.Name, row.Status, row.Amount);
For database storage, map into a parameterized command or an ORM entity. Keep parsing and persistence separate so a schema change does not require changing the downloader.
Free tools Windows power users keep installed
One-click scans. No signup required.
Static HTML versus JavaScript-rendered tables
A server-side HttpClient receives the response body returned by the server. If JavaScript later builds the table in the browser, the response may contain no rows for HAP to find. First inspect the downloaded HTML and log its length or a safe excerpt. If the table is absent, use browser developer tools to identify the data endpoint or rendering mechanism, then request that endpoint if the site’s terms and authentication permit it. Parse the returned JSON or HTML instead of guessing at browser state.
If the data endpoint requires a session, reproduce the necessary cookies or authorization securely. Do not put credentials in source control or logs. Browser automation is a separate solution when rendering, interaction, or client-side execution is essential; it adds startup time, resource use, and operational failure modes.
Why regular expressions are the wrong default
HTML can be malformed, contain nested elements, include quoted attributes in many forms, and place meaningful text across descendants. Microsoft guidance for structured table scraping recommends a parser rather than regex for these reasons. A DOM parser lets you express “cells directly under this row” and still handle nested spans and entity decoding.
Rank #4
Reliability, performance, and safety
- Timeouts: set an explicit
HttpClienttimeout and pass cancellation tokens. - Retries: retry only transient transport or 5xx failures, with bounded exponential backoff. Do not aggressively retry 4xx responses or anti-bot challenges.
- Limits: cap response size where appropriate and reject unexpectedly large documents to protect memory.
- Validation: check status codes, content type, table existence, expected headers, row widths, and numeric/date conversions.
- Logging: record URI, status, elapsed time, selector, and row count without exposing secrets or personal data.
- Caching: cache according to the site’s freshness needs and terms; avoid repeatedly downloading unchanged pages.
- Concurrency: reuse the factory-managed client and limit parallel requests to the target’s published or implied capacity.
- Compliance: respect authentication requirements, rate limits, robots policies, copyright, and privacy obligations.
Troubleshooting checklist
| Symptom | Likely cause | Fix |
|---|---|---|
| “Table not found” | Selector mismatch or table is generated by JavaScript | Save and inspect the response; use a stable id/class or locate the data endpoint. |
| Rows are empty | Only td was selected, or cells contain nested markup |
Select ./th|./td and read InnerText. |
| Text contains odd spacing | Whitespace nodes and line breaks inside cells | Decode entities, split on whitespace, remove empty entries, and join with a single space. |
| Columns shift | Header omitted, colspan, or irregular rows |
Include th, validate counts, and implement explicit span handling. |
| 403, 429, or challenge page | Access control or rate limit | Use permitted authentication, slow requests, and stop retrying challenge responses. |
| Timeout or partial response | Slow origin, large page, or network instability | Set a realistic timeout, use cancellation, bounded retries, and response-size safeguards. |
| Parser throws on broken markup | Unexpected document encoding or severe structure damage | Capture the raw response, verify encoding, and test the selector against the actual document. |
Or skip the browser setup
If you need a rendered page image rather than structured cell values, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie-consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsFor a direct capture, see the ScreenshotNeo documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
It also supports full-page and element captures, device and retina settings, dark mode, custom CSS or JavaScript, waits, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and PDF options. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Sign up for the free plan.
Frequently asked questions
Can HAP execute JavaScript?
No. It parses the HTML it receives. Use an accessible data endpoint or a browser automation solution when client-side rendering is required.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteShould I store the entire HTML response?
Store it only when your privacy, retention, and storage policies allow. Otherwise log a request identifier and targeted diagnostics such as selector status and row count.
What if the table has multiple header rows?
Capture all th cells, then apply a deliberate header-normalization rule that combines or selects the rows your export schema expects.
Frequently Asked Questions
Can HAP execute JavaScript?
No. It parses the HTML received by HttpClient; use the site’s data endpoint or browser automation for client-rendered content.
Should I select rows with .//tr or ./tr?
Use .//tr unless you have verified the exact structure. It tolerates an intervening tbody and other valid nesting.
How do I preserve links in a cell?
Read the cell’s descendant anchor attributes separately (for example, its href) while using InnerText for the displayed label.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




