Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Do not concatenate complete HTML files as plain text. Parse each page, create one destination document, and append the selected nodes from every source body in order. This avoids repeated <html>, <head> and <body> elements and gives you explicit control over styles, scripts, links, IDs and metadata.
In .NET, AngleSharp is a strong default when standards-oriented HTML5 parsing and fragment insertion matter. HtmlAgilityPack is also practical for applications already using its node API. The implementation below uses AngleSharp and shows the decisions you must make when combining real pages.
What “combine” should mean
A valid HTML document has one document element, normally one head and one body. Joining two strings such as File.ReadAllText("a.html") + File.ReadAllText("b.html") creates multiple document shells, duplicate metadata and often invalid resource behavior. Instead, define a destination shell and decide what to import from each source:
- Body content: usually the source
<main>, an article element, or selected children of<body>. - Head resources: stylesheets, fonts, metadata and scripts that the combined page actually needs.
- Order: the sequence in which source pages appear in the output.
- Conflict policy: how to handle duplicate IDs, repeated styles, scripts, titles and base URLs.
The WHATWG HTML Standard defines separate document and fragment parsing algorithms. Use a document parser for complete pages and a fragment-oriented operation when inserting markup into a particular destination element.
#1 Best Overall
Choose a parser
AngleSharp
AngleSharp builds a standards-oriented DOM, supports CSS-style querying and provides document and fragment parsing capabilities. It is a good fit for HTML5 cleanup and predictable insertion contexts. Its project examples and fragment guidance are available in the examples and fragment questions documentation.
HtmlAgilityPack
HtmlAgilityPack loads HTML from files or strings and exposes node manipulation APIs documented at its manipulation guide. The NuGet listing showed version 1.13.0 at the time of research; verify the current package version and target-framework compatibility before installing.
Choose based on malformed-markup behavior, fragment support, API familiarity and your target framework. Neither parser renders JavaScript like a browser. If a page builds its content client-side, you need a browser capture or an API response rather than static-source parsing.
Install AngleSharp and prepare input
Add the package with your normal package manager:
dotnet add package AngleSharp
The following example reads complete files, selects each page’s <main> when present (otherwise the body), clones nodes into one output document, and writes combined.html. It is intentionally explicit so you can change the selection and conflict policy.
using AngleSharp;
using AngleSharp.Dom;
using System.Text;
static async Task<IDocument> ParseFileAsync(IBrowsingContext context, string path)
{
var html = await File.ReadAllTextAsync(path);
return await context.OpenAsync(req => req.Content(html));
}
var config = Configuration.Default;
var context = BrowsingContext.New(config);
var output = await context.OpenAsync(req => req.Content("<!doctype html><html><head><meta charset="utf-8"><title>Combined document</title></head><body></body></html>"));
var sources = new[] { "page1.html", "page2.html", "page3.html" };
foreach (var path in sources)
{
var source = await ParseFileAsync(context, path);
var container = source.QuerySelector("main") ?? source.Body;
if (container is null) continue;
// Clone before insertion so the source DOM remains independent.
foreach (var node in container.ChildNodes.ToArray())
{
output.Body!.AppendChild(node.Clone(true));
}
}
await File.WriteAllTextAsync("combined.html", output.DocumentElement!.OuterHtml, Encoding.UTF8);
Compile this against the AngleSharp version installed in your project and confirm API signatures, because package releases can change overloads. The important operations are parsing each source as a document, selecting content, cloning nodes and appending them to one destination body.
Rank #2
Import only the content you need
Select a stable container
If every page has a semantic <main>, selecting it prevents navigation bars and duplicate footers from being copied. For a known template, use a selector such as .article-content. Check for a missing selector and choose whether to skip the file, fall back to body or fail the build.
Preserve page boundaries
Wrap each imported section so CSS and later processing can identify its origin:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
var section = output.CreateElement("section");
section.ClassList.Add("source-page");
section.SetAttribute("data-source", Path.GetFileName(path));
foreach (var node in container.ChildNodes.ToArray())
section.AppendChild(node.Clone(true));
output.Body!.AppendChild(section);
Use fragment parsing for strings
When inputs are snippets rather than complete documents, parse them as fragments in the context where they will be inserted. A fragment containing table rows, for example, should be parsed for a table body rather than as a standalone document. This follows the standard’s distinction between document and fragment parsing and avoids browser-style reparenting surprises.
Handle head elements, URLs and duplicate IDs
Stylesheets and metadata
Do not blindly copy every source head. Select required <link rel="stylesheet">, inline styles and metadata into the single destination head. Deduplicate identical URLs and decide which title, description and canonical link represent the combined document. Repeated viewport or charset declarations are unnecessary.
Relative URLs and base elements
Moving markup changes the document base used to resolve relative images, links, stylesheets and scripts. A source <base> element cannot represent several origins at once. Convert relative URLs to absolute URLs using each source page’s original URL before insertion, or keep each page’s assets under a known path and rewrite references deliberately.
IDs and fragment links
IDs must be unique in the final document. Prefix imported IDs (for example, page2-) and rewrite matching href="#...", aria-labelledby and for attributes when combining pages. If you do not need cross-section anchors, remove or namespace conflicting IDs instead of leaving broken links.
Scripts and event behavior
Copying a script tag does not safely merge application state. Remove page-specific scripts unless you have audited them for duplicate initialization, global variables and selector collisions. Inline event attributes and assumptions about a page’s original DOM can also fail after composition. A parser manipulates markup; it does not execute the scripts or reproduce browser layout.
Using HtmlAgilityPack instead
HtmlAgilityPack’s document model is suitable when your project already uses it:
var destination = new HtmlAgilityPack.HtmlDocument();
destination.LoadHtml("<!doctype html><html><head><meta charset="utf-8"></head><body></body></html>");
var body = destination.DocumentNode.SelectSingleNode("//body");
foreach (var path in new[] { "page1.html", "page2.html" })
{
var source = new HtmlAgilityPack.HtmlDocument();
source.Load(path);
var container = source.DocumentNode.SelectSingleNode("//main")
?? source.DocumentNode.SelectSingleNode("//body");
if (container == null) continue;
foreach (var child in container.ChildNodes.ToArray())
body.AppendChild(destination.ImportNode(child, true));
}
destination.Save("combined.html");
Check the installed HtmlAgilityPack release for the exact import/clone method available in your target version. The same integration decisions—head ownership, URLs, IDs and scripts—still apply.
Production checklist
- Validate that every input is UTF-8 or decode its declared encoding correctly.
- Set a deterministic source order and record missing or skipped files.
- Choose one title, description, canonical URL and charset.
- Rewrite relative URLs against each source URL.
- Namespace duplicate IDs and update internal references.
- Deduplicate stylesheets and remove unsafe duplicate scripts.
- Inspect the serialized output in the browser or validator used by your application.
- Test pages containing malformed tags, tables, forms, SVG, iframes and custom elements.
Troubleshooting
The output contains nested HTML or several bodies
You appended complete source documents instead of their selected children. Parse each source, select main or body, and append cloned child nodes to the one destination body.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteRank #4
Images or styles are missing
Relative references now resolve from the combined file’s location. Rewrite them to absolute URLs or to paths relative to the output document, and remove conflicting base elements.
Anchors jump to the wrong section
Duplicate IDs are being resolved unpredictably. Prefix IDs per source and rewrite fragment links and ARIA relationships.
Dynamic content is absent
Static HTML does not include content generated after JavaScript runs. Fetch the underlying data endpoint or use a browser-based capture.
Styles leak between pages
Global selectors from one page affect another. Scope imported markup with a wrapper, namespace selectors during a CSS build, or keep only the styles required by the destination template.
Recommended Free Tools
Parser APIs do not compile
Confirm the package version, target framework and namespace imports. AngleSharp and HtmlAgilityPack have different node and cloning APIs; do not mix examples from different libraries without adapting them.
Best Value
Or skip the browser setup
If your real goal is a clean image or PDF of several web pages rather than a merged source file, ScreenshotNeo provides a one-request capture API. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.
For one URL, the cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for options including full-page capture, element selectors, device and retina settings, PDF paper and page ranges, custom CSS or JavaScript, clicks, waits, request blocking, headers, cookies, geolocation, caching, signed links, asynchronous webhooks and bulk capture. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Frequently Asked Questions
Can I merge pages without an external library?
You can use .NET string processing, but safely parsing malformed HTML, selecting fragments and maintaining one document shell is substantially easier with a DOM parser such as AngleSharp or HtmlAgilityPack.
Free tools Windows power users keep installed
One-click scans. No signup required.
Will combining pages preserve their original JavaScript behavior?
No. Scripts may depend on the original URL, global variables and DOM structure. Audit and redesign them for the combined document, or use a browser-rendered workflow for dynamic behavior.
Should I copy every source page’s head element?
No. Keep one destination head and import only the metadata and resources required by the combined output.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

