Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Do not concatenate complete HTML files as plain text. Parse each page, create one destination document, and append the selected nodes from every source body in order. This avoids repeated <html>, <head> and <body> elements and gives you explicit control over styles, scripts, links, IDs and metadata.

In .NET, AngleSharp is a strong default when standards-oriented HTML5 parsing and fragment insertion matter. HtmlAgilityPack is also practical for applications already using its node API. The implementation below uses AngleSharp and shows the decisions you must make when combining real pages.

What “combine” should mean

A valid HTML document has one document element, normally one head and one body. Joining two strings such as File.ReadAllText("a.html") + File.ReadAllText("b.html") creates multiple document shells, duplicate metadata and often invalid resource behavior. Instead, define a destination shell and decide what to import from each source:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Body content: usually the source <main>, an article element, or selected children of <body>.
  • Head resources: stylesheets, fonts, metadata and scripts that the combined page actually needs.
  • Order: the sequence in which source pages appear in the output.
  • Conflict policy: how to handle duplicate IDs, repeated styles, scripts, titles and base URLs.

The WHATWG HTML Standard defines separate document and fragment parsing algorithms. Use a document parser for complete pages and a fragment-oriented operation when inserting markup into a particular destination element.

Choose a parser

AngleSharp

AngleSharp builds a standards-oriented DOM, supports CSS-style querying and provides document and fragment parsing capabilities. It is a good fit for HTML5 cleanup and predictable insertion contexts. Its project examples and fragment guidance are available in the examples and fragment questions documentation.

HtmlAgilityPack

HtmlAgilityPack loads HTML from files or strings and exposes node manipulation APIs documented at its manipulation guide. The NuGet listing showed version 1.13.0 at the time of research; verify the current package version and target-framework compatibility before installing.

Choose based on malformed-markup behavior, fragment support, API familiarity and your target framework. Neither parser renders JavaScript like a browser. If a page builds its content client-side, you need a browser capture or an API response rather than static-source parsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install AngleSharp and prepare input

Add the package with your normal package manager:

dotnet add package AngleSharp

The following example reads complete files, selects each page’s <main> when present (otherwise the body), clones nodes into one output document, and writes combined.html. It is intentionally explicit so you can change the selection and conflict policy.

using AngleSharp;
using AngleSharp.Dom;
using System.Text;

static async Task<IDocument> ParseFileAsync(IBrowsingContext context, string path)
{
    var html = await File.ReadAllTextAsync(path);
    return await context.OpenAsync(req => req.Content(html));
}

var config = Configuration.Default;
var context = BrowsingContext.New(config);

var output = await context.OpenAsync(req => req.Content("<!doctype html><html><head><meta charset="utf-8"><title>Combined document</title></head><body></body></html>"));
var sources = new[] { "page1.html", "page2.html", "page3.html" };

foreach (var path in sources)
{
    var source = await ParseFileAsync(context, path);
    var container = source.QuerySelector("main") ?? source.Body;
    if (container is null) continue;

    // Clone before insertion so the source DOM remains independent.
    foreach (var node in container.ChildNodes.ToArray())
    {
        output.Body!.AppendChild(node.Clone(true));
    }
}

await File.WriteAllTextAsync("combined.html", output.DocumentElement!.OuterHtml, Encoding.UTF8);

Compile this against the AngleSharp version installed in your project and confirm API signatures, because package releases can change overloads. The important operations are parsing each source as a document, selecting content, cloning nodes and appending them to one destination body.

Import only the content you need

Select a stable container

If every page has a semantic <main>, selecting it prevents navigation bars and duplicate footers from being copied. For a known template, use a selector such as .article-content. Check for a missing selector and choose whether to skip the file, fall back to body or fail the build.

Preserve page boundaries

Wrap each imported section so CSS and later processing can identify its origin:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
var section = output.CreateElement("section");
section.ClassList.Add("source-page");
section.SetAttribute("data-source", Path.GetFileName(path));
foreach (var node in container.ChildNodes.ToArray())
    section.AppendChild(node.Clone(true));
output.Body!.AppendChild(section);

Use fragment parsing for strings

When inputs are snippets rather than complete documents, parse them as fragments in the context where they will be inserted. A fragment containing table rows, for example, should be parsed for a table body rather than as a standalone document. This follows the standard’s distinction between document and fragment parsing and avoids browser-style reparenting surprises.

Handle head elements, URLs and duplicate IDs

Stylesheets and metadata

Do not blindly copy every source head. Select required <link rel="stylesheet">, inline styles and metadata into the single destination head. Deduplicate identical URLs and decide which title, description and canonical link represent the combined document. Repeated viewport or charset declarations are unnecessary.

Relative URLs and base elements

Moving markup changes the document base used to resolve relative images, links, stylesheets and scripts. A source <base> element cannot represent several origins at once. Convert relative URLs to absolute URLs using each source page’s original URL before insertion, or keep each page’s assets under a known path and rewrite references deliberately.

IDs and fragment links

IDs must be unique in the final document. Prefix imported IDs (for example, page2-) and rewrite matching href="#...", aria-labelledby and for attributes when combining pages. If you do not need cross-section anchors, remove or namespace conflicting IDs instead of leaving broken links.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scripts and event behavior

Copying a script tag does not safely merge application state. Remove page-specific scripts unless you have audited them for duplicate initialization, global variables and selector collisions. Inline event attributes and assumptions about a page’s original DOM can also fail after composition. A parser manipulates markup; it does not execute the scripts or reproduce browser layout.

Using HtmlAgilityPack instead

HtmlAgilityPack’s document model is suitable when your project already uses it:

var destination = new HtmlAgilityPack.HtmlDocument();
destination.LoadHtml("<!doctype html><html><head><meta charset="utf-8"></head><body></body></html>");
var body = destination.DocumentNode.SelectSingleNode("//body");

foreach (var path in new[] { "page1.html", "page2.html" })
{
    var source = new HtmlAgilityPack.HtmlDocument();
    source.Load(path);
    var container = source.DocumentNode.SelectSingleNode("//main")
                  ?? source.DocumentNode.SelectSingleNode("//body");
    if (container == null) continue;

    foreach (var child in container.ChildNodes.ToArray())
        body.AppendChild(destination.ImportNode(child, true));
}

destination.Save("combined.html");

Check the installed HtmlAgilityPack release for the exact import/clone method available in your target version. The same integration decisions—head ownership, URLs, IDs and scripts—still apply.

Production checklist

  • Validate that every input is UTF-8 or decode its declared encoding correctly.
  • Set a deterministic source order and record missing or skipped files.
  • Choose one title, description, canonical URL and charset.
  • Rewrite relative URLs against each source URL.
  • Namespace duplicate IDs and update internal references.
  • Deduplicate stylesheets and remove unsafe duplicate scripts.
  • Inspect the serialized output in the browser or validator used by your application.
  • Test pages containing malformed tags, tables, forms, SVG, iframes and custom elements.

Troubleshooting

The output contains nested HTML or several bodies

You appended complete source documents instead of their selected children. Parse each source, select main or body, and append cloned child nodes to the one destination body.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Images or styles are missing

Relative references now resolve from the combined file’s location. Rewrite them to absolute URLs or to paths relative to the output document, and remove conflicting base elements.

Anchors jump to the wrong section

Duplicate IDs are being resolved unpredictably. Prefix IDs per source and rewrite fragment links and ARIA relationships.

Dynamic content is absent

Static HTML does not include content generated after JavaScript runs. Fetch the underlying data endpoint or use a browser-based capture.

Styles leak between pages

Global selectors from one page affect another. Scope imported markup with a wrapper, namespace selectors during a CSS build, or keep only the styles required by the destination template.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parser APIs do not compile

Confirm the package version, target framework and namespace imports. AngleSharp and HtmlAgilityPack have different node and cloning APIs; do not mix examples from different libraries without adapting them.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your real goal is a clean image or PDF of several web pages rather than a merged source file, ScreenshotNeo provides a one-request capture API. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

For one URL, the cURL request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for options including full-page capture, element selectors, device and retina settings, PDF paper and page ranges, custom CSS or JavaScript, clicks, waits, request blocking, headers, cookies, geolocation, caching, signed links, asynchronous webhooks and bulk capture. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can I merge pages without an external library?

You can use .NET string processing, but safely parsing malformed HTML, selecting fragments and maintaining one document shell is substantially easier with a DOM parser such as AngleSharp or HtmlAgilityPack.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Will combining pages preserve their original JavaScript behavior?

No. Scripts may depend on the original URL, global variables and DOM structure. Audit and redesign them for the combined document, or use a browser-rendered workflow for dynamic behavior.

Should I copy every source page’s head element?

No. Keep one destination head and import only the metadata and resources required by the combined output.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.