October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Chrome DevTools Protocol

How to Save a Webpage as MHT with Puppeteer (MHTML)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Puppeteer’s Chrome DevTools Protocol (CDP) session and the Page.captureSnapshot command. Set format to mhtml, then write the returned string to a .mhtml (or .mht) file. This captures a serialized page package rather than just the current HTML or a PDF.

The complete Puppeteer method

Puppeteer does not expose an MHTML method directly on Page. Instead, create a CDP session attached to the page and call Chrome’s Page.captureSnapshot protocol method. The method returns serialized page data as a string, which your Node.js process must save.

import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();

  await page.goto('https://example.com', {
    waitUntil: 'networkidle2'
  });

  const cdp = await page.createCDPSession();
  const { data } = await cdp.send('Page.captureSnapshot', {
    format: 'mhtml'
  });

  await writeFile('page.mhtml', data, 'utf8');
} finally {
  await browser.close();
}

Run this as an ES module (for example, save it as save-mhtml.mjs) after installing Puppeteer:

npm install puppeteer
node save-mhtml.mjs

The resulting page.mhtml contains the snapshot data. Rename it to .mht only if your target application specifically expects that extension; the protocol option remains mhtml.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The relevant APIs are documented in the Puppeteer Page API and the Chrome DevTools Protocol Page reference.

What each step does

1. Launch a Chromium browser

puppeteer.launch() starts the browser bundled with, or configured for, your Puppeteer installation. In CI or a server container you may need the system dependencies required by that Chromium build.

2. Navigate and wait for readiness

page.goto() loads the URL. networkidle2 waits until there are no more than two active network connections for a short period. It is a useful baseline, not a guarantee that an application has finished rendering. Single-page apps, analytics streams, advertisements and chat connections can keep requests active or render content later.

Use a readiness condition that matches the site:

await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('main article');

For a known client-rendered component, wait for its selector. For a fixed animation or delayed request, use a deliberate delay sparingly:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
await new Promise(resolve => setTimeout(resolve, 1500));

3. Attach CDP

page.createCDPSession() creates a DevTools Protocol session associated with that page. It gives your script access to protocol commands that are not represented by a high-level Puppeteer method.

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

4. Capture the MHTML snapshot

Page.captureSnapshot accepts { format: 'mhtml' }. The response’s data property is the serialized page package. The protocol documentation says that MHTML serialization includes iframes, shadow DOM, external resources and element-inline styles.

5. Write the returned string

Use writeFile with UTF-8 text. Do not convert the response to JSON or a Buffer containing unrelated metadata. Keep the browser open until the file write completes, then close it in finally so failures do not leave Chromium processes running.

MHTML versus HTML and PDF in Puppeteer

Output API What it is for
MHTML page.createCDPSession() plus Page.captureSnapshot A serialized page package containing document data and captured resources supported by the protocol.
HTML page.content() The current HTML contents of the page. It is not an MHTML archive.
PDF page.pdf() A printable document. It is not a resource-preserving MHTML package.

Puppeteer’s page.content() and page.pdf() are separate APIs; choosing either does not produce an MHT file.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Making captures reliable

Wait for the page’s real completion signal

Prefer a selector that represents useful content, such as an article heading or data table. If the page paginates or lazy-loads, scroll before capture so content that appears only near the viewport has a chance to load:

await page.evaluate(async () => {
  await new Promise(resolve => {
    let y = 0;
    const step = 600;
    const timer = setInterval(() => {
      window.scrollBy(0, step);
      y += step;
      if (y >= document.body.scrollHeight) {
        clearInterval(timer);
        resolve();
      }
    }, 100);
  });
});
await new Promise(resolve => setTimeout(resolve, 500));

This is site-dependent: some pages virtualize old content, and some never settle because of live feeds.

Set a viewport and locale when they matter

await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
await page.setExtraHTTPHeaders({ 'Accept-Language': 'en-US,en;q=0.9' });

The resulting MHTML reflects the browser state you created, including responsive layout and localized responses.

Handle authentication and consent

Log in before navigation or set cookies on the page, then wait for the authenticated UI. Cookie banners and one-time dialogs can become part of the captured state unless your script dismisses them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a deterministic filename

const safeName = new URL(page.url()).hostname.replace(/[^a-z0-9.-]/gi, '_');
await writeFile(`${safeName}-${Date.now()}.mhtml`, data, 'utf8');

Limitations of an MHTML snapshot

MHTML is a snapshot, not a replay of every behavior in a modern web application. The protocol describes what it serializes; it does not promise a perfect offline reconstruction of every dynamic state, service-worker behavior, streaming response or post-capture interaction. Verify the file in the reader you intend to use.

Chrome’s protocol reference currently labels Page.captureSnapshot experimental and uses a moving “tot” document. Match Puppeteer to a compatible Chromium version and check the protocol reference for that version if a deployment depends on this command.

Some viewers impose their own restrictions. Chrome’s extension documentation notes that MHTML files can be loaded only from the file system and only in the main frame when using its page-capture workflow. That is an extension-specific loading rule, but it is a reminder to test the consuming application rather than assuming every viewer handles every package identically.

Troubleshooting

“Cannot find module puppeteer”

Install the dependency in the project that runs the script: npm install puppeteer. If you use a package manager workspace, run the command from the workspace containing the script.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

Chromium fails to launch

On Linux containers, install the libraries required by the Puppeteer Chromium build or configure executablePath to a compatible installed browser. Avoid adding --no-sandbox unless your deployment’s security model specifically requires it and you understand the risk.

The file is empty or truncated

Make sure you await both cdp.send() and writeFile(), and do not close the browser before the write resolves. Log data.length and confirm that the process has permission to write the destination directory.

Content is missing

The capture may have happened before client-side rendering or lazy loading completed. Wait for a meaningful selector, perform the required scroll, dismiss overlays, and capture again. A page that requires login must be authenticated in the same browser context.

The page never reaches network idle

Long-lived sockets and analytics can prevent an idle condition. Use waitUntil: 'domcontentloaded' followed by waitForSelector(), or use a bounded delay after the application’s own ready marker.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The MHT opens differently from the live site

That is expected for a static snapshot. Test with the intended viewer, and remember that scripts, live data, cross-origin behavior and browser security policies can differ offline.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, storage and operational notes

  • Capture cost is dominated by navigation, rendering and resource loading, not the final file write.
  • Use one browser process with separate pages for batches when isolation permits; repeatedly launching Chromium adds startup overhead.
  • Close pages and browsers in error paths to prevent memory leaks.
  • Set navigation and selector timeouts appropriate to your environment, and record the URL, timestamp and Puppeteer/Chromium versions alongside the file for reproducibility.
  • Expect larger files for media-heavy pages. If storage or transfer is constrained, decide whether MHTML is actually required instead of silently substituting HTML or PDF.

Chrome extension alternative

In a Chrome extension, the separate chrome.pageCapture.saveAsMHTML() API saves a tab as MHTML and resolves with a Blob or undefined. It requires the extension permissions and tab context documented by Chrome. This is not the Puppeteer/CDP method and is useful only when the capture runs inside an extension.

Or skip the browser setup

If you need an image or PDF rather than an MHTML archive, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL in one request and can return PNG, JPEG, WebP or PDF. Cookie and consent banners, newsletter popups and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, with the result identified by X-Page-Verdict and X-Billed headers. Its MCP tools—take_screenshot, get_page_info and capture_pdf—work with Claude, Cursor and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all options, including full-page and element capture, device presets, custom CSS/JavaScript, waiting rules, headers and cookies, PDF settings, caching, bulk jobs and signed webhooks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Should I use .mht or .mhtml?

Use .mhtml in your code and example because it names the documented format. Choose .mht only when a consuming tool requires that extension.

Does MHTML save a screenshot?

No. It saves a serialized web page package. Use PDF or an image API when the deliverable is a visual print or raster image.

Can I call Page.captureSnapshot through page.evaluate()?

No. It is a browser DevTools Protocol command, so send it through the CDP session created with page.createCDPSession().

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can MHTML preserve iframes and shadow DOM?

The Chrome DevTools Protocol documentation says its MHTML serialization includes iframes, shadow DOM, external resources and element-inline styles; verify the result in your target viewer.

Is the extension API suitable for a Node.js Puppeteer script?

No. chrome.pageCapture.saveAsMHTML() is a Chrome extension API. A Puppeteer script should use the CDP Page.captureSnapshot command.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.