Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Speed up Pyppeteer on AWS Lambda by measuring three separate phases—Lambda initialization, Chromium startup, and page navigation/readiness—then optimize the slow phase. For navigation, replace Pyppeteer’s default wait for the full load event only when a faster condition, such as a required selector, still guarantees the page state your task needs. There is no reliable universal speed-up figure: results depend on the Lambda configuration, browser build, page, and workload.

First find out which part is slow

A slow invocation does not necessarily mean page.goto() is slow. Lambda has initialization work before the handler runs; Chromium then has to launch; only after that does the browser navigate and wait for the page condition your code requested. AWS describes initialization as code download, runtime startup, and initialization code. It notes package and dependency size, initialization work, and connection setup as contributing factors. See AWS’s Lambda execution environment lifecycle documentation.

Record these intervals separately. Log elapsed time from handler entry through browser launch, page creation, navigation, and the final application-specific readiness check. If your platform provides an initialization duration separately, retain it alongside handler logs. Compare cold and warm invocations over several representative URLs, and record whether the returned data or screenshot is correct as well as how long it took.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import logging
import time

log = logging.getLogger()
log.setLevel(logging.INFO)

def stamp(label, started):
    log.info("%s elapsed_ms=%.1f", label, (time.perf_counter() - started) * 1000)

# In your async Lambda handler, set a timer before each phase:
# t = time.perf_counter()                         # handler entry
# browser = await launch_browser()                 # Chromium launch
# stamp("browser_launch", t)
# page = await browser.newPage()
# t = time.perf_counter()
# response = await page.goto(url, {"waitUntil": "domcontentloaded"})
# stamp("navigation", t)
# t = time.perf_counter()
# await page.waitForSelector("main article", {"timeout": 10_000})
# stamp("readiness", t)

The example shows the measurement boundaries, not a complete Lambda handler: adapt browser launch and cleanup to your installed Pyppeteer and Lambda packaging. Keep timestamps and labels consistent so a change to imports, launch strategy, or navigation condition can be compared fairly. Include errors and timeouts in the sample; looking only at successful, fast pages can hide a change that makes the function less reliable.

Choose a page-ready condition that matches the task

Pyppeteer’s 0.0.25 API reference says goto() defaults to waiting for load. Its navigation options include load, domcontentloaded, networkidle0, and networkidle2. The network-idle conditions require a 500 ms interval with no more than the respective number of active connections. The documented behavior is version-specific; check the API and behavior of the Pyppeteer package you actually deploy. See the Pyppeteer API Reference.

Do not select the condition that merely returns first. Select the earliest signal that establishes the content your job needs. For example, a task that only needs a server-rendered article element might navigate until the DOM is parsed and then wait for that element. A task that depends on images, a client-side render, or data populated by JavaScript may need a different condition or an additional wait. Verify the output rather than assuming that a navigation event means the target data is ready.

Use a selector or function for application readiness

When a known element or JavaScript state marks completion, a targeted waitForSelector() or waitForFunction() can express that requirement more directly than waiting for every network request to settle. A selector only helps if it is a trustworthy signal: a placeholder, skeleton, or empty container can exist before its content is useful. Check the element’s content or state when necessary.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
response = await page.goto(
    url,
    {"waitUntil": "domcontentloaded", "timeout": 30_000},
)
await page.waitForSelector("main article", {"timeout": 10_000})
text = await page.Jeval("main article", "el => el.innerText")
if not text.strip():
    raise RuntimeError("Article element appeared, but contains no text")

This is an illustrative Pyppeteer pattern, not a measured recommendation that domcontentloaded is always best. The API reference documents a 30-second default navigation timeout and allows it to be changed. Increasing the timeout gives a slow navigation more time to finish; it does not make it faster.

Use network idle selectively

Network-idle waits can be useful when the page’s useful content depends on requests completing, but they are not a general synonym for “ready.” Analytics, long polling, streaming, or other continuing requests can prevent an idle condition from occurring. Conversely, a quiet network does not prove that the specific content you need has appeared. Test the condition against the actual target site and prefer a content-specific readiness signal when one is available.

Reduce initialization work and package overhead

AWS identifies initialization code as the largest contributor to latency before function execution in its lifecycle guidance. Keep imports and setup lean: remove dependencies the handler does not use, avoid doing expensive work during module initialization unless every invocation needs it, and defer optional setup to the code path that uses it. A smaller deployment can also reduce code-loading work.

Measure browser binary preparation separately before changing packaging. Whether extraction or setup is costly depends on how the browser is packaged and deployed. Moving work around without timing it can simply shift latency from initialization to the first handler invocation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the browser dependency set compatible

Chromium and the automation library must agree on their protocol and runtime requirements. The chrome-aws-lambda repository provides a Puppeteer-oriented example and runtime guidance, including a recommendation of at least 512 MB and 1600 MB or more for its package use. That is repository guidance for its stated use, not evidence of an optimum for Pyppeteer or proof that its current binary works with every Lambda runtime. Before adopting a package, verify its maintenance status, supported runtime and architecture, Chromium/Pyppeteer compatibility, and deployment-size constraints. Do not copy launch flags from a Node.js/Puppeteer example without validating them for your Python setup.

Reuse warm resources without relying on permanence

Lambda may retain an execution environment and reuse it for a later invocation, but the environment can be frozen and reused only temporarily or terminated. AWS documents that files in /tmp can persist across a freeze and reuse. Treat both the process and temporary files as disposable: startup must work when a cache is missing, and cached data must tolerate staleness.

Where your handler and concurrency model make it safe, consider reusing expensive reusable state rather than rebuilding it for every warm invocation. Keep page contents, cookies, authentication, and other invocation-specific state isolated so one request cannot affect another. Close or reset pages and browser state as appropriate for the library and workload. Pyppeteer’s API documents browser and page operations, but it does not establish one best browser-lifetime strategy for every Lambda deployment. Benchmark the chosen approach under your actual invocation and failure patterns.

Tune memory, timeout, and startup controls from measurements

Browser launch and rendering can consume CPU and memory, while navigation can be dominated by the remote site and network. More memory may change CPU allocation as well as resource headroom, so do not assume that increasing it always speeds up a network-bound page. AWS recommends reviewing the Max Memory Used field, analyzing memory settings, using the open-source AWS Lambda Power Tuning project, and load-testing timeout choices. Compare duration, failure rate, and cost across settings with representative URLs and concurrency.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Memory: check actual peak use and compare measured configurations. A setting that is too low risks memory pressure; a larger setting is not automatically faster or cheaper overall.
  • Timeout: set it to cover realistic browser launch and page behavior, with headroom for variability. A higher timeout prevents premature termination in some cases but does not reduce latency.
  • Concurrency: include simultaneous invocations in tests. Browser resource use and target-site behavior can change under load, so single-request timings alone are not a capacity plan.
  • Cost: evaluate the configured memory and billed duration together rather than optimizing milliseconds in isolation.

For predictable startup, AWS offers Provisioned Concurrency to pre-initialize execution environments. AWS also describes SnapStart as capable of startup performance as low as sub-second in eligible configurations. Those statements concern Lambda startup, not the remote site’s response or page readiness. SnapStart has configuration limitations; AWS’s SnapStart documentation lists constraints including supported-runtime requirements, incompatibility with Provisioned Concurrency, no EFS/S3 Files, and a 512 MB ephemeral-storage ceiling. Check current eligibility and constraints for the exact runtime and deployment before relying on it.

How to tell whether an optimization helped

Use the same package, runtime, architecture, target pages, and output checks before and after a change. Capture cold and warm observations separately; AWS says cold starts typically occur in under 1% of invocations and can range from under 100 ms to over 1 second, but these are general Lambda figures, not predictions for Pyppeteer. A low frequency does not make cold latency irrelevant if your user or downstream job is sensitive to the slowest requests.

Change Phase it may affect What to verify
Reduce dependencies or static initialization Lambda initialization Initialization duration and package behavior on a clean environment
Change browser setup or reuse Chromium launch and warm invocations Cold and warm launch time, state isolation, recovery after termination
Change waitUntil or add a targeted wait Navigation and readiness Elapsed time plus the presence and correctness of required content
Change memory or timeout Resource pressure, runtime and failure behavior Duration, peak memory, errors, concurrency behavior, and cost

No reviewed source supplies a benchmark for the exact combination of Pyppeteer version, Chromium build, Lambda runtime, target page, network, and workload here. Publish numeric claims only when they come from reproducible measurements of that combination.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common slowdowns and failures

  • The handler looks fast, but the invocation is slow: compare Lambda initialization with handler duration. If initialization dominates, inspect package size and module-level setup before changing navigation waits.
  • goto() waits too long: identify the requested waitUntil condition. If the task does not require the full load event, test a narrower navigation event followed by a selector or function that proves the needed content is ready.
  • networkidle0 or networkidle2 never completes: the page may keep requests open. Use a content-specific condition if appropriate, and confirm the output does not depend on unfinished requests.
  • The function times out after navigation: determine whether navigation or the readiness wait consumed the timeout, and log each phase. Reassess the condition and timeout separately; simply raising the timeout can leave an unboundedly slow workflow intact.
  • Warm runs are fast but occasional runs are slow: separate cold from warm measurements and check initialization, browser launch, and page timings. Lambda reuse is temporary, so design and size the workload for cold starts as well.
  • Pages fail after changing the Chromium package or flags: verify runtime, architecture, browser build, Pyppeteer compatibility, and launch configuration together. A Puppeteer-oriented Lambda example does not establish compatibility with Pyppeteer.
  • Results vary under parallel load: load-test at expected concurrency and inspect memory, duration, and target-site response behavior. Avoid sharing mutable page or authentication state across invocations.

Or skip the browser setup

If the task is simply to obtain a screenshot or PDF, a screenshot API can avoid packaging and operating Chromium inside your Lambda function. ScreenshotNeo is a website screenshot API and MCP server; its clean-shot flow accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture, with each step optional. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, this cURL request captures Stripe as WebP. Replace the URL and use your API key; the endpoint and parameters are documented in the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://stripe.com 
  -o shot.webp

ScreenshotNeo also has an MCP server for Claude, Cursor, and other MCP clients, with tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. If those trade-offs fit your workload, sign up for the free plan.

FAQ

Should I always switch from load to domcontentloaded?

No. Switch only if your task can establish readiness through that event plus a suitable check. A page that renders required content later may need another condition.

Does Provisioned Concurrency make a remote page load faster?

No. It pre-initializes Lambda environments to make startup more predictable; it does not control the target website’s response time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is chrome-aws-lambda a drop-in browser package for Pyppeteer?

The cited repository demonstrates a Puppeteer-oriented setup. It does not establish general Pyppeteer compatibility; verify the specific versions, runtime, and architecture you plan to deploy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.