Run Puppeteer reliably on Cloud Run by matching the execution model to the work, setting an explicit deadline, measuring browser memory at realistic concurrency, and cleaning up when requests end. A single Chrome launch flag or memory value is not a reliability strategy: Cloud Run can return a 504 while your container continues running, and an instance is terminated when it exceeds its memory limit.
The reliable Cloud Run pattern
Google documents headless Chrome automation on Cloud Run for scraping, form submissions, UI tests, PDFs and screenshots, with Puppeteer as one of the high-level control libraries. A production design has four parts:
- Choose a service or job. Use a service when a caller needs an HTTP response; use a job when work can run asynchronously and be retried.
- Bound every browser operation. Navigation, selectors and page scripts need their own limits in addition to the platform deadline.
- Control parallelism. Measure peak memory and latency with the pages and assets you actually process.
- Instrument cleanup. Log browser startup, navigation, task completion and cleanup, and close pages even when an operation fails.
There is no official, universal Puppeteer launch-flag set, Chrome/Puppeteer version pairing, or memory-and-concurrency number that works for every workload. Pin the versions used in your image and validate them with representative load tests.
Choose a Cloud Run service or job
Use a service for interactive requests
A Cloud Run service is the natural fit when an API caller expects a screenshot, PDF or extracted result in the HTTP response. Its request timeout defaults to five minutes and can be configured up to 60 minutes. When the deadline expires, Cloud Run closes the connection and returns HTTP 504. The container instance is not necessarily terminated, so Chrome work can continue and consume resources for later requests.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
That behavior makes a long timeout alone insufficient. Set a platform timeout longer than the expected operation where appropriate, but have application code stop work and close resources before the deadline. A client-side timeout does not prove that Chrome stopped.
Use a job for queued or long-running work
Cloud Run Jobs are appropriate when the caller does not need a long-lived HTTP response. The documented default task timeout is 10 minutes; the configurable maximum is 168 hours (7 days). GPU tasks have a one-hour maximum. Retries apply the timeout separately to each task attempt.
These are platform limits, not a promise that an unusually long browser task will succeed. Put a bounded unit of work in each task, persist the result, and make retries safe to repeat. Choose a job when duration, queueing or retry semantics matter more than synchronous response latency.
Build a container with pinned browser dependencies
Pin Node.js and Puppeteer in package.json, then test the exact image you deploy. Puppeteer and the browser it controls must be compatible; official guidance does not establish one canonical pairing.
Example application files
{
"name": "cloud-run-puppeteer",
"private": true,
"type": "module",
"engines": { "node": "20.x" },
"dependencies": {
"express": "^4.19.2",
"puppeteer": "^24.0.0"
}
}
Install dependencies during the image build so the browser is present before a request arrives. The following Dockerfile is a starting point, not a guarantee that every future Puppeteer release uses the same installer behavior:
FROM node:20-bookworm-slim
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY server.js ./
ENV PORT=8080
CMD ["node", "server.js"]
Build and run this image locally, then exercise the same URL and page shape you expect in production. If your selected Puppeteer release requires an explicit browser download or additional Linux packages, follow that release’s installation instructions and keep the resulting image under version control.
Run Puppeteer with bounded work and guaranteed cleanup
This Node.js example creates one browser per request for isolation. That is easy to reason about but slower than reusing a browser. If you later reuse a browser, still create and close a page per request and load-test the shared process.
import express from 'express';
import puppeteer from 'puppeteer';
const app = express();
app.use(express.json());
const port = process.env.PORT || 8080;
const operationLimitMs = Number(process.env.OPERATION_LIMIT_MS || 45000);
function deadline(ms) {
return new Promise((_, reject) =>
setTimeout(() => reject(new Error(`operation exceeded ${ms} ms`)), ms)
);
}
async function runWithLimit(promise, ms) {
return Promise.race([promise, deadline(ms)]);
}
app.post('/capture', async (req, res) => {
const target = req.body?.url;
if (typeof target !== 'string' || !/^https?:///i.test(target)) {
return res.status(400).json({ error: 'body.url must be an http(s) URL' });
}
let browser;
let page;
const started = Date.now();
console.log(JSON.stringify({ stage: 'browser_start', target }));
try {
browser = await puppeteer.launch({
headless: true,
// Verify these flags against your image and security policy.
args: ['--no-sandbox', '--disable-setuid-sandbox']
});
page = await browser.newPage();
page.setDefaultNavigationTimeout(operationLimitMs);
page.setDefaultTimeout(operationLimitMs);
console.log(JSON.stringify({ stage: 'navigation_start', target }));
await runWithLimit(page.goto(target, { waitUntil: 'networkidle2' }), operationLimitMs);
const title = await runWithLimit(page.title(), operationLimitMs);
const image = await runWithLimit(page.screenshot({ type: 'png' }), operationLimitMs);
console.log(JSON.stringify({ stage: 'task_complete', ms: Date.now() - started }));
res.type('png').send(image);
} catch (error) {
console.error(JSON.stringify({ stage: 'task_error', message: error.message, ms: Date.now() - started }));
if (!res.headersSent) res.status(504).json({ error: error.message });
} finally {
console.log(JSON.stringify({ stage: 'cleanup_start' }));
if (page) await page.close().catch(() => {});
if (browser) await browser.close().catch(() => {});
console.log(JSON.stringify({ stage: 'cleanup_complete' }));
}
});
app.listen(port, () => console.log(`listening on ${port}`));
The timeout wrapper prevents a stuck navigation or script from waiting forever, while the finally block closes resources after success, failure or an application-level deadline. The launch arguments shown are common container settings, not a universal prescription; test them with the Chrome build in your image and your security requirements.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsDeploy with an explicit request deadline
Build and deploy the service, setting memory, CPU and timeout deliberately rather than inheriting an accidental default:
gcloud builds submit --tag REGION-docker.pkg.dev/PROJECT_ID/REPO/puppeteer:1
gcloud run deploy puppeteer
--image REGION-docker.pkg.dev/PROJECT_ID/REPO/puppeteer:1
--region REGION
--memory 2Gi
--cpu 2
--timeout 300
--concurrency 1
Replace the placeholders with your project, repository and region. The five-minute timeout in this example matches the documented service default; choose another value only after comparing it with measured browser duration. Keep the application limit below the platform deadline so cleanup can run before Cloud Run closes the request.
Tune memory and concurrency from measurements
Cloud Run terminates an instance that exceeds its configured memory. A useful sizing model is standing process memory plus memory consumed by each simultaneous request multiplied by service concurrency. Chromium, pages, JavaScript heaps, images and downloaded assets can make the per-request portion substantial.
Start conservatively
- Begin with low concurrency, often one for a browser-heavy endpoint.
- Run a load test containing the largest pages, slowest redirects, PDFs or screenshots you will serve.
- Record request latency, peak memory, browser crashes, HTTP 5xx responses and instance restarts.
- Increase concurrency in small steps only while peak memory and failure rates remain stable.
- When raising concurrency, revisit memory and CPU together; more parallel pages increase per-instance demand.
Cloud Run permits configuring maximum concurrent requests per instance. The console default is 80; when a service is first created through CLI or Terraform, the default is 80 times the vCPU count. Those are platform defaults, not recommended Puppeteer settings. If your application cannot process requests in parallel, lower the value.
Free tools Windows power users keep installed
One-click scans. No signup required.
What to measure
- Browser startup time: separates image or process overhead from page navigation.
- Navigation and rendering time: identifies slow sites, redirects and heavy assets.
- Peak resident memory: capture the maximum during simultaneous requests, not only idle usage.
- Cleanup duration: a delayed close can overlap with the next request after a timeout.
- Failure rate by concurrency: compare throughput with 5xx responses and browser errors, not throughput alone.
Handle timeouts and diagnose failures
HTTP 504 after a long navigation
Check the Cloud Run request log and your stage logs. If the request exceeded the service deadline, the caller can receive 504 while Chrome continues. Shorten the page task, split it into smaller units, move it to a job, or increase the service timeout within its documented limit. Also stop application work before the deadline and close the page and browser.
Instance terminated for memory
Correlate system or instance logs with request concurrency. Reduce concurrency, process fewer pages simultaneously, block unnecessary resources in your own code, or increase the memory limit. Repeat the representative load test after each change; an idle smoke test will not reveal production peaks.
Intermittent navigation failures
Log the target, redirect path, navigation duration and error type without recording secrets. Distinguish a target site’s bot check or timeout from a Chrome crash. Set navigation and selector timeouts explicitly, and make retries bounded and idempotent. A retry that launches another browser while the first continues after a 504 can multiply memory pressure.
Requests fail only under load
Compare the same revision at different concurrency values. A stable single-request result does not establish safe parallelism. Lower concurrency first, then inspect CPU and memory pressure before changing browser code.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Production checklist
- Pin Node.js, Puppeteer and the browser image; rebuild and test the exact artifact deployed.
- Choose a service for synchronous responses and a job for queued, retryable work.
- Set a Cloud Run timeout that covers measured work, with an application deadline below it.
- Close every page and browser in a
finallypath. - Start with low concurrency and load-test realistic pages before increasing it.
- Alert on 504s, memory terminations, browser startup failures and cleanup errors.
- Make job output durable and retries safe to repeat.
Or skip the browser setup
If your goal is a dependable website screenshot rather than operating Chromium yourself, ScreenshotNeo provides a single GET request and supports PNG, JPEG, WebP or PDF output. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers identify the page verdict and billing status.
It also offers an MCP server for Claude, Cursor and other MCP clients, with take_screenshot, get_page_info and capture_pdf tools. Every plan includes the features, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, retina scale, PDF controls, custom CSS and JavaScript, pre-capture clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous webhooks, bulk capture for 100 URLs per call, a usage API and an OpenAPI specification.
Use the ScreenshotNeo API documentation for authentication and options. The following calls use the supplied API format:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free. Sign up for the free ScreenshotNeo plan and test a capture without configuring a browser container.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
FAQ
Can I rely on Cloud Run’s default concurrency for Puppeteer?
No. The platform default is a starting setting, not evidence that your pages are safe at that parallelism. Measure the workload and tune until latency, memory and failures remain stable.
Best Value
Does a Cloud Run 504 kill Chromium automatically?
Not necessarily. The connection can close while the container continues processing, so your application must enforce its own deadline and clean up browser resources.
When should a screenshot endpoint become a Cloud Run Job?
Move to a job when the caller can accept asynchronous completion and the task benefits from queueing, task-level timeouts or retries instead of a single HTTP response deadline.
Frequently Asked Questions
Can I rely on Cloud Run’s default concurrency for Puppeteer?
No. The platform default is a starting setting, not evidence that your pages are safe at that parallelism. Measure the workload and tune until latency, memory and failures remain stable.
Does a Cloud Run 504 kill Chromium automatically?
Not necessarily. The connection can close while the container continues processing, so your application must enforce its own deadline and clean up browser resources.
When should a screenshot endpoint become a Cloud Run Job?
Move to a job when the caller can accept asynchronous completion and the task benefits from queueing, task-level timeouts or retries instead of a single HTTP response deadline.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




