Yes—you can automate a daily newsletter with Playwright. A reliable pipeline schedules a run in a defined timezone, opens an isolated browser context, collects only allowlisted pages with bounded waits, saves raw evidence, normalizes and deduplicates stories, applies editorial rules, renders HTML and plain text, validates compliance, sends through an email provider, and records delivery events. Browser automation should gather and verify material; it should not silently publish ambiguous or low-quality items.
What the finished system should do
Think of the newsletter as a repeatable data pipeline rather than a single scraping script. Each run receives a unique run ID, produces an auditable set of source records, and either sends a validated issue or stops safely for review.
- Schedule: start at a fixed timezone and record the scheduled and actual start times.
- Collect: visit an allowlist of source URLs in a fresh Playwright context. Use semantic locators, per-source timeouts, and bounded waits.
- Preserve: save raw HTML, canonical URL, extraction timestamp, and page metadata before changing text.
- Normalize: standardize URLs, titles, publication timestamps, tags, and author fields.
- Deduplicate: collapse repeated links and near-identical titles while retaining the original URL beside every claim.
- Edit: enforce source-quality, recency, topic, and minimum-content rules. Route uncertain records to a human-review queue.
- Render: produce responsive HTML and a plain-text alternative, including a generated sources section.
- Validate and send: check every link, title, image alt attribute, unsubscribe mechanism, sender identity, physical address, and dry-run recipient before delivery.
- Monitor: retain send, delivery, bounce, complaint, and unsubscribe events.
Choose the browser and isolate every run
Playwright’s BrowserType API can launch Chromium, Firefox, or WebKit, or connect to an existing browser server. Microsoft also documents Playwright as a cross-browser API that can automate Edge. For a scheduled newsletter, launch a fresh, headless context for each run so cookies, local storage, and authentication state from one source cannot leak into another.
Install a minimal Node.js project
mkdir daily-newsletter
cd daily-newsletter
npm init -y
npm install playwright
npx playwright install chromium
Keep source definitions in configuration rather than embedding them throughout the crawler. An allowlist makes accidental collection from unapproved domains less likely.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11#1 Best Overall
const SOURCES = [
{
name: 'Example news',
url: 'https://news.example.com/latest',
article: 'article',
title: 'h2 a',
link: 'h2 a',
published: 'time'
}
];
Collect pages with bounded waits and raw snapshots
Do not use an unbounded networkidle wait on pages that maintain analytics or streaming connections. Prefer a short document-ready wait, then wait for the specific selector you need. Give each source its own timeout and continue to the next source when one fails.
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import crypto from 'node:crypto';
const runId = `${new Date().toISOString()}-${crypto.randomUUID()}`;
const timeoutMs = 20000;
function clean(value) {
return (value || '').replace(/\s+/g, ' ').trim();
}
async function collectSource(browser, source) {
const context = await browser.newContext({
userAgent: 'DailyNewsletterBot/1.0 (+contact [email protected])'
});
const page = await context.newPage();
page.setDefaultTimeout(timeoutMs);
try {
const response = await page.goto(source.url, {
waitUntil: 'domcontentloaded',
timeout: timeoutMs
});
await page.locator(source.article).first().waitFor({ state: 'visible', timeout: timeoutMs });
const rawHtml = await page.content();
const records = await page.locator(source.article).evaluateAll((nodes, selectors) =>
nodes.map(node => {
const title = node.querySelector(selectors.title)?.textContent || '';
const href = node.querySelector(selectors.link)?.href || '';
const published = node.querySelector(selectors.published)?.getAttribute('datetime') ||
node.querySelector(selectors.published)?.textContent || '';
return { title, href, published };
}),
{ title: source.title, link: source.link, published: source.published }
);
await mkdir(`snapshots/${runId}`, { recursive: true });
await writeFile(`snapshots/${runId}/${source.name}.html`, rawHtml);
return {
source: source.name,
requestedUrl: source.url,
finalUrl: page.url(),
status: response?.status() ?? null,
fetchedAt: new Date().toISOString(),
records: records.map(item => ({
...item,
title: clean(item.title),
originalUrl: item.href,
published: clean(item.published)
}))
};
} finally {
await context.close();
}
}
const browser = await chromium.launch({ headless: true });
const results = [];
for (const source of SOURCES) {
try {
results.push(await collectSource(browser, source));
} catch (error) {
results.push({ source: source.name, error: String(error), records: [] });
}
}
await browser.close();
console.log(JSON.stringify({ runId, results }, null, 2));
The example assumes each source exposes an article container, a title link, and a time element. Inspect each site and adjust selectors; do not rely on a single CSS pattern across unrelated publishers. If a page is client-rendered, wait for the article selector or a source-specific readiness marker, not an arbitrary long sleep.
Normalize, deduplicate, and apply editorial rules
Canonicalize links
Resolve relative links against the final page URL, remove known tracking parameters, and retain the resulting canonical URL alongside the requested URL. If a page declares a canonical link, prefer it after checking that it remains on an allowed domain. Never discard the original URL needed for attribution or auditing.
Use a deterministic duplicate key
First deduplicate by canonical URL. For syndicated copies with different URLs, compare a normalized title and publication date. Keep the earliest record, merge source names, and preserve every source URL. A simple key can be sha256(canonicalUrl + normalizedTitle); store the title and URL in plain form as well so an editor can inspect collisions.
Enforce a recency and quality window
Convert publication timestamps to UTC, reject records outside the issue’s window, and define what happens when a source omits a date. Require a minimum title and body length, an allowed topic tag, and a source on your allowlist. Mark uncertain dates, paywalled summaries, or conflicting metadata for human review instead of silently guessing.
Rank #2
function normalizeUrl(raw, base) {
const url = new URL(raw, base);
for (const key of ['utm_source', 'utm_medium', 'utm_campaign', 'gclid']) {
url.searchParams.delete(key);
}
return url.toString();
}
function titleKey(title) {
return title.toLowerCase()
.replace(/[^a-z0-9 ]/g, '')
.replace(/\s+/g, ' ')
.trim();
}
function dedupe(items) {
const seenUrl = new Map();
const seenTitle = new Map();
const output = [];
for (const item of items) {
const url = normalizeUrl(item.originalUrl, item.requestedUrl);
const tKey = titleKey(item.title);
if (seenUrl.has(url) || seenTitle.has(tKey)) continue;
const normalized = { ...item, canonicalUrl: url };
seenUrl.set(url, normalized);
seenTitle.set(tKey, normalized);
output.push(normalized);
}
return output;
}
Schedule a predictable morning run
Choose one timezone for the publication promise and document daylight-saving behavior. A server’s local timezone is not a reliable substitute. On a Unix host, cron can start a wrapper script at 07:00 UTC:
0 7 * * * cd /opt/daily-newsletter && /usr/bin/node run.mjs >> logs/cron.log 2>&1
The wrapper should create the run ID before launching the browser, acquire a lock so two runs cannot overlap, and write a final status such as collected, awaiting_review, sent, or failed. If you use a hosted scheduler, configure the same timezone, retry policy, secret storage, and timeout explicitly. Retries must be idempotent: a repeated run should not send duplicate mail.
Render one issue in HTML and plain text
Build the issue from normalized records, not from raw page fragments. Escape title and summary text before inserting it into HTML. Generate a source list from the records so every item can be traced back to its original page.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
function escapeHtml(value) {
return value.replace(/[&<>"']/g, ch => ({ '&': '&', '<': '<', '>': '>', '"': '"', "'": ''' }[ch]));
}
function renderIssue(items, issueDate) {
const htmlItems = items.map(item => `
<article>
<h2><a href="${escapeHtml(item.canonicalUrl)}">${escapeHtml(item.title)}</a></h2>
<p>${escapeHtml(item.summary || '')}</p>
<p>Source: <a href="${escapeHtml(item.canonicalUrl)}">${escapeHtml(item.source)}</a></p>
</article>`).join('');
const text = items.map(item =>
`${item.title}\n${item.summary || ''}\n${item.canonicalUrl}\n`).join('\n');
return {
html: `<!doctype html><html><body><main><h1>Daily briefing — ${escapeHtml(issueDate)}</h1>${htmlItems}</main><p>You are receiving this because you subscribed. <a href="{{unsubscribe_url}}">Unsubscribe</a></p></body></html>`,
text: `Daily briefing — ${issueDate}\n\n${text}\nUnsubscribe: {{unsubscribe_url}}`
};
}
Keep the unsubscribe URL functional in both formats. Add descriptive alt text whenever you include an image, and avoid depending on remote images to convey essential information.
Validate before sending
- Resolve every link and reject missing, malformed, or disallowed destinations.
- Require a non-empty title, source, publication date (or an explicit “date unavailable” review state), and summary.
- Check that the HTML and plain-text versions contain the same essential links.
- Verify the From identity, reply address, physical postal address, and unsubscribe path.
- Send first to a dry-run list and inspect on desktop and mobile clients.
- Abort when the issue is empty, the source count is unexpectedly low, or a review queue is non-empty.
Commercial email must use truthful routing information and a non-deceptive subject, include a valid physical postal address, and provide a clear opt-out path. FTC guidance says opt-outs must be honored within 10 business days, and the opt-out mechanism must remain usable for at least 30 days after the message is sent. Treat these as operational requirements: store the request, suppress the address before the next send, and retain an audit event.
Rank #3
If an item contains a paid recommendation or affiliate relationship, disclose that relationship clearly and conspicuously near the recommendation. The phrase “affiliate link” by itself may not explain the relationship to readers.
Sending, retries, and observability
Use an email provider with an API or SMTP service, but keep provider-specific code behind a small sendIssue() function. Pass an idempotency key based on the run ID and issue date when the provider supports one. Retry transient network errors with exponential backoff; do not retry authentication failures or compliance rejections. After submission, record the provider message ID and subscribe to delivery, bounce, complaint, and unsubscribe events. Alert on repeated source failures, sudden item-count changes, elevated bounces, or complaints.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsControl concurrency and resource use
Run a small number of pages concurrently rather than opening every source at once. Limit navigation and selector waits, close each context in a finally block, and cap raw snapshot retention. Block nonessential resource types only when doing so does not remove the content you need; an image-heavy source may require a different policy than a text-only source.
Common failures and fixes
Selector timeout
Cause: the selector changed, content is client-rendered, or a consent dialog covers the page. Fix: inspect the current DOM, prefer semantic roles or stable attributes, wait for a source-specific readiness selector, and record a screenshot or HTML snapshot for diagnosis.
Empty or partial results
Cause: pagination, lazy loading, an infinite scroll, or a geofenced response. Fix: implement a bounded “load more” loop, wait for newly added nodes, set the intended locale and timezone, and fail the source when the expected minimum count is not met.
Rank #4
Bot check or CAPTCHA
Cause: the site requires an interactive challenge or disallows automated access. Fix: respect the site’s terms and robots policies, remove that source or obtain permission, and route the run to review. Do not attempt to defeat a CAPTCHA.
Duplicate sends after a retry
Cause: the scheduler retried after the provider accepted the message. Fix: persist the run and provider message IDs before acknowledging success, use provider idempotency where available, and check send state before retrying.
Broken links in the delivered issue
Cause: relative URLs, tracking redirects, or links that expired between collection and send. Fix: resolve against the final page URL, validate during the pre-send step, and retain the original URL for editors when a canonical redirect is unavailable.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
When you need a clean image of a source page for an issue, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the outcome with X-Page-Verdict and X-Billed headers.
One GET request returns PNG, JPEG, WebP, or a PDF. See the ScreenshotNeo API documentation for parameters.
Recommended Free Tools
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
For newsletter artwork or evidence captures, relevant options include full-page shots with lazy images loaded, a CSS-selected element, dark mode, 12 device presets or any viewport, retina scale, PDF paper size and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, blocked ads and trackers, custom headers, cookies, user agent and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients, so an AI agent can gather page images without your own browser orchestration.
Best Value
| Plan | Price | Included shots |
|---|---|---|
| Free | $0 | 1,000 per month; no card |
| Starter | $5 | 3,000 |
| Growth | $15 | 15,000 |
| Pro | $39 | 60,000 |
| Scale | $99 | 250,000 |
| Business | $249 | 1,000,000 |
Every feature is available on every plan, and yearly billing provides two months free. Start with 1,000 free screenshots a month—no card required.
Operating-cost and reliability decisions
- Browser runtime: cost grows with navigation time, concurrency, and the number of contexts. Cache stable source pages only when freshness rules allow it.
- Email delivery: budget for provider volume, webhook retention, and list hygiene rather than only the send API.
- Review labor: a human queue is an intentional control for ambiguous dates, duplicates, and policy-sensitive items.
- Failure isolation: one unavailable source should produce a visible warning, not an empty or misleading issue.
- Reproducibility: retain run ID, configuration version, source response status, extracted records, rendered artifacts, and send result long enough to investigate complaints.
A practical launch checklist
- Write the source allowlist, topic rules, recency window, and timezone.
- Implement isolated Playwright contexts with per-source selectors and timeouts.
- Persist raw HTML and metadata before normalization.
- Add canonicalization, duplicate suppression, and a review state.
- Render escaped HTML plus plain text and generate a sources section.
- Run link, content, identity, address, and unsubscribe validation.
- Send to a dry-run list, then enable the scheduler with a lock and idempotency key.
- Capture delivery, bounce, complaint, and unsubscribe events and alert on anomalies.
Frequently Asked Questions
Can Playwright use an existing browser instead of launching one?
Yes. Its BrowserType API supports connecting to an existing browser server as well as launching Chromium, Firefox, or WebKit. Use a dedicated server and isolated context when sharing infrastructure.
What should happen when a source blocks automation?
Stop collecting that source, record the failure, and either obtain permission or replace it. Do not bypass a CAPTCHA or other access control.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →How do I prevent an editor from sending an empty issue?
Make minimum item count, source-health, review-queue, link, and compliance checks hard pre-send gates. A failed gate should leave the run in a review or failed state.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




