The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To convert a JavaScript-rendered page to Markdown, first make its content available, then select the useful part of the page, and finally convert that HTML to Markdown. A plain HTTP fetch is enough when the response already contains the content; when it returns an app shell, render the page in a browser before extracting and converting it. Turndown handles HTML-to-Markdown conversion, but it does not run page JavaScript or decide which content is the main article.
Why a successful fetch can still produce empty Markdown
A web page can change substantially after its initial HTML response arrives. The browser processes HTML, CSS, and JavaScript; page scripts can add or alter DOM content as the application runs. Consequently, an HTTP request may succeed while returning only the shell of a single-page app (SPA), not the content a person sees after the browser finishes rendering. See how browsers work.
Markdown conversion is a separate operation. A converter such as Turndown turns supplied HTML or a DOM node into Markdown. It is not a browser renderer or a main-content classifier. A complete pipeline therefore has three stages: render if necessary, extract the relevant region, then convert it.
Choose static fetch or browser rendering
| Approach | Use it when | Trade-off |
|---|---|---|
| Static fetch plus converter | The HTTP response already contains the text and structure you need. | Simple to operate, but an SPA shell can yield empty or incomplete output. |
| Browser render, extraction, and converter | The route depends on JavaScript, or requires browser interaction. | Can expose client-rendered content, but needs browser setup and page-specific readiness and extraction choices. |
| Hosted rendering and extraction service | You prefer a service to combine some or all of the stages. | Less infrastructure to assemble; verify the service’s capabilities, quality, limits, and price for your workload. |
Start with the static path when appropriate, but inspect the returned text rather than using HTTP success as proof that the page is complete. A static-first request with a Chromium fallback is one documented implementation pattern, not a universal standard: fetch_as_markdown describes this approach.
#1 Best Overall
Build a browser-rendered conversion pipeline
The example below uses Playwright for browser navigation and Turndown for conversion. It deliberately leaves the extraction selector configurable: page structures differ, and there is no one selector or readiness condition that works for every site. Install these packages in a Node.js project with npm install playwright turndown, then install Playwright’s Chromium browser with npx playwright install chromium.
Runnable Node.js example
Save as convert.mjs. Pass the target URL and, optionally, a CSS selector as arguments. The default selector is main; if the site uses another article container, supply its selector explicitly.
import { chromium } from 'playwright';
import TurndownService from 'turndown';
const url = process.argv[2];
const selector = process.argv[3] || 'main';
if (!url) {
console.error('Usage: node convert.mjs <url> [content-selector]');
process.exit(1);
}
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
}
// Replace this with a selector for the actual content on the target site.
await page.locator(selector).waitFor({ state: 'visible', timeout: 15000 });
const html = await page.locator(selector).evaluate(element => element.innerHTML);
const turndown = new TurndownService({ headingStyle: 'atx' });
const markdown = turndown.turndown(html);
console.log(markdown);
} finally {
await browser.close();
}
Run it with node convert.mjs https://example.com/article article, changing both the URL and selector to match the page. The script checks the navigation response and waits for a visible target before converting its inner HTML. That is a practical starting point, not a guarantee: a visible container may still be empty, partially populated, or the wrong region.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsRank #2
Choose a meaningful readiness condition
Playwright’s Page API provides navigation, browser-page interaction, and page content methods. The sample uses a selector wait because it expresses a page-specific condition, but you should confirm that the selected element contains the data you need. Some sites populate content after a click, a scroll, or additional application work. There is no universal wait condition or timeout established for all pages.
A scroll can trigger deferred content on pages that load it as the user moves down the page. The yomi README documents a render-and-scroll option as one tool’s approach; scrolling is not a guarantee that every lazy-loaded item will appear. If the target needs interaction, perform that interaction before extracting the content.
Extract before converting
Convert the article or content container, rather than the entire document, when possible. Whole-page conversion can include navigation, footers, cookie notices, and repeated interface text. Selecting a narrower region reduces that noise, but the right region depends on the site’s markup. Hosted tools advertise content cleanup, but their output quality can vary by page; the cited material does not establish an independent comparison of extraction accuracy.
Turndown accepts HTML strings and DOM nodes and converts them to Markdown. In the sample, the browser obtains the selected element’s HTML and Turndown serializes it. After conversion, inspect whether headings, links, lists, tables, and content revealed only after interaction were preserved. A converter cannot restore information that was absent from the HTML you supplied.
Use a hosted service when you do not want to assemble the stages
Services may combine browser rendering and Markdown output. Firecrawl describes browser-based page processing with Markdown output, while Microlink discusses SPA rendering and page readiness: Firecrawl’s guide and Microlink’s SPA guide. These are vendor descriptions, not independent evidence of comparative extraction quality or performance.
Compare services and self-hosted approaches on JavaScript rendering, readiness controls, content extraction, preservation of structure and links, interactions or authentication, operating cost, and whether you can inspect raw HTML when debugging. The conversion stages explain why these criteria matter; no independent head-to-head quality study is established by the cited material.
Rank #4
Troubleshoot empty or incomplete Markdown
- Output is empty, but the request succeeded: Inspect the original response. If it is an SPA shell without the target text, use browser rendering before conversion.
- The page opened, but your selector timed out: Check the site’s actual DOM and pass a selector for a real content container. The example’s default
mainselector is not guaranteed to exist. - The selector exists, but Markdown is blank: Check whether the selected element has text or child content at extraction time. It may be an empty app container that fills later, or the selector may point to a wrapper without the desired content.
- Some article sections or images are missing: Determine whether they load after scrolling or interaction. Trigger the page behavior needed before extraction; do not assume a single scroll captures all deferred content.
- Navigation reports a failure or no response: Check the URL, network reachability, and response status. The sample throws when navigation fails or returns a non-success response; handle redirects and site-specific access behavior according to your use case.
- The Markdown contains menus, footers, or repeated text: Narrow the extraction selector to the main article region and rerun conversion.
- Headings, links, lists, or tables look wrong: Compare the extracted HTML with the converted output. The loss may come from selecting incomplete HTML or from conversion behavior; validate the result rather than assuming conversion is lossless.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its API returns an image or PDF, not Markdown, so it is not a direct substitute for the render-extract-convert pipeline above. For workflows where a clean rendered capture is useful before your own processing, one GET request can capture a page:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; these steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers indicate the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Free tools Windows power users keep installed
One-click scans. No signup required.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Cost and reliability considerations
A self-hosted browser workflow gives you control over navigation, readiness, extraction, and conversion, but you must operate the browser and account for pages that load inconsistently or depend on interaction. Static fetching avoids browser setup when it returns complete content. Hosted services reduce the number of components you assemble, but check their current limits, price, access requirements, and outputs for the pages you need. The cited sources do not establish universal reliability figures, measured speed, or comparative prices.
Best Value
Frequently Asked Questions
Does Turndown render JavaScript?
No. Turndown converts HTML or DOM content supplied to it; use a browser renderer first when page JavaScript creates the content.
Can I convert every SPA with one generic selector or wait setting?
No universal selector or readiness condition is established. Inspect the page and choose a content-specific target and wait condition.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

