Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Short answer: Use Jina Reader when you need a fast URL-to-Markdown request, Browserless when you need GraphQL and browser-state control, and Firecrawl when you need clean single-page extraction or a domain-wide crawl. The conversion is only as good as the fetch: JavaScript rendering, selector scoping, waiting, access permissions, retries and rate limits matter as much as the HTML-to-Markdown step.

What a website-to-Markdown API actually does

A website-to-Markdown API accepts a URL, retrieves the page, optionally runs it in a browser, removes navigation and other boilerplate, and returns Markdown or structured data. A useful mental model is a five-stage pipeline:

  1. Fetch: Resolve the URL, follow redirects and obtain the response.
  2. Render: Execute JavaScript when the useful content is not present in the initial HTML.
  3. Scope: Keep the article or selected DOM region and exclude menus, ads, consent panels and unrelated widgets.
  4. Convert: Serialize the cleaned DOM as Markdown, HTML, text, JSON or another requested format.
  5. Operate: Apply timeouts, retries, caching, rate limits, deduplication and logging.

If a page is client-rendered, blocked, or still loading when extraction begins, changing the Markdown serializer will not fix the result. The fetch and render stages must succeed first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the right API for the job

Need Best fit Why Trade-off
Prototype one URL with minimal code Jina Reader Prepend https://r.jina.ai/ to the target URL and receive reader-style content. Execution controls are less application-centric than a full browser API.
Control rendered browser state through GraphQL Browserless Its goto operation navigates and its markdown mutation returns converted content; selector, timeout and visibility controls are available. You must manage a GraphQL request and browser timing.
Extract one page as Markdown, JSON, links or a screenshot Firecrawl Scrape It renders pages in a real browser, removes common page chrome and supports multiple output types. Provider-specific API setup and usage limits apply.
Build a corpus from an entire domain Firecrawl Crawl It discovers and processes subpages into a Markdown or JSON corpus. Crawl discovery, deduplication, scope control and rate planning become essential.

Use a single-page workflow for an article, product page or support document. Use crawling only when you genuinely need many pages; a crawl has a different failure and cost profile from one conversion.

Fastest path: Jina Reader

Jina’s basic Reader pattern is a URL prefix. This is a complete command you can run immediately:

curl "https://r.jina.ai/https://www.example.com"

The response is the page’s extracted content in Markdown. In Python:

import requests

reader_url = 'https://r.jina.ai/https://www.example.com'
response = requests.get(reader_url, timeout=90)
response.raise_for_status()
print(response.text)

In Node.js 18 or newer:

const response = await fetch('https://r.jina.ai/https://www.example.com');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const markdown = await response.text();
console.log(markdown);

Jina documents Markdown, HTML, text, screenshot, frontmatter and markdown+frontmatter response modes. Its controls are useful when the default extraction is too broad:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Control Use it when
Browser fetching The initial response lacks content that appears only after JavaScript executes.
x-target-selector You know the CSS selector for the article or main content and want to exclude surrounding chrome.
Wait-for selector A known element appears only after a client-side request finishes.
Exclude selectors Comments, related links, navigation, ads or other regions pollute the result.
Output format You need HTML, plain text, frontmatter, Markdown with frontmatter or a screenshot instead of default Markdown.
Cache controls You want repeat requests to reuse a recent result rather than fetch the page again.

Use the exact parameter or header names from Jina’s current documentation for these controls; they can change independently of the URL-prefix behavior. Jina’s 2026 rate-limit table lists 20 requests per minute without an API key, 500 requests per minute with a free key and up to 5,000 requests per minute with a premium key. The same table lists 7.9 seconds average latency. Those are provider-published, time-sensitive operational figures, not guarantees; verify the current limits and latency before launch.

Browser-level control with Browserless GraphQL

Browserless is appropriate when your application already uses GraphQL or needs explicit control over the rendered page. Its documented pattern navigates first, then converts the resulting page:

mutation Markdownify {
  goto(url: "https://example.com") { status }
  markdown { markdown }
}

The markdown operation accepts a CSS selector, a timeout and a visible setting. The documented default timeout is 30,000 milliseconds. A selector limits conversion to the content region; visible is useful when the DOM contains hidden duplicate templates. Set a longer timeout only when the page is known to need it, because waiting on every request reduces throughput. Check the returned navigation status before accepting the Markdown, and treat an empty result as a failed extraction rather than a valid document.

Single-page extraction or a whole-site corpus with Firecrawl

Firecrawl Scrape

Firecrawl Scrape is designed for one URL. It renders the page in a real browser, removes navigation, footers, ads and tracking elements, and can return clean Markdown, structured data, links or screenshots. Select this pattern for a controlled queue of URLs where you want one result per input and can record the source URL alongside the output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Firecrawl Crawl

Firecrawl Crawl discovers and processes subpages on a domain, returning a Markdown or JSON corpus. Before starting, define an allowed host, path rules, a maximum page count and a deduplication key. Store each page’s canonical URL, retrieval time, HTTP result and content hash. Without those controls, faceted navigation, print views and tracking parameters can create a surprisingly large and repetitive corpus.

Make the Markdown useful for search and RAG

Scope to the content you need

Selector scoping is usually the highest-impact quality improvement. Target the article body, documentation container or product description instead of accepting the entire page. Exclude cookie notices, sticky headers, recommendation rails, comments and chat widgets. Keep headings and lists; they provide structure for chunking and citations.

Wait for the real content

For a JavaScript application, wait for a meaningful selector or network completion rather than sleeping for an arbitrary number of seconds. A fixed delay can be too short on a slow run and wasteful on a fast run. Record the selector, timeout and final HTTP status with every document so a later quality check can explain an empty or partial result.

Preserve provenance

Store the original URL, final URL after redirects, retrieval timestamp, provider, response status and extraction options next to the Markdown. If you request frontmatter or markdown+frontmatter, retain it instead of stripping it before indexing. Provenance lets you refresh stale pages and trace an answer back to its source.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Normalize before indexing

Normalize line endings, remove repeated whitespace and reject pages that contain only a navigation shell or an error message. Do not blindly remove every link or code block: links and code are often the information a developer is trying to retrieve. Chunk by headings where possible, and keep the source URL in each chunk’s metadata.

Designing a reliable conversion pipeline

  1. Queue URLs: Deduplicate exact URLs and, for crawls, normalize tracking parameters according to your policy.
  2. Fetch with bounded retries: Retry transient network failures and 5xx responses with exponential backoff. Do not retry a deterministic 4xx access denial indefinitely.
  3. Render only when necessary: Start with a normal fetch or direct Reader request; enable browser rendering for pages whose content depends on JavaScript.
  4. Validate output: Require a minimum amount of meaningful text, expected headings or a known selector. Mark validation failures for review.
  5. Cache intentionally: Cache immutable documentation longer than frequently changing pages. Include the extraction options in the cache key.
  6. Monitor provider behavior: Track latency, status codes, empty-output rate, retry count and rate-limit responses by provider.

For a whole-site ingestion, process pages in bounded batches, honor the provider’s limits and the source site’s access rules, and stop when the crawl leaves the approved host or path. A queue that can pause is safer than firing thousands of parallel requests.

Or skip the browser setup

ScreenshotNeo is a screenshot and PDF API, not a Markdown converter. It is useful when your pipeline also needs a visual record of the rendered page, a regression artifact or an image for an AI agent. One GET request returns a PNG, JPEG, WebP or PDF; the service accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture.

cURL (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', data);

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the outcome with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Symptom Likely cause Fix
Markdown is empty or only contains a menu The page is client-rendered or the extractor selected the wrong region. Enable browser fetching, wait for a content selector and set x-target-selector or an equivalent article selector.
Content stops before the article ends Lazy loading, an early timeout or a selector that matches only the first block. Wait for the final content marker, increase the timeout within provider limits and inspect the rendered DOM.
Requests slow down or return rate-limit errors Your concurrency exceeds the provider’s current allowance. Use a bounded queue, exponential backoff and a cache; re-check the provider’s published limits.
Repeated pages appear in a crawl Tracking parameters, print URLs or faceted navigation create distinct URLs. Restrict paths, normalize URLs and deduplicate by canonical URL and content hash.
HTTP success but unusable text The response is a login wall, bot challenge, error template or consent screen. Validate the text and status, handle authentication where permitted, and do not present an access-control bypass as a feature.
Characters are corrupted Incorrect encoding handling in your storage or post-processing layer. Preserve the provider’s UTF-8 response, avoid lossy byte conversions and test pages containing non-Latin scripts.

Access, robots and rights

A fetcher should respect robots directives, authentication boundaries, rate limits and the source site’s terms. Jina states that Reader does not actively circumvent or bypass anti-bot systems, anti-bot defenses or access controls. You remain responsible for having permission to retrieve, store, transform and redistribute third-party material. For private documentation, use an authenticated workflow only when the site owner authorizes it, and protect any tokens or cookies supplied to the browser.

FAQ

Can I switch providers without rewriting my data model?

Yes. Store provider-neutral fields such as source URL, final URL, retrieval time, status, Markdown, metadata and extraction options. Keep provider-specific response headers and diagnostics in a separate object so a later migration does not discard useful operational history.

Should I save Markdown, HTML or both?

Save Markdown for indexing and human review, and retain the raw or rendered representation when you may need to re-extract with a different selector. Keeping both increases storage use but avoids refetching a page solely to repair an extraction rule.

How do I know whether a page changed?

Compare a normalized-content hash and the retrieval timestamp. If the hash changes, re-index the page; if it does not, keep the existing chunks and update only the freshness metadata.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I switch providers without rewriting my data model?

Yes. Keep provider-neutral fields for URLs, status, timestamps, Markdown and metadata, while storing provider-specific diagnostics separately.

Should I save Markdown, HTML or both?

Save Markdown for indexing and review; retain the raw or rendered representation when future re-extraction may be needed.

How do I detect page changes reliably?

Hash normalized content and compare the hash plus retrieval timestamp on each refresh.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.