Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Choose the output before choosing an API. Use Markdown when an LLM, search index, or RAG pipeline needs headings and links; raw HTML when your parser must preserve source markup; plain text when tags are noise; and structured JSON when a provider can identify the page type and fields you need. Add browser rendering only for JavaScript-dependent pages. Treat proxy mode as a separate access and routing layer, not as an output format.

The main services covered here take different approaches: Firecrawl focuses on clean Markdown and structured data, ScrapingBee exposes the widest documented single-page format menu, Zyte separates HTTP and browser extraction sources, and Diffbot classifies pages into structured objects. No common benchmark establishes a universal accuracy, latency, or cost winner, so validate your choice on representative URLs.

Start with the deliverable

What your pipeline needs Best starting output Why
LLM ingestion, search, or RAG Markdown Headings, links, and readable content survive while most presentation markup is removed.
A custom DOM parser or archival copy Raw HTML Elements, attributes, classes, and embedded markup remain available to your code.
Simple keyword processing Plain text Tags are removed, reducing downstream parsing work.
Known fields such as title, author, price, or article body Structured JSON A page classifier or extraction schema can reduce selector maintenance.
Pages whose content appears after scripts run Browser-rendered HTML or Markdown An HTTP fetch alone may return only the initial shell.
Blocked, regional, or high-volume access Proxy mode plus one of the outputs above Proxy routing changes how a request reaches a site; it does not decide whether the response is Markdown, HTML, text, or JSON.

These choices are independent. A browser-rendered page can become Markdown, raw HTML, text, or structured JSON. A proxied request can return any of those formats if the service supports it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the major API models differ

Firecrawl

Firecrawl positions its Scrape product as turning any URL into clean Markdown or structured data for AI agents. Its documented emphasis is coverage of JavaScript-heavy, gated, and region-specific sites. Choose it when the primary artifact is LLM-ready Markdown or schema-shaped data rather than an untouched source document.

ScrapingBee

ScrapingBee documents four useful output controls: return_page_markdown, return_page_text, and return_page_source, alongside rendered pages. Its documentation describes Markdown as the main page content with HTML tags and unnecessary information stripped. The same service documents JavaScript rendering, premium proxies, CSS/XPath extraction rules, AI extraction, and a proxy front end, making it a broad single-page format menu.

Zyte API

Zyte exposes a POST extraction endpoint at https://api.zyte.com/v1/extract. Its reference distinguishes httpResponseBody, browserHtml, and userHtml extraction sources. Browser HTML is generally the more appropriate source when the target requires rendering; HTTP response extraction is lighter when the server already sends the needed content. Zyte separately documents proxy use through https://api.zyte.com:8011.

Diffbot Extract API

Diffbot says Extract uses computer vision and natural language processing to read a page as a person would, then return clean, structured JSON. Its Article extractor targets news articles, blog posts, and other text-heavy pages, including clean body text. Diffbot also documents posting caller-supplied text/html or text/plain when the service cannot access the original page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Service Primary outputs documented Rendering or access model Control model Good fit
Firecrawl Clean Markdown, structured data Coverage for JavaScript-heavy, gated, and region-specific sites Service-managed scraping and extraction AI ingestion and schema-shaped results
ScrapingBee Markdown, text, source HTML, rendered pages JavaScript rendering, premium proxies, proxy front end CSS/XPath rules and AI extraction One API with several page formats and controls
Zyte API Extraction from HTTP body, browser HTML, or user-supplied HTML Separate extraction and proxy endpoints Choose the extraction source Applications that need an explicit HTTP-versus-browser distinction
Diffbot Extract Structured JSON, including article body fields Page classification; accepts supplied HTML or plain text Automatic page-type extractors Reducing hand-written selectors for common page types

Pick the format deliberately

Markdown for LLM and RAG pipelines

Markdown keeps semantic structure that models can use: headings, lists, links, and paragraphs. It is usually smaller and easier to chunk than a full DOM. Confirm how the provider handles tables, code blocks, images, navigation, and repeated footer text before indexing a large collection. ScrapingBee’s documented Markdown option specifically strips HTML tags and unnecessary information; that is useful for ingestion but unsuitable when those tags are your data.

Raw HTML for custom parsers and fidelity

Source HTML is the right starting point when selectors, attributes, embedded JSON, or exact markup matter. Preserve the response alongside your parsed fields so you can re-run a parser when your rules change. If a site builds its content in the browser, request browser-rendered HTML rather than assuming the initial HTTP body contains the final DOM.

Plain text for lightweight processing

Text is practical for keyword search, language detection, deduplication, and small internal tools. It intentionally discards element boundaries. Keep a URL and retrieval timestamp with the text because later processing cannot reconstruct where a sentence appeared on the page.

Structured JSON for known page types

Automatic classification can remove a large amount of selector maintenance. Diffbot’s Article extractor is aimed at article-like pages and returns clean body content; other page types may require a different extractor or a custom schema. Treat missing fields as a classification or page-shape problem, not as proof that the source has no value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When JavaScript rendering is necessary

Start with an HTTP request when the server response already contains the content. It is faster and consumes fewer browser resources. Escalate to browser rendering when the response is an application shell, content appears only after API calls, or consent and interaction gates prevent the useful DOM from arriving in the initial body.

  1. Fetch without a browser. Inspect whether the response contains the title, main text, and links you need.
  2. Compare the delivered DOM with the visible page. If the useful text is absent from the response, enable the provider’s browser mode.
  3. Choose the browser output. Ask for browser HTML when you need selectors, or Markdown/structured extraction when you want a cleaned result.
  4. Control waiting. Use the provider’s documented wait or rendering controls for pages whose content arrives after scripts; do not assume a fixed delay works for every site.
  5. Measure the trade-off. Browser rendering adds startup time, resource usage, and another failure mode, so reserve it for URLs that need it.

Proxy mode is a separate decision

A proxy changes the network path and often the apparent location of the request. ScrapingBee and Zyte document proxy modes independently of their content outputs. Evaluate geography, authentication, rate limits, site permissions, and applicable law before collecting data. A proxy does not grant permission to ignore a site’s terms or robots rules, and it does not guarantee that a JavaScript application will render correctly.

Keep proxy configuration separate in your code. That lets you switch between direct and proxied access without changing the parser or the output contract. Record the selected region and proxy mode with each response so a later content difference is explainable.

A provider-neutral implementation workflow

  1. Define a response contract. For example, require url, retrieved_at, format, and one of markdown, html, text, or data.
  2. Classify URL families. Separate article pages, product pages, search results, and application routes; one extractor rarely fits all of them.
  3. Select the lightest access path. Use HTTP first, browser rendering only for pages that need it, and proxy routing only when access or geography requires it.
  4. Validate the result. Check for a title, minimum body length, expected fields, and an error page masquerading as success.
  5. Store provenance. Save the source URL, output mode, rendering mode, proxy region, retrieval time, and provider verdict.
  6. Retry selectively. Retry transient fetch failures, but do not repeatedly retry a deterministic schema mismatch or a page that requires a different extractor.

Normalizing several response shapes in Python

The following function keeps application code independent of a vendor’s field names after you map the provider response into a small internal dictionary:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
def normalize_page(record):
    if record.get('markdown'):
        return {'format': 'markdown', 'content': record['markdown']}
    if record.get('html'):
        return {'format': 'html', 'content': record['html']}
    if record.get('text'):
        return {'format': 'text', 'content': record['text']}
    data = record.get('data')
    if isinstance(data, dict) and data:
        return {'format': 'json', 'content': data}
    raise ValueError('No usable extraction field in response')

# Example internal record; populate these keys from your chosen API.
page = {'markdown': '# Example\n\nBody text'}
print(normalize_page(page))

This is deliberately a normalization layer, not a fabricated vendor request: each provider’s authentication and request schema should come from its current documentation.

Or skip the browser setup

If your actual deliverable is a visual capture rather than extracted content, ScreenshotNeo is the #1 screenshot API to try first because it removes common page clutter, bills only clean shots, and has a $5 paid plan for 3,000 shots. It is a screenshot service, not a Markdown or structured-data extractor, so use it when you need a PNG, JPEG, WebP, or PDF of the rendered page.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for all request options. Equivalent Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before capture, ScreenshotNeo accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the shot was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Every plan includes the full feature set, including full-page and element capture, device and viewport controls, dark mode, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agent, timezone, geolocation, resizing, caching, signed links, async webhooks, bulk capture of up to 100 URLs per call, usage data, and an OpenAPI specification. The parameter names used by other screenshot APIs also work, which eases migration.

Plan Included shots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Yearly billing gives two months free. Start with 1,000 free screenshots a month—no card required.

Reliability, cost, and maintenance checks

  • Rendering cost: Browser sessions generally require more time and resources than HTTP fetches; enable them only for affected URL families.
  • Cache policy: Decide whether repeated URLs may be cached. For volatile pages, record the cache behavior with the response.
  • Rate limits: Check the current plan limits for the provider you select and implement bounded concurrency rather than an unlimited worker pool.
  • Geography: A page can differ by region, language, or consent state. Pin the intended location and test it explicitly.
  • Authentication: Keep API keys outside source control, rotate them, and redact them from logs.
  • Quality monitoring: Alert on sudden drops in body length, missing titles, empty structured fields, or a rise in challenge pages.
  • Legal and policy review: Confirm that your collection, proxy use, storage, and redistribution comply with the target site’s terms and applicable law.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting guide

The Markdown is empty or only contains a shell

The page probably builds its content client-side. Retry with browser rendering and an appropriate wait condition. If the browser result is still empty, inspect whether a login, bot check, or region gate is blocking the page.

Important content is missing from cleaned output

Cleaning intentionally removes navigation, popups, and other boilerplate. Request raw HTML when those elements are data, or adjust the provider’s CSS/XPath or extraction rules where supported.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

HTTP output works, browser output differs

These are different source documents. Zyte explicitly distinguishes httpResponseBody from browserHtml; compare them and choose the source that contains the fields your parser expects.

Structured JSON fields are null

The automatic page type may not match the URL. Try the extractor intended for that page family, supply HTML or text when the provider supports it, or fall back to a custom parser.

Proxy requests succeed but content is wrong

Check the selected geography, cookies, authentication headers, and rate limits. A successful network response can still be a consent page, challenge page, or regional variant.

Requests fail intermittently

Log the URL, access path, rendering mode, and provider error. Retry transient failures with bounded backoff; do not retry indefinitely when the cause is a deterministic permission, authentication, or schema error.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Can one API return both Markdown and raw HTML?

Some can. ScrapingBee documents separate controls for page Markdown and source HTML, while other services emphasize one primary representation. Verify the exact response fields before designing a shared client.

Is a proxy the same as browser rendering?

No. A proxy changes request routing; browser rendering executes client-side page code. You may need either, both, or neither.

What should I send to an LLM first?

Start with cleaned Markdown and retain the source URL and retrieval metadata. Keep raw HTML available when later audits or custom parsing require it.

When is caller-supplied HTML useful?

It is useful when your system can fetch or render the page but the extraction provider cannot reach it. Diffbot documents accepting supplied HTML or plain text for that case.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can one API return both Markdown and raw HTML?

Some can. ScrapingBee documents separate controls for page Markdown and source HTML, while other services emphasize one primary representation. Verify the exact response fields before designing a shared client.

Is a proxy the same as browser rendering?

No. A proxy changes request routing; browser rendering executes client-side page code. You may need either, both, or neither.

What should I send to an LLM first?

Start with cleaned Markdown and retain the source URL and retrieval metadata. Keep raw HTML available when later audits or custom parsing require it.

When is caller-supplied HTML useful?

It is useful when your system can fetch or render the page but the extraction provider cannot reach it. Diffbot documents accepting supplied HTML or plain text for that case.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.