Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convert each URL independently, store its Markdown under a deliberate cache key, and return one success or failure record per input. For small batches, stream results as they finish; for large batches, submit a background job. The important distinction is that a provider’s cache switch is not automatically an application-owned, durable per-URL cache with your chosen freshness rules.

What a production-ready converter must do

A useful bulk converter has three separate layers:

  1. Batch orchestration: accept a list, bound concurrency, retry transient failures, pace requests per host, and report every input.
  2. Fetching and conversion: choose a lightweight HTTP fetch or browser rendering, extract useful content, and produce Markdown. JavaScript-heavy pages, access controls and unusual layouts may need browser handling.
  3. Per-URL caching: map each request to a stable identity, apply freshness rules, and provide an explicit refresh or bypass path.

Keep the submitted URL, normalized cache key, final redirect URL, Markdown, status, fetch time and error details. Update a successful record only after conversion completes; otherwise a temporary outage can overwrite good content.

Choose the batch execution model

Streaming for small and moderate lists

Crawl4AI’s hosted API documents a streaming batch endpoint accepting up to 50 URLs per call and emitting one newline-delimited JSON result as each URL completes. This lets downstream work begin without waiting for the slowest page. Treat that limit as a hosted-API limit documented at Crawl4AI’s API documentation, not as a universal limit of its open-source library.

Background jobs for large lists

The same hosted documentation describes background jobs for lists up to 10,000 URLs. Submit the list, retain the job ID, poll its status, then retrieve results. This avoids keeping a client connection open for a long crawl and makes resumability easier.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Self-hosted execution

The open-source Crawl4AI library gives you ownership of browser runtimes, storage, proxies and monitoring. It does not automatically inherit the hosted service’s limits. Jina Reader can also be self-hosted; its project documentation describes a stateless default and optional S3-compatible caching.

Define URL identity before writing code

Do not casually strip query strings: they can select different content. Decide and document how your system treats:

  • Host-name casing and default ports.
  • Trailing slashes.
  • Fragments. They may be irrelevant to server HTML, but can select client-rendered sections.
  • Query parameters, including tracking parameters and content selectors.
  • Redirects. Keep the original URL for auditability and store the final URL separately.

Use a standards-compliant URL parser and a versioned canonicalization policy. A policy change should produce a new cache-key version rather than silently mixing old and new records.

A complete Python implementation

The example below uses SQLite for durable per-URL records, bounded asynchronous concurrency, freshness checks and independent failures. Replace fetch_markdown with Crawl4AI, Jina Reader, or your own HTTP/browser extractor.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio, hashlib, json, sqlite3, time
from urllib.parse import urlsplit, urlunsplit

DB = "url_markdown.db"
TTL = 3600
CONCURRENCY = 8

def canonicalize(url):
    p = urlsplit(url.strip())
    host = (p.hostname or "").lower()
    port = p.port
    netloc = host
    if port and not ((p.scheme == "http" and port == 80) or (p.scheme == "https" and port == 443)):
        netloc += f":{port}"
    # Query strings are retained; fragments are removed for this server-content policy.
    return urlunsplit((p.scheme.lower(), netloc, p.path or "/", p.query, ""))

def key_for(url):
    return hashlib.sha256(canonicalize(url).encode()).hexdigest()

def init_db():
    with sqlite3.connect(DB) as c:
        c.execute("""CREATE TABLE IF NOT EXISTS pages(
          cache_key TEXT PRIMARY KEY, submitted_url TEXT, canonical_url TEXT,
          final_url TEXT, markdown TEXT, status TEXT, error TEXT,
          fetched_at REAL, updated_at REAL)""")

def cached(key, now):
    with sqlite3.connect(DB) as c:
        row = c.execute("SELECT * FROM pages WHERE cache_key=?", (key,)).fetchone()
    if row and row[7] and now - row[7] < TTL and row[5] == "ok":
        return row
    return None

async def fetch_markdown(url):
    # Call your chosen converter here. Return (markdown, final_url).
    raise NotImplementedError

async def one(url, sem, refresh=False):
    key, canonical, now = key_for(url), canonicalize(url), time.time()
    if not refresh:
        hit = cached(key, now)
        if hit:
            return {"url": url, "status": "cache_hit", "markdown": hit[4], "final_url": hit[3]}
    async with sem:
        try:
            markdown, final_url = await fetch_markdown(url)
            with sqlite3.connect(DB) as c:
                c.execute("INSERT OR REPLACE INTO pages VALUES (?,?,?,?,?,?,?,?,?)",
                    (key,url,canonical,final_url,markdown,"ok",None,now,time.time()))
            return {"url":url,"status":"ok","markdown":markdown,"final_url":final_url}
        except Exception as e:
            with sqlite3.connect(DB) as c:
                c.execute("INSERT OR REPLACE INTO pages VALUES (?,?,?,?,?,?,?,?,?)",
                    (key,url,canonical,None,None,"error",str(e),None,time.time()))
            return {"url":url,"status":"error","error":str(e)}

async def convert(urls, refresh=False):
    init_db(); sem = asyncio.Semaphore(CONCURRENCY)
    return await asyncio.gather(*(one(u, sem, refresh) for u in urls))

# results = asyncio.run(convert(["https://example.com", "https://example.org"]))
# force a refetch: asyncio.run(convert(urls, refresh=True))

In a real service, replace SQLite with a database that supports concurrent workers, add a unique constraint on the cache key, and encrypt sensitive cookies or authorization data. Store failures separately if you want a short negative-cache interval; otherwise retry them normally.

Freshness, retries and politeness

Freshness policy

A time-to-live is only one choice. You can also invalidate by deployment, content version, webhook, or an operator-triggered refresh. Expose refresh=true (or an equivalent queue action) so users never need to alter URLs to bypass stale content.

Retries

Retry connection resets, 408, 429 and selected 5xx responses with exponential backoff and jitter. Do not repeatedly retry authentication failures, invalid URLs, or deterministic 4xx responses. Set per-request connect and total timeouts.

Concurrency and host pacing

Use a global worker limit plus per-host limits. Respect provider quotas and target-site capacity. Crawl4AI documents concurrency and delay controls, and exposes a robots.txt setting whose documented default is false; enable and configure it deliberately rather than implying automatic compliance. See its parameter documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Using hosted readers and cache controls

Jina Reader converts a URL to LLM-friendly representations, including Markdown; its documentation describes the simple https://r.jina.ai/ prefix. The project notes that Reader may select a browser or a lightweight curl-based engine. Its self-hosted deployment is stateless unless you configure an S3-compatible bucket. The project also documents x-cache-tolerance and x-no-cache headers for freshness and bypass behavior; these are provider controls, not a replacement for your own cache-key and audit model. Read the Reader project documentation.

Jina’s hosted rate limits are tier-dependent and expressed through requests-per-minute and tokens-per-minute enforcement. They change, so check the live Reader API page before hard-coding a number or price.

Result format and observability

Return one object per submitted URL, even when a fetch fails:

{"url":"https://example.com?a=1","status":"ok","cache":"miss","canonical_url":"https://example.com/?a=1","final_url":"https://www.example.com/","fetched_at":"2026-09-29T12:00:00Z","markdown":"# Example"}

For failures include a stable error class, human-readable detail, attempt count and elapsed time. Track cache hit rate, conversion latency, response sizes, status-code distribution, retry counts and per-host throttling. Never log authorization headers or private cookies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failure modes

Every request is a miss

Compare the stored canonical key with the incoming one, check that the database is persistent, and verify that the freshness clock uses UTC consistently.

Distinct pages collapse into one record

Your canonicalizer probably removed meaningful query parameters or fragments. Restore them, or make the content-selection rule explicit.

Markdown is empty or incomplete

The page may require JavaScript, consent interaction, authentication, or scrolling. Switch to browser rendering, wait for a selector or network idle, and record the final URL and rendered status.

429 or repeated timeouts

Lower concurrency, add per-host delay, honor Retry-After when present, and raise timeouts only after checking page size and rendering behavior.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One bad URL aborts the batch

Use per-task exception handling and persist each result independently. A batch aggregate should report counts, not hide individual errors.

When screenshots are part of the pipeline

If your workflow also needs visual evidence, ScreenshotNeo is the first screenshot API to try: it removes consent banners, newsletter popups and chat widgets before capture, and bills only clean shots. It is separate from Markdown extraction, but useful for QA records or pages whose rendered state must be inspected.

Or skip the browser setup

One GET request returns PNG, JPEG, WebP or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for options. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

How to choose an architecture

Need Best fit What to verify
Immediate progress for up to 50 URLs Hosted streaming batch NDJSON handling, current limit and account quota
Thousands of URLs Hosted background job 10,000-URL documented cap, polling and result retention
Strict data control Self-hosted crawler or Reader Browser updates, storage, proxies, monitoring and S3 cache setup
Simple URL-to-Markdown conversion Jina Reader Current RPM/TPM limits and rendering behavior

Estimate cost from expected URL count, cache-hit rate, rendered-browser share, retries and storage—not just the number of submitted URLs. Recheck provider limits and pricing immediately before deployment because the cited hosted values are time-sensitive.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Should failed pages be cached?

Usually only briefly, with a separate negative-cache expiry; otherwise a transient outage can persist until the normal TTL ends.

Is a URL fragment part of the cache key?

Only if the target application uses it to select client-rendered content. Make that behavior an explicit, documented policy.

Can streaming and background jobs share storage?

Yes. Use the same per-URL record schema and job identifier, while allowing either execution mode to write independently.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.