Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Perplexity does not automatically crawl a website in this workflow. Your Python program fetches the page (for example, through the Crawlbase Crawling API), removes irrelevant markup, and sends the resulting text to Perplexity for interpretation. Keeping collection and interpretation separate makes failures easier to diagnose and lets you choose static or JavaScript rendering deliberately.

This guide builds that pipeline end to end: environment setup, HTML trimming, Markdown conversion, schema-constrained extraction, validation, retries, JavaScript-heavy pages, and production concerns.

How the fetch-then-interpret architecture works

The pipeline has five distinct stages:

  1. Collect: request the target URL with a crawling service.
  2. Trim: select the article, product, or other useful DOM region with BeautifulSoup.
  3. Normalize: convert the selected HTML to Markdown with markdownify, reducing navigation and markup noise.
  4. Interpret: send the cleaned text and an explicit extraction instruction to Perplexity.
  5. Validate: parse the response as JSON and check required types and missing values before storing it.

In the implementation described by Crawlbase, Crawlbase is the collection layer and Perplexity reads only the text your application supplies. As Hassan Rehan of Crawlbase puts it, “Perplexity does not crawl the site in this flow. It reads the text you give it.” See the Crawlbase guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That boundary matters. A timeout, proxy challenge, consent wall, or empty JavaScript shell is a collection failure; an incorrect field or malformed JSON is an interpretation or validation failure. Treating them as separate stages prevents you from trying to fix a crawler problem by changing the prompt.

Prerequisites and safe configuration

Install the Python packages

The example uses Python 3.10 or newer, matching the currently documented requirements of the official Perplexity Python SDK.

python -m venv .venv
source .venv/bin/activate       # Windows: .venvScriptsactivate
python -m pip install crawlbase beautifulsoup4 markdownify openai pydantic

The perplexityai package is the official SDK and supports synchronous and asynchronous clients, Search API calls, chat completions, and typed responses. This tutorial uses the OpenAI-compatible endpoint documented for Perplexity Agent workflows, so the openai client keeps the request explicit and easy to inspect.

Keep credentials out of source control

Create environment variables in your shell, secret manager, or deployment platform:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
export CRAWLBASE_TOKEN="your-crawlbase-token"
export PERPLEXITY_API_KEY="your-perplexity-api-key"

Do not print these values, commit a .env file containing real keys, or put them in a prompt. Rotate a key immediately if it appears in logs or a repository.

Complete Python implementation

The following script fetches a page, extracts a likely content region, converts it to Markdown, asks Perplexity for a fixed JSON object, and validates the result. Replace the example URL and selectors with those appropriate for the site you are allowed to access.

import json
import os
import re
from typing import Optional

import requests
from bs4 import BeautifulSoup
from crawlbase import CrawlingAPI
from markdownify import markdownify as html_to_markdown
from openai import OpenAI
from pydantic import BaseModel, ConfigDict, Field, ValidationError

TARGET_URL = "https://example.com/product"

class Product(BaseModel):
    model_config = ConfigDict(extra="forbid")
    name: Optional[str] = None
    description: Optional[str] = None
    price: Optional[str] = None
    currency: Optional[str] = None
    specifications: list[str] = Field(default_factory=list)


def fetch_html(url: str) -> str:
    """Use Crawlbase's normal token for server-rendered HTML."""
    api = CrawlingAPI({"token": os.environ["CRAWLBASE_TOKEN"]})
    response = api.get(url)
    # Crawlbase responses expose the fetched body as response.body.
    body = response.body
    if isinstance(body, bytes):
        body = body.decode("utf-8", errors="replace")
    if not body or len(body.strip()) < 200:
        raise RuntimeError("The crawler returned an empty or suspiciously small body")
    return body


def trim_html(html: str) -> str:
    soup = BeautifulSoup(html, "html.parser")
    for tag in soup(["script", "style", "noscript", "svg", "nav", "footer", "form"]):
        tag.decompose()

    # Prefer semantic content, then common article/product containers.
    content = (
        soup.find("main")
        or soup.find("article")
        or soup.select_one(".product, .product-page, .entry-content, .post-content")
        or soup.body
        or soup
    )
    return str(content)


def to_markdown(html_fragment: str) -> str:
    markdown = html_to_markdown(html_fragment, heading_style="ATX")
    markdown = re.sub(r"n{3,}", "nn", markdown)
    return markdown.strip()


def extract_product(markdown: str) -> Product:
    client = OpenAI(
        api_key=os.environ["PERPLEXITY_API_KEY"],
        base_url="https://api.perplexity.ai/v1",
    )
    instruction = """Extract the product fields from the supplied page text.nReturn only a JSON object with exactly these keys: name, description, price, currency, specifications.nUse null when a scalar field is not present and [] when no specifications are present.nNever infer a price, currency, name, or specification that is absent from the text.nKeep price as the displayed string and preserve the page's currency notation.nnPAGE TEXT:n""" + markdown

    completion = client.chat.completions.create(
        model="sonar",
        messages=[
            {"role": "system", "content": "You extract facts conservatively from supplied text."},
            {"role": "user", "content": instruction},
        ],
        temperature=0,
        response_format={"type": "json_schema", "json_schema": {
            "name": "product",
            "schema": {
                "type": "object",
                "additionalProperties": False,
                "properties": {
                    "name": {"type": ["string", "null"]},
                    "description": {"type": ["string", "null"]},
                    "price": {"type": ["string", "null"]},
                    "currency": {"type": ["string", "null"]},
                    "specifications": {"type": "array", "items": {"type": "string"}},
                },
                "required": ["name", "description", "price", "currency", "specifications"],
            },
        }},
    )
    raw = completion.choices[0].message.content
    return Product.model_validate(json.loads(raw))


def main() -> None:
    html = fetch_html(TARGET_URL)
    markdown = to_markdown(trim_html(html))
    if not markdown:
        raise RuntimeError("No usable text remained after HTML trimming")
    product = extract_product(markdown)
    print(product.model_dump_json(indent=2))


if __name__ == "__main__":
    main()

The JSON Schema makes the response shape predictable, while Pydantic rejects unexpected keys or wrong types. If your installed SDK version exposes a different structured-output parameter, follow its current documentation and retain the same schema and post-response validation. Never assume that a successful HTTP response means the extracted facts are correct.

Choosing the crawler token

Normal token for server-rendered HTML

Use Crawlbase’s normal token when the useful content is present in the initial HTML response. This is common for traditional server-rendered pages. Confirm by saving the response and searching it for a visible heading or distinctive product string.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JavaScript token for client-rendered pages

Modern applications may return an almost empty shell and populate the page only after JavaScript runs. Crawlbase describes its JavaScript token for those client-rendered pages. If your saved HTML contains the root element but not the visible text, switch to the JavaScript-capable token before changing the extraction prompt. Rendering adds latency and operational cost, so do not enable it for every URL without evidence that it is needed.

Even browser rendering cannot guarantee access to a page protected by a CAPTCHA, login requirement, geofence, or terms that prohibit automated collection. Obtain permission, authenticate through supported mechanisms, or skip the URL.

Making extraction accurate and repeatable

Trim before you prompt

Raw HTML includes navigation, scripts, tracking attributes, repeated links, and hidden elements. Selecting main, article, or a site-specific container and converting it to Markdown lowers token use and removes distractions. Keep headings, lists, tables, and labels because they often carry the relationship between a field and its value.

State the missing-data policy

Tell the model exactly what to do when a field is absent: null for a scalar and an empty array for a list. Explicitly prohibit inference. For prices, preserve the displayed string and currency rather than converting it unless your application has a separate, documented conversion step.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use selectors for stable layouts, schemas for variable ones

Fixed CSS selectors are fast and deterministic when every page follows the same template. Schema-directed extraction is more useful when labels, ordering, or wording vary across pages. A practical hybrid is to select the product container with CSS, then let Perplexity map its text to a validated schema.

Control input size

Set a maximum character or token budget after trimming. If a page contains reviews, recommendations, and comments unrelated to the target fields, remove those sections before the model call. If truncation is unavoidable, preserve the beginning and the sections where the required fields normally appear, and record that truncation in your job metadata.

Perplexity Agent and Search options

Perplexity’s current API Platform separates Agent and Search capabilities. Agent workflows include web search, URL fetching, and reasoning controls; Search provides ranked results, domain filtering, multi-query search, and content extraction. The Agent API announcement documents web_search, fetch_url, JSON Schema structured outputs, and the OpenAI-compatible base URL used above.

Those capabilities can complement a custom fetcher, but they do not change the central rule of this tutorial: when your program fetches a page and sends its text, the model interprets the supplied bytes. If you instead ask an Agent workflow to fetch a URL, document that as a different architecture with its own access, freshness, and failure behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliability, retries, and operational controls

Retry only transient failures

  • Retry network resets, gateway timeouts, and HTTP 5xx responses with exponential backoff and a maximum attempt count.
  • Do not blindly retry 401/403 responses; check credentials, permissions, robots or site policy, and challenge pages.
  • Do not retry malformed model JSON indefinitely. Log the response safely, tighten the schema or prompt, and cap repair attempts.

Record stage-level metadata

For each URL, store the fetch status, final URL, rendering mode, response length, extraction selector, Markdown length, model identifier, latency, and validation result. Hashing the cleaned text lets you detect unchanged pages and avoid unnecessary interpretation calls. Keep page content and credentials in separate access-controlled stores.

Respect site rules and rate limits

Check the target site’s terms, robots directives, authentication requirements, and applicable law before collecting. Throttle requests per host, honor provider limits, and use a queue for large jobs. Cache fetched content only for a period that fits the page’s update frequency and your legal obligations.

Troubleshooting common failures

Symptom Likely cause Fix
HTML is an empty app shell Content is rendered client-side Use Crawlbase’s JavaScript token, then verify visible text exists before extraction.
Markdown contains menus but no article Wrong container selector or a consent wall Inspect the saved HTML, choose a site-specific selector, and handle the consent flow lawfully.
Model invents a value Prompt permits inference or context is missing Require null/empty values, prohibit inference, preserve labels, and validate against source text.
JSON parse or schema error Free-form output, unsupported response-format option, or truncated response Use the documented structured-output interface, lower input size, retry once, and reject rather than silently accepting invalid data.
401 from Perplexity Missing, revoked, or mis-scoped API key Check PERPLEXITY_API_KEY, endpoint, and account access without logging the secret.
429 or repeated timeouts Rate limit, overloaded provider, or too much input Back off, reduce concurrency and page size, and queue work.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF rather than text extraction, ScreenshotNeo provides a one-request website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

For a screenshot, use the documented endpoint and options:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for full-page and element capture, 12 device presets plus custom viewports, retina scale, dark mode, PDF controls, custom CSS and JavaScript, clicks, waits, blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data, and the OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account.

When this pattern is the right choice

  • Choose fetch-then-interpret when you need auditable control over the exact page bytes sent to the model.
  • Use a normal crawler token for static HTML and a JavaScript token when the initial response lacks the content.
  • Prefer deterministic selectors for uniform templates and schema-constrained Perplexity extraction for varied layouts.
  • Validate every response and retain enough metadata to reproduce a failed job.

Frequently Asked Questions

Does Perplexity scrape the target website by itself in this Python workflow?

No. Your crawler fetches the page and your application supplies cleaned text; Perplexity interprets that text. Agent workflows can provide URL-fetching tools, but that is a different architecture.

When should I use a JavaScript-rendering crawler token?

Use it when the initial HTML is an app shell without the visible content and the page populates after JavaScript executes. Verify the saved response before switching.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I extract fields without asking for free-form prose?

Yes. Request a JSON Schema response, require every key, represent missing scalars as null and missing lists as [], then validate the parsed object in Python.

Is Markdown conversion mandatory?

No, but trimming HTML and converting the relevant fragment to Markdown usually removes navigation and markup noise, reducing irrelevant input and improving consistency.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.