Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Perplexity does not automatically crawl a website in this workflow. Your Python program fetches the page (for example, through the Crawlbase Crawling API), removes irrelevant markup, and sends the resulting text to Perplexity for interpretation. Keeping collection and interpretation separate makes failures easier to diagnose and lets you choose static or JavaScript rendering deliberately.
This guide builds that pipeline end to end: environment setup, HTML trimming, Markdown conversion, schema-constrained extraction, validation, retries, JavaScript-heavy pages, and production concerns.
How the fetch-then-interpret architecture works
The pipeline has five distinct stages:
- Collect: request the target URL with a crawling service.
- Trim: select the article, product, or other useful DOM region with BeautifulSoup.
- Normalize: convert the selected HTML to Markdown with markdownify, reducing navigation and markup noise.
- Interpret: send the cleaned text and an explicit extraction instruction to Perplexity.
- Validate: parse the response as JSON and check required types and missing values before storing it.
In the implementation described by Crawlbase, Crawlbase is the collection layer and Perplexity reads only the text your application supplies. As Hassan Rehan of Crawlbase puts it, “Perplexity does not crawl the site in this flow. It reads the text you give it.” See the Crawlbase guide.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteThat boundary matters. A timeout, proxy challenge, consent wall, or empty JavaScript shell is a collection failure; an incorrect field or malformed JSON is an interpretation or validation failure. Treating them as separate stages prevents you from trying to fix a crawler problem by changing the prompt.
#1 Best Overall
Prerequisites and safe configuration
Install the Python packages
The example uses Python 3.10 or newer, matching the currently documented requirements of the official Perplexity Python SDK.
python -m venv .venv
source .venv/bin/activate # Windows: .venvScriptsactivate
python -m pip install crawlbase beautifulsoup4 markdownify openai pydantic
The perplexityai package is the official SDK and supports synchronous and asynchronous clients, Search API calls, chat completions, and typed responses. This tutorial uses the OpenAI-compatible endpoint documented for Perplexity Agent workflows, so the openai client keeps the request explicit and easy to inspect.
Keep credentials out of source control
Create environment variables in your shell, secret manager, or deployment platform:
export CRAWLBASE_TOKEN="your-crawlbase-token"
export PERPLEXITY_API_KEY="your-perplexity-api-key"
Do not print these values, commit a .env file containing real keys, or put them in a prompt. Rotate a key immediately if it appears in logs or a repository.
Complete Python implementation
The following script fetches a page, extracts a likely content region, converts it to Markdown, asks Perplexity for a fixed JSON object, and validates the result. Replace the example URL and selectors with those appropriate for the site you are allowed to access.
Rank #2
import json
import os
import re
from typing import Optional
import requests
from bs4 import BeautifulSoup
from crawlbase import CrawlingAPI
from markdownify import markdownify as html_to_markdown
from openai import OpenAI
from pydantic import BaseModel, ConfigDict, Field, ValidationError
TARGET_URL = "https://example.com/product"
class Product(BaseModel):
model_config = ConfigDict(extra="forbid")
name: Optional[str] = None
description: Optional[str] = None
price: Optional[str] = None
currency: Optional[str] = None
specifications: list[str] = Field(default_factory=list)
def fetch_html(url: str) -> str:
"""Use Crawlbase's normal token for server-rendered HTML."""
api = CrawlingAPI({"token": os.environ["CRAWLBASE_TOKEN"]})
response = api.get(url)
# Crawlbase responses expose the fetched body as response.body.
body = response.body
if isinstance(body, bytes):
body = body.decode("utf-8", errors="replace")
if not body or len(body.strip()) < 200:
raise RuntimeError("The crawler returned an empty or suspiciously small body")
return body
def trim_html(html: str) -> str:
soup = BeautifulSoup(html, "html.parser")
for tag in soup(["script", "style", "noscript", "svg", "nav", "footer", "form"]):
tag.decompose()
# Prefer semantic content, then common article/product containers.
content = (
soup.find("main")
or soup.find("article")
or soup.select_one(".product, .product-page, .entry-content, .post-content")
or soup.body
or soup
)
return str(content)
def to_markdown(html_fragment: str) -> str:
markdown = html_to_markdown(html_fragment, heading_style="ATX")
markdown = re.sub(r"n{3,}", "nn", markdown)
return markdown.strip()
def extract_product(markdown: str) -> Product:
client = OpenAI(
api_key=os.environ["PERPLEXITY_API_KEY"],
base_url="https://api.perplexity.ai/v1",
)
instruction = """Extract the product fields from the supplied page text.nReturn only a JSON object with exactly these keys: name, description, price, currency, specifications.nUse null when a scalar field is not present and [] when no specifications are present.nNever infer a price, currency, name, or specification that is absent from the text.nKeep price as the displayed string and preserve the page's currency notation.nnPAGE TEXT:n""" + markdown
completion = client.chat.completions.create(
model="sonar",
messages=[
{"role": "system", "content": "You extract facts conservatively from supplied text."},
{"role": "user", "content": instruction},
],
temperature=0,
response_format={"type": "json_schema", "json_schema": {
"name": "product",
"schema": {
"type": "object",
"additionalProperties": False,
"properties": {
"name": {"type": ["string", "null"]},
"description": {"type": ["string", "null"]},
"price": {"type": ["string", "null"]},
"currency": {"type": ["string", "null"]},
"specifications": {"type": "array", "items": {"type": "string"}},
},
"required": ["name", "description", "price", "currency", "specifications"],
},
}},
)
raw = completion.choices[0].message.content
return Product.model_validate(json.loads(raw))
def main() -> None:
html = fetch_html(TARGET_URL)
markdown = to_markdown(trim_html(html))
if not markdown:
raise RuntimeError("No usable text remained after HTML trimming")
product = extract_product(markdown)
print(product.model_dump_json(indent=2))
if __name__ == "__main__":
main()
The JSON Schema makes the response shape predictable, while Pydantic rejects unexpected keys or wrong types. If your installed SDK version exposes a different structured-output parameter, follow its current documentation and retain the same schema and post-response validation. Never assume that a successful HTTP response means the extracted facts are correct.
Choosing the crawler token
Normal token for server-rendered HTML
Use Crawlbase’s normal token when the useful content is present in the initial HTML response. This is common for traditional server-rendered pages. Confirm by saving the response and searching it for a visible heading or distinctive product string.
Recommended Free Tools
JavaScript token for client-rendered pages
Modern applications may return an almost empty shell and populate the page only after JavaScript runs. Crawlbase describes its JavaScript token for those client-rendered pages. If your saved HTML contains the root element but not the visible text, switch to the JavaScript-capable token before changing the extraction prompt. Rendering adds latency and operational cost, so do not enable it for every URL without evidence that it is needed.
Even browser rendering cannot guarantee access to a page protected by a CAPTCHA, login requirement, geofence, or terms that prohibit automated collection. Obtain permission, authenticate through supported mechanisms, or skip the URL.
Making extraction accurate and repeatable
Trim before you prompt
Raw HTML includes navigation, scripts, tracking attributes, repeated links, and hidden elements. Selecting main, article, or a site-specific container and converting it to Markdown lowers token use and removes distractions. Keep headings, lists, tables, and labels because they often carry the relationship between a field and its value.
State the missing-data policy
Tell the model exactly what to do when a field is absent: null for a scalar and an empty array for a list. Explicitly prohibit inference. For prices, preserve the displayed string and currency rather than converting it unless your application has a separate, documented conversion step.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use selectors for stable layouts, schemas for variable ones
Fixed CSS selectors are fast and deterministic when every page follows the same template. Schema-directed extraction is more useful when labels, ordering, or wording vary across pages. A practical hybrid is to select the product container with CSS, then let Perplexity map its text to a validated schema.
Control input size
Set a maximum character or token budget after trimming. If a page contains reviews, recommendations, and comments unrelated to the target fields, remove those sections before the model call. If truncation is unavoidable, preserve the beginning and the sections where the required fields normally appear, and record that truncation in your job metadata.
Perplexity Agent and Search options
Perplexity’s current API Platform separates Agent and Search capabilities. Agent workflows include web search, URL fetching, and reasoning controls; Search provides ranked results, domain filtering, multi-query search, and content extraction. The Agent API announcement documents web_search, fetch_url, JSON Schema structured outputs, and the OpenAI-compatible base URL used above.
Those capabilities can complement a custom fetcher, but they do not change the central rule of this tutorial: when your program fetches a page and sends its text, the model interprets the supplied bytes. If you instead ask an Agent workflow to fetch a URL, document that as a different architecture with its own access, freshness, and failure behavior.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Reliability, retries, and operational controls
Retry only transient failures
- Retry network resets, gateway timeouts, and HTTP 5xx responses with exponential backoff and a maximum attempt count.
- Do not blindly retry 401/403 responses; check credentials, permissions, robots or site policy, and challenge pages.
- Do not retry malformed model JSON indefinitely. Log the response safely, tighten the schema or prompt, and cap repair attempts.
Record stage-level metadata
For each URL, store the fetch status, final URL, rendering mode, response length, extraction selector, Markdown length, model identifier, latency, and validation result. Hashing the cleaned text lets you detect unchanged pages and avoid unnecessary interpretation calls. Keep page content and credentials in separate access-controlled stores.
Respect site rules and rate limits
Check the target site’s terms, robots directives, authentication requirements, and applicable law before collecting. Throttle requests per host, honor provider limits, and use a queue for large jobs. Cache fetched content only for a period that fits the page’s update frequency and your legal obligations.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| HTML is an empty app shell | Content is rendered client-side | Use Crawlbase’s JavaScript token, then verify visible text exists before extraction. |
| Markdown contains menus but no article | Wrong container selector or a consent wall | Inspect the saved HTML, choose a site-specific selector, and handle the consent flow lawfully. |
| Model invents a value | Prompt permits inference or context is missing | Require null/empty values, prohibit inference, preserve labels, and validate against source text. |
| JSON parse or schema error | Free-form output, unsupported response-format option, or truncated response | Use the documented structured-output interface, lower input size, retry once, and reject rather than silently accepting invalid data. |
| 401 from Perplexity | Missing, revoked, or mis-scoped API key | Check PERPLEXITY_API_KEY, endpoint, and account access without logging the secret. |
| 429 or repeated timeouts | Rate limit, overloaded provider, or too much input | Back off, reduce concurrency and page size, and queue work. |
Or skip the browser setup
If your goal is a clean image or PDF rather than text extraction, ScreenshotNeo provides a one-request website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
For a screenshot, use the documented endpoint and options:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallcurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for full-page and element capture, 12 device presets plus custom viewports, retina scale, dark mode, PDF controls, custom CSS and JavaScript, clicks, waits, blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data, and the OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account.
Best Value
When this pattern is the right choice
- Choose fetch-then-interpret when you need auditable control over the exact page bytes sent to the model.
- Use a normal crawler token for static HTML and a JavaScript token when the initial response lacks the content.
- Prefer deterministic selectors for uniform templates and schema-constrained Perplexity extraction for varied layouts.
- Validate every response and retain enough metadata to reproduce a failed job.
Frequently Asked Questions
Does Perplexity scrape the target website by itself in this Python workflow?
No. Your crawler fetches the page and your application supplies cleaned text; Perplexity interprets that text. Agent workflows can provide URL-fetching tools, but that is a different architecture.
When should I use a JavaScript-rendering crawler token?
Use it when the initial HTML is an app shell without the visible content and the page populates after JavaScript executes. Verify the saved response before switching.
Can I extract fields without asking for free-form prose?
Yes. Request a JSON Schema response, require every key, represent missing scalars as null and missing lists as [], then validate the parsed object in Python.
Is Markdown conversion mandatory?
No, but trimming HTML and converting the relevant fragment to Markdown usually removes navigation and markup noise, reducing irrelevant input and improving consistency.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

