Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To get started with Crawlee for Python, install Python 3.10 or newer, install the crawlee package (plus the extra for your chosen crawler), create a crawler with a request handler, run it against a URL, and read the JSON dataset saved under ./storage/datasets/default/. Use an HTTP crawler when the HTML already contains the data; use PlaywrightCrawler when JavaScript or browser interaction is required.

What Crawlee does

Crawlee is a Python framework that organizes web-crawling work around requests, queues, handlers, retries, concurrency, sessions and storage. You describe what should happen when a page is reached; Crawlee handles the repeated request-processing cycle. A request identifies a URL. A request queue holds starting URLs and any additional URLs discovered during the crawl. A request handler receives the current request and crawler-specific page data, extracts information, and can enqueue more work.

The official introductory documentation describes the workflow as visiting a page, opening it, doing work, saving results, continuing to another page, and repeating until the job is complete. That model is useful for a beginner because it separates navigation and orchestration from extraction logic.

Prerequisites and installation

Check Python first

The current setup documentation requires Python 3.10 or newer. Verify the interpreter that will run your crawler:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python --version

If your system has multiple Python installations, use the same executable for both the version check and package installation.

Install the core package

python -m pip install crawlee
python -c 'import crawlee; print(crawlee.__version__)'

The core package is enough to learn the request-handler workflow, but each crawler class has an optional dependency set:

Need Install Important limitation or requirement
HTML returned directly by HTTP python -m pip install "crawlee[beautifulsoup]" Does not execute client-side JavaScript.
HTTP HTML with CSS-selector extraction python -m pip install "crawlee[parsel]" Does not execute client-side JavaScript.
Rendered pages or browser interaction python -m pip install "crawlee[playwright]"
playwright install
Requires Playwright browser dependencies; use a browser context for rendered content.

You can install all extras if a project genuinely needs several integrations, but selecting only the extra for your first crawler keeps the environment smaller and makes failures easier to diagnose.

Optional project scaffolding

The documented CLI can create a prepared project:

uvx 'crawlee[cli]' create my-crawler

When Crawlee is already installed, the equivalent command is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
crawlee create my_crawler

Activate the generated environment as instructed by the scaffold, then run its module with:

python -m my_crawler

Which Crawlee crawler should you use?

Choose according to where the data exists, not according to the site’s visual appearance. Open the page source or inspect the HTTP response when possible. If the desired text is already in the response HTML, a browser adds setup and runtime overhead without solving a problem. If JavaScript creates the content after load, an HTTP parser will see only the pre-rendered shell.

BeautifulSoupCrawler

Start with BeautifulSoupCrawler for ordinary server-rendered HTML and familiar parser methods. It is the straightforward, fast-to-start option for titles, headings, links and metadata that arrive in the response. It does not run client-side JavaScript.

ParselCrawler

Use ParselCrawler when CSS-selector-oriented extraction is the clearest fit. It also fetches over HTTP and therefore has the same JavaScript limitation. Parsel can make selectors concise when the page has repeated structures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PlaywrightCrawler

Use PlaywrightCrawler when the target requires JavaScript execution, interaction, scrolling, a browser session or other rendered-page behavior. Install the Crawlee Playwright extra and run playwright install. The quick-start documentation describes Chromium, Firefox and WebKit support. During development, headful mode can make navigation visible so you can see what the browser is doing.

Question HTTP crawler Playwright crawler
Does it execute JavaScript? No Yes, in a controlled browser
Setup and runtime Smaller dependency footprint; no browser launch Browser binaries and a heavier runtime are required
Extraction style BeautifulSoup parser or Parsel selectors Rendered page and browser context
Best first use Server-rendered HTML Client-rendered or interactive pages

The main crawler classes share a common interface, so starting with an HTTP crawler does not lock the project permanently to that fetching method.

Make your first Crawlee crawler

Minimal BeautifulSoup example

Create main.py:

import asyncio

from crawlee.beautifulsoup_crawler import BeautifulSoupCrawler, BeautifulSoupCrawlingContext


async def main() -> None:
    crawler = BeautifulSoupCrawler()

    @crawler.router.default_handler
    async def request_handler(context: BeautifulSoupCrawlingContext) -> None:
        title = context.soup.title.get_text(strip=True) if context.soup.title else None
        context.log.info(f"URL: {context.request.url} | title: {title}")
        await context.push_data({
            "url": context.request.url,
            "title": title,
        })

    await crawler.run(["https://example.com"])


if __name__ == "__main__":
    asyncio.run(main())

Run it with:

python main.py

crawler.run([...]) is the short form. Crawlee still uses a request queue internally; the list supplies its starting requests. The handler runs once for each successfully processed page and writes a record with the URL and title.

Explicit queue form

An explicit queue is useful when you want to add requests before the crawl starts or pass the queue to other code:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio

from crawlee.beautifulsoup_crawler import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
from crawlee import Request
from crawlee.request_queue import RequestQueue


async def main() -> None:
    queue = await RequestQueue.open()
    await queue.add_request(Request.from_url("https://example.com"))
    crawler = BeautifulSoupCrawler(request_queue=queue)

    @crawler.router.default_handler
    async def request_handler(context: BeautifulSoupCrawlingContext) -> None:
        title = context.soup.title.get_text(strip=True) if context.soup.title else None
        await context.push_data({"url": context.request.url, "title": title})

    await crawler.run()


if __name__ == "__main__":
    asyncio.run(main())

Use the exact imports exposed by the Crawlee version installed in your environment; package APIs can change between releases.

Enqueue links to turn one page into a crawl

Once the one-page example works, discover links and add them to the queue. Keep a domain or URL policy in your handler so a test does not unexpectedly expand into an unrestricted crawl:

from urllib.parse import urljoin, urlparse

# inside request_handler, after saving the current record
base_host = urlparse(context.request.url).netloc
for link in context.soup.select("a[href]"):
    next_url = urljoin(context.request.url, link["href"])
    if urlparse(next_url).netloc == base_host:
        await context.add_requests([next_url])

Before using this pattern in production, add limits for depth, page count, URL patterns and duplicate handling. Respect the target site’s terms and applicable rules.

Where does Crawlee save the results?

By default, Crawlee writes JSON dataset files below ./storage/datasets/default/. In the minimal example, open that directory after the process exits and inspect the generated JSON records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To move storage, set CRAWLEE_STORAGE_DIR before starting the program:

CRAWLEE_STORAGE_DIR=./crawl-data python main.py

On Windows PowerShell:

$env:CRAWLEE_STORAGE_DIR=".crawl-data"
python main.py

Keeping storage outside the source tree can simplify cleanup and separate crawl output from code. Dataset records are the durable hand-off from the handler to later processing, such as loading into a database or exporting another format.

Using Playwright when JavaScript is required

Install the browser-enabled extra and browser binaries:

python -m pip install "crawlee[playwright]"
playwright install

A browser crawler follows the same conceptual pattern: create the crawler, register a handler, run starting URLs, and push structured data. The handler receives rendered-page context instead of a BeautifulSoup object. Use the page API to read the title or other DOM content after scripts have run. During debugging, configure headful operation so the browser window is visible; switch back to headless operation for unattended runs.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not choose Playwright merely because a page contains some JavaScript. First confirm that the field you need is absent from the HTTP response. Browser automation costs more setup and resources, while an HTTP crawler is usually the simpler choice for static content.

Common errors and fixes

ModuleNotFoundError for Crawlee or a parser

The package was installed into a different interpreter or the optional extra is missing. Run installation through the interpreter used to launch the script: python -m pip install crawlee or the relevant extra, then rerun the version check.

Playwright browser executable is missing

Installing the Python extra does not always install browser binaries. Run playwright install in the same environment. In restricted systems, verify that the required browser download is permitted.

The title or content is empty

You may be using an HTTP crawler against a JavaScript-rendered page, selecting the wrong element, or receiving an error page. Inspect the response HTML. If the data appears only after scripts execute, move to PlaywrightCrawler; otherwise correct the parser or selector.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

No files appear in the expected directory

Check the process log for an exception and confirm that CRAWLEE_STORAGE_DIR is not pointing elsewhere. The default location is relative to the process’s working directory, not necessarily the directory containing main.py.

The crawl expands unexpectedly

Link discovery can enqueue navigation, fragments, logout links or external hosts. Restrict hosts, URL patterns, depth and total requests before enabling broad discovery.

A request repeatedly fails

Check the URL, DNS and network access first. Crawlee’s orchestration includes retries and request processing, but retries cannot fix a permanently invalid URL, a blocked environment or a page that requires browser rendering. Log the failing URL and response details, then choose the appropriate crawler or adjust the request policy.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability and project growth

Start with one URL and a small handler. Add concurrency only after extraction is correct and the target can tolerate the request rate. Crawlee provides orchestration for retries, sessions, concurrency and storage, so you do not need to rebuild those mechanisms around every script.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For repeatable runs, make records explicit and stable: include the source URL, extraction timestamp if your application needs it, and only the fields downstream consumers require. Keep the storage directory configurable. Test selectors against representative pages, including missing fields and error responses.

When built-in components are insufficient, the extension guidance describes points for custom parsers, HTTP backends, databases and browser integrations. Treat those as a second-stage design decision. A custom extension is justified when a project requirement cannot be met by the standard crawler, handler and storage flow.

Or skip the browser setup

If your immediate goal is a clean image or PDF of a page rather than a data crawl, ScreenshotNeo provides a single screenshot API request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response reports the page verdict and billing status in headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

See the ScreenshotNeo API documentation for parameters and response details. cURL:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently asked questions

Can Crawlee crawl a site without JavaScript?

Yes. BeautifulSoupCrawler and ParselCrawler fetch and parse HTTP HTML without launching a browser. They are appropriate when the required content is present in the response.

Can I change crawler type later?

Usually. Crawlee’s main crawler classes share an interface, so the handler and surrounding workflow can often remain similar while the fetching component changes.

Is the CLI required?

No. You can install the package and write a Python module directly. The CLI is an optional way to scaffold a prepared project.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I learn after the first page?

Learn controlled link discovery, then add limits and structured records. After that, study sessions, retries, concurrency and custom extensions as the project requires them.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.