What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To build a first crawler with Crawlee for Python, install the crawler integration that fits the page, create a crawler with a request handler, extract a field, and save it with context.push_data(). You need Python 3.10 or newer. For a static or server-rendered page, start with BeautifulSoupCrawler; use PlaywrightCrawler when the content appears only after client-side JavaScript runs.
What this tutorial builds
The example below visits a page, reads its title, and writes one record to Crawlee’s default dataset storage. It uses BeautifulSoupCrawler, which fetches pages over HTTP and parses their returned HTML rather than opening a browser. The example uses https://example.com as a simple starting URL; change it to a page you are permitted to crawl and whose title is present in the returned HTML.
Crawlee’s basic pattern is reusable: a request tells the crawler where to go, a request handler says what to do with the response, and crawler.run() starts processing. You can add links to the queue later, after confirming that the first page produces the data you expect. See the official Crawlee Python quick start and first crawler guide.
Check prerequisites and install Crawlee
Verify Python and pip
The official quick start requires Python 3.10 or newer. Check which interpreter and pip your shell will use:
#1 Best Overall
python --version
python -m pip --version
If your system uses python3 instead of python, substitute that command consistently. Using python -m pip ties installation to the interpreter invoked by python, which helps avoid installing into a different Python environment.
Install the integration you plan to use
Crawlee is distributed as the crawlee Python package. Its crawler integrations are optional extras, so installing the base package does not mean every parser and browser integration is present. For this tutorial’s BeautifulSoup crawler, install:
python -m pip install 'crawlee[beautifulsoup]'
Other documented install choices include python -m pip install crawlee for the core package, python -m pip install 'crawlee[parsel]' for Parsel, and python -m pip install 'crawlee[playwright]' for Playwright. The official setup guide covers installing multiple extras and project setup. The optional extras matter: choose the integration your code imports rather than assuming it came with a minimal install.
Rank #2
Choose the crawler by how the page gets its content
| Crawler | Best fit | Parsing or rendering | Setup and resource trade-off |
|---|---|---|---|
BeautifulSoupCrawler |
Static or server-rendered HTML and a beginner-friendly parsing API | Fetches returned HTML over HTTP; does not execute client-side JavaScript | No browser dependency; generally lighter than browser rendering |
ParselCrawler |
Pages where CSS or XPath selectors suit the extraction task, or where you already know Parsel | HTTP response parsing with CSS/XPath selection; Parsel also supports regex | No browser dependency; the HTTP crawler guide notes performance characteristics, but does not establish a universal speed ranking for every target |
PlaywrightCrawler |
Content that depends on JavaScript execution or browser behavior | Controls a browser and can inspect the rendered page | Install the Playwright extra and browser dependencies; typically takes more time and resources than HTTP fetching |
The decision is about page behavior, not which crawler is universally best. If the title, product details, or other target data are in the HTML returned by the server, an HTTP crawler is often sufficient. If the page fills that data in only after JavaScript runs, an HTTP response may not contain it; switch to PlaywrightCrawler. The official guides explain HTTP crawlers and the Playwright crawler.
Build and run a first crawler
Save this as main.py
This example follows the quick-start pattern: register a default handler, extract a title from the response’s parsed soup, save a record, then run the crawler on a starting URL. The URL is intentionally a small, stable example target; replace it with the page relevant to your task.
import asyncio
from crawlee.crawlers import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
async def main() -> None:
crawler = BeautifulSoupCrawler()
@crawler.router.default_handler
async def handle_page(context: BeautifulSoupCrawlingContext) -> None:
title = context.soup.title.string if context.soup.title else None
await context.push_data({
"url": context.request.url,
"title": title.strip() if title else None,
})
await crawler.run(["https://example.com"])
if __name__ == "__main__":
asyncio.run(main())
The handler receives a crawling context. Here, context.request.url records the URL being processed, while context.soup is the parsed HTML. The title check matters because some pages have no <title> element; the script stores null for that field rather than failing on a missing title. context.push_data() persists the record in Crawlee’s dataset storage.
Run the script and inspect its output
python main.py
The quick start documents JSON dataset files under ./storage/datasets/default/, relative to the current working directory. If the output is not where expected, check the directory from which you launched the script. The CRAWLEE_STORAGE_DIR environment variable can change the storage directory; set it before running the program if you want storage elsewhere. The quick start also shows link enqueuing and the await context.enqueue_links() pattern, which you can add when you intentionally want a crawl to follow links rather than process only the starting URL.
Use Parsel or Playwright when the page calls for it
Parsel: CSS and XPath-oriented extraction
Choose ParselCrawler when its selector interface fits your extraction work. Install the integration with python -m pip install 'crawlee[parsel]', then follow the Parsel-specific examples in the official examples index. It remains an HTTP crawler: selectors can locate elements in returned HTML, but they cannot reveal content that exists only after client-side JavaScript executes.
Playwright: browser-rendered content
For a JavaScript-dependent page, install both Crawlee’s Playwright extra and the browser dependencies:
python -m pip install 'crawlee[playwright]'
playwright install
The official quick start’s Playwright pattern reads a title with await context.page.title(), rather than reading context.soup.title. Browser setup is an additional prerequisite, and browser-based work typically uses more time and resources than an HTTP request. Use it when rendering is needed, not as a default upgrade for every page. Consult the Playwright crawler guide for the current crawler API and examples.
Or skip the browser setup
Crawlee is for crawling and extracting page data; ScreenshotNeo is an alternative when the task is to obtain a screenshot or PDF rather than build a crawler. One GET request returns an image or PDF. For example, this cURL request saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners are accepted and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; those cleanup steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card.
Generate a starter project or deploy later
If you prefer a generated project skeleton to a single script, the setup guide documents these Crawlee CLI options:
Best Value
uvx 'crawlee[cli]' create my-crawler
# or, if the CLI is installed:
crawlee create my_crawler
The generated project’s setup instructions explain how to run it as a Python module. This is optional: for learning the request-handler pattern, one script is easier to inspect. The official Crawlee for Python project page also describes converting a project into an Apify Actor and deploying it there; consider that path if you need hosted execution rather than a local process.
Troubleshooting common first-run problems
- Python is older than 3.10: install or select a supported Python interpreter, then repeat the version check and package installation with that interpreter.
ModuleNotFoundErrorfor Crawlee or an integration: activate the environment used to run the script and install the matching package or extra with that environment’spython -m pip. For this example, usepython -m pip install 'crawlee[beautifulsoup]'.- Playwright launches without a browser: install the Playwright extra and run
playwright installso browser dependencies are available. - The saved title is
nullor missing: inspect whether the response HTML contains a title. If the desired content is injected by JavaScript, an HTTP crawler cannot execute that code; use Playwright and inspect the rendered page instead. - No dataset file appears where expected: look under
./storage/datasets/default/relative to the process’s working directory, and check whetherCRAWLEE_STORAGE_DIRpoints storage to another location. - The crawler sees an unexpected page or no useful content: confirm the starting URL and what the page returns to an HTTP client. A site may require a browser-rendered view or may limit automated requests; use the crawler type that matches the page and follow the site’s access rules.
Make the example a useful crawler
Once the one-page script saves the expected title, change one thing at a time: choose fields that exist on your target, test extraction against its actual markup, and add link enqueuing only if the task requires visiting additional pages. Use an HTTP crawler for pages whose needed data is in the response, and switch to Playwright when rendering is a real requirement. For custom datasets and crawler-specific patterns, the official Crawlee examples index links to dataset storage, BeautifulSoup, Parsel, Playwright, and adaptive crawling examples.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Frequently Asked Questions
Does Crawlee for Python work on Windows, macOS, and Linux?
The cited quick-start and setup pages establish the Python version requirement and installation methods, but do not provide a platform-specific compatibility matrix here. Check the current Crawlee setup documentation for platform details relevant to your environment.
Can a Crawlee HTTP crawler scrape a page that uses JavaScript?
It can scrape data present in the returned HTML, but it does not execute client-side JavaScript. If the target data appears only after browser-side rendering, use a browser crawler such as PlaywrightCrawler.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

