Recommended Free Tools
To get started with Crawlee for Python, install Python 3.10 or newer, install the crawlee package (plus the extra for your chosen crawler), create a crawler with a request handler, run it against a URL, and read the JSON dataset saved under ./storage/datasets/default/. Use an HTTP crawler when the HTML already contains the data; use PlaywrightCrawler when JavaScript or browser interaction is required.
What Crawlee does
Crawlee is a Python framework that organizes web-crawling work around requests, queues, handlers, retries, concurrency, sessions and storage. You describe what should happen when a page is reached; Crawlee handles the repeated request-processing cycle. A request identifies a URL. A request queue holds starting URLs and any additional URLs discovered during the crawl. A request handler receives the current request and crawler-specific page data, extracts information, and can enqueue more work.
The official introductory documentation describes the workflow as visiting a page, opening it, doing work, saving results, continuing to another page, and repeating until the job is complete. That model is useful for a beginner because it separates navigation and orchestration from extraction logic.
Prerequisites and installation
Check Python first
The current setup documentation requires Python 3.10 or newer. Verify the interpreter that will run your crawler:
#1 Best Overall
python --version
If your system has multiple Python installations, use the same executable for both the version check and package installation.
Install the core package
python -m pip install crawlee
python -c 'import crawlee; print(crawlee.__version__)'
The core package is enough to learn the request-handler workflow, but each crawler class has an optional dependency set:
| Need | Install | Important limitation or requirement |
|---|---|---|
| HTML returned directly by HTTP | python -m pip install "crawlee[beautifulsoup]" |
Does not execute client-side JavaScript. |
| HTTP HTML with CSS-selector extraction | python -m pip install "crawlee[parsel]" |
Does not execute client-side JavaScript. |
| Rendered pages or browser interaction | python -m pip install "crawlee[playwright]"playwright install |
Requires Playwright browser dependencies; use a browser context for rendered content. |
You can install all extras if a project genuinely needs several integrations, but selecting only the extra for your first crawler keeps the environment smaller and makes failures easier to diagnose.
Optional project scaffolding
The documented CLI can create a prepared project:
uvx 'crawlee[cli]' create my-crawler
When Crawlee is already installed, the equivalent command is:
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →crawlee create my_crawler
Activate the generated environment as instructed by the scaffold, then run its module with:
python -m my_crawler
Which Crawlee crawler should you use?
Choose according to where the data exists, not according to the site’s visual appearance. Open the page source or inspect the HTTP response when possible. If the desired text is already in the response HTML, a browser adds setup and runtime overhead without solving a problem. If JavaScript creates the content after load, an HTTP parser will see only the pre-rendered shell.
BeautifulSoupCrawler
Start with BeautifulSoupCrawler for ordinary server-rendered HTML and familiar parser methods. It is the straightforward, fast-to-start option for titles, headings, links and metadata that arrive in the response. It does not run client-side JavaScript.
Rank #2
ParselCrawler
Use ParselCrawler when CSS-selector-oriented extraction is the clearest fit. It also fetches over HTTP and therefore has the same JavaScript limitation. Parsel can make selectors concise when the page has repeated structures.
PlaywrightCrawler
Use PlaywrightCrawler when the target requires JavaScript execution, interaction, scrolling, a browser session or other rendered-page behavior. Install the Crawlee Playwright extra and run playwright install. The quick-start documentation describes Chromium, Firefox and WebKit support. During development, headful mode can make navigation visible so you can see what the browser is doing.
| Question | HTTP crawler | Playwright crawler |
|---|---|---|
| Does it execute JavaScript? | No | Yes, in a controlled browser |
| Setup and runtime | Smaller dependency footprint; no browser launch | Browser binaries and a heavier runtime are required |
| Extraction style | BeautifulSoup parser or Parsel selectors | Rendered page and browser context |
| Best first use | Server-rendered HTML | Client-rendered or interactive pages |
The main crawler classes share a common interface, so starting with an HTTP crawler does not lock the project permanently to that fetching method.
Make your first Crawlee crawler
Minimal BeautifulSoup example
Create main.py:
import asyncio
from crawlee.beautifulsoup_crawler import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
async def main() -> None:
crawler = BeautifulSoupCrawler()
@crawler.router.default_handler
async def request_handler(context: BeautifulSoupCrawlingContext) -> None:
title = context.soup.title.get_text(strip=True) if context.soup.title else None
context.log.info(f"URL: {context.request.url} | title: {title}")
await context.push_data({
"url": context.request.url,
"title": title,
})
await crawler.run(["https://example.com"])
if __name__ == "__main__":
asyncio.run(main())
Run it with:
python main.py
crawler.run([...]) is the short form. Crawlee still uses a request queue internally; the list supplies its starting requests. The handler runs once for each successfully processed page and writes a record with the URL and title.
Explicit queue form
An explicit queue is useful when you want to add requests before the crawl starts or pass the queue to other code:
import asyncio
from crawlee.beautifulsoup_crawler import BeautifulSoupCrawler, BeautifulSoupCrawlingContext
from crawlee import Request
from crawlee.request_queue import RequestQueue
async def main() -> None:
queue = await RequestQueue.open()
await queue.add_request(Request.from_url("https://example.com"))
crawler = BeautifulSoupCrawler(request_queue=queue)
@crawler.router.default_handler
async def request_handler(context: BeautifulSoupCrawlingContext) -> None:
title = context.soup.title.get_text(strip=True) if context.soup.title else None
await context.push_data({"url": context.request.url, "title": title})
await crawler.run()
if __name__ == "__main__":
asyncio.run(main())
Use the exact imports exposed by the Crawlee version installed in your environment; package APIs can change between releases.
Enqueue links to turn one page into a crawl
Once the one-page example works, discover links and add them to the queue. Keep a domain or URL policy in your handler so a test does not unexpectedly expand into an unrestricted crawl:
from urllib.parse import urljoin, urlparse
# inside request_handler, after saving the current record
base_host = urlparse(context.request.url).netloc
for link in context.soup.select("a[href]"):
next_url = urljoin(context.request.url, link["href"])
if urlparse(next_url).netloc == base_host:
await context.add_requests([next_url])
Before using this pattern in production, add limits for depth, page count, URL patterns and duplicate handling. Respect the target site’s terms and applicable rules.
Where does Crawlee save the results?
By default, Crawlee writes JSON dataset files below ./storage/datasets/default/. In the minimal example, open that directory after the process exits and inspect the generated JSON records.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →To move storage, set CRAWLEE_STORAGE_DIR before starting the program:
CRAWLEE_STORAGE_DIR=./crawl-data python main.py
On Windows PowerShell:
$env:CRAWLEE_STORAGE_DIR=".crawl-data"
python main.py
Keeping storage outside the source tree can simplify cleanup and separate crawl output from code. Dataset records are the durable hand-off from the handler to later processing, such as loading into a database or exporting another format.
Using Playwright when JavaScript is required
Install the browser-enabled extra and browser binaries:
python -m pip install "crawlee[playwright]"
playwright install
A browser crawler follows the same conceptual pattern: create the crawler, register a handler, run starting URLs, and push structured data. The handler receives rendered-page context instead of a BeautifulSoup object. Use the page API to read the title or other DOM content after scripts have run. During debugging, configure headful operation so the browser window is visible; switch back to headless operation for unattended runs.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Do not choose Playwright merely because a page contains some JavaScript. First confirm that the field you need is absent from the HTTP response. Browser automation costs more setup and resources, while an HTTP crawler is usually the simpler choice for static content.
Common errors and fixes
ModuleNotFoundError for Crawlee or a parser
The package was installed into a different interpreter or the optional extra is missing. Run installation through the interpreter used to launch the script: python -m pip install crawlee or the relevant extra, then rerun the version check.
Playwright browser executable is missing
Installing the Python extra does not always install browser binaries. Run playwright install in the same environment. In restricted systems, verify that the required browser download is permitted.
The title or content is empty
You may be using an HTTP crawler against a JavaScript-rendered page, selecting the wrong element, or receiving an error page. Inspect the response HTML. If the data appears only after scripts execute, move to PlaywrightCrawler; otherwise correct the parser or selector.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteNo files appear in the expected directory
Check the process log for an exception and confirm that CRAWLEE_STORAGE_DIR is not pointing elsewhere. The default location is relative to the process’s working directory, not necessarily the directory containing main.py.
The crawl expands unexpectedly
Link discovery can enqueue navigation, fragments, logout links or external hosts. Restrict hosts, URL patterns, depth and total requests before enabling broad discovery.
A request repeatedly fails
Check the URL, DNS and network access first. Crawlee’s orchestration includes retries and request processing, but retries cannot fix a permanently invalid URL, a blocked environment or a page that requires browser rendering. Log the failing URL and response details, then choose the appropriate crawler or adjust the request policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Performance, reliability and project growth
Start with one URL and a small handler. Add concurrency only after extraction is correct and the target can tolerate the request rate. Crawlee provides orchestration for retries, sessions, concurrency and storage, so you do not need to rebuild those mechanisms around every script.
Best Value
For repeatable runs, make records explicit and stable: include the source URL, extraction timestamp if your application needs it, and only the fields downstream consumers require. Keep the storage directory configurable. Test selectors against representative pages, including missing fields and error responses.
When built-in components are insufficient, the extension guidance describes points for custom parsers, HTTP backends, databases and browser integrations. Treat those as a second-stage design decision. A custom extension is justified when a project requirement cannot be met by the standard crawler, handler and storage flow.
Or skip the browser setup
If your immediate goal is a clean image or PDF of a page rather than a data crawl, ScreenshotNeo provides a single screenshot API request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response reports the page verdict and billing status in headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
See the ScreenshotNeo API documentation for parameters and response details. cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Frequently asked questions
Can Crawlee crawl a site without JavaScript?
Yes. BeautifulSoupCrawler and ParselCrawler fetch and parse HTTP HTML without launching a browser. They are appropriate when the required content is present in the response.
Can I change crawler type later?
Usually. Crawlee’s main crawler classes share an interface, so the handler and surrounding workflow can often remain similar while the fetching component changes.
Is the CLI required?
No. You can install the package and write a Python module directly. The CLI is an optional way to scaffold a prepared project.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesWhat should I learn after the first page?
Learn controlled link discovery, then add limits and structured records. After that, study sessions, retries, concurrency and custom extensions as the project requires them.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

