Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Crawlee is an open-source library for building web scrapers and browser automation workflows in JavaScript and Python. Choose CheerioCrawler when the information is already in fetched HTML; choose PlaywrightCrawler or PuppeteerCrawler when the page depends on JavaScript or browser behavior. Crawlee provides the crawl and data-handling framework, while Playwright and Puppeteer are separate dependencies. This guide focuses on the JavaScript implementation and shows where each approach fits.

What is Crawlee?

Crawlee is an open-source library for scraping websites and automating browser tasks. Its JavaScript and Python implementations provide crawler patterns and tools for processing requests and storing results. The project repository states that Crawlee is licensed under the Apache License 2.0: Crawlee README.

This is a library, not a hosted scraping service that you must use through one vendor. You can run a crawler locally or deploy it on cloud infrastructure. Apify is one optional deployment path, not a requirement for using Crawlee.

Should I use CheerioCrawler or PlaywrightCrawler?

Start by checking where the page’s data comes from. If it is present in the HTML returned by an ordinary HTTP request, parse that HTML with CheerioCrawler. If the page needs client-side JavaScript, browser events, or other browser behavior to expose the content, use a browser crawler such as PlaywrightCrawler. If you already use Puppeteer or prefer its workflow, Crawlee also provides PuppeteerCrawler.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need Starting point What to know
Read server-rendered or otherwise directly available HTML CheerioCrawler Uses HTTP and HTML parsing; it does not render client-side JavaScript.
Load a JavaScript-driven page or interact with browser controls PlaywrightCrawler Uses Playwright to control a browser. Install Playwright separately.
Keep an existing Puppeteer-based workflow PuppeteerCrawler Uses Puppeteer to control a browser. Install Puppeteer separately.

Crawlee describes CheerioCrawler as fast and efficient, but the official material cited here does not establish a general performance benchmark. Browser crawlers usually involve more dependencies and browser resource use than an HTTP-and-HTML approach; use the simplest crawler that can reliably obtain the content you need.

What do I need to install?

The JavaScript quick start requires Node.js 16 or later and gives npm install crawlee as the general installation command. Browser automation packages are not bundled with Crawlee. Install the engine you intend to use as a separate dependency. The API documentation also describes smaller packages, including @crawlee/cheerio and @crawlee/playwright.

  • For a basic JavaScript crawler: npm install crawlee
  • For Playwright browser crawling: npm install crawlee playwright
  • For Puppeteer browser crawling: npm install crawlee puppeteer

For a starter project, the documented scaffold command is npx crawlee create my-crawler; follow the prompts to select a template. Check the current JavaScript quick start and API documentation for version-specific setup details.

How do I scrape a website with Crawlee?

This small example fetches HTML with CheerioCrawler, extracts each page title, limits the crawl to five requests, and writes the results to Crawlee’s default dataset. Save it as main.js in a project where you have installed Crawlee.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const { CheerioCrawler } = require('crawlee');

const crawler = new CheerioCrawler({
    maxRequestsPerCrawl: 5,
    async requestHandler({ request, $, pushData, log }) {
        const title = $('title').first().text().trim();
        await pushData({ url: request.url, title });
        log.info(`Saved title for ${request.url}`);
    },
});

await crawler.run(['https://example.com/']);

Run it with node main.js. The request limit keeps the example bounded; the handler reads the first title element, associates the extracted value with the requested URL, and sends the object to the dataset. Crawlee’s quick start demonstrates this general pattern. Use a page you are allowed to access, and adapt the URL and extraction selectors to that site’s HTML.

When the HTML crawler is enough

Use this pattern for pages where a regular HTTP response contains the content you need. Inspect the returned markup when selectors produce empty values: a title or data visible on screen may have been inserted by JavaScript and therefore be absent from the HTML Cheerio receives.

When a browser crawler is necessary

For a JavaScript-dependent page, install Playwright separately and substitute PlaywrightCrawler for the HTTP crawler. The handler then receives a browser page rather than a Cheerio document; use browser locators or page evaluation to wait for and read rendered content. Follow the current Crawlee quick start for the browser crawler’s current API and starter example, since browser integrations can change between releases. Do not install a browser engine when the fetched HTML already contains what you need.

What Crawlee’s request and storage features do

A crawler is more than a selector. The request limit in the example is a practical guardrail: it bounds a run while you develop and helps prevent an accidental crawl from expanding beyond its intended scope. The handler processes each request, and pushData sends extracted records to a dataset so they can be consumed after the run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a real project, decide which pages should enter the crawl, what fields constitute a valid record, and how to handle missing or malformed fields before scaling up. Keep a stable source URL with each record so an extraction can be checked against its page. The quick start’s request-and-dataset flow is a minimal starting point, not a guarantee that any site’s structure will remain unchanged.

Proxy configuration and sessions

Crawlee includes proxy configuration and session management. Its proxy guide describes integrating ProxyConfiguration with HTTP and browser crawler classes. Its session management guide explains how a SessionPool can associate a session with cookies and proxy details; session-specific settings can persist across requests, and proxy IPs can be rotated.

These features are mechanisms for managing crawl requests, not promises of anonymity, access, or success against a site’s controls. A proxy does not make collection permitted, and session handling does not ensure a site will serve the desired page. Follow the target site’s rules and applicable requirements; do not treat Crawlee configuration as a way to override them. The proxy documentation retrieved for this article is versioned at Crawlee’s 3.13 proxy-management guide, so verify the matching guidance for the Crawlee version in your project.

Version and deployment considerations

The JavaScript documentation version shown in the official materials is 3.18. The changelog lists Crawlee 3.18.1, dated August 12, 2026, and 3.18.0, dated August 4, 2026. The listed 3.18.1 fixes include an update to Playwright Cloudflare challenge handling for changed markup; the 3.18.0 entries include link-clicking options, dependency declarations, and type-safe router labels. These changelog notes describe release entries, not a guarantee that a crawler will pass a site’s challenge or load successfully. Check the JavaScript changelog when choosing or upgrading a version.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can run Crawlee on your own machine or deploy it on cloud infrastructure. Apify is an optional platform for people who want managed deployment and related tooling; using Crawlee does not require an Apify account. The repository README describes the project’s local and other-cloud use as well.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common Crawlee problems

The selector finds no text

First verify that the selector matches the returned HTML and that the selected element contains text. If the content appears only after client-side code runs, Cheerio cannot render it; switch to a browser crawler. If the content is in the HTML, adjust the selector and handle empty or absent values explicitly.

The browser crawler package is missing

Crawlee and its browser automation engine are separate installs. Add the matching dependency, such as npm install crawlee playwright or npm install crawlee puppeteer, then follow the current quick-start instructions for that integration.

The crawl makes more requests than intended

Set a request limit such as maxRequestsPerCrawl while developing, and review how new requests enter the crawl before removing or raising the limit. Keep the start URLs and crawl scope narrow until the handler and stored records are correct.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A request is blocked or a page does not load

Check the URL, network access, the site’s behavior, and whether the content requires browser execution. Crawlee’s proxy and session features can configure request identity and session state, but they do not guarantee access or authorize collection. Stop or revise a crawl if the site’s rules or applicable requirements do not permit it.

Results look stale or repeat unexpectedly

Inspect the requested URL and the fields saved for each record, then check whether the page itself returned the content you expected. A successful handler execution does not validate the meaning or freshness of extracted data; verify the output against representative pages.

Or skip the browser setup

If your task is to capture a page as an image or PDF rather than build a custom crawl, ScreenshotNeo offers a one-request screenshot API and an MCP server. Here is a cURL example; replace the target URL as needed. See the ScreenshotNeo documentation for API options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots per month with no card.

Frequently Asked Questions

Does Crawlee work only with Apify?

No. Crawlee can run locally or on other cloud infrastructure; Apify is an optional deployment platform.

Does CheerioCrawler execute JavaScript?

No. It fetches and parses HTML; use a browser crawler for content that requires client-side JavaScript.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.