Free tools Windows power users keep installed
One-click scans. No signup required.
Crawlee is an open-source library for building web scrapers and browser automation workflows in JavaScript and Python. Choose CheerioCrawler when the information is already in fetched HTML; choose PlaywrightCrawler or PuppeteerCrawler when the page depends on JavaScript or browser behavior. Crawlee provides the crawl and data-handling framework, while Playwright and Puppeteer are separate dependencies. This guide focuses on the JavaScript implementation and shows where each approach fits.
What is Crawlee?
Crawlee is an open-source library for scraping websites and automating browser tasks. Its JavaScript and Python implementations provide crawler patterns and tools for processing requests and storing results. The project repository states that Crawlee is licensed under the Apache License 2.0: Crawlee README.
This is a library, not a hosted scraping service that you must use through one vendor. You can run a crawler locally or deploy it on cloud infrastructure. Apify is one optional deployment path, not a requirement for using Crawlee.
Should I use CheerioCrawler or PlaywrightCrawler?
Start by checking where the page’s data comes from. If it is present in the HTML returned by an ordinary HTTP request, parse that HTML with CheerioCrawler. If the page needs client-side JavaScript, browser events, or other browser behavior to expose the content, use a browser crawler such as PlaywrightCrawler. If you already use Puppeteer or prefer its workflow, Crawlee also provides PuppeteerCrawler.
#1 Best Overall
| Need | Starting point | What to know |
|---|---|---|
| Read server-rendered or otherwise directly available HTML | CheerioCrawler |
Uses HTTP and HTML parsing; it does not render client-side JavaScript. |
| Load a JavaScript-driven page or interact with browser controls | PlaywrightCrawler |
Uses Playwright to control a browser. Install Playwright separately. |
| Keep an existing Puppeteer-based workflow | PuppeteerCrawler |
Uses Puppeteer to control a browser. Install Puppeteer separately. |
Crawlee describes CheerioCrawler as fast and efficient, but the official material cited here does not establish a general performance benchmark. Browser crawlers usually involve more dependencies and browser resource use than an HTTP-and-HTML approach; use the simplest crawler that can reliably obtain the content you need.
What do I need to install?
The JavaScript quick start requires Node.js 16 or later and gives npm install crawlee as the general installation command. Browser automation packages are not bundled with Crawlee. Install the engine you intend to use as a separate dependency. The API documentation also describes smaller packages, including @crawlee/cheerio and @crawlee/playwright.
- For a basic JavaScript crawler:
npm install crawlee - For Playwright browser crawling:
npm install crawlee playwright - For Puppeteer browser crawling:
npm install crawlee puppeteer
For a starter project, the documented scaffold command is npx crawlee create my-crawler; follow the prompts to select a template. Check the current JavaScript quick start and API documentation for version-specific setup details.
How do I scrape a website with Crawlee?
This small example fetches HTML with CheerioCrawler, extracts each page title, limits the crawl to five requests, and writes the results to Crawlee’s default dataset. Save it as main.js in a project where you have installed Crawlee.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11const { CheerioCrawler } = require('crawlee');
const crawler = new CheerioCrawler({
maxRequestsPerCrawl: 5,
async requestHandler({ request, $, pushData, log }) {
const title = $('title').first().text().trim();
await pushData({ url: request.url, title });
log.info(`Saved title for ${request.url}`);
},
});
await crawler.run(['https://example.com/']);
Run it with node main.js. The request limit keeps the example bounded; the handler reads the first title element, associates the extracted value with the requested URL, and sends the object to the dataset. Crawlee’s quick start demonstrates this general pattern. Use a page you are allowed to access, and adapt the URL and extraction selectors to that site’s HTML.
When the HTML crawler is enough
Use this pattern for pages where a regular HTTP response contains the content you need. Inspect the returned markup when selectors produce empty values: a title or data visible on screen may have been inserted by JavaScript and therefore be absent from the HTML Cheerio receives.
When a browser crawler is necessary
For a JavaScript-dependent page, install Playwright separately and substitute PlaywrightCrawler for the HTTP crawler. The handler then receives a browser page rather than a Cheerio document; use browser locators or page evaluation to wait for and read rendered content. Follow the current Crawlee quick start for the browser crawler’s current API and starter example, since browser integrations can change between releases. Do not install a browser engine when the fetched HTML already contains what you need.
What Crawlee’s request and storage features do
A crawler is more than a selector. The request limit in the example is a practical guardrail: it bounds a run while you develop and helps prevent an accidental crawl from expanding beyond its intended scope. The handler processes each request, and pushData sends extracted records to a dataset so they can be consumed after the run.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
For a real project, decide which pages should enter the crawl, what fields constitute a valid record, and how to handle missing or malformed fields before scaling up. Keep a stable source URL with each record so an extraction can be checked against its page. The quick start’s request-and-dataset flow is a minimal starting point, not a guarantee that any site’s structure will remain unchanged.
Proxy configuration and sessions
Crawlee includes proxy configuration and session management. Its proxy guide describes integrating ProxyConfiguration with HTTP and browser crawler classes. Its session management guide explains how a SessionPool can associate a session with cookies and proxy details; session-specific settings can persist across requests, and proxy IPs can be rotated.
These features are mechanisms for managing crawl requests, not promises of anonymity, access, or success against a site’s controls. A proxy does not make collection permitted, and session handling does not ensure a site will serve the desired page. Follow the target site’s rules and applicable requirements; do not treat Crawlee configuration as a way to override them. The proxy documentation retrieved for this article is versioned at Crawlee’s 3.13 proxy-management guide, so verify the matching guidance for the Crawlee version in your project.
Version and deployment considerations
The JavaScript documentation version shown in the official materials is 3.18. The changelog lists Crawlee 3.18.1, dated August 12, 2026, and 3.18.0, dated August 4, 2026. The listed 3.18.1 fixes include an update to Playwright Cloudflare challenge handling for changed markup; the 3.18.0 entries include link-clicking options, dependency declarations, and type-safe router labels. These changelog notes describe release entries, not a guarantee that a crawler will pass a site’s challenge or load successfully. Check the JavaScript changelog when choosing or upgrading a version.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
You can run Crawlee on your own machine or deploy it on cloud infrastructure. Apify is an optional platform for people who want managed deployment and related tooling; using Crawlee does not require an Apify account. The repository README describes the project’s local and other-cloud use as well.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common Crawlee problems
The selector finds no text
First verify that the selector matches the returned HTML and that the selected element contains text. If the content appears only after client-side code runs, Cheerio cannot render it; switch to a browser crawler. If the content is in the HTML, adjust the selector and handle empty or absent values explicitly.
The browser crawler package is missing
Crawlee and its browser automation engine are separate installs. Add the matching dependency, such as npm install crawlee playwright or npm install crawlee puppeteer, then follow the current quick-start instructions for that integration.
The crawl makes more requests than intended
Set a request limit such as maxRequestsPerCrawl while developing, and review how new requests enter the crawl before removing or raising the limit. Keep the start URLs and crawl scope narrow until the handler and stored records are correct.
Best Value
A request is blocked or a page does not load
Check the URL, network access, the site’s behavior, and whether the content requires browser execution. Crawlee’s proxy and session features can configure request identity and session state, but they do not guarantee access or authorize collection. Stop or revise a crawl if the site’s rules or applicable requirements do not permit it.
Results look stale or repeat unexpectedly
Inspect the requested URL and the fields saved for each record, then check whether the page itself returned the content you expected. A successful handler execution does not validate the meaning or freshness of extracted data; verify the output against representative pages.
Or skip the browser setup
If your task is to capture a page as an image or PDF rather than build a custom crawl, ScreenshotNeo offers a one-request screenshot API and an MCP server. Here is a cURL example; replace the target URL as needed. See the ScreenshotNeo documentation for API options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots per month with no card.
Frequently Asked Questions
Does Crawlee work only with Apify?
No. Crawlee can run locally or on other cloud infrastructure; Apify is an optional deployment platform.
Does CheerioCrawler execute JavaScript?
No. It fetches and parses HTML; use a browser crawler for content that requires client-side JavaScript.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

