Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Start with ScrapeWebsiteTool for a single, mostly static page. Install CrewAI’s tools extra, pass a URL, and call run(). If the page needs a CSS selector, use ScrapeElementFromWebsiteTool; if it requires a real browser, use the documented Selenium integration (with its current-development caveat); and for bounded multi-page work, use Firecrawl’s scrape or crawl tools. The examples below show each path, how to attach tools to an agent, how to validate results, and how to avoid common failures.

1. Install CrewAI’s scraping tools

The official basic example installs the tools extra:

pip install 'crewai[tools]'

Documentation pages currently resolve to CrewAI versions 1.15.18, 1.15.22 and 1.15.23, while the overview is unversioned. Match the documentation to the CrewAI version installed in your project before copying an example.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Scrape one page with ScrapeWebsiteTool

Use this when you need the readable content of one URL and the content is present in the returned HTML.

from crewai_tools import ScrapeWebsiteTool

scraper = ScrapeWebsiteTool(website_url="https://example.com")
text = scraper.run()
print(text)

CrewAI documentation describes this tool as being designed to extract and read the content of a specified website (official ScrapeWebsiteTool documentation). You can also initialize it without a URL so an agent supplies the target at call time:

from crewai_tools import ScrapeWebsiteTool

scraper = ScrapeWebsiteTool()

For a known target, fixing the URL in code is safer: it makes the workflow auditable and prevents an unconstrained agent from visiting arbitrary sites. The documented fetch path uses CrewAI’s SSRF-safe HTTP helper, checks the requested URL and every redirect against private or reserved ranges, and pins the connection to the checked IP. That property applies to this documented tool, not automatically to custom or third-party tools.

Give an agent a narrow extraction contract

from crewai import Agent, Task, Crew
from crewai_tools import ScrapeWebsiteTool

scraper = ScrapeWebsiteTool(website_url="https://example.com/news")

researcher = Agent(
    role="News extractor",
    goal="Extract only the article headlines from the supplied page",
    backstory="You return one headline per line and never invent missing text.",
    tools=[scraper],
    verbose=True,
)

task = Task(
    description=(
        "Use the scraping tool on the configured page. Return each headline "
        "as a separate bullet. If no headlines are found, return exactly "
        "'NO_HEADLINES'."
    ),
    expected_output="A bullet list of headlines or NO_HEADLINES.",
    agent=researcher,
)

result = Crew(agents=[researcher], tasks=[task]).kickoff()
print(result)

A direct run() is simpler for a one-off fetch. An agent and task are useful when scraping is one step in a larger workflow, such as extracting fields and then classifying or summarizing them. Keep the output contract specific (“return each headline as a separate bullet”) instead of asking the agent to “scrape everything.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Extract a known section with a CSS selector

ScrapeElementFromWebsiteTool is the better fit when you know the element or field you need. Its documented inputs include a URL, a CSS selector and optional cookies; matching text is returned joined by newlines. The page describes an implementation based on requests and BeautifulSoup (selector-tool documentation).

from crewai_tools import ScrapeElementFromWebsiteTool

headlines = ScrapeElementFromWebsiteTool(
    website_url="https://example.com/news",
    css_selector="article h2",
).run()
print(headlines)

Install the dependencies listed by the matching documentation if your environment does not already include them:

pip install requests beautifulsoup4

Selector failure modes

  • No matches: inspect the current HTML and update the selector. A class used for styling can change without notice.
  • Unexpected text: narrow the selector (for example, main article h2) and exclude navigation or footer elements.
  • Empty result on a modern app: the data may be inserted by JavaScript after the initial response; switch to a browser-based option.
  • Login-protected content: use the tool’s documented cookie input only where you have authorization, and keep session values out of source control.

4. Handle JavaScript-rendered pages with Selenium

Use SeleniumScrapingTool when the required content appears only after browser execution, scrolling, or a wait. The documented inputs include a URL, CSS selector, optional cookies, wait time, and text or HTML output (SeleniumScrapingTool documentation).

CrewAI’s page currently labels this tool “in development” and warns that unexpected behavior may occur. Treat it as a documented option to test against your target, not as a guaranteed production-ready browser service.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The page calls for Selenium, Chrome and Chrome WebDriver (and lists webdriver-manager in its setup). Follow the version-matched installation instructions, then use a bounded wait:

from crewai_tools import SeleniumScrapingTool

scraper = SeleniumScrapingTool(
    website_url="https://example.com/app",
    css_selector="main article",
    wait_time=5,
    return_html=False,
)
print(scraper.run())

Prefer a selector wait or the shortest reliable delay rather than a large fixed sleep. Test browser and WebDriver versions together, and record failures because browser automation is more sensitive to environment changes than an HTTP fetch.

5. Use Firecrawl for extraction or a bounded crawl

Firecrawl adds an external extraction service and API-key configuration. CrewAI documents two integrations:

FirecrawlScrapeWebsiteTool: one URL with optional structured extraction

The scrape integration supports the URL, main-content filtering, raw HTML, and an LLM extraction prompt or schema (Firecrawl scrape documentation). Store the key in the environment variable shown by the documentation:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
pip install 'crewai[tools]' firecrawl-py
export FIRECRAWL_API_KEY="your-key"
from crewai_tools import FirecrawlScrapeWebsiteTool

scraper = FirecrawlScrapeWebsiteTool(
    website_url="https://example.com/product",
    params={
        "onlyMainContent": True,
        "formats": ["markdown"],
    },
)
print(scraper.run())

Use the exact parameter names supported by the installed integration. An extraction prompt or schema is useful when downstream code needs fields such as name, price and availability, but you still need to validate that required fields are present.

FirecrawlCrawlWebsiteTool: multiple pages from a starting URL

The crawl integration supports include and exclude patterns, crawl depth, page limits, timeouts and other scope controls (Firecrawl crawl documentation).

from crewai_tools import FirecrawlCrawlWebsiteTool

crawler = FirecrawlCrawlWebsiteTool(
    website_url="https://example.com/docs",
    params={
        "maxDepth": 2,
        "limit": 50,
        "includePaths": ["/docs/.*"],
        "excludePaths": ["/docs/archive/.*"],
    },
)
print(crawler.run())

Always set a depth and page limit. A starting URL without boundaries can expand into a much larger job than intended, especially on sites with calendars, search parameters or faceted navigation.

6. Choose the smallest tool that fits

Requirement Start with Important trade-off
One static page ScrapeWebsiteTool Simple HTTP/HTML extraction; it does not promise client-side JavaScript rendering.
Known field or section ScrapeElementFromWebsiteTool Fast and focused, but the selector must match the current HTML.
Browser interaction or delayed content SeleniumScrapingTool Needs Chrome/WebDriver and is currently marked in development.
One page with external extraction FirecrawlScrapeWebsiteTool Requires Firecrawl credentials and an external service.
Many linked pages FirecrawlCrawlWebsiteTool Define depth, limits and path rules to keep the crawl bounded.

CrewAI’s overview also points to Browserbase for cloud browser infrastructure and Stagehand for complex interactions. Those are selection guidance, not independent speed, accuracy, reliability or price benchmarks; the official material does not establish comparable performance figures for these options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. A production-minded scraping checklist

  1. Define the target and fields. Decide whether you need one URL, selected elements, browser interaction or a crawl. Write the expected output and required fields before running an agent.
  2. Install only what you need. Start with crewai[tools]; add the documented selector, Selenium or Firecrawl dependencies for that path.
  3. Fix inputs where possible. Configure a known URL and selector in code. Put Firecrawl credentials in FIRECRAWL_API_KEY, never in prompts or committed files.
  4. Respect the site. Check robots.txt, terms and applicable law; identify your client appropriately; add delays and rate limits; and do not bypass access controls.
  5. Validate before downstream use. Reject empty results, check required fields and types, remove duplicates, and preserve the source URL and retrieval time with each record.
  6. Handle transient failures. Add bounded retries with backoff for network errors, but do not retry indefinitely or turn a blocked request into a denial-of-service pattern.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

8. Troubleshooting

Import or installation errors

Confirm that the virtual environment running your script is the one where crewai[tools] was installed. For Selenium, verify Chrome and WebDriver compatibility and follow the current tool page’s dependency instructions. For Firecrawl, install firecrawl-py and check that FIRECRAWL_API_KEY is available to the process.

The output is blank or missing the visible text

Inspect the raw response. If the text is injected after page load, ScrapeWebsiteTool or the selector tool may not see it; move to Selenium or Firecrawl. If only a region is needed, verify the selector against the current DOM rather than guessing from an old class name.

Requests are blocked or redirected

Check the site’s policy, reduce request frequency, use an honest identifying user-agent where appropriate, and stop if access is refused. The SSRF checks documented for ScrapeWebsiteTool protect against private/reserved destinations; they do not make a blocked public site accessible.

A crawl runs too long or returns unrelated pages

Lower maxDepth and limit, add include/exclude patterns, and remove tracking or search paths. Keep each crawl’s scope explicit and log the URLs returned for review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The agent invents fields

Make the task’s expected output machine-checkable, require a sentinel such as NO_VALUE for missing data, and validate the result in Python before storing it. Scraping retrieves text; it does not guarantee that an agent’s interpretation is factual.

Or skip the browser setup

If your immediate need is a clean visual capture rather than parsed text, ScreenshotNeo is a complementary website screenshot API. It accepts cookie and consent banners like a visitor, removes 60-plus known consent platforms plus newsletter popups and chat widgets before capture, and bills only clean shots. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed; each response reports the result with X-Page-Verdict and X-Billed headers. It also offers an MCP server for AI agents with take_screenshot, get_page_info and capture_pdf.

One GET request returns PNG, JPEG, WebP or PDF. The complete option set includes full-page and CSS-element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper and page controls, custom CSS/JavaScript, clicks, selector or network-idle waits, request/resource blocking, headers, cookies, user-agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, async webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for response formats and options. The same request in Python:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account to try it.

9. Cost, reliability and operational boundaries

CrewAI’s official scraping pages do not publish comparable benchmarks for speed, accuracy, success rate, adoption or cost savings. Choose based on page behavior, scope and dependency tolerance rather than an unverified performance claim. A local HTTP fetch avoids an external extraction-service credential; a browser introduces Chrome/WebDriver maintenance; Firecrawl adds service configuration but can simplify rendered extraction and crawling.

For repeatable jobs, pin package versions, log tool and browser versions, keep selectors under tests, and store failures separately from valid empty results. Re-run a small fixture set after upgrades because the documentation is versioned and the tools’ behavior can change.

Frequently Asked Questions

Can ScrapeWebsiteTool scrape a site that requires JavaScript?

Not reliably. It is documented as an HTTP/HTML scraper, so use Selenium or a Firecrawl integration when the required content is rendered or revealed in a browser.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use a CrewAI agent for every scrape?

No. Call a tool directly for a one-off fetch. Add it to an agent when scraping is one bounded step in a workflow that needs reasoning or transformation.

How do I crawl only a documentation section?

Use FirecrawlCrawlWebsiteTool with an explicit starting URL, include pattern, exclude pattern, maximum depth and page limit, then validate the returned URLs.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.