What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To crawl a website with an API, give a crawler a starting URL, define which paths and domains it may visit, choose how it discovers and renders pages, then retrieve and audit the results. “Entire website” is a target scope, not a guarantee: page and depth limits, sitemap coverage, inaccessible pages, JavaScript rendering and filtering can all affect what comes back.

Firecrawl documents an asynchronous crawl workflow at its crawl API documentation: submit a job, receive an ID, and retrieve its status and results. Its documented default crawl limit is 10,000 pages, so set a deliberate limit and check the job’s reported URLs, errors and skipped pages rather than assuming the crawl is exhaustive.

What a website crawl API does—and what “entire” means

A crawl API starts with one or more seed URLs, discovers additional URLs, fetches pages, and returns content in formats you can use in a documentation corpus, migration, knowledge base or other downstream system. Firecrawl describes sitemap and link discovery, and its product page lists outputs including Markdown, JSON, HTML, links, screenshots, images and metadata (Firecrawl Web Crawling API).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A crawler can only return pages it discovers and is configured to include, and that it can fetch under the service’s access rules and limits. A page may be absent because it is not in a sitemap or reachable link path, falls outside the configured domain or path boundary, exceeds a page or depth limit, fails to load, or depends on rendering the service does not perform. Treat “entire site” as the scope you intend to crawl, then verify coverage against an inventory.

Plan the crawl boundary before submitting a job

Choose the seed and host scope

Decide whether the seed is a homepage, a section such as https://example.com/docs/, or several known starting URLs. Also decide whether the crawl should stay on that host, include subdomains, or follow external links. Firecrawl documents crawlEntireDomain, allowSubdomains and allowExternalLinks; the latter two and crawlEntireDomain are documented as false by default. Set these intentionally rather than assuming that a homepage seed means every related host is included.

Set page, depth and path limits

Firecrawl documents a default crawl limit of 10,000 pages when limit is omitted, and supports maxDiscoveryDepth to bound link traversal. Choose a limit appropriate to the project and include/exclude patterns for the sections that matter. One subtle failure mode: Firecrawl notes that the starting URL itself is checked against include-path patterns. If the seed does not match, the crawl can return zero pages; ensure the seed and intended paths are compatible.

Account for query strings and duplicate URLs

Firecrawl documents similar-URL deduplication as enabled by default and ignoring query parameters as disabled by default. Do not enable query-parameter ignoring without checking the site’s URL behavior: two URLs differing only in a query string may represent distinct content, such as filtered records or language variants. Conversely, parameterized tracking URLs may create redundant variants. Decide based on the target site, then audit the resulting URL set.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose how the crawler discovers pages

Firecrawl’s documented default combines sitemap and link discovery. Its sitemap setting supports include, skip and only modes (Firecrawl crawl configuration).

  • Use sitemap discovery when the site maintains a useful sitemap and you want URLs listed there included. Sitemap-only discovery can miss pages that the sitemap omits.
  • Use link discovery when pages are reachable by following links from the seed. Link-only discovery can miss pages that are not linked from reachable pages.
  • Use both when you want the two discovery routes to complement each other, then compare the results with the sitemap or a separate URL inventory.

Neither route alone proves completeness. A sitemap can be stale or selective; link traversal is constrained by reachable paths, filters and depth. The useful question is not whether the API says “crawl,” but which URL sources it checked and what the job reports it fetched or skipped.

Choose rendering and output for the downstream task

Firecrawl’s crawl documentation says Markdown is the default and allows per-page scrape options. Its product page lists Markdown, JSON, HTML, links, screenshots, images and metadata. Select the output your next system can consume: Markdown is convenient for text-oriented knowledge bases, while HTML may preserve markup useful in migrations and structured formats may suit extraction workflows. The available formats and options are provider-specific.

JavaScript-heavy pages need particular attention. Firecrawl’s product page says each page is rendered in Chromium, which is a Firecrawl service description—not a guarantee for every crawler API. The Apify Website Crawler listing describes automatic, raw HTTP and browser rendering choices, as well as sitemap use and configurable page and depth limits; these are features of that listing, not category-wide defaults (Apify Website Crawler). Check the selected service’s current rendering mode and test representative pages before committing to a large crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Submit a Firecrawl crawl and retrieve the results

The endpoint documented for Firecrawl v2 is POST https://api.firecrawl.dev/v2/crawl. The job is asynchronous: the initial response returns an ID, and you then request status/results. The example below shows the request shape and the scope options to consider; use the authentication and complete request requirements in Firecrawl’s current documentation for your account.

curl -X POST "https://api.firecrawl.dev/v2/crawl" 
  -H "Content-Type: application/json" 
  -H "Authorization: Bearer $FIRECRAWL_API_KEY" 
  -d '{
    "url": "https://example.com/docs/",
    "limit": 2000,
    "maxDiscoveryDepth": 5,
    "crawlEntireDomain": false,
    "allowExternalLinks": false,
    "allowSubdomains": false,
    "sitemap": "include"
  }'

Replace the example URL and limits with the project’s actual boundary. The documented options include path filters, a delay, and the sitemap modes described above. Firecrawl says setting delay forces concurrency to one, so a delay can reduce request concurrency when that is appropriate. Consult the current API documentation for the exact status-request procedure and any additional request fields your account or workflow requires.

Do not stop after the job ID. Retrieve the job status and consume all result pages. Firecrawl documents that a result response may include a next URL when a job is still running or when content exceeds 10 MB. Follow that URL and continue until the API indicates there is no further page of results; a single response may not contain everything.

Audit coverage before using the crawl output

  1. Record the intended scope. Keep the seed URLs, allowed hosts, path patterns, page limit and depth limit with the crawl output so you can explain what the run was meant to include.
  2. Compare discovered URLs with an inventory. Use the site’s sitemap or a maintained URL list as a comparison point. Differences are leads to investigate, not automatic proof that the crawler failed: the two sources may represent different scopes.
  3. Inspect errors and skips. Review the provider’s status and reported errors or skipped pages. A successful job response does not itself establish that every intended page was fetched.
  4. Check representative content. Open a sample from each important section, including pages known to rely on client-side rendering, to confirm the chosen format contains the content your downstream system needs.
  5. Rerun a narrow crawl for gaps. If a section is missing, check filters, seed matching, depth and discovery source, then submit a deliberately scoped follow-up rather than blindly increasing the limit.

Respect site rules and manage crawl load

Before running a crawl, review the target site’s access rules and the provider’s handling of them. Firecrawl says it reads robots.txt for rules applying to FirecrawlAgent and *. The Apify Website Crawler listing says its robots.txt option is enabled by default. These statements apply to those named services; behavior varies by provider, so verify current settings and the site’s rules before a run.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Limits, depth, concurrency and delay affect both the scope and operational impact of a crawl. A lower page limit may stop a run before all desired pages are discovered; a configured delay can slow throughput, and Firecrawl says its delay setting forces concurrency to one. For a large crawl, begin with a representative, bounded run, inspect results and errors, and expand only when the output and site rules support it.

Costs and how to compare crawler services

Firecrawl’s product page states a price of one credit per page crawled. Its displayed plans and prices can change, so confirm the current pricing and what counts as a page before estimating project cost. The page also states a default crawl limit of 10,000; Firecrawl’s documentation describes that as the default limit when omitted. These are provider-published terms, not independent measurements of completeness or extraction quality.

Compare services against the job you need done rather than assuming one API is universally best. Relevant differences include:

  • Discovery: sitemap support, link traversal and sitemap-only controls.
  • Scope: page and depth limits, path filters, subdomain and external-link behavior, and URL deduplication.
  • Rendering: whether the service fetches raw HTML or uses a browser for JavaScript-rendered pages.
  • Site instructions: documented robots.txt handling and request-delay controls.
  • Results: output formats, asynchronous job retrieval, pagination and error reporting.
  • Operations: pricing, concurrency, retries, infrastructure and maintenance needs.

The available provider documentation does not establish a controlled comparison of completeness, speed or extraction quality across services. A feature listing can tell you what a provider says it supports; it cannot guarantee how well a particular site will crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

ScreenshotNeo is a website screenshot API, not a crawler for discovering and exporting every page on a domain. Use a crawler for site-wide URL discovery; use a screenshot when you need a visual capture of a known URL. A single request can return an image or PDF. See the ScreenshotNeo site and API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo removes cookie banners, newsletter popups and chat widgets before capture. Bot checks, blank pages and failed loads are not billed. Its MCP server lets AI agents take screenshots, and the Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month with no card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common crawl problems

The job returns zero pages

Check that the starting URL matches the include-path patterns. Firecrawl documents that it checks the starting URL against those patterns, so a mismatch can produce zero pages. Also verify the URL is valid and the selected discovery mode can find pages within the configured boundary.

Pages from a section are missing

Compare path filters, discovery depth and the sitemap mode. The missing URLs may not be in the sitemap or reachable from crawled links; a path rule may exclude them, or the depth limit may stop traversal before reaching them. Check provider-reported skips and errors, then try a narrowly scoped crawl of the section.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The crawl stops before the expected number of pages

Inspect the configured limit, depth and job status rather than equating a completed job with full coverage. Firecrawl documents a default limit of 10,000 pages when limit is omitted. If the intended scope is larger, verify the current service limits and split the scope into deliberate sections if appropriate.

Only part of the output appears in one response

Check for the response’s next URL. Firecrawl says it can appear while a job is ongoing or when result content exceeds 10 MB. Follow the pagination path until there are no more results.

Text from a dynamic page is absent

Check whether the selected service renders JavaScript and whether the page needs more time or a specific interaction before content appears. Rendering behavior differs across services. Firecrawl describes Chromium rendering on its product page; do not assume a different API uses the same approach.

Two seemingly similar URLs produce unexpected duplicates or omissions

Review query-parameter handling and similar-URL deduplication. Firecrawl documents similar-URL deduplication as enabled by default and query-parameter ignoring as disabled by default. If query strings carry distinct content, ignoring them may merge meaningful pages; if they are tracking-only, retaining all variants may add noise.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently asked questions

When should I use crawl rather than scrape or map?

Use a crawl when you want to follow a site’s pages and retrieve their content. A single-page scrape is for a known URL; mapping is for discovering URLs without necessarily retrieving each page’s full content. Firecrawl frames these as separate product workflows on its crawl page.

Can I process pages as they are crawled?

That depends on the provider’s result-delivery options. The Firecrawl details covered here establish asynchronous job retrieval and result pagination; consult its current API documentation for any streaming or webhook workflow you require.

Can I crawl only part of a site?

Yes. Use a section seed and appropriate path, domain, depth and page-limit settings. Check that the starting URL satisfies any include-path pattern.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.