Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Cloudflare AI Search—not Vectorize by itself—when you want an MCP client to search your website. Create an AI Search instance, connect a site you own (or upload files), let AI Search maintain its Vectorize-backed index, enable its public endpoint and MCP, then give your client the endpoint URL ending in /mcp. Vectorize alone is a database: a direct Vectorize implementation requires your own Worker, embedding pipeline, metadata, and query logic.

Choose the right Cloudflare product first

The phrase “Vectorize MCP” combines two Cloudflare components that have different jobs. AI Search is the managed indexing and search layer. It automatically indexes connected sources, provides semantic, keyword, and hybrid retrieval, supports metadata filters, and includes an MCP endpoint. Its index is powered by Vectorize and is created and maintained for you.

Cloudflare Vectorize is the general-purpose vector database for Workers applications. It stores embeddings and returns similar vectors, but it does not crawl a website or publish an MCP endpoint automatically. Use it directly when you need to control ingestion, chunking, embedding models, metadata, authorization, and application behavior.

Route Best for What you build Content source
AI Search + MCP Website or knowledge-base search for agents Instance configuration, source connection, endpoint security, and client setup Owned website crawler or uploaded files
Direct Vectorize Custom retrieval inside a Worker Vector index, embeddings, ingestion Worker, metadata, query API, and any MCP layer Data supplied by your application; website crawling is your responsibility

AI Search is available on all Cloudflare plans according to its overview. The direct Vectorize tutorial lists a Workers Free or Paid plan as a prerequisite. Current usage limits and pricing depend on workload, so check Cloudflare’s current limits before budgeting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prerequisites and decisions

  • A Cloudflare account and a domain onboarded to that account if you plan to crawl a site. The crawler is restricted to sites the account owner owns.
  • Node.js and Wrangler. The setup guide for the documented Wrangler version lists Node.js 16.17.0 or later; verify the current requirement before installation because it can change.
  • An MCP-compatible client that supports a remote HTTP server. Client configuration and header names vary, so use that client’s current documentation.
  • A decision about exposure. The default public endpoint is unauthenticated; anyone who obtains its URL can query the indexed content.
  • An embedding model choice. AI Search’s selected model determines vector dimensions and cannot be changed after the instance is created.

For public documentation, the default endpoint may be acceptable. For internal or customer content, plan a custom hostname protected by Cloudflare Access before you index anything.

Create and monitor an AI Search instance

The documented Wrangler flow creates a web-crawler instance in one command:

npx wrangler ai-search create docs-search --type web-crawler --source developers.cloudflare.com

Replace docs-search with your instance name and the source with a domain you own. If crawling is unsuitable, choose AI Search’s built-in storage and upload files instead. The crawler and file-storage options are managed by AI Search; neither requires you to create a separate Vectorize index.

Check progress with:

npx wrangler ai-search stats docs-search

Do not connect an agent until the content you expect is indexed. Review the indexed pages, remove content that should not be searchable, and select the embedding model deliberately because its dimensions are fixed for the life of the instance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Enable the endpoint and MCP

  1. Open the Cloudflare dashboard and select the AI Search instance.
  2. Go to Settings > Public Endpoint.
  3. Enable the public endpoint, then enable MCP.
  4. Copy the generated endpoint host and append /mcp. That complete URL is the remote MCP server address.
  5. Set a useful tool description. Explain what the indexed material covers and the kinds of questions it can answer. A clear description helps an agent decide when to call the search tool.

The MCP reference exposes a search tool that queries the indexed content. MCP is only the discovery and invocation interface; it does not crawl pages or create vectors.

Connect an MCP client safely

Many clients use an mcpServers object for remote servers. A generic shape is:

{
  "mcpServers": {
    "docs-search": {
      "url": "https://YOUR-ENDPOINT-HOST/mcp"
    }
  }
}

This is illustrative, not universal. Some clients require a transport field such as "type": "http", while others place remote servers in a different settings file. Confirm the current format for Claude, Cursor, or your MCP client, then restart or reload its MCP connections. Ask the client a question whose answer appears in the indexed site and inspect the tool call to confirm that it used search.

Secure a public endpoint before indexing private data

By default, the generated endpoint accepts queries without authentication. Treat the URL as a capability: anyone who has it can search the indexed corpus. Do not put customer records, internal runbooks, unpublished product information, or other sensitive material behind an unauthenticated endpoint.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a custom domain with Cloudflare Access

  1. Attach a custom domain to the AI Search endpoint.
  2. Protect that hostname with a Cloudflare Access application and service-token headers.
  3. Configure your MCP client to send the required Access headers. The exact header configuration differs by client.
  4. Set default_domain_enabled to false as documented by Cloudflare. Otherwise the generated default hostname can continue responding without Access protection.

Allowed-origin settings control browser clients; they are not general server-side authentication. Rate limiting can reduce abuse on a public endpoint, but it does not make private content private. Test both the custom hostname and the old default hostname after changing the setting.

Use search modes and metadata deliberately

AI Search supports semantic, keyword, and hybrid search. Semantic search is useful when a question uses different wording from the source. Keyword search is important for exact API names, error codes, product identifiers, and version strings. Hybrid search combines both signals. Choose based on your content and test representative questions rather than assuming one mode always wins.

Metadata filters let you restrict results by fields such as category, version, or language. Add consistent metadata during ingestion or file upload, then use filters when an agent must stay within a product version or locale. Keep values normalized—for example, use v2 everywhere instead of mixing 2, 2.0, and version-two.

When direct Vectorize is the better architecture

Choose direct Vectorize when AI Search’s managed crawler and MCP surface do not provide enough control. A typical Worker implementation has these stages:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Fetch or receive source documents and split them into meaningful chunks.
  2. Generate embeddings with your selected model.
  3. Create a Vectorize index with dimensions matching that model.
  4. Insert vectors with stable IDs and metadata such as URL, title, language, and version.
  5. Query the index from a Worker, optionally combining vector similarity with keyword logic.
  6. Build your own API or MCP server, authentication, rate limits, and citation handling.

This route gives you control over re-indexing, tenant isolation, deletion, ranking, and response formatting. It also makes you responsible for crawl scheduling, robots and access rules, embedding failures, stale documents, and every security boundary. A Vectorize index is not a substitute for those application services.

Common failures and fixes

The crawler indexes nothing

Confirm that the domain is onboarded to the same Cloudflare account and that you are not trying to crawl a site you do not own. Check the instance statistics, verify that the source URL is correct, and use uploaded files if the site cannot be crawled.

The MCP client cannot connect

Check that the URL ends in /mcp, not only the endpoint host. Confirm that the client supports remote HTTP MCP and that its configuration uses the required transport field. Look for JSON syntax errors, restart the client, and test the endpoint from the client’s own connection diagnostics.

The agent returns irrelevant results

Try hybrid or keyword search for exact terms, improve document titles and chunk boundaries, and add consistent metadata filters. Test with real user questions and include version or language constraints when those distinctions matter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Access protection appears to be bypassed

Test the generated default hostname directly. If it still answers, disable it with default_domain_enabled=false. Also verify that the MCP client actually sends the Access service-token headers; a browser-origin allowlist will not add server authentication.

Embedding or index creation errors occur

Check that the Vectorize index dimensions match the selected embedding model. In AI Search, the model cannot be swapped after instance creation, so create a new instance if you need a different dimensionality and re-index the source.

Performance, reliability, and operating practice

  • Wait for indexing to complete before evaluating search quality; newly changed pages may not appear immediately.
  • Keep source URLs, titles, versions, and languages in metadata so agents can narrow results and explain where an answer came from.
  • Schedule content reviews and deletions. Search quality degrades when obsolete pages remain indexed.
  • Use rate limits and monitoring on public endpoints, and log failed client connections without recording sensitive query text unnecessarily.
  • Measure your own representative questions. The documentation describes hybrid search as improving accuracy, but no universal accuracy, latency, or cost benchmark is established for every site.
  • Recheck Cloudflare’s current plan limits, pricing, Wrangler requirements, and MCP-client transport guidance before production rollout.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your immediate need is a clean image or PDF of a page—not a searchable MCP index—ScreenshotNeo is a separate website screenshot API. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo documentation for all options. A direct call looks like this:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month without a card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account.

FAQ

Does Vectorize automatically crawl my website?

No. Website crawling and the MCP endpoint are provided by AI Search. Direct Vectorize requires you to supply and maintain the content pipeline.

Can I leave the default MCP hostname enabled with Access?

Not if all access must be authenticated. Disable the default hostname with default_domain_enabled=false after configuring the protected custom domain.

Can I change the AI Search embedding model later?

No. The selected model determines index dimensions and cannot be changed after instance creation; use a new instance to migrate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is an allowed-origin list an authentication mechanism?

No. It affects browser clients. Use a custom domain and Cloudflare Access service-token headers for authenticated server-side access.

Frequently Asked Questions

Does Vectorize automatically crawl my website?

No. Website crawling and the MCP endpoint are provided by AI Search. Direct Vectorize requires you to supply and maintain the content pipeline.

Can I leave the default MCP hostname enabled with Access?

Not if all access must be authenticated. Disable the default hostname with default_domain_enabled=false after configuring the protected custom domain.

Can I change the AI Search embedding model later?

No. The selected model determines index dimensions and cannot be changed after instance creation; use a new instance to migrate.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is an allowed-origin list an authentication mechanism?

No. It affects browser clients. Use a custom domain and Cloudflare Access service-token headers for authenticated server-side access.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.