Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Cloudflare’s 2024 Year in Review shows that AI crawlers became an important new category of automated web traffic, but it does not show that they generated the largest share of all Cloudflare traffic. Cloudflare identified Googlebot as the highest-volume request source overall. Meanwhile, ByteDance’s Bytespider declined sharply during the year, and Anthropic’s ClaudeBot became consistently active in late April before declining after an early peak in May and June.

The distinction matters for publishers and site owners: AI crawling may be strategically significant even when it is not the largest traffic category, because it can consume resources, reuse content without sending an immediate referral, and force difficult decisions about access, search visibility and compensation.

The short answer

Cloudflare added AI bot and crawler traffic as a new metric in its 2024 Year in Review, published December 9, 2024. The review covered Cloudflare-observed trends from January 1 through December 1, 2024.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Its findings support four conclusions:

  • AI crawlers had become a visible and strategically important source of automated requests.
  • Googlebot generated the highest request volume identified by Cloudflare.
  • Bytespider activity was approximately 80–85% lower at the end of November than at the beginning of the year.
  • ClaudeBot became consistently active in late April, peaked around May and June, and then declined through the rest of the measured period.

So the accurate interpretation is not “AI crawlers were the biggest source of Internet traffic.” It is that AI-specific crawling had become important enough to measure separately—and important enough for website owners to decide whether to allow, monitor, block or eventually charge for it.

What Cloudflare actually measured

Cloudflare Radar reports activity observed across Cloudflare’s network and related data sources. That provides a large and useful vantage point, but it is not a census of every request on the Internet. The figures describe traffic reaching Cloudflare customers during the stated period, not traffic from every website, hosting provider or CDN.

The AI metric was based on known AI crawler user agents associated with the ai.robots.txt project. That means the graph tracks recognized crawlers; it cannot reliably include bots that disguise themselves, rotate identities or fail to identify their purpose honestly.

Several kinds of activity should be kept separate:

  • Human visits: requests made by people using browsers, apps or other interactive clients.
  • General bot traffic: automated requests for indexing, monitoring, scraping, security research, testing and other purposes.
  • Search crawlers: bots such as Googlebot that help search engines discover and index pages.
  • AI training crawlers: agents associated with collecting data for model development.
  • AI search and retrieval crawlers: agents that may fetch content to answer a search or assistant query.
  • User-action crawlers: systems that retrieve a page because a user has asked an AI service to use it.

A request count is not the same as human traffic, bandwidth, unique pages, origin cost, content value or referrals. A crawler can make many inexpensive cached requests, while another may retrieve fewer but much larger pages. Neither count alone reveals whether a publisher benefited.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which crawlers stood out in 2024?

Crawler Operator Cloudflare’s observed pattern What it suggests
Googlebot Google Highest request volume identified overall Search indexing remained a major automated use of the Web.
Bytespider ByteDance Approximately 80–85% lower in late November than at the beginning of the year AI crawler activity can be volatile and specific to an operator or development phase.
ClaudeBot Anthropic Consistent activity began in late April; an early peak appeared around May and June, followed by a decline New model or product activity can arrive in sharp waves rather than as a steady trend.

Cloudflare described Bytespider as being used to download training data for ByteDance’s large language models. ClaudeBot was associated with collecting training data for models used by Claude. Those descriptions should not be expanded into claims about exactly how much content entered a model, or what either company did with every retrieved page.

Nor does a decline prove that an operator stopped crawling or stopped training. Crawl schedules, infrastructure, blocked requests and changes to crawler identities can all affect an observed series. A decline in one named agent also does not prove that AI content collection declined overall.

Was AI crawler traffic really a “major source”?

That depends on what “major” means:

  • Major strategic issue: Yes. Cloudflare created a dedicated AI crawler metric and introduced tools for customers to manage this traffic.
  • Major new category of automated traffic: Yes, especially for publishers, documentation sites and content-heavy businesses.
  • Largest overall source of Cloudflare requests: No. Cloudflare identified Googlebot as the highest-volume request source in the review.
  • Largest source of human visits: Not established by this report. Crawler requests are not human referrals.
  • Largest source of infrastructure cost: Not established. Request totals do not show bandwidth, cache behavior or origin expenditure.

The report therefore supports a governance and business story more strongly than a simple traffic-ranking story. AI crawlers became important because they challenged the traditional exchange between access and discovery: a search crawler can help a page appear in results that send visitors back, while an AI system may retrieve information and present an answer without producing an immediate visit to the source.

General bot traffic provides context—but is not AI traffic

Cloudflare also reported that the United States accounted for more than one-third of global bot traffic, while the top 10 countries generated 68.5% of observed bot traffic. By source network, AWS accounted for 12.7% and Google for 7.8%; Microsoft, Hetzner, DigitalOcean and OVH each contributed more than 1%.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These are figures for global bot traffic, not for AI crawlers specifically. Combining them with the AI-crawler graph would create a misleading comparison. The same caution applies to other findings in the Year in Review, such as overall Internet traffic growth, DNS-based popularity rankings and security trends: they are useful background, but they do not prove anything about the share or value of AI crawling.

Why publishers are concerned

AI crawling creates several different costs and risks:

  • Requests consume edge capacity, bandwidth, origin resources and monitoring attention.
  • Training access may use a publisher’s work without producing a direct referral.
  • AI-generated answers may compete with the page that supplied the information.
  • Blocking all automation can damage search indexing, uptime checks, accessibility tools or legitimate integrations.
  • User-agent labels are not perfect evidence of identity or intent.

Later Cloudflare analysis, published after the 2024 review, examined the imbalance between AI crawling and referrals. It is useful context for why the 2024 data mattered, but it should not be treated as a measurement of 2024 results. Similarly, Cloudflare’s later analysis of the 2025 crawler landscape showed that the mix of GPTBot, Meta-ExternalAgent and user-action crawling changed over time. Later crawler rankings should not be imported into a 2024 ranking.

A practical policy for site owners

1. Measure before changing access

Review request logs, CDN analytics, origin load, response sizes, cache status and referral data. Identify which agents are making requests and which paths they access. Separate search bots, AI training bots, AI retrieval bots, monitoring services and suspicious unidentified traffic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Decide what outcome matters

A publisher may want search visibility, citations, licensing revenue, lower infrastructure cost or simply control over reuse. Those goals can conflict. A blanket “allow” policy may maximize exposure but increase cost; a blanket block may reduce unwanted use but also reduce discovery.

3. Use robots.txt as a policy signal

robots.txt is a widely understood way to communicate crawler preferences. It is not authentication, payment or guaranteed network enforcement. It works only when a crawler chooses to comply, so teams that need reliable prevention must add controls at the CDN, WAF, server or application layer.

4. Allow useful crawlers separately

Do not treat Googlebot, AI training crawlers, AI search crawlers and user-action crawlers as interchangeable. A site may want to permit search indexing while restricting training crawlers, or permit a retrieval agent for selected public documentation while blocking bulk collection.

5. Block only after testing

Blocking is reasonable when a crawler has no business value, creates unwanted load or conflicts with the site’s content policy. Check for false positives, search-engine access, APIs, monitoring tools and existing WAF rules. User-agent-only blocking can also be bypassed by bots that disguise themselves.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

6. Consider charging as an experiment, not a guaranteed revenue stream

Charging may make sense for content with clear commercial or training value, but it requires operator support and a way to assess whether the income offsets lost access or referrals. It should not be assumed that a high request count translates into meaningful revenue.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Cloudflare’s current AI crawler controls

Cloudflare’s current documentation uses the name AI Crawl Control; the product was previously referred to as AI Audit. It is documented as available on all Cloudflare plans and provides visibility into AI crawler requests, per-crawler decisions and robots.txt compliance monitoring.

Detection differs by plan. Basic detection uses user-agent strings. More advanced detection uses a Bot Management detection ID and is associated with Enterprise plans that include Cloudflare Bot Management. This can improve classification, but it should not be understood as a perfect guarantee against disguised automation.

AI Crawl Control supports three actions—Allow, Block and Charge—as described in Cloudflare’s crawler-management documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What Pay Per Crawl does

Pay Per Crawl is documented as closed beta or private beta in the current 2026 documentation, not as a universally available mature marketplace. Cloudflare documents charging for successful content retrieval, generally an HTTP 200 response, and a minimum price of $0.01 per crawl.

There are important limitations:

  • One configured price applies to all crawlers assigned the Charge action, although individual crawlers can be assigned Allow, Block or Charge.
  • Repeated crawls can create repeated charges, so spending and crawl frequency need monitoring.
  • /robots.txt, /sitemap.xml, /security.txt, /.well-known/security.txt and /crawlers.json are documented as free paths.
  • HTTP errors are not treated as successful billable retrievals.
  • Charging or blocking search-engine crawlers may prevent proper indexing, so search bots need separate treatment.

Rule order is another common failure point. According to Cloudflare’s WAF integration documentation, WAF blocks occur before bot solutions and Pay Per Crawl. If an earlier WAF or bot rule blocks the request, Pay Per Crawl may never process it.

Cloudflare also documents advanced configuration involving URI exceptions and dynamic pricing through origin or Worker response headers. Those options do not change the central business question: whether a particular type of access is worth allowing, restricting or monetizing.

What the 2024 data cannot answer

Cloudflare’s review does not establish:

  • How many AI requests became training data.
  • How many produced citations or referrals.
  • How much bandwidth or origin cost they imposed on publishers.
  • How often each crawler violated robots.txt.
  • How much unidentified or disguised automation escaped classification.
  • Whether the overall value of AI access exceeded the value of search discovery.

Those questions require site-level logs, commercial agreements, referral measurements and content-specific economics. Cloudflare’s network-level view is valuable for showing that AI crawling was a growing policy concern, but it cannot by itself tell every publisher what access decision is profitable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The defensible takeaway

Cloudflare Radar’s 2024 review was an early signal that AI crawlers had become a meaningful force in the web ecosystem. It was not proof that AI crawlers collectively dominated all traffic, surpassed search crawlers or delivered the most value to publishers.

The important comparison is not simply AI versus Googlebot. It is automated content consumption versus the benefits returned to the sites being crawled. Site owners should measure those benefits, keep search access distinct from training access, enforce policies in more than one layer when necessary, and treat payment experiments as experiments—not as an automatic consequence of high crawler volume.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.