Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Cloudflare’s 2024 Year in Review shows that AI crawlers became an important new category of automated web traffic, but it does not show that they generated the largest share of all Cloudflare traffic. Cloudflare identified Googlebot as the highest-volume request source overall. Meanwhile, ByteDance’s Bytespider declined sharply during the year, and Anthropic’s ClaudeBot became consistently active in late April before declining after an early peak in May and June.
The distinction matters for publishers and site owners: AI crawling may be strategically significant even when it is not the largest traffic category, because it can consume resources, reuse content without sending an immediate referral, and force difficult decisions about access, search visibility and compensation.
The short answer
Cloudflare added AI bot and crawler traffic as a new metric in its 2024 Year in Review, published December 9, 2024. The review covered Cloudflare-observed trends from January 1 through December 1, 2024.
Recommended Free Tools
Its findings support four conclusions:
- AI crawlers had become a visible and strategically important source of automated requests.
- Googlebot generated the highest request volume identified by Cloudflare.
- Bytespider activity was approximately 80–85% lower at the end of November than at the beginning of the year.
- ClaudeBot became consistently active in late April, peaked around May and June, and then declined through the rest of the measured period.
So the accurate interpretation is not “AI crawlers were the biggest source of Internet traffic.” It is that AI-specific crawling had become important enough to measure separately—and important enough for website owners to decide whether to allow, monitor, block or eventually charge for it.
#1 Best Overall
What Cloudflare actually measured
Cloudflare Radar reports activity observed across Cloudflare’s network and related data sources. That provides a large and useful vantage point, but it is not a census of every request on the Internet. The figures describe traffic reaching Cloudflare customers during the stated period, not traffic from every website, hosting provider or CDN.
The AI metric was based on known AI crawler user agents associated with the ai.robots.txt project. That means the graph tracks recognized crawlers; it cannot reliably include bots that disguise themselves, rotate identities or fail to identify their purpose honestly.
Several kinds of activity should be kept separate:
- Human visits: requests made by people using browsers, apps or other interactive clients.
- General bot traffic: automated requests for indexing, monitoring, scraping, security research, testing and other purposes.
- Search crawlers: bots such as Googlebot that help search engines discover and index pages.
- AI training crawlers: agents associated with collecting data for model development.
- AI search and retrieval crawlers: agents that may fetch content to answer a search or assistant query.
- User-action crawlers: systems that retrieve a page because a user has asked an AI service to use it.
A request count is not the same as human traffic, bandwidth, unique pages, origin cost, content value or referrals. A crawler can make many inexpensive cached requests, while another may retrieve fewer but much larger pages. Neither count alone reveals whether a publisher benefited.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Which crawlers stood out in 2024?
| Crawler | Operator | Cloudflare’s observed pattern | What it suggests |
|---|---|---|---|
| Googlebot | Highest request volume identified overall | Search indexing remained a major automated use of the Web. | |
| Bytespider | ByteDance | Approximately 80–85% lower in late November than at the beginning of the year | AI crawler activity can be volatile and specific to an operator or development phase. |
| ClaudeBot | Anthropic | Consistent activity began in late April; an early peak appeared around May and June, followed by a decline | New model or product activity can arrive in sharp waves rather than as a steady trend. |
Cloudflare described Bytespider as being used to download training data for ByteDance’s large language models. ClaudeBot was associated with collecting training data for models used by Claude. Those descriptions should not be expanded into claims about exactly how much content entered a model, or what either company did with every retrieved page.
Nor does a decline prove that an operator stopped crawling or stopped training. Crawl schedules, infrastructure, blocked requests and changes to crawler identities can all affect an observed series. A decline in one named agent also does not prove that AI content collection declined overall.
Was AI crawler traffic really a “major source”?
That depends on what “major” means:
- Major strategic issue: Yes. Cloudflare created a dedicated AI crawler metric and introduced tools for customers to manage this traffic.
- Major new category of automated traffic: Yes, especially for publishers, documentation sites and content-heavy businesses.
- Largest overall source of Cloudflare requests: No. Cloudflare identified Googlebot as the highest-volume request source in the review.
- Largest source of human visits: Not established by this report. Crawler requests are not human referrals.
- Largest source of infrastructure cost: Not established. Request totals do not show bandwidth, cache behavior or origin expenditure.
The report therefore supports a governance and business story more strongly than a simple traffic-ranking story. AI crawlers became important because they challenged the traditional exchange between access and discovery: a search crawler can help a page appear in results that send visitors back, while an AI system may retrieve information and present an answer without producing an immediate visit to the source.
General bot traffic provides context—but is not AI traffic
Cloudflare also reported that the United States accounted for more than one-third of global bot traffic, while the top 10 countries generated 68.5% of observed bot traffic. By source network, AWS accounted for 12.7% and Google for 7.8%; Microsoft, Hetzner, DigitalOcean and OVH each contributed more than 1%.
These are figures for global bot traffic, not for AI crawlers specifically. Combining them with the AI-crawler graph would create a misleading comparison. The same caution applies to other findings in the Year in Review, such as overall Internet traffic growth, DNS-based popularity rankings and security trends: they are useful background, but they do not prove anything about the share or value of AI crawling.
Why publishers are concerned
AI crawling creates several different costs and risks:
- Requests consume edge capacity, bandwidth, origin resources and monitoring attention.
- Training access may use a publisher’s work without producing a direct referral.
- AI-generated answers may compete with the page that supplied the information.
- Blocking all automation can damage search indexing, uptime checks, accessibility tools or legitimate integrations.
- User-agent labels are not perfect evidence of identity or intent.
Later Cloudflare analysis, published after the 2024 review, examined the imbalance between AI crawling and referrals. It is useful context for why the 2024 data mattered, but it should not be treated as a measurement of 2024 results. Similarly, Cloudflare’s later analysis of the 2025 crawler landscape showed that the mix of GPTBot, Meta-ExternalAgent and user-action crawling changed over time. Later crawler rankings should not be imported into a 2024 ranking.
Rank #3
A practical policy for site owners
1. Measure before changing access
Review request logs, CDN analytics, origin load, response sizes, cache status and referral data. Identify which agents are making requests and which paths they access. Separate search bots, AI training bots, AI retrieval bots, monitoring services and suspicious unidentified traffic.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall2. Decide what outcome matters
A publisher may want search visibility, citations, licensing revenue, lower infrastructure cost or simply control over reuse. Those goals can conflict. A blanket “allow” policy may maximize exposure but increase cost; a blanket block may reduce unwanted use but also reduce discovery.
3. Use robots.txt as a policy signal
robots.txt is a widely understood way to communicate crawler preferences. It is not authentication, payment or guaranteed network enforcement. It works only when a crawler chooses to comply, so teams that need reliable prevention must add controls at the CDN, WAF, server or application layer.
4. Allow useful crawlers separately
Do not treat Googlebot, AI training crawlers, AI search crawlers and user-action crawlers as interchangeable. A site may want to permit search indexing while restricting training crawlers, or permit a retrieval agent for selected public documentation while blocking bulk collection.
5. Block only after testing
Blocking is reasonable when a crawler has no business value, creates unwanted load or conflicts with the site’s content policy. Check for false positives, search-engine access, APIs, monitoring tools and existing WAF rules. User-agent-only blocking can also be bypassed by bots that disguise themselves.
Free tools Windows power users keep installed
One-click scans. No signup required.
6. Consider charging as an experiment, not a guaranteed revenue stream
Charging may make sense for content with clear commercial or training value, but it requires operator support and a way to assess whether the income offsets lost access or referrals. It should not be assumed that a high request count translates into meaningful revenue.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Cloudflare’s current AI crawler controls
Cloudflare’s current documentation uses the name AI Crawl Control; the product was previously referred to as AI Audit. It is documented as available on all Cloudflare plans and provides visibility into AI crawler requests, per-crawler decisions and robots.txt compliance monitoring.
Detection differs by plan. Basic detection uses user-agent strings. More advanced detection uses a Bot Management detection ID and is associated with Enterprise plans that include Cloudflare Bot Management. This can improve classification, but it should not be understood as a perfect guarantee against disguised automation.
AI Crawl Control supports three actions—Allow, Block and Charge—as described in Cloudflare’s crawler-management documentation.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →What Pay Per Crawl does
Pay Per Crawl is documented as closed beta or private beta in the current 2026 documentation, not as a universally available mature marketplace. Cloudflare documents charging for successful content retrieval, generally an HTTP 200 response, and a minimum price of $0.01 per crawl.
There are important limitations:
- One configured price applies to all crawlers assigned the Charge action, although individual crawlers can be assigned Allow, Block or Charge.
- Repeated crawls can create repeated charges, so spending and crawl frequency need monitoring.
/robots.txt,/sitemap.xml,/security.txt,/.well-known/security.txtand/crawlers.jsonare documented as free paths.- HTTP errors are not treated as successful billable retrievals.
- Charging or blocking search-engine crawlers may prevent proper indexing, so search bots need separate treatment.
Rule order is another common failure point. According to Cloudflare’s WAF integration documentation, WAF blocks occur before bot solutions and Pay Per Crawl. If an earlier WAF or bot rule blocks the request, Pay Per Crawl may never process it.
Cloudflare also documents advanced configuration involving URI exceptions and dynamic pricing through origin or Worker response headers. Those options do not change the central business question: whether a particular type of access is worth allowing, restricting or monetizing.
What the 2024 data cannot answer
Cloudflare’s review does not establish:
- How many AI requests became training data.
- How many produced citations or referrals.
- How much bandwidth or origin cost they imposed on publishers.
- How often each crawler violated robots.txt.
- How much unidentified or disguised automation escaped classification.
- Whether the overall value of AI access exceeded the value of search discovery.
Those questions require site-level logs, commercial agreements, referral measurements and content-specific economics. Cloudflare’s network-level view is valuable for showing that AI crawling was a growing policy concern, but it cannot by itself tell every publisher what access decision is profitable.
The defensible takeaway
Cloudflare Radar’s 2024 review was an early signal that AI crawlers had become a meaningful force in the web ecosystem. It was not proof that AI crawlers collectively dominated all traffic, surpassed search crawlers or delivered the most value to publishers.
The important comparison is not simply AI versus Googlebot. It is automated content consumption versus the benefits returned to the sites being crawled. Site owners should measure those benefits, keep search access distinct from training access, enforce policies in more than one layer when necessary, and treat payment experiments as experiments—not as an automatic consequence of high crawler volume.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

