# AI Crawler Statistics 2026: Which Numbers to Quote, Dated

> AI crawler statistics with each figure dated to its panel and month: which share to quote, training vs search, referral ratios, plus one site's 20-day log.

- URL: https://missiongrowth.io/blog/ai-crawler-statistics
- Published: 2026-10-06
- Author: Furkan Aktaş, Co-Founder, Mission Growth
- Publisher: Mission Growth. Company facts: https://missiongrowth.io/llms.txt

AI crawler statistics circulate as one-line shares: Meta at 52%, AI crawlers at almost 80% of AI bot traffic, GPTBot at 23.71%. Each is a claim about one panel of sites, one list of bots, one denominator and one month, which is why credible sources put the same company 6.9 times apart and GPTBot 3.9 times apart.

This page dates every figure in the AI bot traffic statistics and GPTBot statistics you will meet to its panel, and gives the number you can quote for each question: how much traffic is AI, which bot leads, what share is training, who blocks.

It also adds one site's 20-day request log, with its limits stated. The how-to side, what [GPTBot and the other AI crawlers](https://missiongrowth.io/blog/gptbot-ai-crawlers) do and how to set rules for them, lives on its own page.

The newest figures on this page are a provisional pull of Cloudflare Radar's Q3 2026 data (through September 2026, with September still open when it was read) and our own request log through October 4, 2026.

## AI crawler statistics at a glance

Fastly's almost 80% AI-crawler share, Meta's 52% and OpenAI's 98% of fetcher requests describe mid-April to mid-July 2025, so they are a 2025 snapshot that is now more than a year old.

The almost 80% is the crawler share of all AI bot traffic Fastly observed, and 39,000 requests per minute is a level that fetcher volume exceeds in some cases in that same sample (Fastly press release, August 19, 2025).

The key numbers, each with its scope and month:

- AI crawlers made up almost 80% of AI bot traffic on Fastly's network, mid-April to mid-July 2025.
- Meta generated 52% of AI crawler traffic in that Fastly sample, against Google at 23% and OpenAI at 20%.
- OpenAI's bots made 98% of fetcher requests in the Fastly sample, with fetcher volume above 39,000 requests per minute in some cases.
- Human traffic was 47% of Cloudflare's HTML requests on December 2, 2025, and non-AI bots were 44%.
- More than 50% of Internet traffic is non-human, Cloudflare said in July 2026.
- Training was 47.6% of AI crawler requests in Q3 2026 on Cloudflare Radar data pulled October 1, 2026, with September still provisional.
- Cloudflare put 52% of crawler requests down to AI training in June 2026, up from 22% in spring 2025.
- Apple Bot (27.47%) led Radware's April 2026 sample of AI crawler requests on its protected customer applications.
- 312 of 3,816 top-domain robots.txt files (8.2%) disallowed GPTBot in June 2025, per Cloudflare Radar.

One summary of this topic in Google's AI Overview packs three of the Fastly figures into present-tense sentences. It reads:

> AI crawlers make up nearly 80% of total automated AI bot traffic. Meta leads AI crawler traffic generation at 52%, outperforming Google (23%) and OpenAI (20%). OpenAI/ChatGPT drives up to 98% of real-time fetcher bot requests, scaling past 39,000 requests per minute during peak windows.

The wording drops the window. The table below dates each circulating figure and sets a newer or different reading beside it. Our [AI SEO statistics](https://missiongrowth.io/blog/ai-seo-statistics) hub holds the wider set of AI SEO impact numbers. How often Google shows an Overview at all is covered in [Google AI Overviews statistics](https://missiongrowth.io/blog/ai-overviews-statistics).

::dataset{key="circulating-figures-dated" name="Circulating AI crawler figures, what each measured and a newer reading"}

| Circulating figure | What it measured and when | Newer or different reading |
|---|---|---|
| AI crawlers almost 80% of AI bot traffic | Fastly: crawlers' share of all AI bot traffic observed, mid-April to mid-July 2025 | Our log, single-purpose AI requests: 89.4% crawlers, 10.6% fetchers (one site, September 15 to October 4, 2026) |
| Meta 52% | Fastly: Meta's share of AI crawler traffic, same window | Cloudflare: Meta 7.5% of AI and search crawler traffic, July 2025 |
| OpenAI 98% of fetcher requests, 39,000 per minute | Fastly: OpenAI's share of fetcher requests and a rate that fetcher volume exceeds in some cases, same window | Our log: OpenAI's fetchers made 80.0% of user-triggered requests |
| North America nearly 90% | Fastly: share of observed AI crawler activity, same window | Our log: 90.2% of requests came from the United States (where requests came from, not which sites were crawled) |
| Apple Bot 27.47%, GPTBot 23.71% | Radware: shares of AI crawler requests on its protected customer applications, April 2026 | Census of two properties: GPTBot 26.2% of AI crawl, January 11 to July 3, 2026 |
| Training and user-action shares for Cloudflare Radar | An earlier reading from July 2026, before Cloudflare relabelled the series | Radar data pulled October 1, 2026: training 47.6% in Q3 2026, September provisional |
| The news-site block rate | A sample of news sites | GPTBot disallowed by 8.2% of top domains with a robots.txt file, June 2025 |

The circulating figures above come from different panels and months, and the oldest are 15 to 16 months old. For the fetcher figure, one site gives a first-party comparison:

OpenAI's fetchers made 375 of the 469 user-triggered requests in our log (80.0%).

That is one site, Worker-logged requests only, and it is no confirmation of Fastly's 98%: the two samples differ in window, size and what counts as a fetcher.

## Which AI crawlers lead? The ranking depends on the panel

Radware's April 2026 sample puts Apple Bot (27.47%) first, then GPTBot (23.71%), OAI-SearchBot (17.13%), ChatGPT-User (11.50%) and meta-externalagent (11.49%) of AI crawler requests on its protected customer applications, and no other panel in the table below ranks Apple Bot first.

Radware's list is one vendor's customers in one month, and Applebot is mixed-use: Cloudflare lists it with Bingbot and Googlebot as a mixed-use crawler in September 2026, and Apple documents it for search in Spotlight, Siri and Safari, with data that may also help train Apple foundation models.

::dataset{key="panel-ranking" name="Which AI crawler leads, by panel, window and denominator"}

| Panel | Window | The share is of | Leaders |
|---|---|---|---|
| Fastly | Mid-April to mid-July 2025 | AI crawler traffic | Meta 52%, Google 23%, OpenAI 20% |
| Cloudflare | May 2025 | AI-only crawling | GPTBot 30%, Meta-ExternalAgent 19% |
| Cloudflare | May 2025 | All AI and search crawler traffic | GPTBot 7.7% |
| Cloudflare | July 2025 | All AI and search crawler traffic | Googlebot 39%, Meta 7.5% |
| Radware | April 2026 | AI crawler requests on its protected customer applications | Apple Bot 27.47%, GPTBot 23.71% |
| Two-property census | January 11 to July 3, 2026 | AI crawl | GPTBot 26.2% |
| Our log (one site) | September 15 to October 4, 2026 | All 20,180 logged requests, AI or not | Meta family 25.7%, Googlebot 9.6%, ClaudeBot 5.5% |

Read the last row with care, because it counts all 20,180 requests in our log.

Meta's crawler family made 5,186 of those requests (25.7%), more than any other family; Googlebot made 1,940 (9.6%) and ClaudeBot 1,108 (5.5%). Our Worker files Meta's training crawler, fetcher, web indexer and link-preview agents under one family name, and the Amazonbot family also counts Amzn-User and Amzn-SearchBot.

That 25.7% is a labelled bundle of several tokens, so don't read it as one bot's share.

### Why Meta runs from 7.5% to 52%

Meta's share of AI crawling runs 6.9 times apart across panels, from 7.5% to 52%, and three causes explain it.

Cloudflare's 7.5% for July 2025 is a share of AI and search crawler traffic, a denominator that includes Googlebot at 39%. Cloudflare's 19% for May 2025 is a share of AI-only crawling, and Radware's 11.49% for meta-externalagent comes from its own customer applications in April 2026. Fastly's 52% sits on a different panel and grouping again.

So 7.5% to 19% is mostly a denominator effect. The step from 19% to 52% happens on a similar denominator and window, which leaves panel population or bot grouping as the cause. That split is our reading of the published figures; none of the panels states it.

::figure{src="/blog/figures/ai-crawler-statistics-1.svg" alt="Dot-range chart of Meta's crawler share across published panels, from 7.5% on Cloudflare in July 2025 to 52% on Fastly, with our own 25.7% marked as a bundle." caption="Meta's share of AI crawling runs from 7.5% to 52% depending on the panel and denominator, so a single number is never the answer. By Cloudflare's and Fastly's counts." width="720" height="368"}

### One bot, several classes

The same bot lands in different classes depending on who counts it. Cloudflare's September 2026 post lists Amazon among the operators with training-only crawlers without naming the token, the two-property census files Amazonbot under AI search indexing, and Amazon documents it as improving its products and services, with data that may be used to train Amazon AI models.

Cloudflare classifies bots by behavior, and one bot can have more than one.

::dataset{key="class-by-source" name="One bot, different classes by source"}

| Bot or family | What the operator documents | Cloudflare, September 2026 | Census | Our log class |
|---|---|---|---|---|
| GPTBot | Crawls content that may be used to train OpenAI's foundation models | OpenAI's training-only crawler | - | Training |
| ClaudeBot | Collects web content that could contribute to model training | Anthropic's training-only crawler | - | Training |
| Claude-SearchBot | Improves search result quality | - | - | AI search |
| ChatGPT-User, Claude-User | Visit pages when a user asks | - | - | User fetch |
| Amazonbot family | Improves Amazon products and services, may train Amazon AI models | Amazon's training-only crawler, token not named | AI search indexing | Bundled family |
| Meta-ExternalAgent family | Not documented here; Cloudflare is the source | Meta's training-only crawler | - | Bundled family |
| Applebot | Search in Spotlight, Siri and Safari, may also help train; Applebot-Extended is the opt-out | Mixed-use | - | Mixed-use |
| Googlebot, Bingbot | Classic search crawlers | Mixed-use | - | Mixed-use |

Our rule has four parts. Class a token by what its operator documents it for: training, AI search or user-triggered fetch. Keep Googlebot, Bingbot and Applebot as mixed-use crawlers, as Cloudflare does. Keep bundled families such as Meta and Amazon out of every single class. Then report the share together with the rule.

## How much web traffic is AI crawler traffic?

AI crawler traffic share is 4.2% of HTML requests on Cloudflare in 2025, 20% of verified-bot traffic and 7.2% of all requests on a two-property census, so each denominator gives a different percentage. Ask which denominator a share uses before you quote it, because the same crawling can be a single digit or a fifth of the total.

::dataset{key="share-denominators" name="AI crawler share of traffic under five denominators"}

| Panel | Window | Denominator | AI share |
|---|---|---|---|
| Cloudflare | 2025 | Share of HTML requests | Other AI bots 4.2%, against Googlebot alone at 4.5% |
| Cloudflare | 2025 | Verified Bot traffic | AI crawlers 20%, search engine crawlers 40% |
| Two-property census | January 11 to July 3, 2026 | All 1,348,706 requests | 7.2% |
| Radware | April 2026 | All legitimate bot traffic (legitimate bots, crawlers and AI crawlers) | Around 30% |
| Fastly | Q2 2025 | Observed activity on its network | Automated bots 37%, the bot share rather than AI |

Googlebot belongs in the comparison too. Counting it as AI would put Googlebot at about 52% of the HTML requests from Googlebot and the other AI bots together (4.5% against 4.2%, our arithmetic from Cloudflare's 2025 figures), which is why Cloudflare's 2025 review, the census and Radware each report it apart from the AI bots.

The human-versus-bot answer also depends on the date and the denominator: Cloudflare's December 2, 2025 reading counts HTML requests, while its July 2026 statement covers all Internet traffic, so treat them as two separate measurements. Sector matters as well: Retail and Computer Software together drew just over 40% of AI crawler activity on Cloudflare in October 2025.

### What one site's log looks like

The same question on one site gives a sixth denominator, the site's own requests.

In our own request log, 20,180 requests reached our Worker between September 15 and October 4, 2026 (one site, 79 named bot families plus three catch-all labels, requests that hit the Worker only, user agents as declared); 1,436 of them, or 7.1%, were robots.txt and sitemap fetches.

We sort families by the purpose their operator documents for the token: single-purpose AI training crawlers made 1,880 requests (9.3%), AI search crawlers 2,076 (10.3%), user-triggered fetchers 469 (2.3%), mixed-use or bundled families (Meta, Amazonbot, Applebot, Googlebot, Bingbot) 9,821 (48.7%), search-only bots 2,124 (10.5%) and other automated clients 3,810 (18.9%).

::dataset{key="log-classes" name="Our request log by class under a stated family rule"}

| Class | Families | Requests | Share of 20,180 |
|---|---|---|---|
| Mixed-use or bundled families | Meta-ExternalAgent family, Amazonbot family, Applebot, Googlebot, Bingbot | 9,821 | 48.7% |
| Other automated clients | Everything else, including SEO-Crawler, Google-Fetcher, other-bot, http-client and unclassified | 3,810 | 18.9% |
| Search-only bots | DuckDuckBot, PetalBot, Search-Crawler, Bravebot, GoogleOther, Google-InspectionTool | 2,124 | 10.5% |
| AI search crawlers | OAI-SearchBot, PerplexityBot, Claude-SearchBot, ExaSearchBot, DuckAssistBot, Kimi-SearchBot | 2,076 | 10.3% |
| Single-purpose AI training crawlers | GPTBot, ClaudeBot, Bytespider, cohere-ai, DeepSeekBot | 1,880 | 9.3% |
| User-triggered fetchers | ChatGPT-User, Claude-User, Perplexity-User, Claude-Code, ChatGPT-Agent, Gemini-Deep-Research, GoogleAgent-URLContext, Google-GeminiNotebook, Shap-User | 469 | 2.3% |

::figure{src="/blog/figures/ai-crawler-statistics-2.svg" alt="Stacked bar for AI crawler statistics on our own site: of 20,180 logged requests, mixed-use families made 48.7% and training crawlers 9.3%." caption="On our own site, mixed-use or bundled families made 48.7% of logged requests and single-purpose training crawlers 9.3%." width="720" height="504"}

The training share on this site rests on how the mixed-use families are classed. If the Meta and Amazon families (6,686 requests, 33.1%) are counted as training instead, our training requests rise from 1,880 to 8,566 (42.4% of all 20,180).

So the same log gives a training share of 9.3% or 42.4%, and the swing comes from two classification calls. On a single site, name the families whose class is a judgment before you quote a training share, and pick the denominator before the number.

## Training, search and fetch: what AI crawlers are for

Cloudflare Radar data put training at 47.6% of AI crawler requests in Q3 2026 (data through September 2026, September provisional), and the training and user-action figures in Google's AI Overview were an earlier July 2026 reading, before Cloudflare relabelled the series again.

We read the Q3 2026 figure on a third-party page updated October 1, 2026, which pulled it from Cloudflare Radar's crawl-purpose summary that day and also reports search crawling at 8.3%. Its denominator includes a mixed-purpose bucket, and the page itself calls September and the Q3 totals provisional.

Cloudflare has published no post with this figure that we could find, and we could not read the live Radar dashboard, so treat it as a third-party pull and the page's account of Cloudflare relabelling the series as unconfirmed.

The earlier reading is the one in circulation:

> Cloudflare Radar data indicates roughly 44.54% of AI crawler requests target model training, while live user fetches (RAG / real-time retrieval) account for just 2.66%.

Cloudflare's own June 2026 report puts 52% of crawler requests down to AI training, up from 22% in spring 2025, and says mixed-use crawlers, which blend search, agent use and training, represent over 36% of activity. That mixed-use bucket changes what the 52% divides by.

Do not line the 22% up against Cloudflare's July 2025 reading, which put training at 79% of AI crawling and search at 17%. The 2026 report's population is crawler requests with a mixed-use bucket, so the two are not like for like; that is our reading of the two posts. The 2025 post also reported user actions growing modestly from 2% to 3.2%.

In the two-panel split below, both sources divide AI crawling three ways and land close together. A single site does not have to.

::dataset{key="purpose-splits" name="Training, AI search and user fetch shares across panels"}

| Split | Training | AI search | User fetch |
|---|---|---|---|
| Cloudflare, July 2025 | 79% | 17% | 3.2% |
| Two-property census, January 11 to July 3, 2026 | 68.9% | 21.6% | 9.5% |
| Our log, single-purpose families only (one site) | 42.5% | 46.9% | 10.6% |

::figure{src="/blog/figures/ai-crawler-statistics-3.svg" alt="Grouped bars of training, AI search and user-fetch shares: Cloudflare 79%, 17%, 3.2%; a census 68.9%, 21.6%, 9.5%; our site 42.5%, 46.9%, 10.6%." caption="Two panels that split AI crawling three ways put training at 68.9% and 79%; our single-purpose split is 42.5% because mixed-use families are left out. By Cloudflare's count for July 2025." width="720" height="342"}

Our single-purpose split differs because its denominator excludes the 48.7% of requests from mixed-use families. Training sits 36.5 points below Cloudflare's July 2025 figure, AI search 29.9 points above it and user fetch 7.4 points above it, so on this site agreement with either panel is no safe assumption.

Counting only the single-purpose families, 4,425 requests, our split is 42.5% training, 46.9% AI search and 10.6% user fetch; crawlers (training plus search) are 89.4% of them and fetchers 10.6%.

The trend over 2025 is steep from a small base. Cloudflare reported AI user-action crawling up over 15x in 2025, and crawling for model training reached as much as 7-8x search crawling and 32x user-action crawling at peak.

Compare purpose shares only inside one taxonomy and one period. The 79% and 68.9% of the two three-way panels agree within about 10 points; any figure from a different taxonomy, a different month or a mixed-use bucket answers a different question.

## Crawl-to-referral ratios: how many crawls buy one visit

Anthropic crawled 38,000 pages for every referred visit in July 2025, about 26 referred visits per million crawls, while Perplexity's 194 crawls per visit is about 5,155 per million. Cloudflare defines the crawl-to-referral ratio as HTML crawl requests from a platform's crawler divided by HTML requests that platform referred, which measures human visits.

A ratio is easier to compare once you invert it to referred visits per million crawls:

| Platform | Crawls per referred visit | Window | Referred visits per million crawls |
|---|---|---|---|
| Anthropic | 38,000 | July 2025 | About 26 |
| Perplexity | 194 | July 2025 | About 5,155 |

In that month Perplexity's ratio was about 196 times lower than Anthropic's. Cloudflare's 2025 year-in-review gives the wider ranges:

- Anthropic reached as much as 500,000:1 early in 2025, then ranged from about 25,000:1 to about 100,000:1.
- OpenAI reached as much as 3,700:1 in March 2025.
- Perplexity started the year below 100:1 and spiked above 700:1 in late March.
- Google started just over 3:1 and rose to as high as 30:1.

These are 2025 readings kept as history. Cloudflare itself describes the Anthropic series as erratic from January through May 2025 and OpenAI's as spiky, and Cloudflare Radar's live series shows current values that this page does not reproduce.

Invert the ratio before you compare, and read any current ratio from the live series. To tie crawl traffic to the visits it does or does not produce on your own site, see how to [measure AI search traffic](https://missiongrowth.io/blog/ai-search-analytics). Who gets cited is a separate question, which our [AI citation statistics](https://missiongrowth.io/blog/ai-citation-statistics) page covers.

## Who blocks AI crawlers, and who ignores the block

GPTBot was disallowed by 312 of 3,816 top-domain robots.txt files (8.2%) in June 2025, so the news-site block rate in circulation describes a sector, not the web, and the figure moves with the sample. Google's AI Overview states the news-site claim like this:

> Roughly 49.4% of top news sites completely block OpenAI's GPTBot via robots.txt, with over 54% blocking at least one major AI scraper.

That is a statement about news sites, a sector sample, and the top-domain figure below sits far under it.

::dataset{key="block-rate-samples" name="AI crawler block rates by sample"}

| Sample | What was measured | Rate |
|---|---|---|
| Top 10,000 domains with a robots.txt file (3,816), June 6, 2025 | GPTBot disallowed in any form | 312 domains (8.2%) |
| Same 3,816 domains | GPTBot disallowed fully | 250 domains (6.6%) |
| Same 3,816 domains | Any allow or disallow directive aimed at AI bots | 546 domains (about 14%) |
| Cloudflare sites, post of September 15, 2026 | Enable some mechanism to block training | 17% |
| Cloudflare sites, post of September 15, 2026 | Block search bots | Under 1% |

The pattern in Cloudflare's own data is to block training and allow search: 17% of its sites enable some training block while fewer than 1% block search bots. For the per-role rules behind that pattern, see our guide on how to [allow or block AI crawlers](https://missiongrowth.io/blog/gptbot-ai-crawlers).

A robots.txt honeypot shows the other half, who ignores the block.

On one of the census's two properties, 125 distinct IPs of declared AI crawlers violated the honeypot between February 7 and June 29, 2026, and 124 of them identified as Bytespider. One operator accounts for nearly all declared violations. The census never checked source IPs against the ranges operators publish, so its user agents are self-declared and a scraper impersonating a crawler counts as that crawler.

A block rate means nothing without its sample, and a compliance rate means nothing without knowing who declared what.

## What AI crawlers fetch on missiongrowth.io

In our log, 10.3% of the 20,180 requests asked for a markdown copy of a page, ranging from 0.1% of PerplexityBot's requests to 89.5% of ExaSearchBot's.

A markdown copy means a request for a page's markdown version, a .md URL or a request that prefers text/markdown. The log does not say why a bot asks for it. Everything in this section is one site, Worker-logged requests only.

2,075 of the 20,180 requests (10.3%) asked for a markdown copy of a page: ExaSearchBot did so on 681 of its 761 requests (89.5%), ShapBot on 386 of 452 (85.4%) and Applebot on 217 of 409 (53.1%), while the Meta family did on 76 of 5,186 (1.5%) and PerplexityBot on 1 of 755 (0.1%).

::dataset{key="markdown-by-bot" name="Markdown-copy requests by bot in our request log"}

| Bot | Requests | Markdown requests | Share |
|---|---|---|---|
| ExaSearchBot | 761 | 681 | 89.5% |
| ShapBot | 452 | 386 | 85.4% |
| Applebot | 409 | 217 | 53.1% |
| GPTBot | 760 | 125 | 16.4% |
| ClaudeBot | 1,108 | 142 | 12.8% |
| Amazonbot | 1,500 | 166 | 11.1% |
| Meta-ExternalAgent family | 5,186 | 76 | 1.5% |
| PerplexityBot | 755 | 1 | 0.1% |

The llms files tell a quieter story. missiongrowth.io publishes its own llms.txt and llms-full.txt, a curated plain-text knowledge base for AI crawlers, and our log records every request for them from any client.

Of 118 requests for our llms.txt and llms-full.txt files (73 and 45), 8 came from single-purpose AI training, search or fetcher families and 15 from the mixed-use families; 53 came from clients with no recognised bot user agent, which includes browsers and is not bot traffic.

The other 42 came from other automated clients. To check your own file, our [llms.txt checker](https://missiongrowth.io/tools/llms-txt-checker) tests it, and the [llms.txt file](https://missiongrowth.io/blog/llms-txt-guide) guide covers what belongs in it.

Two more first-party frames. The first is geography:

90.2% of the requests in our log came from the United States (18,211 of 20,180).

That measures where requests came from and says nothing about which sites were crawled, so it can't confirm or refute the North America figure above. The second is volume. Daily request counts ranged from 162 to 3,228 over the 20 days (median day 657.5, average 1,009), so the window is not a stable baseline.

JavaScript is the older caveat. A Vercel and MERJ analysis found in December 2024 that the crawlers of OpenAI, Anthropic, Meta, ByteDance and Perplexity do not render JavaScript, while Gemini (through Googlebot's infrastructure) and Applebot do. We migrated our own React single-page app to prerendered static HTML for 20 marketing pages because AI crawlers do not execute JavaScript.

A bot's request mix tells you which version of your page it wants, so read your own log bot by bot. Single-purpose AI families made 8 of the 118 llms-file requests on this site, so check your own log before you assume AI crawlers read llms.txt.

## How to count AI crawlers in your own log

To count AI crawlers in a server log, group requests by user-agent token, assign each token to training, AI search, user fetch, mixed-use or other, and divide each single-purpose class by the single-purpose total instead of by all requests. The steps:

1. Pull one fixed window and write down its length in days, since one day can be 20 times another.
2. Group requests by user-agent token and sum them.
3. Assign each token to training, AI search, user fetch, classic search, mixed-use or other, using what its operator documents.
4. Keep mixed-use and bundled families out of every single class.
5. Divide each single-purpose class by the single-purpose total, then show the mixed-use total beside it.

Our worked example: 1,880 training, 2,076 AI search and 469 user-fetch requests make 4,425 requests of single-purpose AI traffic. Dividing each by 4,425 gives the split in the purpose table above. The 9,821 requests from mixed-use families stay out of all three, and we print them beside the result so a reader can see how much of the log the split leaves aside.

### How we counted ours

Our log only sees requests that reach our Worker and match a known bot family or a generic bot pattern, plus every request for llms.txt or llms-full.txt from any client; blocked or edge-cached requests and human page views are not in it.

The window is September 15 to October 4, 2026, which is 20 days with hits. Family labels bundle tokens: the Meta family also counts meta-externalfetcher, meta-webindexer, FacebookBot and facebookexternalhit, and the Bingbot family also counts BingPreview and AzureAI-SearchBot.

### Why your numbers will differ

- Panels count verified bots or declared user agents; a raw log counts only the second. Treat self-declared user agents as claims: a scraper impersonating GPTBot inflates GPTBot.
- Denominators differ: all requests, HTML requests, verified-bot traffic or AI traffic only.
- Family labels in your own tooling may bundle tokens of different purposes, as ours does.
- A short window on a small site swings widely, as our 162 to 3,228 requests a day showed.

A user-agent count is a claim that can be wrong. A share of AI crawler traffic is a property of a denominator, and on one site it also hinges on how the families that mix purposes (Meta, Amazon, Apple, Google, Bing) are classed. Split your own log this week into training, AI search, user fetch, mixed-use and other, and publish the share with the rule that produced it.

## FAQ

### Is 50% of internet traffic bots?

On some panels, yes. Cloudflare counted 47% of HTML requests as human and 44% as non-AI bots on December 2, 2025, and said in July 2026 that more than 50% of Internet traffic is now non-human. The denominators differ, so the two readings are not a trend line.

### How many bot requests does a small site get?

One site's Worker logged 20,180 requests in 20 days, counting known bot families, generic bot patterns and every llms.txt and llms-full.txt request, with a median of 657.5 a day and a range of 162 to 3,228. That is one site, Worker-logged requests only, and the window is not a stable baseline, so treat it as an order of magnitude and measure your own.

### What percentage of AI crawler traffic is for training?

It depends on the panel and taxonomy. A provisional pull of Cloudflare Radar data shows 47.6% for Q3 2026, and Cloudflare reported 52% for June 2026 with a mixed-use bucket. Two panels that split three ways give 79% (Cloudflare, July 2025) and 68.9% (a two-property census).

### Why don't my server logs match these AI crawler numbers?

Published panels count verified bots or declared user agents over different denominators and months. A server log counts user agents as declared, and family labels can bundle several tokens. Our own log is user-agent based, which is why we show the class rule beside every share.

### What is an AI crawler?

An AI crawler is a program that requests web pages on behalf of an AI system. It does one of three jobs: collecting content for model training, indexing pages for AI search, or fetching a page on a user's request. Mixed-use bots such as Googlebot sit between those jobs.
