TrendingThe thresholds you restructure your account around were never published.
SEO

Everyone reports AI traffic from a channel grouping that does not exist.

The slide says AI traffic is up and to the right. Underneath it, some assistants send a referrer and some send nothing, the biggest AI surface of all arrives labeled as ordinary organic search, and the crawler that decides whether you get cited at all is a different bot from the one everybody blocked. Here is what is actually measurable, and what to stop claiming.

MSMikołaj Salecki, portrait
Editor-in-chief
Aug 14, 2026·7 min read
Several thin ribbons converging on a hollowed plaster head, most of them bone-white and unlabeled while one is solid brand blue, with a dotted measurement matrix behind and a fine registration rule beneath them
Plenty of it arrives. Almost none of it arrives wearing a name tag.Illustration: Mediovsky · generated with AI
TL;DR
  • There is no default AI channel. Referrer-sending assistants land in Referral, the rest land in Direct.
  • Clicks from AI Overviews are clicks on Google results, so they arrive as ordinary organic.
  • Google documents four Performance report search types: web, image, video, and news. [4]
  • OpenAI splits GPTBot for training from OAI-SearchBot for search citations, plus an ads validator. [1]
  • Anthropic splits ClaudeBot, Claude-SearchBot, and Claude-User the same way. [2]
  • Perplexity states that Perplexity-User generally ignores robots.txt, because a person asked. [3]
  • Ahrefs published 0.5% of visitors driving 12.1% of signups, a 23x rate. One company’s funnel. [5]
  • Report absolute numbers with denominators, and say what you are not counting.

There is a slide going around every marketing department right now. AI traffic, up several hundred percent, with a line going up and to the right. It is built on a channel grouping nobody defined, counting sources nobody listed, against a denominator nobody mentions.

Nothing on it is deliberately dishonest. It is just that almost every part of it is harder to measure than the chart implies.

The underlying interest is entirely legitimate. Something real is happening and it deserves measurement. The problem is that the measurement is harder than the slide implies, in four specific ways that are worth knowing before you present anything.

What actually arrives

Some assistants send a referrer header when a user clicks through to you. Those visits land in your Referral channel, mixed in with every blog and directory that ever linked to you, identifiable only if you know which hostnames to look for.

Some send nothing, and those visits land in Direct, indistinguishable from someone typing your URL. There is no way to recover them after the fact. They are simply not labeled.

And the largest AI surface most sites are exposed to is not a referral at all. A click from an AI Overview is a click on a Google search result, which arrives as ordinary organic search with no marker distinguishing it from a click on a blue link. Google’s documentation for the Search Console Performance report lists the available search types as web, image, video, and news. [4] There is no separate AI surface documented there to filter on.

So the honest starting position is that you can see some of it, none of the biggest part of it, and you will never know the size of what you are missing.

The bots, which are three different things wearing similar names

This is the part with real operational consequence, and it is where most sites have quietly made a decision they did not intend to make.

Every major assistant runs distinct crawlers for distinct purposes, and the vendors document them clearly.

Vendor Training Search and citation User-initiated Other
OpenAI GPTBot OAI-SearchBot ChatGPT-User OAI-AdsBot, which validates pages submitted as ads
Anthropic ClaudeBot Claude-SearchBot Claude-User
Perplexity none stated for foundation models PerplexityBot Perplexity-User
Google Google-Extended, a control token rather than a crawler Googlebot, the same one as Search Google-CloudVertexBot, for owner-requested crawls

The distinctions are the vendors’ own. OpenAI describes GPTBot as crawling “content that may be used in training our generative AI foundation models,” while OAI-SearchBot is “used to surface websites in search results in ChatGPT’s search features,” and blocking that one means sites will not appear in ChatGPT search answers. It also runs OAI-AdsBot to “validate the safety of web pages submitted as ads on ChatGPT.” [1] Anthropic splits the same three ways: ClaudeBot for content that could contribute to training, Claude-SearchBot to “improve search result quality for users,” and Claude-User for when an individual asks Claude a question and it accesses a site. [2] Perplexity is explicit that PerplexityBot “is not used to crawl content for AI foundation models” and exists to surface and link sites in its results. [3]

Google is the odd one out and the difference is worth understanding, because it is the case people get wrong most often. It does not run a separate AI crawler. Google-Extended is a control token rather than a bot, letting publishers manage whether already-crawled content may be used for training Gemini models and for grounding, and Google states that it “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search.” [6] Which means the crawler feeding AI Overviews is Googlebot, the same one that has always been there, and there is no way to accept Search while declining that surface.

The decision most sites made by accident

What happened: in the first wave of blocking, a great many sites added the training crawler to robots.txt, left the search crawler alone, saw no drop in anything, and concluded that blocking AI bots costs nothing.

What that actually did: it declined to contribute training data, which was the intent, and left citation eligibility untouched. Both outcomes were reasonable. Neither was chosen deliberately, and the reverse mistake, blocking the search crawler and keeping the training one, is the same error pointed at your visibility instead.

One more asymmetry matters. User-initiated fetchers behave differently on purpose, because a person asked for the page. Perplexity states that Perplexity-User “generally ignores robots.txt rules” since a user initiated the request, and OpenAI notes the same distinction for ChatGPT-User. [3][1] Robots.txt is a lever on crawling and citation. It was never a wall against a human clicking through.

Three near-identical plaster keys lying in a row on a dotted grid, each cut to a slightly different profile, the middle one rendered in solid brand blue beside a single closed stone door
Three keys that look alike. Only one of them opens the door marked citation.Illustration: Mediovsky · generated with AI

Building a view that admits what it does not know

The referrer side is a custom channel group, and the work is entirely in maintaining the list rather than in writing the rule.

Source matches regex
  chatgpt\.com|openai\.com|perplexity\.ai
OR Source matches regex
  copilot\.microsoft\.com|gemini\.google\.com
OR Source matches regex
  claude\.ai|you\.com|phind\.com

Three rules keep this from decaying into the thing it was meant to replace. Keep the list in one documented place with a date on it, because new surfaces appear constantly and a stale regex silently undercounts. Name the channel for what it measures, something like “AI assistant referrals,” not “AI traffic,” so nobody reads it as the total. And record what is excluded, in the same place, so the next person does not have to rediscover that AI Overviews are absent.

The log side is the other half, and it is the half most teams skip. Referrer data tells you who arrived. Server logs tell you who is reading you at all, which crawlers are visiting, how often, and how deep. For a question like whether you are eligible to be cited, the log is the primary evidence and the analytics view has nothing to say.

The quality question, honestly

Every published dataset points the same direction: visitors arriving from assistants convert better than organic search visitors. The direction is consistent enough to believe.

The magnitude is not. Ahrefs published its own first-party numbers showing AI search visitors making up 0.5% of its traffic while driving 12.1% of signups, which it works out as converting 23 times better than traditional organic. [5] That is a real figure from a real funnel, and it describes a company selling software to exactly the audience most likely to be using AI assistants heavily in the first place. Other published datasets report multiples that differ from that one by more than an order of magnitude, in both directions.

The direction is a finding. The multiple is someone else’s funnel.

Two things are going on underneath the spread, and both are worth naming when someone quotes a number at you. Selection: people who arrive after an assistant has already answered their question and recommended you are far down the consideration path, so a high conversion rate is partly a description of who they were before they arrived. And attribution: because a chunk of assistant traffic arrives as Direct, whichever slice you can identify is a biased sample of the whole.

None of that means the effect is fake. It means the number is yours to measure, and the honest version of the claim is that this traffic converts well for us, at this volume, on this definition.

What to actually report

  • Absolute sessions and conversions, with the denominator visible. A percentage on a base of 200 sessions is noise wearing a suit.
  • The explicit list of sources counted, with the date the list was last updated.
  • A one-line statement of what is not counted, naming AI Overviews and referrer-stripped visits.
  • Crawler activity from server logs, separated into training, search, and user-initiated, because they answer different questions.
  • A deliberate robots.txt position per crawler type, written down with the reasoning, rather than one inherited rule.
  • Conversion rate compared against your own organic baseline, never against a published multiple.
  • The trend over months rather than weeks, since the base is small enough that a single week is mostly variance.

The reason to do this properly is not the traffic, which is still small for most sites. It is that the measurement is the feedback loop for everything else you might do about AI search. Getting cited by these systems is a strategy with real work attached, and without an honest view of who arrives and which crawlers can reach you, that work has no scoreboard. The teams that will get this right are not the ones with the biggest AI traffic number this quarter. They are the ones whose number still means the same thing next quarter, because somebody wrote down what it counts.

Sources

  1. OpenAI · OpenAI crawlers and user agentsGPTBot for training, OAI-SearchBot for surfacing sites in ChatGPT search features, OAI-AdsBot for validating pages submitted as ads, and ChatGPT-User for user-initiated actions
  2. Anthropic · Does Anthropic crawl data from the web, and how can site owners block the crawlerClaudeBot, Claude-SearchBot, and Claude-User, and the robots.txt directives honored
  3. Perplexity · Perplexity bots and user agentsPerplexityBot for surfacing and linking sites and explicitly not for foundation-model training, and Perplexity-User which generally ignores robots.txt because a user initiated the request
  4. Google Search Console Help · Search performance reportthe documented search types: web, image, video, and news
  5. Ahrefs · Does AI search traffic convert better than traditional search? For Ahrefs, yesfirst-party data from one company: 0.5% of visitors driving 12.1% of signups, reported as a 23x conversion rate against organic search
  6. Google Search Central · Google crawlers and user agentsGoogle-Extended as a standalone control token for Gemini training and grounding, explicitly not affecting inclusion in Google Search nor used as a ranking signal, plus Google-CloudVertexBot for owner-requested crawls

Frequently asked questions

Does GA4 have an AI channel by default?

No. Assistants that send a referrer land in Referral alongside every other site that links to you, and assistants that send nothing land in Direct. There is no default grouping that isolates them, which is why almost every AI traffic number in circulation was assembled by hand from a list somebody wrote once and never updated.

Can I see clicks from AI Overviews?

Not as a distinct source in your analytics. A click from an AI Overview is a click on a Google search result, so it arrives as ordinary organic search. Google’s documentation for the Search Console Performance report lists web, image, video, and news as the search types, so there is no separate AI surface documented there either.

Which AI crawlers should I actually care about?

Three kinds, and conflating them is the expensive mistake. Training crawlers collect content that may be used to train models. Search crawlers decide whether you can be surfaced and cited in the assistant’s answers. User-initiated fetchers visit a page because a person just asked about it. They have different names and different consequences.

If I block GPTBot, do I disappear from ChatGPT?

No. OpenAI documents GPTBot as the crawler for content that may be used in training its foundation models, while OAI-SearchBot is what surfaces websites in ChatGPT’s search features. Blocking OAI-SearchBot is what removes you from those answers. Many sites blocked the training crawler, kept the search one, and concluded incorrectly that blocking had no cost.

Do user-initiated fetchers respect robots.txt?

Often not, by design, because a person asked for the page. Perplexity states plainly that its Perplexity-User fetcher generally ignores robots.txt rules since a user initiated the request. OpenAI notes the same distinction for ChatGPT-User. Your robots.txt is a lever on crawling and citation, not a wall against a human clicking through an assistant.

How do I build an AI traffic view that is honest?

Create a custom channel group matching the known assistant hostnames, keep the list in one documented place, and accept that it is incomplete. Then read server logs for the bot side, because referrer data tells you who arrived and log data tells you who is reading you at all.

Is AI traffic really higher quality?

The direction is consistent and the size of the effect is not. Ahrefs published its own first-party figures showing AI search visitors at 0.5% of traffic driving 12.1% of signups, a 23x conversion rate against organic. Other published datasets report multiples that differ by more than an order of magnitude, so treat any specific multiple as one company’s funnel rather than a benchmark.

What should I report to leadership?

Absolute numbers with their denominators, the list of sources you are counting, and a plain statement of what is not counted. A percentage growth figure on a base of two hundred sessions is not a finding, and reporting it as one is how a measurement program loses its credibility in a single meeting.

Found this useful?
MSMikołaj Salecki, portrait
Editor-in-chief

Mikołaj Salecki

Writes about media, tech, and AI business for people who actually run digital. Former agency lead. Skeptic of frameworks that read better than they perform.

More articles →