Published
What attribution asks that visibility does not
AI search attribution is the problem of establishing, from your own analytics and server logs, that a visitor arrived because of an AI answer — which surface sent them, for what, and what they did next. It is a strictly narrower question than AI visibility. Visibility asks whether you appeared; attribution asks whether appearing produced anything.
The plumbing is ordinary and decades old. A browser sends a Referer header naming the page the link sat on, analytics software maps that hostname to a source, and some AI products additionally append a UTM parameter — a query string such as ?utm_source=chatgpt.com read by the destination site as an explicit tag. Where either signal survives the journey, the visit is attributable.
The difficulty is that a large and unmeasured share of AI-influenced visits arrive with neither signal and land in analytics as direct traffic, the bucket for sessions with no discoverable origin. The reasons are structural rather than fixable, and the most important of them is not a technical failure at all: when an AI answer names a brand without linking it and the person then searches for that brand or types the domain, there is no referral event to lose. That is a visit with no referrer to break, and no instrument anywhere can see it.
What that produces is not a measurement but a floor. Every AI referral figure you have is a lower bound of unknown tightness, and should be labeled that way on every report built from it.
The signals, surface by surface
Only one operator has made a published commitment. Everything else in this list is a convention observed in the wild.
- ChatGPT — referrer
chatgpt.comand the UTM parameterutm_source=chatgpt.com. Documented by OpenAI. - Perplexity — referrer
perplexity.ai. Observed, no published convention. - Gemini app — referrer
gemini.google.com. Observed, and it appears in no Google reporting surface at all. - Microsoft Copilot — referrer
copilot.microsoft.com, observed since 29 March 2024 and never announced by Microsoft in the two years since. - Claude — referrer
claude.ai, with a path form ofclaude.ai/referral/YYYY-MM-DDreported by a practitioner. Anthropic documents nothing. - Google AI Overviews and AI Mode — referrer
google.com, indistinguishable from organic search. - Bing, DuckDuckGo and Brave AI answers — the engine's own hostname, indistinguishable from that engine's organic traffic.
- Grok — referrer
grok.comorx.com, the second of which cannot be separated from ordinary social traffic.
Two structural facts fall out of that list: the surface with the most reach is the one you cannot identify, and the surface you identify best is identifiable only because one vendor chose to make it so.
The largest surface is the least attributable
Google AI Overviews and AI Mode send clicks with a google.com referrer. In analytics they are Google organic traffic and nothing distinguishes them. This is not a gap that better tagging or a smarter segment closes; there is no signal to segment on, and any dashboard presenting AI Overview traffic from analytics data is presenting an inference whose method should be demanded and read.
Google's own reporting does not close it either, and this is the most consequential measurement fact on the subject. Google counts those clicks — its documentation states that clicking an external link in an AI Overview counts as a click, and says the same of AI Mode — and pools them into the Web search type with no filter. The generative AI performance report that launched on 3 June 2026 reports impressions and nothing else: no clicks, no click-through rate, no queries, and no dimension separating AI Overviews from AI Mode. The surface with by far the most reach is therefore one where neither your analytics nor Google's reporting isolates the traffic.
One partial exception is worth knowing, with its limits attached. John Mueller stated on 6 August 2026 that "if a user asks a follow-up question within AI Mode, they are essentially performing a new query" and that all impression, position and click data for the new response count as coming from that new query. Those queries are therefore already in Search Console — in the main Performance report, unlabeled and unfilterable. Practitioners spot them by their conversational shape, the short continuations nobody types into a search box, which is a heuristic rather than a filter and will both miss real ones and collect false positives.
Five ways an AI visit becomes direct traffic
These are in rough order of how much of the loss each probably explains, and that ordering is inference rather than measurement: no study decomposing the loss was found.
- The brand was named, not linked. The answer recommends you, the person searches your name or types your domain, and the visit surfaces later as direct or as branded organic. There is genuinely no referral event, and this is plausibly the largest component because it is exactly the outcome AI answers are built to produce.
- Native apps do not send referrers reliably. ChatGPT, Perplexity, Copilot, Claude and Gemini all have mobile and desktop applications, and links opened from an app commonly arrive with no Referer header. A UTM parameter survives this, which is why OpenAI's tag matters more than OpenAI's referrer.
- Referrer policy strips the header. A restrictive policy on the sending page suppresses the header entirely. It is a per-surface implementation detail operators do not publish and can change without notice.
- Copy and paste. Answers are read, URLs are copied, links are shared into messaging and email. All of it lands as direct.
- Cross-device and delayed action. Read on a phone, acted on days later from a laptop, with no identifiable source on the first touch.
On the size of the gap, almost nothing is known. The one quantified estimate located is practitioner research rather than platform data: Lawrence Hitches, working from more than 100 brands in an AI ecommerce dataset on 17 June 2026, estimated true Claude-influenced visits at roughly two to three times what referral data shows, with the rest landing as direct. That is one consultant, one platform, one undisclosed method. No study measuring the leakage rate with a published method was found, and after the Search Console facts that absence is the most important thing on this page.
How much it is, and how differently it behaves
The largest dataset available is SE Ranking's study of 101,574 websites across 250 countries with analytics integration, covering 16 months to April 2026. AI referral traffic was 0.02% of all website traffic in 2024, 0.24% in 2025 and 0.32% in 2026, which the study characterizes as 16-fold growth from 2024 to 2026 (SE Ranking, June 2026).
Both halves belong in the same sentence and are almost always separated. Quoting 16-fold growth without a third of one percent overstates the present; quoting a third of one percent as proof that AI search does not matter ignores a sixteen-fold move in two years. Within that traffic, ChatGPT accounted for 74.78% of AI referrals in 2026, ahead of Gemini at 11.56%, Perplexity at 7.23%, Copilot at 3.51% and Claude at 2.62%.
The behavioral differences are larger than the volume. The same dataset found AI visitors spending 67.7% longer on site than organic visitors, 9 minutes 19 seconds against 5 minutes 33 seconds. SE Ranking's July 2026 follow-up found 60% of AI-referred traffic landing on homepages, against 17% from organic search — consistent with brand-level rather than page-level recommendation, and with a direct analytical consequence: landing-page analysis of AI traffic concentrates on the homepage and says little about which content earned the recommendation.
Two widely quoted figures need their labels kept on. Semrush's finding that an AI search visit is 4.4 times as valuable as an organic one covers 500 or more digital marketing and SEO topics, uses Semrush's own composite of value rather than a conversion rate, and carries Semrush's caveat that its projections are extrapolations. On the other side of the click, Pew Research's panel of 900 US adults sharing browsing data — 68,879 unique Google searches, 12,593 of which produced AI summaries — found a result link clicked in 8% of visits with an AI summary against 15% without, links inside the summary clicked in 1% of visits, and 26% of sessions ending after the search against 16% without. Google's position is that click volume is relatively stable year over year with slightly more quality clicks, defined as those where users do not quickly click back, published with no sample size, figures or method. One side published a specified method; the other published an assertion.
Server logs measure the other side of the transaction
Analytics sees people arriving. Logs see machines fetching, and the two answer different questions. AI operators publish user agents and IP ranges, so identification can be verified rather than assumed. OpenAI documents four agents with four distinct meanings and publishes an address list for each (OpenAI): GPTBot for training foundation models, OAI-SearchBot for the index behind ChatGPT's search features, ChatGPT-User for user-initiated fetches inside a conversation, and OAI-AdsBot for ad page safety validation.
That distinction is the payoff, and treating GPTBot hits as a visibility metric — which is common — confuses training harvest with answer-time retrieval. A ChatGPT-User hit is the closest thing to a live signal available in this field: a person, in a conversation, caused your specific URL to be fetched at that moment. It proves neither citation nor a visit, but no other instrument shows retrieval at answer time on a real request. Logs also separate a retrieval failure from a selection failure: a page that was never fetched cannot be cited, and that is a different problem with a different fix.
Three limits belong with every log-derived figure. Crawling is not citing, and no log line indicates that a fetched page appeared in an answer. Google's AI crawling cannot be isolated: Googlebot covers all Google Search features, AI Overviews and AI Mode among them, while Google-Extended governs Gemini training and grounding and does not affect Search inclusion — so anything claiming to show you a Google AI crawler is showing you Googlebot. And a user-agent count is a lower bound rather than a census: Cloudflare's Cloudforce One team reported on 4 August 2025 that Perplexity ran 3 to 6 million daily requests from undeclared agents against 20 to 25 million from its declared crawler, a characterization Perplexity disputed. Either way, header matching under-reports by an unknown margin, and crawl-to-refer ratios inherit the problem.
Building a segment that will not quietly lie to you
The workable approach is a custom channel group or segment matching on both referrer hostname and UTM source, with the hostname list enumerated explicitly and maintained by hand. Match a prefix on the UTM value rather than one exact string: OpenAI's own documentation gives utm_source=chatgpt.com while secondary coverage of the June 2025 extension to the additional sources panel renders it as utm_source=chatgpt, and a segment built on one exact string silently drops traffic (OpenAI publisher FAQ).
Avoid the regex that has circulated since 2024, which matches any hostname containing the characters of the .ai namespace. It catches an entire country-code namespace unrelated to AI and misses chatgpt.com and gemini.google.com entirely. Enumerating hostnames is more work and the only version that can be audited.
Three properties of any such segment should be stated on every report built from it. It is a floor, because it cannot see mention without a link and loses app traffic and stripped referrers. It excludes Google's own AI surfaces completely, since they are indistinguishable from organic, and for most sites that is the majority of AI exposure. And it will break without warning, because every hostname in it except ChatGPT's tag is an undocumented convention no operator has committed to keeping.
One last check. No Google Analytics documentation establishing a default AI or generative-AI channel group was found as of 29 August 2026, and every practitioner account describes hand-built groups. That is recorded as unverified rather than as a confirmed absence — but if a dashboard shows you an AI channel you did not build, find out what defines it before reporting from it.
Frequently asked questions
Can I see AI Overview traffic in Google Analytics?
No. AI Overviews and AI Mode send visitors with a google.com referrer, identical to organic search, so no signal exists to segment on. Google counts those clicks internally and pools them into the Web search type with no filter, and its generative AI performance report contains impressions only. Any analytics view labeled AI Overview traffic is an inference, and its method should be examined.
Why does so much AI traffic show up as direct?
Five mechanisms, and the largest is not a tracking failure. When an answer names a brand without linking it, the person searches or types the domain and there is no referral event to capture. The rest is native apps that omit the referrer, restrictive referrer policies, copy and paste, and cross-device delay. The mechanisms are well documented; the magnitude has never been measured.
Is ChatGPT traffic reliably tracked?
It is the best-instrumented AI surface by a wide margin and still incomplete. OpenAI documents that ChatGPT includes utm_source=chatgpt.com in referral URLs from its search results, which is the only publisher attribution commitment any AI operator has made. That scoping matters: it is not a statement that every outbound link in the product carries the tag. Segment on the hostname and the tag together.
How much traffic does AI search actually send?
In the largest dataset available, 101,574 sites tracked over 16 months, AI referrals were 0.32% of all website traffic in 2026, up from 0.24% in 2025 and 0.02% in 2024. That is a sixteen-fold rise on a very small base, and both facts belong together. It also excludes Google's own AI surfaces, which are unattributable, so the real figure is higher by an unknown amount.
Do AI visitors behave differently once they land?
Measurably so. In the same 101,574-site dataset, AI-referred visitors spent 67.7% longer on site than organic visitors, and 60% of AI-referred traffic landed on homepages against 17% from organic search. The homepage skew is a real behavioral difference rather than a tracking artifact, and it means page-level analysis of AI traffic reveals little about which content earned the recommendation.
Does GA4 have a built-in AI traffic channel?
No documentation establishing a default AI or generative-AI channel group was found as of 29 August 2026, and every practitioner account describes hand-built custom channel groups and segments. This is recorded as unverified rather than as a confirmed absence. If your reporting shows an AI channel you did not define, establish what rules produce it before drawing conclusions from it.