Published
Why these belong on one page
An AI assistant, in the sense this page uses, is a conversational product that answers a question by fetching live web pages and, usually, showing the reader where the answer came from. This page covers the ones that are not ChatGPT, not Google and not Microsoft: Claude from Anthropic, Meta AI, DuckDuckGo's AI-assisted answers, Brave Search, Grok from xAI, and Le Chat from Mistral.
They are grouped because none of them individually justifies a page, and pretending otherwise would misrepresent the evidence. On SE Ranking's panel of 101,574 websites covering January 2025 to April 2026, Claude accounted for 2.62% of AI referral traffic and everything outside the top five rounded to approximately nothing.
What makes the group worth writing about is the comparison. Between them these operators give four incompatible answers to the same question — how much a company will tell publishers about what it takes and what it sends back. Anthropic publishes the clearest crawler documentation in the field. Brave publicly refuses to be identifiable. Meta AI is the largest assistant here by reach and frequently does not link out at all. xAI publishes nothing whatsoever. Read side by side, that spread says more about the state of AI search than any one of them says alone.
Anthropic documents the most and measures the least
Claude's web search launched on 20 March 2025 for paid users in the United States and went global on all plans on 27 May 2025. Anthropic states that when Claude incorporates web information into a response it provides direct citations so the reader can fact-check the sources, and the links are clickable.
Anthropic's crawler documentation, last updated 7 April 2026, names three agents and separates them cleanly: ClaudeBot gathers content for training, Claude-User accesses websites when individuals ask Claude questions, and Claude-SearchBot crawls to improve search result quality for users. Anthropic says its bots respect do-not-crawl signals by honoring industry standard directives in robots.txt, supports Crawl-delay, publishes an IP range file, and warns that blocking by IP address may not work reliably because it prevents the crawler from reading robots.txt in the first place.
One detail deserves more attention than it gets. Anthropic is the only operator on this page — and one of very few anywhere — that states no exception for user-initiated fetches. OpenAI says robots.txt "may not apply" to ChatGPT-User. Perplexity says Perplexity-User "generally ignores robots.txt rules." Meta says its external fetcher "may bypass robots.txt." Anthropic's page makes no such carve-out for Claude-User. That is a documented difference in posture, not an inference, and it is worth stating plainly.
Then the reversal. Anthropic documents nothing about measurement: no UTM parameter, no publisher help page, no citation reporting. Practitioners report clicks arriving with a claude.ai referrer in the path form claude.ai/referral/YYYY-MM-DD, the date being the conversation date — practitioner research across roughly 100 brands in early 2026, which ranks below platform documentation because it is not documentation.
Nor does Anthropic document its search backend. In March 2025 TechCrunch reported evidence that Claude's web search ran on Brave Search: Brave appeared on Anthropic's subprocessor list, spotted by engineer Antonio Zugaldia, and Simon Willison found identical citations between the two plus a BraveSearchParams parameter in Claude's search function. Anthropic gave no comment. Well-evidenced, never confirmed, possibly changed since.
Claude's size depends on whose panel you read, and the disagreement is the finding. SE Ranking's 101,574-site panel puts it at 2.62% of AI referral traffic; the 100-brand dataset puts it at 8 to 12%. Cite the range. Growth is better attested: Similarweb has Claude's share of traffic to generative AI sites rising from 1.6% in June 2025 to 8.9% in May 2026, and SE Ranking has referral volume up 320% year over year — different metrics pointing the same way.
Meta AI has the reach and almost none of the attribution
Meta AI is embedded in Facebook, Instagram, WhatsApp and Messenger and available as a standalone app. By installed reach it is the largest assistant on this page. By publisher-referral value it is close to the smallest, and the reason is straightforward: it frequently does not link out.
Meta documents its crawlers well — five of them, with stated purposes and stated robots.txt behavior. meta-webindexer/1.1 crawls to improve Meta AI search result quality and honors robots.txt, making it the structural equivalent of OAI-SearchBot and Claude-SearchBot. meta-externalagent/1.1 crawls for training and product indexing. meta-externalads/1.1 crawls for advertising products. Two are exceptions: meta-externalfetcher/1.1, which fetches individual links at a user's request and "may bypass robots.txt because it performs fetches that were requested by the user," and facebookexternalhit/1.1, which "might bypass robots.txt when performing security or integrity checks."
The citation behavior is where it falls down. Meta publishes no policy on when Meta AI links to a source, and reporting recorded since May 2024 describes it summarizing news from various outlets without linking to the originals. Behavior varies by surface. The safe model for a publisher is a zero-click surface unless you can demonstrate otherwise on your own site — and you probably cannot, because there is no referrer convention, no UTM and no publisher tooling, and Meta app traffic has always been hard to separate from ordinary social referral.
Reach is not attribution. Treating Meta AI's install base as latent referral potential is the standard error here, and nothing published supports it.
DuckDuckGo: the most specific opt-out anyone has published
DuckDuckGo runs two AI surfaces: AI-assisted answers shown above search results, and Duck.ai, an anonymized front end to third-party chatbots. DuckAssist launched in March 2023 as a Wikipedia-grounded summarizer; the broader feature set was described in March 2025, including user controls over how often AI answers appear — often, sometimes as the default, on-demand, or never. Answers cite sources below the text and readers can click through.
The crawler is DuckAssistBot, and its help page is the most operationally specific opt-out documentation any operator in this field has published. DuckDuckGo states that it crawls pages in real time for AI-assisted answers, that the data is not used to train AI models, and — the part nobody else matches — that once you disallow the user agent, "the change will take effect after 72 hours and DuckAssistBot will stop crawling your site." Opting out does not affect organic visibility. An IP list and an opt-out contact address are published alongside.
A stated effect and a stated timeframe should be unremarkable and instead are unique. Every other opt-out here is a request with no promised latency and no stated consequence.
Measurement, though, is structurally impossible. Clicks arrive with a duckduckgo.com referrer and cannot be separated from classic DuckDuckGo organic, and the product's privacy design actively works against attribution. Anyone showing you a DuckDuckGo AI referral figure is inferring it.
Brave publishes a refusal to be identifiable
Brave Search runs an independent index with its own AI answer layer, and sells a search API to other AI companies — which makes it plausibly upstream of assistants you do measure. TechCrunch's March 2025 reporting placed Brave behind Claude's web search, and Brave announced a partnership with Mistral in February 2025 under which Le Chat's live web results run on the Brave Search API.
Its crawler help page contains the single most consequential sentence anyone writing about AI crawler control needs to read: "The Brave Search crawler does not advertise a differentiated user agent because we must avoid discrimination from websites that allow only Google to crawl them." And separately: "Note that robots.txt is not used to prevent a page from being indexed." Brave directs publishers to the noindex robots meta directive instead, and re-crawls to apply it.
This is a documented decision not to be blockable by user agent, published openly, with a stated rationale. Agree with it or not, the consequence is unambiguous: no user-agent blocklist can touch Brave, and the "block all AI bots" snippets circulating since 2023 do nothing here.
Le Chat makes the general problem concrete. Mistral publishes no crawler documentation relevant to publishers, because the fetching is Brave's — so the applicable control is Brave's, and Brave's crawler is deliberately anonymous. The operator whose name is on the interface is often not the operator fetching your page. Publisher control follows the retrieval layer, not the brand.
xAI publishes nothing, and the silence is the finding
Grok launched in November 2023 and retrieves from both the web and posts on X. As of 29 August 2026 no xAI-published crawler documentation, user-agent list, IP range file or robots.txt guidance could be located, and every GrokBot or xAI-Grok user-agent string in circulation traces to a commercial bot directory rather than to xAI. This is the largest documentation gap among assistants with meaningful usage, and it sits oddly beside X's terms of service, which have banned crawling and scraping since September 2023.
What is documented is not the mechanism but the output. Ahrefs' Brand Radar analysis of over 1.9 million US queries, published 1 June 2026, reports Reddit at 16.3% of mention share, YouTube 15.1%, Facebook 13.9%, Instagram 5.9% and Quora 5.5%. Quote those with the qualifier attached: they are mention shares within the top 50 domains, not shares of all Grok citations, and the article is auto-refreshed monthly, so the figures move. Restating them as "most of Grok's citations are social media" misreports the method.
Measurement is the usual story. Referrers are grok.com or x.com, nothing is documented, and the X case cannot be separated from ordinary social traffic.
The column of no
Line the operators up and one column is empty for every row. Anthropic documents three crawlers and no measurement. Meta documents five crawlers and no citation policy. DuckDuckGo documents a 72-hour opt-out and offers no way to attribute a click. Brave documents its refusal to be identified. xAI documents nothing. Apple documents Applebot and Applebot-Extended well, and is explicit that Applebot-Extended is a training opt-out only: it "does not crawl webpages" and pages disallowing it "can still be included in search results". Apple publishes no citation guidance either.
Not one operator on this page documents how a publisher should measure its referral traffic, and not one publishes citation counts. OpenAI's referral parameter and Microsoft's citation-impression reporting are the only two exceptions in the entire field, and neither of them appears here. Everything practitioners know about measuring Claude, Meta AI, DuckDuckGo, Brave, Grok or Le Chat traffic is reverse-engineered from referrer strings that could change tomorrow without an announcement.
Which leads to the advice this page exists to give. Pasting a "block all AI bots" list into robots.txt does materially less than the people copying it believe: Brave is unlisted by design, Meta's fetcher may bypass, xAI publishes no name to block, and OpenAI and Perplexity both exempt user-initiated fetches. What settles it for any site is server-log analysis against published IP ranges, not a text file. And the reason to read about these engines is control and measurement — what you can block and what you can prove — not campaign planning, because the combined volume does not justify a campaign.
Frequently asked questions
Does blocking AI crawlers in robots.txt stop AI systems reading my site?
Only partly, and less than the popular snippets imply. Brave publishes no differentiated user agent by design, Meta's external fetcher may bypass robots.txt, xAI publishes no agent name at all, and OpenAI and Perplexity both exempt user-initiated fetches. The only way to know what actually reaches your pages is server-log analysis against operators' published IP ranges.
Which of these operators documents its crawlers best?
Anthropic. Its crawler documentation, updated 7 April 2026, names ClaudeBot for training, Claude-User for user requests and Claude-SearchBot for search quality, publishes IP ranges, and states no exception for user-initiated fetches — unlike OpenAI, Perplexity and Meta, each of which carves one out in writing.
How much traffic does Claude actually send?
Sources disagree by an order of magnitude, and that disagreement is worth reporting. SE Ranking's panel of 101,574 sites puts Claude at 2.62% of AI referral traffic for January 2025 to April 2026. A practitioner dataset covering about 100 brands in early 2026 puts it at 8 to 12%. Different panels, different site mixes, different periods. Cite the range, never one figure.
Can I measure DuckDuckGo AI answer traffic separately?
No. Clicks arrive with a duckduckgo.com referrer that cannot be distinguished from classic DuckDuckGo organic, and the product's privacy design works against attribution by intent. DuckDuckGo documents no separate convention. Any tool reporting a DuckDuckGo AI referral number is estimating it, and should say so; treat that traffic as unmeasurable at the publisher end.
Does Meta AI link to the sources it summarizes?
Not reliably. Meta publishes no citation policy, and reporting recorded since May 2024 describes Meta AI summarizing news from various outlets without linking to the originals. Behavior varies by surface and has changed over time. Model Meta AI as a zero-click surface unless you can demonstrate otherwise from your own data, which is difficult because no referrer convention exists.
Is it worth optimizing for these engines individually?
Usually no, on the volume evidence. Combined, everything on this page is a low-single-digit share of AI referral traffic, which is itself a fraction of a percent of all referral traffic. The reason to understand them is control and measurement: which crawler you can block, which opt-out has a stated effect, and which traffic you can prove arrived.