Published
What ChatGPT search actually is
ChatGPT search is the web retrieval layer inside OpenAI's ChatGPT assistant. When a prompt looks as though it needs current or external information, ChatGPT issues queries against search providers, reads a small set of the pages that come back, and writes an answer carrying inline links to the sources it used. OpenAI launched it on 31 October 2024 for Plus and Team subscribers, folding in the SearchGPT prototype it had announced that July, extended it to all logged-in users on 16 December 2024, and made it available to everyone without an account on 5 February 2025.
It is not a search engine in the older sense. There is no ranked list a publisher can inspect, no query-by-query index, and nothing from OpenAI resembling Google Search Console. What it is, for a site owner, is a surface — a place where a link to your page may appear inside somebody else's answer, and from which a small, measurable flow of clicks arrives. A surface is not a results page, and most of the measurement problems here follow from that difference.
Two things are confused constantly and should not be. ChatGPT search is live retrieval, performed when somebody asks. GPTBot is OpenAI's training crawler, with no documented relationship to whether you appear in an answer.
Four user agents, four different jobs
OpenAI's bot documentation — which moved from platform.openai.com to the developers domain during 2025 and still answered there when checked on 29 August 2026 — names four fetching agents and gives each of them one job:
- OAI-SearchBot builds and refreshes what can be shown in ChatGPT search. OpenAI states that sites opted out of it "will not be shown in ChatGPT search answers", while adding that such sites can still appear as plain links.
- ChatGPT-User fetches a specific page because a user's request caused it. OpenAI writes that it "is not used for crawling the web in an automatic fashion" and that, because the action is user-initiated, "robots.txt rules may not apply."
- GPTBot is the training crawler. Disallowing it "indicates a site's content should not be used in training generative AI foundation models."
- OAI-AdsBot safety-checks pages submitted as ads, and its data "is not used to train generative AI foundation models."
OpenAI publishes the full user-agent string for each agent and a JSON file of its IP ranges, which is what makes any of this checkable in your own server logs rather than taken on trust. As of 29 August 2026 the search agent identifies as OAI-SearchBot/1.4 and the training crawler as GPTBot/1.4. Verify against the published ranges: user-agent strings are trivially forged.
Read the ChatGPT-User sentence carefully, because the hedge sits in the source itself. OpenAI does not say the agent obeys robots.txt and does not say it ignores it; it says the rules may not apply. That is permissive language rather than a commitment, and it is the same posture Perplexity takes with its own on-demand fetcher.
The GPTBot mistake, and what it costs
The most expensive misreading in this field is the belief that disallowing GPTBot governs whether ChatGPT cites you. It does not, on OpenAI's own account. GPTBot is assigned to training and OAI-SearchBot to search, as separate agents with separate opt-outs, and OpenAI has never published a statement connecting a GPTBot disallow to search visibility.
The belief spread for a reason. In 2023 the advice to "block the AI bots" was written as one instruction, in a year when GPTBot was the only OpenAI agent there was; OAI-SearchBot did not exist until 2024. The snippet outlived the situation it was written for and is still being copied between posts three years later.
The consequence runs both ways. Blocking GPTBot to keep content out of training while staying visible in answers is coherent and, on the documentation, works. Blocking it in the belief that this controls citation is not, and neither is watching citations fall afterwards and concluding the two are connected. What would settle it is an OpenAI statement, or a controlled test on paired domains differing only in that robots.txt line. Neither exists.
Measurement is the one thing OpenAI does document
OpenAI's publisher FAQ states that ChatGPT automatically appends utm_source=chatgpt.com to referral URLs, so inbound clicks can be identified in ordinary analytics. That single sentence makes OpenAI the only major assistant operator that has documented a tagging convention of its own, and it is the reason ChatGPT is the easiest AI surface to measure. The same FAQ states that "Any public website can appear in ChatGPT search," and directs publishers who want out entirely to a noindex tag — noting the crawler must still reach the page to read it. OpenAI's publishers and developers FAQ is the primary source for both.
Documented is not complete. The tagging changed once without announcement: around June 2025, Glenn Gabe of G-Squared Interactive observed the parameter extending to links in the "More" sources panel, not inline citations alone. Clicks passing through an in-app browser, or copied and pasted, arrive untagged as direct. Segment on the referrer host and the UTM together, and assume leakage.
The deeper limit is structural. Only clicks are counted. A citation read and not clicked leaves no trace you can see, and OpenAI operates no console that would show it. Every ChatGPT visibility figure a publisher holds is a floor, not a measure.
The silence about source selection is itself the finding
Set the documentation beside the questions practitioners actually ask and the gap takes shape. OpenAI publishes nothing on how candidate pages are scored, how many are retrieved, how they are ordered, or what makes one page preferred over another. There is no quality guideline and no equivalent of Google's Search Essentials. Generative engine optimization — the practice of trying to influence what an answer engine cites — has, for this surface, no documented target.
Three further silences deserve recording. OpenAI has never expanded "third-party search providers" into a list of names. It publishes no impression or citation data to publishers, unlike Microsoft, which shipped an AI Performance report in Bing Webmaster Tools in February 2026. And it has released no aggregate figures on outbound clicks or cited domains, so every number in circulation comes from third-party panels rather than the operator.
The vacuum fills with inference, and the inference gets quoted as fact. The common example is the assertion that ChatGPT runs on Bing, so being indexed in Bing gets you cited. SearchGPT was widely reported as Bing-backed in 2024, and precise-sounding percentages for how much of ChatGPT's citation set comes from Bing still circulate; the ones traceable at all lead to undated commercial posts with no method, and are unverified. Optimizing for Bing may help. It is not documented that it does.
Citation does not follow organic rank
The strongest evidence about ChatGPT citation is not about how to earn one, but about what does not. Two independent studies, different vendors, different query sets, point the same way.
Semrush, on 21 July 2025, reported that pages cited by ChatGPT search rank in traditional organic positions 21 or worse for related queries almost 90% of the time. The sample was over 500 digital marketing and SEO topics — one vertical, and the vertical Semrush sells into, so do not generalize it without saying so.
The Ahrefs citation overlap study of 11 August 2025, by Louise Linehan with Xibeijia Guan, ran 15,000 long-tail queries across Google, Bing, ChatGPT, Gemini, Copilot and Perplexity, extracted the citations and compared them with Google's results. Only 12% of AI-cited links appeared in Google's top 10 for the same query, and roughly 80% ranked nowhere in the top 100. ChatGPT was at the low end: 8.0% overlap for in-text citations and 6.1% for references, against Perplexity's 28.6%.
Two studies converging from different directions is the strongest evidence in this subject, and what they converge on is that ChatGPT citation is not a by-product of organic rank. Neither claims rankings are irrelevant. Both mean that a budget built on the assumption that ranking work automatically buys citation is misallocated.
Scale, and the two numbers people mix up
ChatGPT is the largest single source of AI referral traffic to websites. SE Ranking's June 2026 study of anonymized analytics from 101,574 websites, covering January 2025 to April 2026, put ChatGPT at 74.78% of AI referral traffic, ahead of Gemini at 11.56%, Perplexity at 7.23%, Copilot at 3.51% and Claude at 2.62%. The same panel showed 79.74% a year earlier while referral volume grew 27% year over year. Falling share and rising volume are not contradictory: the category grew faster than ChatGPT did.
Now the error. Similarweb's widely quoted series measures something else: traffic to the assistants' own websites, which is usage share, not referral share. On that measure ChatGPT was 52.7% in May 2026, down from 76.4% a year earlier. Both series are real. Posts quoting 52.7% as ChatGPT's share of referrals to publishers have swapped one for the other, and the mistake is common enough to check for whenever a share figure arrives without a denominator.
Keep the base rate in view too. ChatGPT's share of all referral traffic across that panel was 0.32% in May 2026, up from 0.23% in April and above the previous peak of 0.29% in October 2025 — a 36.7% month-on-month rise on a very small base. Quoting the growth without the base, or the base without the growth, produces two misleading stories from one dataset.
One figure travels much further than its evidence: Semrush's finding that an AI search visitor is 4.4 times as valuable as an organic visitor, by conversion rate. It is one vendor study, in one vertical, a conversion-rate ratio rather than a revenue measure, with no published sample size, time window or significance test. It has since been re-quoted as "4-5x", "5-9x" and as a revenue claim. Use the original sentence with its method attached, or leave it alone.
What changed in 2026, and what to stop citing
Ads began testing in ChatGPT in the United States on 9 February 2026. OpenAI's help documentation states that they appear below the end of a response, that they are clearly labeled as sponsored and visually separated from it, and that ads do not influence ChatGPT's answers. No independent replication exists either way. What would settle it is a controlled comparison of citation sets for identical prompts in categories with and without an active advertiser. Until somebody runs that, OpenAI's statement is the operator's own claim, not verified behavior.
ChatGPT Atlas, the browser launched on 21 October 2025, is finished: OpenAI announced on 9 July 2026 that browser-based agentic work was moving into the desktop app and a Chrome extension, and Atlas stopped working on 9 August 2026. Guidance written around Atlas is now historical, which matters because much 2025 and early 2026 advice had that browser in frame.
The housekeeping point generalizes: documentation moves without announcement, and advice ages in months. Date every claim you record about this surface, re-read the bot documentation before acting on a robots.txt change, and treat any undated statement about ChatGPT's behavior as probably stale.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT search?
No. OpenAI's bot documentation assigns GPTBot to training and OAI-SearchBot to search, as separate agents with separate opt-outs, and has published nothing linking a GPTBot disallow to search visibility. Blocking OAI-SearchBot is the documented control, and even then OpenAI notes that opted-out sites can still appear as plain links. The confusion dates from 2023, before OAI-SearchBot existed.
How do I track ChatGPT referral traffic in analytics?
OpenAI documents that referral URLs carry utm_source=chatgpt.com, so ordinary analytics will pick them up. Segment on both the UTM parameter and the referrer host chatgpt.com, because tagging has changed once without announcement and some clicks arrive untagged as direct traffic. Remember you are counting clicks only: citations that are read but not clicked are invisible, so the number is a floor.
Does ChatGPT search run on Bing?
OpenAI says only “third-party search providers” and has never named them. SearchGPT was widely reported as Bing-backed in 2024, and specific percentages circulating in agency posts trace to undated sources with no stated method. Treat the claim as an untested inference. An OpenAI statement, or a large paired study of Bing-indexed versus Bing-absent URLs appearing in ChatGPT, would settle it.
Do I need to rank in Google to be cited by ChatGPT?
The evidence says no. Semrush found in July 2025 that ChatGPT-cited pages rank in organic positions 21 or worse almost 90% of the time for related queries. Ahrefs, testing 15,000 long-tail queries in August 2025, found only 8.0% of ChatGPT in-text citations appeared in Google's top 10. Two independent studies, one conclusion: citation is not a function of rank.
Can I see how often ChatGPT cites my site without a click?
Not from OpenAI. There is no impression report, no citation count and no publisher console of any kind, so uncited-but-read appearances leave no trace you can access. Microsoft is the exception in this field: its AI Performance report in Bing Webmaster Tools, in preview since February 2026, shows citations without clicks for Microsoft's own surfaces. Nothing equivalent exists for ChatGPT.
Does ChatGPT-User obey robots.txt?
OpenAI's own wording is that because the fetch is user-initiated, “robots.txt rules may not apply.” That is deliberately permissive: it is neither a promise of compliance nor a statement that the file is ignored. Perplexity takes the same posture with its on-demand fetcher. Plan on the basis that a user-triggered fetch of a public page will succeed, whatever your robots.txt says.