Over the previous few months, I’ve been paying shut consideration to how ChatGPT’s fan-out queries have been evolving, as some fascinating developments have taken place that I imagine are altering the standard of ChatGPT’s responses. I’ve been engaged on this text for some time, however this one has been significantly tough to put in writing, as a result of as with all issues in AI search, the data modifications extra shortly than I can end writing about it. That’s positively been true for the way OpenAI seems to be tweaking and refining its means of retrieving info by way of internet search (RAG), and particularly for the way closely ChatGPT has began counting on web site: searches in fan-out queries, doubtlessly utilizing them to curate outcomes from higher-quality sources.
The TL;DR: I feel OpenAI is utilizing fan-out queries, and the positioning: operator specifically, as one technique of decreasing spammy outputs of their solutions derived from web content material. I imagine it’s their effort to enhance the standard of retrieved sources whereas taking early steps to fight spam and low-quality info of their outcomes. It jogs my memory of what Google has examined with E-E-A-T, however ChatGPT model.
I feel analyzing fan-out queries issues as a result of ChatGPT is utilizing engines like google to retrieve the outcomes it makes use of to formulate a solution, which suggests that is essentially an search engine marketing drawback. Each phrase the mannequin chooses to place right into a fan-out question, and each search operator it makes use of, tells us one thing about what the mannequin is searching for and the place within the search outcomes it expects to search out it. When it scopes a search with web site:, provides the phrase “official,” or factors at a particular subreddit, the mannequin is telling us what sort of content material it believes will greatest reply the consumer’s query. These queries are the closest factor we’ve got to understanding why ChatGPT pulls within the info that it does, and I feel there’s a lot we will be taught from unpacking them.
There are a couple of of us within the trade who’ve accomplished nice work sharing their findings on the internal workings of ChatGPT: studying the uncooked site visitors, scraping the dialog recordsdata, and pulling fan-outs out of the API. Their datasets have began to converge on the identical findings, and this text combines the learnings from their analysis with my very own observations from watching fan-out conduct over time, utilizing a mix of Peec AI, Profound, the Resoneo plugin, FanoutFox, and Google Search Console. So I purpose to do two issues with this text: First, lay out what everybody has really discovered, in a single place, with the numbers attributed to whoever ran them. Then provide you with my learn on what’s evolving and why I feel it’s essential.
The Mechanics Of How ChatGPT Fan-Out Queries Work
Not all searches on ChatGPT use internet search. Something contained in its coaching knowledge could be answered shortly with out utilizing RAG, and OpenAI’s free or cheaper fashions usually tend to depend on coaching knowledge to reply questions shortly, because it prices them much less cash to generate. When the query requires up-to-date information, ChatGPT will use internet search (retrieval-augmented era or RAG) to tug info from a wide range of sources, together with exterior search engine knowledge and its personal inside index (Labrador).
Once you do ask ChatGPT a query that triggers an online search, it begins by deconstructing your immediate right into a set of its personal background searches (fan-out queries), runs them in parallel, and synthesizes a solution from no matter comes again. Monitoring how fan-out queries change over time tells us loads about how OpenAI is tweaking the mannequin’s search conduct to attempt to produce higher outcomes. I feel watching this house offers us an enormous clue about what they had been hoping to realize with every mannequin replace.
It’s essential to begin by defining two phrases generally thrown round in our house, with out it all the time being understood what the nuance is between them. Retrieved means a web page ChatGPT fetched whereas working its fan-out queries. Cited means a web page that made it into the seen reply as a hyperlink. Quotation and retrieval can behave in a different way, and proper now they seem like transferring in reverse instructions: the variety of “retrieved” URLs in ChatGPT’s responses is rising, whereas unbiased measurements present the variety of distinctive domains cited per response falling over the identical interval. Whereas extra pages are being thought-about for the reply, fewer pages get cited and credited.
Whereas studying any research about AI search, it’s value asking whether or not the article refers to retrieved URLs or cited URLs, and if it’s discussing citations, for which prompts these citations are showing. In lots of circumstances, the discrepancies I’ve seen between fan-out research come down to at least one measuring retrieval and the opposite measuring citations.
To learn the precise fan-out queries (not simply the ultimate reply), I’ve been utilizing a mix of some instruments: Peec AI and Profound each supply fan-out queries for tracked prompts, and the free Resoneo ChatGPT Chrome plugin and FanoutFox (proven beneath) each floor the queries the mannequin runs and the sources it pulls. These plugins make it simpler to observe the mannequin “suppose” by its searches in actual time.
Timing issues right here too: it’s important to think about how and when ChatGPT releases new fashions, and which fashions and tiers are mostly utilized by nearly all of its customers. ChatGPT 5.6 ships in a couple of variant: Sol is the usual model, and the cheaper “Luna” variant is what rolled out across the begin of August as the brand new default mannequin for Free and Go customers. That distinction is essential once you learn the research beneath, as a result of they aren’t all measuring the identical mannequin or tier. And in response to Olivier de Segonzac’s breakdown in Search Engine Land, greater than 90% of ChatGPT’s weekly customers are on the free plan. So regardless of the free default does when it searches, it’s now more than likely what the massive majority of ChatGPT customers get.
Half 1: What Trade Analysis At the moment Exhibits
What I Noticed In My Personal Analysis
Just a few patterns stood out to me early on, earlier than I went anybody else’s knowledge:
- The higher fashions search extra, and lean on web site: extra. In my testing, ChatGPT 5.4 Considering would hearth 10-plus searches for a single immediate, together with a number of web site: queries, whereas 5.3 Prompt usually did simply two to a few fan-outs.
- The kind of web site: search seems to depend upon the question. For opinions and product opinions, it leans closely on Reddit, together with querying particular subreddits. For YMYL-style questions, it seems to want authoritative sources, usually .gov domains or typically .org. In a single batch of roughly 20 prompts within the authorized house, each single one returned citations solely from .gov domains. The beneath screenshot from the Resoneo plugin reveals this course of at work:

- For specs and pricing, the mannequin goes to the model. Searches like web site:sephora.com or web site:costco.com confirmed up usually. And for those who ask a few services or products, ChatGPT will usually run web site: searches in opposition to essentially the most well-established manufacturers in that house, even once you didn’t title them in your immediate. It additionally usually names particular merchandise and does web site: searches for particular person product pages from the model inside the fan-out queries.
The online impact, at the very least in what I checked out, seems to be extra ChatGPT visibility for high-authority websites and trusted manufacturers, and fewer citations for everybody else.
Beneath, I’ll spotlight a couple of latest research on this subject and what they discovered.
Helpful Findings From search engine marketing & AI Search Trade Specialists On ChatGPT Fan-Outs
A number of superior of us in our house have been measuring this independently, with totally different instruments and totally different assortment strategies. David Konitzny at Peec AI ran the numbers the day ChatGPT 5.6 turned the default: the share of prompts with solely a single fan-out question dropped from 94.0% to 43.5%, common retrieved sources roughly doubled from about 12 to 24, prompts needing a second fan-out iteration went from about 5% to 33.5%, and the positioning: operator went from showing in roughly 0.3% of fan-outs to about 23%. Mainly, ChatGPT is turning into extra exact and sturdy in its looking course of.
Chris Lengthy at Nectiv, evaluating roughly 4,000 prompts on 5.6 Sol in opposition to his personal 2025 baseline, discovered common fan-out queries per immediate went from 2.17 to 7.61, the longest question chain went from 4 searches to 29, and “web site:,” “official,” and “gov” all landed as high unigrams, with web site: in 64% of queries. There’s an enormous hole between the positioning: search figures in David’s and Chris’s articles, and it might be defined by the fan-out question assortment technique: Chris’s consultancy pulls fan-outs from OpenAI’s API, whereas a number of of the opposite instruments on this house extract them from the ChatGPT shopper interface. Whereas the API is clear and repeatable, the UI technique could also be nearer to what actual customers really get (with sure limitations like personalization, which no software can observe successfully). It’s value checking which technique a research used to know why totally different research might present discrepancies. In each research, nevertheless, the share of web site: searches in question fan-outs elevated considerably.
David additionally discovered product pages now make up 16.39% of retrieved pages, transferring forward of listicles, which tracks with what I’d been seeing anecdotally: ChatGPT 5.6 seems to go straight to manufacturers and producers for specs and pricing moderately than routing all the things by roundup articles and different listicles, which could be self-serving and susceptible to manipulation. The web page varieties shedding share of retrievals (listicles, how-to guides, and comparability pages specifically) additionally occur to be the codecs most closely spammed for GEO over the previous 12 months or two. In my talks all through this 12 months, I’ve shared how these actual web page varieties trigger search engine marketing and AI search issues.
Olivier de Segonzac and the Resoneo crew, who’ve accomplished among the most detailed reverse-engineering of the retrieval structure on the market, discovered that the distinctive domains cited per response dropped from 19 to fifteen after the 5.3 replace, confirming the sample from the opposite path: retrieval counts grew whereas quotation counts dropped. Ahrefs’ most-cited-domains knowledge reveals the place the surviving citations land most regularly: Reddit, Wikipedia, Forbes, Merriam-Webster, Client Studies, Healthline, and Walmart.
Suganthan Mohanadasan has been coming at it from the network-traffic aspect, and his findings are helpful to know: ChatGPT decides earlier than it searches. Throughout 57 conversations and three,554 retrieved pages, 21 of 27 preliminary queries contained model names the consumer by no means talked about of their immediate, throughout 11 of 13 product classes. For instance, when looking “the perfect AI note-taking app,” the primary fan-out question already contained the phrases “Granola,” “Notion AI,” “Otter,” “Fireflies,” “Fathom,” “Mem,” and “Limitless.”
Beneath is a snippet from Suganthan’s latest article, shared with permission:

Being the model talked about in that first question is the true objective: manufacturers named within the fan-out had been cited 68.9% of the time, whereas pages that had been solely fetched had been cited 2.1% of the time, and solely 110 of the three,554 retrieved pages, about 3.1%, made it into a solution in any respect. Suganthan additionally defined ChatGPT’s routing logic inside fan-out queries: information path to official pages, and opinions path to opinions and Reddit. So ChatGPT will look to various kinds of websites, and distinct pages inside them, relying on which info helps to greatest reply the consumer’s query.
As with all research, it’s useful to learn his methodology alongside these numbers, which he states himself: a single account, based mostly in Dubai, sampled in July 2026, weighted towards software program and AI instruments, with the brand-injection sample measured throughout 27 preliminary queries. That’s a comparatively small pattern dimension, however it’s value citing as a result of the path additionally matches what the bigger datasets present, and what I see in my very own testing.
The Consensus Between Current ChatGPT Fan-Out Question Research
Though the above researchers used totally different instruments, totally different fashions, and totally different assortment strategies, their findings nonetheless line up throughout 4 frequent patterns:
- ChatGPT is working considerably extra searches per immediate than it was a 12 months in the past, together with at no cost and cheaper tiers.
- A meaningfully massive share of these searches at the moment are scoped with the positioning: operator.
- The domains it scopes to skew high-authority: official and producer pages, established evaluate platforms (G2, Clutch, Capterra, Client Studies, Wirecutter), Wikipedia, .gov and regulatory sources, main retailers, and Reddit.
- Retrieval goes up whereas the variety of distinctive domains cited goes down.
The skew is towards surfacing higher-quality, recognizable websites, and that development is continuous with newer ChatGPT fashions.
Half 2: My idea On One Purpose Why This Is Taking place
I discover all of this fascinating as a result of it aligns intently with one of many elements of Google’s search rating techniques I’ve spent essentially the most time finding out: E-E-A-T (expertise, experience, authoritativeness, and trustworthiness). I feel OpenAI could also be utilizing fan-out queries as a way of making certain customers get high-quality, reliable, authoritative info, whereas suppressing spammy and extremely manipulated articles (corresponding to self-promotional listicles and different varieties of self-serving content material).
I feel the positioning: operator is functioning (at the very least partially) as a spam filter.
For some time, ChatGPT would regularly retrieve and cite the precise type of spammy, self-serving, manipulative content material that we’ve already turn out to be fairly aware of within the search engine marketing world. Evaluating the standard of a random open-web web page in actual time is genuinely onerous and costly, and at ChatGPT’s quantity you’d should do it tons of of hundreds of thousands of instances a day. Narrowing fan-out queries to .gov domains, established evaluate platforms, massive recognizable manufacturers, and official sources is presumably a less expensive method to get a lot of the identical end result. As an alternative of judging whether or not an unknown web page is reliable, the mannequin seems to be sidestepping the issue totally by solely searching for info in locations it already trusts.
Google spent years constructing techniques to reply “is that this supply authoritative sufficient for this question?” and what OpenAI seems to be doing is a distilled model of this identical course of, relying closely on the positioning: operator and different key phrases that affect which sources get retrieved. For instance: pure information go to the model’s personal area. Well being and authorized questions go to .gov websites. Opinions go to Reddit and established evaluate platforms. Whereas Google makes use of advanced rating algorithms to floor high-quality, reliable, and authoritative pages, ChatGPT can refine its fan-out queries to make sure the search outcomes pull from these trusted sources.
Including “.gov” to sure queries (YMYL, maybe?) is what I discover intriguing, as a result of it jogs my memory of how Google elevates .gov (and different high-authority) websites in its outcomes for sure queries or throughout instances of disaster. For instance, through the Covid period, I shared a whole lot of analysis displaying how Google elevated the official FDA, CDC, and different high-authority and authorities websites within the search outcomes for health-related queries:

The .gov-only sample in that batch of authorized prompts (proven beneath) appears like an enormous change on ChatGPT’s half, because it separates essentially the most “common” or “optimized” pages on the web from the official authority websites. On nearly any shopper authorized subject, the favored, top-ranking pages proven in search come from regulation agency blogs, not authorities businesses. Mixed with the web page varieties shedding retrieval share, that strikes me as a method for ChatGPT to chop by the noise and elevate official info as a substitute.

The refinement additionally seems to be transferring down throughout the worth tiers. The 5.4 Considering mannequin was already doing a model of this: a lot of fan-out queries, web site: operators geared toward trusted domains, a transparent choice for official and authoritative sources. With 5.6 turning into the free default, these search behaviors appear to have additionally prolonged to free tier customers, not simply the customers paying for the reasoning mannequin.
The Use Of The Phrase “Official” In Fan-Out Queries
I discovered it very fascinating that the frequency of the phrase “official” in fan-out queries seems to be climbing, and it reveals up as a high unigram in Chris’s knowledge alongside web site: and gov. Comparable analysis by Conductor, proven beneath, additionally reveals the rise in web site: searches and searches together with “official.” ChatGPT appears to be actively looking for the official supply for a model or product, which is strictly what you’d anticipate if the objective is to keep away from citing unverifiable third-party pages.

This might current an honest motive to revisit whether or not “Official” or “Official Web site” belongs in your homepage title tag or meta description, particularly for firms that share a reputation with different manufacturers or entities, or the place the search outcomes might in any other case be complicated. Whereas this has already been a reasonably commonplace search engine marketing suggestion for a very long time, I feel there are circumstances the place making that language specific can keep away from confusion. And given how a lot ChatGPT now seems to be searching for official info within the search outcomes, it might doubtlessly assist the mannequin reconcile which model you really are, or which area is formally yours.
The Dangers Of Relying On Web site: Searches
Heavy use of web site: searches in fan-out queries can work nicely when the mannequin is aware of which websites to tug from. However articles by Malte Landwehr and Netcraft discovered that if there’s confusion concerning the appropriate area for a given model, ChatGPT seems to guess the area, and typically it guesses incorrectly.
Malte Landwehr discovered examples the place ChatGPT constructed web site: queries in opposition to the fallacious model area, and in some circumstances, a website that doesn’t belong to the corporate in any respect. In a single take a look at, it repeatedly searched web site:census.com for the startup Census (not the federal government bureau), when the startup’s official web site is getcensus.com, and census.com is parked and that can be purchased. He noticed the identical sample pointing at different parked domains like lago.io, persona.id, lightfield.ai, and mesa.com.
A Netcraft research from 2025 additionally discovered that roughly a 3rd of brand name login hyperlinks generated by LLMs pointed to domains the model didn’t personal, and about 29% pointed to unregistered, inactive, or parked domains, with smaller manufacturers essentially the most uncovered.
As Malte mentioned in his article, when the mannequin misidentifies the trusted web site for a lesser-known model, somebody might purchase that parked area, host content material matching what ChatGPT is searching for, and quietly feed it fallacious pricing or a pretend assist quantity. This might be problematic for manufacturers the place the official area isn’t strongly encoded but, or the place another motive prevents ChatGPT from figuring out the suitable area title.
It’s value acknowledging that this narrowing might be as a lot about value and latency as it’s about high quality. Checking a brief listing of identified domains is cheaper and sooner than evaluating the open internet at ChatGPT’s quantity, and from the surface, that might look an identical to a deliberate high quality filter. Malte’s discovering additionally reveals there’s nonetheless room to enhance: a mannequin that was genuinely assessing whether or not a supply is reliable and authoritative could be much less more likely to level a search at a parked area. So there might be a couple of motive why OpenAI is evolving the way it constructs fan-out queries.
Web site: Searches Showing In Google Search Console
One web site: question on the positioning beneath generated roughly 197,000 impressions and precisely one click on. That’s not the one instance: once I checked out web site: searches for a number of main manufacturers in Google Search Console, I discovered 1000’s of impressions for web site: queries with nearly zero clicks, and the amount of those queries is usually growing over time. A click-through charge that near zero is difficult to elucidate with human searchers, and factors as a substitute to bots (LLMs, monitoring instruments, scrapers, and so forth). You’ll be able to see comparable web site: search patterns in Bing Webmaster Instruments’ AI search question reporting.

This bought me pondering: ChatGPT is working much more web site: searches than ever, and if it (or the companions it makes use of to scrape engines like google for RAG) remains to be leaning on a significant search index, it appears believable that at the very least a few of these searches would floor in our question reporting. There’s no official “ChatGPT Search Console,” so that is hypothesis based mostly on varied accounts in each Google and Bing’s search console instruments, and Google has had its personal Search Console impression-related knowledge points muddying this actual type of knowledge, so it’s tough to get a conclusive reply as to the place these searches are coming from.
The sensible tip: Examine each Google Search Console and Bing Webmaster Instruments queries for web site: searches and different repeated patterns. In Search Console, use the Efficiency report and filter Question by “Queries containing” web site: (or a regex). In Bing Webmaster Instruments, use the Search Efficiency report and search the question desk for web site: strings.
Whereas ChatGPT search has traditionally run on Bing (the OpenAI-Microsoft partnership), unbiased testing over the previous 12 months suggests it’s additionally pulling from Google’s index, probably by third-party scrapers, so it genuinely isn’t clear which engine logged a given retrieval. This info can also be evolving on a regular basis, particularly given Google’s lawsuit in opposition to SerpApi, which was believed to be the supplier of scraped Google knowledge to OpenAI. ChatGPT additionally seems to be more and more constructing and leveraging its personal index, which might give OpenAI extra management over freshness and protection with out relying on its largest competitor.
Bing has made it simpler to observe AI search efficiency: in February 2026, it launched an AI Efficiency Report inside Webmaster Instruments that separates AI citations from conventional search and surfaces “grounding queries,” that are the matters derived from search phrases Copilot generates internally to retrieve content material. That was massive information: a search engine selecting to report fan-out sub-queries as their very own class, which tells you this entire conduct is actual sufficient that Microsoft is now constructing reporting round it (whereas Google continues to cover it from us).
What I’d Truly Do About It
I see a whole lot of this as a sign that search engine marketing issues greater than ever. As a result of ChatGPT seems to lean on main search indices (instantly or by RAG companions), rating nicely in search remains to be what feeds the fan-out, together with making certain your model has the suitable info when its pages are instantly looked for. Right here are some things I feel are value doing:
- Construct your model to the purpose the place it’s the one ChatGPT thinks of first. Whereas this one appears fairly apparent, that is the place all the information factors: you need your model to be synonymous with its class, and to be naturally advisable sufficient to enter into the coaching knowledge as one of the crucial trusted and well-known manufacturers within the house. With out good branding, it’s doubtless that your model gained’t even enter the dialog. That is additionally unimaginable to win long-term by way of manipulative optimization ways.
- Put your key information in plain, crawlable HTML by yourself area. Info corresponding to pricing, specs, mannequin numbers, and assist particulars must be readable to the mannequin. If ChatGPT is working web site:yourbrand.com searching for these, they should be included as textual content it may well really learn, not locked inside a picture or client-side JavaScript-rendered content material.
- Nail down which area is “official.” Contemplate “Official Web site” (or comparable) in title tags or meta descriptions the place it reads naturally and helps customers (don’t spam this), and ensure your actual area is persistently strengthened throughout the net, significantly for those who’re a smaller model or share a reputation with different entities. That is the only greatest protection in opposition to the wrong-domain drawback.
- Don’t play whack-a-mole with making an attempt to focus on particular person fan-out queries. They’re long-tail, low-volume, and differ between prompts and customers. Producing many pages focusing on the long-tail is an effective method to get caught up in Google’s scaled content material abuse spam entice. It’s extra helpful to mixture the queries the mannequin retains working and discover the core matters it persistently prioritizes. That is what I imagine Microsoft does within the Bing Webmaster Instruments AI search question report. I additionally suppose that aggregating fan-out queries and distilling them down into core key phrases to make use of for conventional search engine marketing rank-tracking remains to be a very good technique of approximating AI search efficiency: for those who rank nicely for the matters typically requested about in fan-out queries (ideally tracked throughout each Google and Bing), you have got a better likelihood of finally being a part of the AI response.
- Examine whether or not you’re within the question in any respect. Suganthan’s take a look at is an effective one: Run the prompts you care about 5 separate instances and take a look at whether or not your model reveals up in ChatGPT’s preliminary searches, not simply within the reply. If it by no means does, you doubtless have work to do to construct up the notoriety of your model in its area of interest, and that will get accomplished with opinions, comparisons, digital PR, and media protection over an extended stretch of time, not with technical fixes alone. In case your model does present up, the objective is to transform retrievals into citations and to affect the AI response, and that’s the place on-page optimization can rely: state claims clearly and early on the web page, and guarantee essential numbers and enterprise particulars are retrievable for LLMs.
- Take Reddit severely. If the mannequin is pulling opinions from particular subreddits, that’s a part of your model’s visibility now, whether or not you just like the discussions or not. However retrieval and quotation diverge right here: Dan Petrovic discovered ChatGPT discards the Reddit pages it pulls roughly 99% of the time. That reality stood out to me subsequent to Reddit topping Ahrefs’ most-cited-domains listing, and each are true directly: Reddit will get retrieved so relentlessly that even a ~1% survival charge produces extra citations than nearly anybody else in absolute phrases. It’s additionally the identical phenomenon as Suganthan’s 3.1% general quotation charge, seen from a single area’s perspective. So deal with Reddit as one thing that shapes how the mannequin understands your class moderately than as a dependable quotation path. And don’t use any of this as an excuse to spam Reddit with synthetic suggestions of your model: Reddit lately cracked down on this actual kind of GEO spam, utilizing its personal LLMs to flag round 25,000 spammy posts and feedback per day, and taking part in with hearth could cause your account to be banned.
- Watch your web site: impressions in each Search Console and Bing Webmaster Instruments. That is one other instance of bot exercise affecting our Search Console knowledge, on high of the latest discussions round Google Search Console displaying conversational prompts from AI Mode as particular person queries in GSC. That is one thing to think about so the noise doesn’t distort your reporting, and it might doubtlessly maintain some details about how ChatGPT (and possibly different LLMs) is retrieving details about your model, assuming that’s what that GSC knowledge reveals.
I imagine a lot of it is a response to the clear flaws in ChatGPT over the previous 12 months, corresponding to referencing biased and promotional content material, like listicles and comparisons that suggest the identical model writing the article. These modifications seem like doing the alternative, narrowing onerous to .gov domains, massive manufacturers, Reddit, and official sources, and lengthening that method down from the premium fashions to the free default that a lot of the world makes use of. Whereas I don’t see this as anyplace close to as refined as Google’s mechanisms for reaching comparable objectives, it reveals that OpenAI is innovating towards higher-quality outcomes and dealing on options to the issue of manipulated content material showing in its responses.
Extra Assets:
This publish was initially revealed on Lily Ray NYC Substack.
Featured Picture: CineVI/Shutterstock
