More than half of these ChatGPT answers never searched the web at all
Sources: Gaetano DiNardi, Ayomide Joseph, and Shopify Enterprise.
The short version
- ChatGPT searched the live web on only 44% of 270 commercial comparison queries, and on the other 56% the study's instrumentation observed no fresh retrieval step.
- In one network log, ChatGPT's first search already named six vendors, so the candidate set was decided before any page was read.
- The result is extreme concentration, with one brand taking the top slot in 71% of a competitor's alternatives queries.
- Eligibility is the slow lever and it is built before the query: review depth, community presence, original research and being named alongside the category.
Ask ChatGPT for the best tools in a category and it does not go looking for the category. Gaetano DiNardi pulled the network log behind one such answer and found sixteen web search queries, with the very first one already naming six specific vendors.1 The model chose the candidates, then went searching for support. That behavior explains the concentration everyone has noticed in AI answers, and it sets up a harder finding from a separate 270-query run: ChatGPT searched the live web on only 44% of those queries, answering the other 56% from what it already held.2 For more than half of these commercial questions there is no retrieval step to optimize and no page fetch to earn. Whether the model already associates your brand with the category is the lever that remains, and that association was built long before anyone typed the question.
ChatGPT searched the live web on 44% of 270 queries. Gemini searched on 74%. Perplexity searched every time.
Does ChatGPT search your category, or the brands already in it?
ChatGPT searches the brands. Asked for “best AI agent builders”, ChatGPT returned a shortlist and its network stream contained sixteen unique web search queries. The first query already carried Microsoft Copilot Studio, Salesforce Agentforce, Google Vertex AI Agent Builder, ServiceNow AI Agents, IBM watsonx Orchestrate and AWS Bedrock Agents inside it.1 The vendors came from the model, and the searches went out to confirm them.
What does the network log show about how the candidate list is built?
The log shows retrieval being used for verification rather than discovery. Beyond that pre-loaded first query, many of the remaining fifteen are site: searches aimed at the vendors’ own domains, one per brand, looking for official documentation on each.1 DiNardi asked the model directly how it knew to search for those brands, and the answer was that it already held them as strong entities for that category.1
The mechanism is reproducible, which is genuinely useful, since anyone can run the same check on their own category in ten minutes. The shortlist is also decided before a single page is fetched, so on-page work operates on a list you were either already on or not.
The pre-seeded query is the transferable idea here: a retrieval step whose search terms already contain the candidate set, so the retrieval confirms a shortlist instead of discovering one. This is one query rather than a measured rate, so treat it as a demonstration of the mechanism and not as a claim about how ChatGPT always behaves.
How often does an AI answer skip the web entirely?
Often enough to change where you spend money. Across 270 query variations asking three assistants for “alternatives to” Kustomer, Okta and LinearB, ChatGPT searched the live web on 44% of them and answered 56% from training data with no fresh retrieval step. Gemini searched on 74% of queries. Perplexity searched on all of them.2

Ayomide Joseph, guest post on Gaetano DiNardi’s Marketing Advice newsletter; 270 queries, three products, one author, no published protocol.
On more than half of those answers, your newest page did not exist as far as the model was concerned, and whatever description of your product it absorbed months ago is what got repeated. No amount of publishing velocity reaches that half within a quarter. The study’s own reading is that ChatGPT leans harder on its training data to build summaries, so it may be referencing outdated information about your brand, while Gemini and Perplexity are more likely to be working from a fresh read of the current web.2
One caveat on the unit, because it matters. “No fresh retrieval step” is what the study’s instrumentation observed, and a system can also reuse cached or in-session context, so I would stop short of “the answer came purely from pre-training”. The practical consequence is the same either way: nothing you published this week was read.
What does concentration look like once the candidate set is fixed?
Concentration looks like a handful of brands owning nearly everything. In the same 270-query run, Microsoft Entra ID took the top slot in 71% of Okta-alternatives queries, and Zendesk took it in 55% of Kustomer-alternatives queries.2 If the candidates are chosen before retrieval, then AI search is behaving as an eligibility filter first and a ranking system second, which is a real shift in what technical work can accomplish alone.
An eligibility filter and a ranking system fail in different ways, and that difference is the whole practical point. A ranking system rewards incremental improvement, so a page that is slightly better moves up slightly. A filter has a threshold, so effort below it produces nothing visible and effort above it produces a step change. Teams organized around ranking read a flat month as insufficient effort and add more of the same. Teams organized around eligibility read the same flat month as evidence they are working on the wrong axis.
That is where two disciplines stop being separable. Category positioning has historically been a brand and marketing question, measured in awareness and preference, while technical SEO has been an access question, measured in crawls and impressions. If the candidate set is assembled from what the model associates with a category, then positioning determines whether you are retrievable at all and technical work determines whether you survive being retrieved. Neither one is sufficient, and a team that owns only one of them cannot see the half of the problem that belongs to the other. In my opinion, the reporting split most agencies use, brand over here and technical over there, hides exactly this failure.
This is 270 queries across three B2B SaaS products, run by one author with no published protocol, which makes it a strong signal on thin evidence. Nothing here generalizes to consumer retail without separate work, and the piece it comes from is a guest post by Ayomide Joseph on DiNardi’s newsletter, which DiNardi says he reviewed and edited for accuracy.2
Does any of this show up in commerce data?
Yes, and Shopify’s commerce data points at feeds instead of at pages. Shopify Enterprise reports that AI-referred shoppers who reached a product page converted about 80% better than shoppers referred by organic search, that AI search introduced net-new customers at about 1.3 times the rate of organic, and that in Q2, 50% of AI-referred sessions landed directly on a product page.3 Watches convert about 2.4 times better from AI-referred traffic than from organic.3
The figure I would put in front of a merchandising team is a different comparison, and it is easy to misread. When AI search used structured Shopify Catalog data to find and recommend products, the shoppers it referred converted at twice the rate of shoppers from AI sessions that relied on scraped or third-party product feeds.3 That is AI against AI rather than AI against organic. What it says is that the same channel pays double when it is grounded in your catalog instead of in a scrape of your storefront, which is a feed argument and connects directly to this site’s earlier piece on the Shopify catalog land grab.
Shopify’s data is Shopify’s platform, self-reported, with no method published alongside it. That makes it first-party and directional, which is a genuine strength and a genuine limit at once.
Which lever is slow, which is fast, and which one are you pulling?
Eligibility is slow and retrieval readiness is fast, and most teams are pulling the fast one. The slow lever is what the source names as the layer that decides which brands the models converge on: original research programs, review-aggregator depth on G2 and Capterra, genuine community and Reddit presence built by participating, and category-defining content that becomes the canonical reference over time. The source is explicit that these do not produce quarterly visibility movement.2 That honesty is worth preserving rather than smoothing over, because a lever with no quarterly signal is one most reporting cycles will quietly drop.
The fast lever still matters on the 44% where retrieval happens. Being fetchable, parseable and current decides whether you survive that half, and this site has covered how selection works on the retrieval side in how ChatGPT chooses its sources.
How do you run this check on your own category?
Ask the assistant your category question, then ask it which searches it ran, and read whether your brand appears in the sub-queries or only in the sources it cites. That single check separates the two problems. Appearing in the sub-queries means the model already holds you as a candidate, and your work is retrieval readiness. Appearing only in the cited sources, or not at all, means the model reached for other brands first, and no amount of on-page improvement fixes that inside a quarter.
Run it three or four times rather than once, because a single answer is one sample of a stochastic system, and write down the date.
There is a second question worth asking on the 56% where nothing was fetched, and it is uncomfortable enough that most tooling avoids it. If the model answered from memory, was the fact about your brand missing from that memory, or present and not retrieved? Those look identical from the outside, they need opposite responses, and Google Research published a benchmark this month that measures the difference. That is the subject of the next piece in this series. In my opinion, the teams who log these checks monthly will be the first to notice their category consolidating, and consolidation is the part that is hard to reverse once it has happened.
Terms defined here
- The pre-seeded query. A retrieval step whose search terms already contain the answer's candidate set, so the retrieval confirms a shortlist rather than discovering one.
Sources
Recent developments
Related reading
This piece elsewhere