cloro

ChatGPT Grounding Study: 86% Commercial, 1% Informational

Ricardo Batista
Founder, cloro
12 min read
On this page

Our January 2026 study reported one figure: ChatGPT grounded 80.5% of responses. That prompt set was entirely commercial, and the pipeline that collected it opted into web search. Neither fact was stated clearly enough at the time. This update re-runs the study without forcing search, on a prompt set split evenly by intent, which decomposes the old number rather than correcting it.

The headline numbers

MetricValue
Successful responses634 (of 865 attempted)
Prompts100 (50 commercial, 50 informational)
Countries10
Test windowAugust 2026
Commercial grounding rate86.5% (268 of 310)
Informational grounding rate0.9% (3 of 324)
Average sources per grounded commercial answer19.3
Web search forcedNo

Prompt intent decides grounding, and almost nothing else does

“How often does ChatGPT ground?” is not a well-formed question. It is two questions with two different answers, separated by what the user wants.

Ask ChatGPT which project management tool suits a remote team of twenty and it searches. Ask it what a raccoon is and it does not, because the answer is stable, well represented in training data, and no live page will improve it. Grounding buys freshness and specificity, and prompts vary enormously in how much of either they need.

A blended rate is therefore close to meaningless. Move the prompt mix and the headline moves with it, with nothing changing in ChatGPT’s behaviour.

Prompt typeResponsesGroundedRateStd. error
Commercial31026886.5%±1.9 pts
Informational32430.9%±0.5 pts

The gap is roughly 85 points against a standard error under 2, so no sample size makes this a close call.

Six categories never grounded at all

The category breakdown shows how sharp the boundary is.

CategoryIntentResponsesGrounding rate
Social media toolingCommercial4797.9%
Legal and complianceCommercial4493.2%
TechnologyCommercial5689.3%
Health productsCommercial5383.0%
EcommerceCommercial6080.0%
Consumer servicesCommercial5078.0%
Health informationInformational506.0%
NatureInformational540.0%
ScienceInformational450.0%
HistoryInformational530.0%
DefinitionsInformational430.0%
How things workInformational420.0%
GeographyInformational370.0%

The takeaway: six informational categories produced 274 responses and not one web search. The only one that grounded at all was health information, at 6.0%, and health questions carry a freshness angle that “why is the ocean salty” does not. For a brand, that means there is no citation to chase in those six categories: a search that never runs cannot cite anyone.

Geography barely matters, which overturns our own earlier framing

Our January write-up made a lot of the country spread, reporting Ukraine at 92% and Serbia at 46%, and reached for explanations about training-data density. That story does not survive this run.

Prompt typeCountriesMean rateStd. deviationRange
Commercial1086.7%3.0 pts81.3% to 90.9%
Informational100.9%1.6 pts0.0% to 3.9%

Across ten countries chosen to span the old study’s extremes, including Ukraine, Serbia, India, Canada and the United States, the commercial rate moves by three points. The original study’s country variation was most likely prompt mix and sampling noise.

The takeaway: if you are prioritising markets for AI SEO, grounding rate is not the variable to sort on. Pick markets on demand, not on how often ChatGPT searches there.

Why 7%, 58% and 86% are all correct

This update started with Graphite’s Do Not Force Search in Prompt Tracking and Ethan Smith’s post about it. Their argument: tools that force web search produce citations no real user sees, so let the model decide and report a per-prompt grounding rate. We think that is right, and this run adopts it.

Their piece also puts three very different numbers next to each other:

SourceReported ratePrompt population
Similarweb, 20267% of U.S. responses carry citationsAll prompt types
Nectiv, October 202531%Mixed
Graphite, August 202658%Entity-comparison prompts
cloro (this study), August 202686.5%Commercial prompts
cloro (this study), August 20260.9%Informational prompts

As competing estimates of one quantity, these look like a field that cannot measure a basic fact. As measurements of different prompt populations, they line up. Similarweb’s 7% averages everything people type into ChatGPT, most of which is not commercial, and our 0.9% shows what the bottom of that distribution looks like. A population dominated by prompts that never ground produces a low average however reliably the commercial slice searches.

Graphite’s 58% and our 86.5% are both commercial-leaning rates from the same month, so the gap between them is real and we cannot resolve it from the outside. Their prompt set is specifically entity-comparison questions where ours spans six commercial categories, their detector requires a fan-out query or a visible Sources panel where ours also counts citation markers, and infrastructure quality decides how often a session is quietly throttled.

What this means for the “AI visibility tools are wrong” argument

Graphite warns that forcing search on a prompt that never searches naturally produces phantom citations, shifting visibility values by about 20 points. Our data says which prompts they mean: the informational half. Forcing search there does not uncover sources that were influencing the answer all along, it manufactures a result no user will ever see.

On the commercial half the concern is smaller. At 86.5%, forcing search changes the answer for a minority of runs, and our January 80.5% with search opted in sits close to the 86.5% we measure without it. The two runs differ in prompt set, country mix and date, so this is not a controlled comparison, but it is consistent with forcing search mattering least where brands actually compete.

For brands, the commercial rate is the only one that matters

An all-prompt average is the wrong number to plan against, because you cannot be cited in an answer that never runs a search. Under a 1% grounding rate there is no citation surface to optimise for and no outreach that puts you in the answer. Those answers come from model weights, and the only way in is being represented widely enough that the next training run absorbs it, which works on the timescale of model releases rather than anything citation tracking measures.

So when a study reports that 93% of ChatGPT answers have no citations, the honest reading is not “citation optimisation is mostly futile”. It is “most ChatGPT traffic is questions brands were never going to appear in”. The commercial share is the addressable market, and inside it grounding is close to the default.

The same logic has an uncomfortable half. A tracking prompt set full of informational questions is measuring a surface you cannot influence. Ask of each prompt you track whether a real buyer would type it and whether it grounds. If it does not ground, it does not belong in a citation report.

What to do with this

  1. Track the grounding rate per prompt, not just citations. A citation list means little without knowing what share of runs produced any citation at all.
  2. Drop informational prompts from citation reporting. Keep them for brand-mention tracking if you like, since a model can mention you from parametric knowledge, but they will never yield a citation you can act on.
  3. Expect grounding on commercial prompts. At 86.5%, not being cited on a commercial query is a competitive result rather than a structural one.
  4. Do not sort markets by grounding rate. Three points of variance across ten countries is not a strategy input.

How we measured it

Every query ran in a fresh browser session with its own residential exit IP in the target country, no shared cookies or cache, and an organic fingerprint. That is unchanged from the original study, and it keeps sessions from being silently throttled, which is what drags measured grounding rates down.

Three things changed:

We stopped forcing search. The original run used our production ChatGPT pipeline, which opts into web search through a URL hint. This run drops it, so ChatGPT lands on a default composer and decides for itself. That is why the two runs are not directly comparable.

The prompt set is split by intent. 100 prompts, 50 commercial and 50 informational, across thirteen categories. The commercial half reconstructs the six categories of the original study; the informational half is new.

Ungrounded answers are recorded, not discarded. Production treats an answer with no web search as a failure and retries it, which is correct for a scraping product and fatal for this measurement.

A response counts as grounded when its payload carries a web-search signal or citation markers. Graphite instead requires a fan-out query or a visible Sources panel, so ours counts some responses theirs would not.

What this study cannot tell you

  • The prompt set is ours, not a sample of real traffic. The 50/50 ratio is a design decision, not a measurement of what people ask ChatGPT, so the blended rate across this run carries no meaning and we have deliberately not headlined it.
  • The original 100-prompt list is gone. The January prompts were not preserved, so the commercial half here reconstructs its categories rather than repeating the set. The two runs are not a controlled before-and-after.
  • Ten countries, unevenly sampled. Infrastructure failures left per-country counts between 24 and 96 responses. The by-intent split is robust to this; per-country rates are not, which is why we publish no country table this time.
  • August 2026, logged out. Logged-in behaviour may differ. Graphite’s logged-in comparison found no large difference, which is reassuring but is not our own measurement.
  • Single vendor. Read these as the rates cloro’s measurement layer observes, not as universal constants.

Where this leaves AI SEO planning

January’s number was underspecified rather than wrong. ChatGPT does ground commercial prompts at a high rate, and that is what the 80.5% measured, on a prompt set that happened to be entirely commercial. What we had not measured is that the same model grounds informational prompts almost never.

Both halves matter. The high commercial rate means citation optimisation has a large surface, and absence from it is a competitive failure rather than bad luck. The near-zero informational rate means much of the published grounding statistics describe a surface no brand can act on, so treating those averages as a verdict on AI SEO is a category error.

If you are measuring your own ChatGPT visibility, separate your prompts the way we separated ours. How often ChatGPT cites anyone depends entirely on which pile you are looking at.

This study used cloro’s infrastructure and independent session testing. For access to our ChatGPT endpoint, visit our page.

Ricardo Batista

About the author

Founder, cloro

Ricardo is one of the founders and engineers behind its SERP and AI-search scraping infrastructure. Before cloro he scaled a financial comparison site to $7M ARR and ran the full-country operations of a unicorn to $65M ARR, then went back to building. He writes about search engine scraping, generative-engine optimization, and turning live search and AI-answer data into something teams can act on.

Frequently asked questions

How often does ChatGPT use web search?

It depends almost entirely on what you ask. In our August 2026 run of 634 successful responses, ChatGPT used web search in 86.5% of commercial prompts (product comparisons, tool recommendations, buying questions) and 0.9% of informational prompts (definitions, science, history). A single blended figure hides that split.

Why do published grounding rates disagree so much?

Because they measure different prompt populations. Similarweb reports about 7% of U.S. ChatGPT responses carry citations across all prompt types. Graphite measures 58% for entity-comparison prompts. We measure 86.5% for commercial prompts. Once you condition on intent, the three are not in conflict.

Which grounding rate should a brand plan against?

The commercial one. Informational prompts almost never ground, and a brand cannot be cited in an answer that never runs a search. The all-prompt average answers a question brands are not asking.

Does forcing web search change the answer?

Yes, and it matters most for prompts that never search on their own. Graphite found visibility values differ by about 20 points for those prompts. That is why this run let ChatGPT decide instead of forcing search.