ChatGPT Grounding Study: 86% Commercial, 1% Informational
On this page
Our January 2026 study reported one figure: ChatGPT grounded 80.5% of responses. That prompt set was entirely commercial, and the pipeline that collected it opted into web search. Neither fact was stated clearly enough at the time. This update re-runs the study without forcing search, on a prompt set split evenly by intent, which decomposes the old number rather than correcting it.
The headline numbers
| Metric | Value |
|---|---|
| Successful responses | 634 (of 865 attempted) |
| Prompts | 100 (50 commercial, 50 informational) |
| Countries | 10 |
| Test window | August 2026 |
| Commercial grounding rate | 86.5% (268 of 310) |
| Informational grounding rate | 0.9% (3 of 324) |
| Average sources per grounded commercial answer | 19.3 |
| Web search forced | No |
Prompt intent decides grounding, and almost nothing else does
“How often does ChatGPT ground?” is not a well-formed question. It is two questions with two different answers, separated by what the user wants.
Ask ChatGPT which project management tool suits a remote team of twenty and it searches. Ask it what a raccoon is and it does not, because the answer is stable, well represented in training data, and no live page will improve it. Grounding buys freshness and specificity, and prompts vary enormously in how much of either they need.
A blended rate is therefore close to meaningless. Move the prompt mix and the headline moves with it, with nothing changing in ChatGPT’s behaviour.
| Prompt type | Responses | Grounded | Rate | Std. error |
|---|---|---|---|---|
| Commercial | 310 | 268 | 86.5% | ±1.9 pts |
| Informational | 324 | 3 | 0.9% | ±0.5 pts |
The gap is roughly 85 points against a standard error under 2, so no sample size makes this a close call.
Six categories never grounded at all
The category breakdown shows how sharp the boundary is.
| Category | Intent | Responses | Grounding rate |
|---|---|---|---|
| Social media tooling | Commercial | 47 | 97.9% |
| Legal and compliance | Commercial | 44 | 93.2% |
| Technology | Commercial | 56 | 89.3% |
| Health products | Commercial | 53 | 83.0% |
| Ecommerce | Commercial | 60 | 80.0% |
| Consumer services | Commercial | 50 | 78.0% |
| Health information | Informational | 50 | 6.0% |
| Nature | Informational | 54 | 0.0% |
| Science | Informational | 45 | 0.0% |
| History | Informational | 53 | 0.0% |
| Definitions | Informational | 43 | 0.0% |
| How things work | Informational | 42 | 0.0% |
| Geography | Informational | 37 | 0.0% |
The takeaway: six informational categories produced 274 responses and not one web search. The only one that grounded at all was health information, at 6.0%, and health questions carry a freshness angle that “why is the ocean salty” does not. For a brand, that means there is no citation to chase in those six categories: a search that never runs cannot cite anyone.
Geography barely matters, which overturns our own earlier framing
Our January write-up made a lot of the country spread, reporting Ukraine at 92% and Serbia at 46%, and reached for explanations about training-data density. That story does not survive this run.
| Prompt type | Countries | Mean rate | Std. deviation | Range |
|---|---|---|---|---|
| Commercial | 10 | 86.7% | 3.0 pts | 81.3% to 90.9% |
| Informational | 10 | 0.9% | 1.6 pts | 0.0% to 3.9% |
Across ten countries chosen to span the old study’s extremes, including Ukraine, Serbia, India, Canada and the United States, the commercial rate moves by three points. The original study’s country variation was most likely prompt mix and sampling noise.
The takeaway: if you are prioritising markets for AI SEO, grounding rate is not the variable to sort on. Pick markets on demand, not on how often ChatGPT searches there.
Why 7%, 58% and 86% are all correct
This update started with Graphite’s Do Not Force Search in Prompt Tracking and Ethan Smith’s post about it. Their argument: tools that force web search produce citations no real user sees, so let the model decide and report a per-prompt grounding rate. We think that is right, and this run adopts it.
Their piece also puts three very different numbers next to each other:
| Source | Reported rate | Prompt population |
|---|---|---|
| Similarweb, 2026 | 7% of U.S. responses carry citations | All prompt types |
| Nectiv, October 2025 | 31% | Mixed |
| Graphite, August 2026 | 58% | Entity-comparison prompts |
| cloro (this study), August 2026 | 86.5% | Commercial prompts |
| cloro (this study), August 2026 | 0.9% | Informational prompts |
As competing estimates of one quantity, these look like a field that cannot measure a basic fact. As measurements of different prompt populations, they line up. Similarweb’s 7% averages everything people type into ChatGPT, most of which is not commercial, and our 0.9% shows what the bottom of that distribution looks like. A population dominated by prompts that never ground produces a low average however reliably the commercial slice searches.
Graphite’s 58% and our 86.5% are both commercial-leaning rates from the same month, so the gap between them is real and we cannot resolve it from the outside. Their prompt set is specifically entity-comparison questions where ours spans six commercial categories, their detector requires a fan-out query or a visible Sources panel where ours also counts citation markers, and infrastructure quality decides how often a session is quietly throttled.
What this means for the “AI visibility tools are wrong” argument
Graphite warns that forcing search on a prompt that never searches naturally produces phantom citations, shifting visibility values by about 20 points. Our data says which prompts they mean: the informational half. Forcing search there does not uncover sources that were influencing the answer all along, it manufactures a result no user will ever see.
On the commercial half the concern is smaller. At 86.5%, forcing search changes the answer for a minority of runs, and our January 80.5% with search opted in sits close to the 86.5% we measure without it. The two runs differ in prompt set, country mix and date, so this is not a controlled comparison, but it is consistent with forcing search mattering least where brands actually compete.
For brands, the commercial rate is the only one that matters
An all-prompt average is the wrong number to plan against, because you cannot be cited in an answer that never runs a search. Under a 1% grounding rate there is no citation surface to optimise for and no outreach that puts you in the answer. Those answers come from model weights, and the only way in is being represented widely enough that the next training run absorbs it, which works on the timescale of model releases rather than anything citation tracking measures.
So when a study reports that 93% of ChatGPT answers have no citations, the honest reading is not “citation optimisation is mostly futile”. It is “most ChatGPT traffic is questions brands were never going to appear in”. The commercial share is the addressable market, and inside it grounding is close to the default.
The same logic has an uncomfortable half. A tracking prompt set full of informational questions is measuring a surface you cannot influence. Ask of each prompt you track whether a real buyer would type it and whether it grounds. If it does not ground, it does not belong in a citation report.
What to do with this
- Track the grounding rate per prompt, not just citations. A citation list means little without knowing what share of runs produced any citation at all.
- Drop informational prompts from citation reporting. Keep them for brand-mention tracking if you like, since a model can mention you from parametric knowledge, but they will never yield a citation you can act on.
- Expect grounding on commercial prompts. At 86.5%, not being cited on a commercial query is a competitive result rather than a structural one.
- Do not sort markets by grounding rate. Three points of variance across ten countries is not a strategy input.
How we measured it
Every query ran in a fresh browser session with its own residential exit IP in the target country, no shared cookies or cache, and an organic fingerprint. That is unchanged from the original study, and it keeps sessions from being silently throttled, which is what drags measured grounding rates down.
Three things changed:
We stopped forcing search. The original run used our production ChatGPT pipeline, which opts into web search through a URL hint. This run drops it, so ChatGPT lands on a default composer and decides for itself. That is why the two runs are not directly comparable.
The prompt set is split by intent. 100 prompts, 50 commercial and 50 informational, across thirteen categories. The commercial half reconstructs the six categories of the original study; the informational half is new.
Ungrounded answers are recorded, not discarded. Production treats an answer with no web search as a failure and retries it, which is correct for a scraping product and fatal for this measurement.
A response counts as grounded when its payload carries a web-search signal or citation markers. Graphite instead requires a fan-out query or a visible Sources panel, so ours counts some responses theirs would not.
What this study cannot tell you
- The prompt set is ours, not a sample of real traffic. The 50/50 ratio is a design decision, not a measurement of what people ask ChatGPT, so the blended rate across this run carries no meaning and we have deliberately not headlined it.
- The original 100-prompt list is gone. The January prompts were not preserved, so the commercial half here reconstructs its categories rather than repeating the set. The two runs are not a controlled before-and-after.
- Ten countries, unevenly sampled. Infrastructure failures left per-country counts between 24 and 96 responses. The by-intent split is robust to this; per-country rates are not, which is why we publish no country table this time.
- August 2026, logged out. Logged-in behaviour may differ. Graphite’s logged-in comparison found no large difference, which is reassuring but is not our own measurement.
- Single vendor. Read these as the rates cloro’s measurement layer observes, not as universal constants.
Where this leaves AI SEO planning
January’s number was underspecified rather than wrong. ChatGPT does ground commercial prompts at a high rate, and that is what the 80.5% measured, on a prompt set that happened to be entirely commercial. What we had not measured is that the same model grounds informational prompts almost never.
Both halves matter. The high commercial rate means citation optimisation has a large surface, and absence from it is a competitive failure rather than bad luck. The near-zero informational rate means much of the published grounding statistics describe a surface no brand can act on, so treating those averages as a verdict on AI SEO is a category error.
If you are measuring your own ChatGPT visibility, separate your prompts the way we separated ours. How often ChatGPT cites anyone depends entirely on which pile you are looking at.
This study used cloro’s infrastructure and independent session testing. For access to our ChatGPT endpoint, visit our page.

About the author
Ricardo Batista
Founder, cloro
Ricardo is one of the founders and engineers behind its SERP and AI-search scraping infrastructure. Before cloro he scaled a financial comparison site to $7M ARR and ran the full-country operations of a unicorn to $65M ARR, then went back to building. He writes about search engine scraping, generative-engine optimization, and turning live search and AI-answer data into something teams can act on.
Frequently asked questions
How often does ChatGPT use web search?
It depends almost entirely on what you ask. In our August 2026 run of 634 successful responses, ChatGPT used web search in 86.5% of commercial prompts (product comparisons, tool recommendations, buying questions) and 0.9% of informational prompts (definitions, science, history). A single blended figure hides that split.
Why do published grounding rates disagree so much?
Because they measure different prompt populations. Similarweb reports about 7% of U.S. ChatGPT responses carry citations across all prompt types. Graphite measures 58% for entity-comparison prompts. We measure 86.5% for commercial prompts. Once you condition on intent, the three are not in conflict.
Which grounding rate should a brand plan against?
The commercial one. Informational prompts almost never ground, and a brand cannot be cited in an answer that never runs a search. The all-prompt average answers a question brands are not asking.
Does forcing web search change the answer?
Yes, and it matters most for prompts that never search on their own. Graphite found visibility values differ by about 20 points for those prompts. That is why this run let ChatGPT decide instead of forcing search.
Related reading

ChatGPT Ads Penetration Study: 51% in the US
ChatGPT ads over the last 7 days measure 51.0% of US responses, 53.6% CA, 49.8% AU, and, new since May, 18.8% JP. Everywhere else remains near zero. The surface reverted then re-spiked in June and now looks structural, not transient.

Query Fan Out: How AI Search Expands One Prompt Into Many
Query fan-out turns one AI search prompt into many sub-queries. Learn how ChatGPT and AI Mode retrieve sources and how to optimize for it in 2026.
ChatGPT Visibility Tracker: Track Your Brand Mentions
Learn how a ChatGPT visibility tracker works, run a free manual baseline, and compare the best tools to monitor your brand mentions in AI search.