"How often does AI cite you?" is a harder question than it looks. The answer depends on which questions you ask — change the query mix and the numbers move with it, so a citation rate says as much about the sample as it does about the engine.
So we fixed the questions. The AI Search Index runs the same frozen basket of 735 queries — 15 in every combination of eight industries and six kinds of query, plus a local-intent set — through five engines, every cycle. The goal is a straight answer to one thing, tracked over time: how is AI citation behavior actually changing? It builds on our one-off State of AI Search snapshot, whose prompt mix skewed commercial. This edition sets the baseline the next one is measured against.
Edition 01 (baseline) · Published July 16, 2026 · cloro API · a recurring index, measured quarterly
86 / 100
The Citation Presence Index: the mean share of answers that cite a source, across the five engines. This edition sets the baseline, so Edition 02 shows whether AI answers are citing more or less.
Gemini: 33%
Gemini cites a source on barely a third of its answers — even sparser than our July read (41%). Every other engine cites on 98–100%. On Gemini, most of the time there's no citation to win.
91% vs 81%
Query type moves citation presence more than topic does: "best X" queries get cited on 91% of answers, definitions on just 81%. The page you publish shifts your odds more than your industry does.
AI Mode: 19%
Engines split sharply on the big user-generated platforms: 19% of Google AI Mode's citations and 12% of Perplexity's point at YouTube, Reddit and their peers — versus under 2% for ChatGPT, which leans publisher and official instead.
A recurring measurement of how the major AI answer engines cite their sources. Each edition runs the same frozen basket of 735 queries — 15 in every one of the eight verticals × six query archetypes, plus a 15-prompt local-intent set — through ChatGPT, Perplexity, Copilot, Gemini and Google AI Mode, so no topic or intent dominates the sample.
Because the questions never change, the numbers are directly comparable between editions: movement reflects the engines rather than a change in what we asked. Published quarterly; this page always carries the current edition.
Freshness check · re-measured August 19, 2026
Edition 01 was measured in July 2026 and the next basket run is what turns it into a trend. In the meantime we re-check the headline figures against cloro's production monitoring corpus, which runs continuously. Two engines have moved far enough that the baseline no longer describes today.
| Engine | Index basket, July | Corpus, same July window | Corpus, 30d to Aug 19 | Corpus, last 7d |
|---|---|---|---|---|
| ChatGPT | 98.4% | 99.6% | 79.6% | 91.5% |
| Perplexity | 100% | 99.2% | 99.2% | 100% |
| Copilot | 98.6% | 97.2% | 92.4% | 92.4% |
| Google AI Mode | 100% | 99.4% | 90.7% | 92.2% |
| Gemini | 32.9% | 49.5% | 70.2% | 65.1% |
The first two columns cover the same fortnight and still disagree, most sharply on Gemini (32.9% against 49.5%). That gap is method, not movement: the index runs a topic-balanced basket, while the production corpus is weighted toward what customers actually track. So compare the three corpus columns with each other, and treat the basket column as its own separate series.
Within the corpus, Gemini rose from 49.5% to 70.2%, a genuine change in the engine rather than in our sample. Gemini citing on roughly a third of answers was the most quotable number in this edition, and it is the one that has aged worst. ChatGPT went the other way and is unstable with it: 99.6% in July, 79.6% across the 30 days to August 19, 91.5% across the last seven. No single ChatGPT citation-presence figure survives long enough to plan around, which is the finding, not a caveat about the finding.
Everything below this box is Edition 01, as measured in July 2026. Edition 02 re-runs the basket and restates it properly.
Before "how many sources," there's a prior question: does the answer cite anything? For four of the five engines the answer is almost always yes. Gemini is the exception. It returns a source on barely a third of its answers, and on this balanced basket it comes in even lower than the July snapshot's 41%. On four engines there is nearly always a source list to join; on Gemini, two answers in three offer nothing to win. One blended rate would hide that, so we report presence per engine.
When they do cite, engines differ nearly 4×: Google AI Mode backs an answer with about twenty sources, Copilot with five. Gemini is the split personality. It rarely cites, but when it does it's citation-dense (13.4), close to ChatGPT. Each source is a slot, and a page has far better odds of taking one of AI Mode's twenty than one of Copilot's five.
| Engine | Per answer (all) | Per cited answer |
|---|---|---|
| Google AI Mode | 19.5 | 19.5 |
| ChatGPT | 13.3 | 13.6 |
| Gemini | 4.4 | 13.4 |
| Perplexity | 10.6 | 10.6 |
| Copilot | 5.1 | 5.2 |
"Per answer (all)" averages over every answer including uncited ones; "per cited answer" averages only answers that cited something, which is why Gemini's columns diverge. Quote 4.4 and Gemini barely cites; quote 13.4 and it reads like ChatGPT. A Gemini citation is rare but lands in a crowded list.
Because the basket allots the same 15 prompts to every vertical × query-type cell, we can ask something a snapshot can't: does citation behavior depend on what you ask? The answer is that topic barely matters. Presence sits in a tight 82–88% band across every vertical, so citation presence is engine-driven, not subject-driven. A low citation count says something about your pages and the engine, not about a hard category.
Query type, though, moves it more. Commercial intents ("best X," buying guidance) get cited most; definitional and how-to queries least. If you want a citation, the shape of the question you answer does more work than your sector.
The local-intent set sits outside this grid, and we've left it off the chart on purpose: it's a single 15-prompt cell, so its 85.3% carries a margin of error of roughly ±8 points — wide enough to land almost anywhere among the six above. Read it as "local queries get cited about as often as everything else," not as a ranking.
Pool the cited domains across all five engines and, as in July, the top of the list is user-generated platforms, not publishers. YouTube leads by a wide margin, then Reddit and Google, followed by a spread of publishers, finance sites, and official (.gov) sources. The balanced basket strips out the commercial skew of the old snapshot, and the pattern holds anyway. Much of the citable surface is not your own site: a video, a Reddit thread, or a category publisher's roundup counts as much as a page on your domain.
| # | Domain | Type | Citations |
|---|---|---|---|
| 1 | youtube.com | video / UGC | 2,598 |
| 2 | reddit.com | forum / UGC | 771 |
| 3 | google.com | platform | 695 |
| 4 | indeed.com | jobs board | 552 |
| 5 | forbes.com | publisher | 466 |
| 6 | nerdwallet.com | finance publisher | 463 |
| 7 | healthline.com | health publisher | 401 |
| 8 | nih.gov | official (.gov) | 336 |
| 9 | state.gov | official (.gov) | 326 |
| 10 | experian.com | credit bureau | 278 |
One of the sharpest differences between engines is how much they lean on the major user-generated platforms, and it decides where a UGC push pays: a Reddit thread or YouTube video routes into Google AI Mode and Perplexity, and near a dead end on ChatGPT. To keep this exact, we measure the share of each engine's citations pointing at a fixed, named list of them:YouTube, Reddit, Wikipedia, Quora, Stack Overflow, Stack Exchange, GitHub, X/Twitter, Facebook, Instagram, LinkedIn, TikTok, Pinterest, Tripadvisor, Yelp, Medium, Substack. Because UGC sites outside that list aren't counted, read each figure as a floor, not a ceiling.
Concentration is the share of an engine's citations that comes from just its top 10 domains. On this broader, more neutral basket every engine spreads wider than the commercially-skewed snapshot suggested, though Google AI Mode leans hardest on a narrow set. Even there the top ten take a minority, so a specialist page is not shut out.
Of the five engines, only ChatGPT returns a publish date on its cited pages, so freshness is a ChatGPT-only cut this edition. Where it does report dates, the median cited page is about five months old, and two-thirds were published within the last year. On ChatGPT an older page competes against a mostly recent pool. The cut shows what gets cited, not why: no proof that re-dating a page earns a citation.
150 days
Median age of a cited page (ChatGPT, ~6,000 distinct dated citations).
67%
Share of ChatGPT's cited pages published within the last 12 months.
Measured on a query set that isn't skewed toward any one topic, the headline differences hold: four engines cite almost always, Gemini rarely; the citable surface is led by platforms, not publisher homepages; and each engine reaches for a different kind of source. What the balanced basket adds is a cleaner read on where the variation lives: almost entirely in the engine and the query type, barely at all in the subject.
This edition is the baseline. Run the same 735 prompts again and Edition 02 turns each of these numbers into a trend, up or down since the quarter before.
Every figure comes from running a fixed 735-prompt basket through cloro's API across five engines (ChatGPT, Perplexity, Copilot, Gemini, Google AI Mode), US, in July 2026. Google AI Overview, which the July snapshot covered as a sixth surface, is not measured in this edition. The basket allots 15 prompts to each cell of an 8 verticals × 6 query archetypes grid (720), plus a 15-prompt local-intent set carried by Home/local (735 total). So every vertical is sampled with the same 90 prompts and every query type with the same 120 — no topic or intent is over-represented, the deliberate fix for the commercial skew noted in the July snapshot. Home/local is the one asymmetry: its local set makes it marginally larger than the other seven. About 86% of the prompts are drawn from real search-volume data; the remainder are curated to fill cells where genuine demand is too sparse to sample. Citation presence is the share of answers returning ≥1 source; density averages sources per answer; concentration is the top-10 domain share. The platform share counts citations pointing at a fixed, named list of major UGC platforms rather than attempting to classify all 8,823 distinct cited domains — it is exact for that list and a floor for UGC overall. Freshness uses the publish dates ChatGPT returns (the only engine that does), deduplicated by URL so a page carried in both the source list and its citation pill counts once. Perplexity grounds intermittently, so empty responses were re-run until grounded before counting. This is Edition 01, a dated baseline; the recurring trend begins when the same basket runs again.
Citations-per-answer looks different across cloro's research, and it's worth saying why. Gemini reads 2.5 in the July snapshot, 6.6 (US) in AI Overviews Around the World, and 4.4 here. All three are correct. Averaging over every answer multiplies an engine's cited-depth by how often it cites at all — so for a low-presence engine like Gemini the figure moves with the query mix and the window it was measured in. ChatGPT, which cites on ~98% of answers, reads 12–13 in all three. That's exactly why we publish both columns above: the cited-only figure is the stable cross-study comparison, and on it Gemini (13.4) sits close to ChatGPT (13.6) rather than five times below it.
cloro. (July 16, 2026). AI Search Index: How the Major AI Engines Cite. cloro Research. https://cloro.dev/research/ai-search-index/
This Index builds on the one-off State of AI Search snapshot. For which query intents trigger an AI answer at all, see the AI Overview Trigger Index; for how each engine formats its citations, see LLM Citations. Where this Index is a dated edition, the AI Visibility Leaderboard runs the same measurement weekly on a fixed brand roster, so it tracks which brands the engines name and cite over time rather than at one moment.
Nearly all of them, with one exception. On a fixed 735-prompt basket, Google AI Mode and Perplexity carried at least one source on 100% of answers, Copilot on 98.6%, and ChatGPT on 98.4%. Gemini cited a source on only 32.9% of answers, so a brand can be invisible in Gemini's citations while being cited everywhere else.
Google AI Mode cites the most at 19.5 sources per answer, ChatGPT 13.3, Perplexity 10.6, and Copilot 5.1. Gemini averages 4.4 across all answers but 13.4 across the answers that cite anything, because most of its answers cite nothing. Quote the two Gemini figures together; either one alone misleads.
YouTube leads with 2,598 citation instances across the five conversational engines in the basket, followed by Reddit at 771 and Google's own properties at 695. Video and forum content dominate the top of the list ahead of any publisher or brand site.
It is recurring and controlled. The index runs the same topic-balanced 735-prompt basket, eight verticals by six query archetypes in the US, through cloro's API each edition, so editions are comparable over time. The State of AI Search study was a one-off snapshot of cloro's live monitoring corpus. Edition 01 is the baseline; the trend line starts at Edition 02.
cloro reports live whether ChatGPT, Perplexity, Gemini, Copilot, and Google's AI surfaces cite you, how often, and from where, so you can see which engines you're missing. Start with 500 free credits a month.