"How often does AI cite you?" is a harder question than it looks. The answer depends on which questions you ask — change the query mix and the numbers move with it, so a citation rate says as much about the sample as it does about the engine.
So we fixed the questions. The AI Search Index runs the same frozen basket of 735 queries — 15 in every combination of eight industries and six kinds of query, plus a local-intent set — through five engines, every cycle. The goal is a straight answer to one thing, tracked over time: how is AI citation behavior actually changing? It builds on our one-off State of AI Search snapshot, whose prompt mix skewed commercial. This edition sets the baseline the next one is measured against.
Edition 01 (baseline) · Published July 16, 2026 · cloro API · a recurring index, measured quarterly
86 / 100
The Citation Presence Index — the mean share of answers that cite a source, across the five engines. A single number this edition establishes as the baseline; every future edition moves it up or down.
Gemini: 33%
Gemini cites a source on barely a third of its answers — even sparser than our July read (41%). Every other engine cites on 98–100%. On Gemini, most of the time there's no citation to win.
91% vs 81%
Query type moves citation presence more than topic does: "best X" queries get cited on 91% of answers, definitions on just 81%. The balanced basket is what makes this cut possible.
AI Mode: 19%
Engines split sharply on the big user-generated platforms: 19% of Google AI Mode's citations and 12% of Perplexity's point at YouTube, Reddit and their peers — versus under 2% for ChatGPT, which leans publisher and official instead.
A recurring measurement of how the major AI answer engines cite their sources. Each edition runs the same frozen basket of 735 queries — 15 in every one of the eight verticals × six query archetypes, plus a 15-prompt local-intent set — through ChatGPT, Perplexity, Copilot, Gemini and Google AI Mode, so no topic or intent dominates the sample.
Because the questions never change, the numbers are directly comparable between editions: movement reflects the engines rather than a change in what we asked. Published quarterly; this page always carries the current edition.
Before "how many sources," there's a prior question: does the answer cite anything? For four of the five engines the answer is almost always yes. Gemini is the exception — it returns a source on barely a third of its answers, and on this balanced basket it comes in even lower than the July snapshot's 41%.
When they do cite, engines differ nearly 4×: Google AI Mode backs an answer with about twenty sources, Copilot with five. Gemini is the split personality — it rarely cites, but when it does it's citation-dense (13.4), close to ChatGPT.
| Engine | Per answer (all) | Per cited answer |
|---|---|---|
| Google AI Mode | 19.5 | 19.5 |
| ChatGPT | 13.3 | 13.6 |
| Gemini | 4.4 | 13.4 |
| Perplexity | 10.6 | 10.6 |
| Copilot | 5.1 | 5.2 |
"Per answer (all)" averages over every answer including uncited ones; "per cited answer" averages only answers that cited something — which is why Gemini's two columns diverge so far.
Because the basket allots the same 15 prompts to every vertical × query-type cell, we can ask something a snapshot can't: does citation behavior depend on what you ask? The answer is that topic barely matters — presence sits in a tight 82–88% band across every vertical, so citation presence is engine-driven, not subject-driven.
Query type, though, moves it more. Commercial intents — "best X," buying guidance — get cited most; definitional and how-to queries least. If you want a citation, the query shape you're answering matters more than the industry you're in.
The local-intent set sits outside this grid, and we've left it off the chart on purpose: it's a single 15-prompt cell, so its 85.3% carries a margin of error of roughly ±8 points — wide enough to land almost anywhere among the six above. Read it as "local queries get cited about as often as everything else," not as a ranking.
Pool the cited domains across all five engines and, as in July, the top of the list is user-generated platforms, not publishers. YouTube leads by a wide margin, then Reddit and Google — followed by a spread of publishers, finance sites, and official (.gov) sources. On the balanced basket the commercial skew of the old snapshot is gone, and this durable pattern remains.
| # | Domain | Type | Citations |
|---|---|---|---|
| 1 | youtube.com | video / UGC | 2,598 |
| 2 | reddit.com | forum / UGC | 771 |
| 3 | google.com | platform | 695 |
| 4 | indeed.com | jobs board | 552 |
| 5 | forbes.com | publisher | 466 |
| 6 | nerdwallet.com | finance publisher | 463 |
| 7 | healthline.com | health publisher | 401 |
| 8 | nih.gov | official (.gov) | 336 |
| 9 | state.gov | official (.gov) | 326 |
| 10 | experian.com | credit bureau | 278 |
One of the sharpest differences between engines is how much they lean on the major user-generated platforms. To keep this exact rather than a judgement call, we measure the share of each engine's citations pointing at a fixed, named list of them —YouTube, Reddit, Wikipedia, Quora, Stack Overflow, Stack Exchange, GitHub, X/Twitter, Facebook, Instagram, LinkedIn, TikTok, Pinterest, Tripadvisor, Yelp, Medium, Substack. Because UGC sites outside that list aren't counted, read each figure as a floor, not a ceiling.
Concentration is the share of an engine's citations that comes from just its top 10 domains. On this broader, more neutral basket every engine spreads wider than the commercially-skewed snapshot suggested — but Google AI Mode still leans hardest on a narrow set.
Of the five engines, only ChatGPT returns a publish date on its cited pages, so freshness is a ChatGPT-only cut this edition. Where it does report dates, the median cited page is about five months old — and two-thirds were published within the last year.
150 days
Median age of a cited page (ChatGPT, ~6,000 distinct dated citations).
67%
Share of ChatGPT's cited pages published within the last 12 months.
Measured on a query set that isn't skewed toward any one topic, the headline differences hold: four engines cite almost always, Gemini rarely; the citable surface is led by platforms, not publisher homepages; and each engine reaches for a different kind of source. What the balanced basket adds is a cleaner read on where the variation lives — overwhelmingly in the engine and the query type, barely at all in the subject.
This edition is the baseline. Its value compounds at the next one: with the same 735 prompts run again, Edition 02 turns each of these numbers into a trend — up or down since the quarter before.
Every figure comes from running a fixed 735-prompt basket through cloro's API across five engines (ChatGPT, Perplexity, Copilot, Gemini, Google AI Mode), US, in July 2026. Google AI Overview, which the July snapshot covered as a sixth surface, is not measured in this edition. The basket allots 15 prompts to each cell of an 8 verticals × 6 query archetypes grid (720), plus a 15-prompt local-intent set carried by Home/local (735 total). So every vertical is sampled with the same 90 prompts and every query type with the same 120 — no topic or intent is over-represented, the deliberate fix for the commercial skew noted in the July snapshot. Home/local is the one asymmetry: its local set makes it marginally larger than the other seven. About 86% of the prompts are drawn from real search-volume data; the remainder are curated to fill cells where genuine demand is too sparse to sample. Citation presence is the share of answers returning ≥1 source; density averages sources per answer; concentration is the top-10 domain share. The platform share counts citations pointing at a fixed, named list of major UGC platforms rather than attempting to classify all 8,823 distinct cited domains — it is exact for that list and a floor for UGC overall. Freshness uses the publish dates ChatGPT returns (the only engine that does), deduplicated by URL so a page carried in both the source list and its citation pill counts once. Perplexity grounds intermittently, so empty responses were re-run until grounded before counting. This is Edition 01, a dated baseline; the recurring trend begins when the same basket runs again.
Citations-per-answer looks different across cloro's research, and it's worth saying why. Gemini reads 2.5 in the July snapshot, 6.6 (US) in AI Overviews Around the World, and 4.4 here. All three are correct. Averaging over every answer multiplies an engine's cited-depth by how often it cites at all — so for a low-presence engine like Gemini the figure moves with the query mix and the window it was measured in. ChatGPT, which cites on ~98% of answers, reads 12–13 in all three. That's exactly why we publish both columns above: thecited-only figure is the stable cross-study comparison, and on it Gemini (13.4) sits close to ChatGPT (13.6) rather than five times below it.
cloro. (July 16, 2026). AI Search Index: How the Major AI Engines Cite. cloro Research. https://cloro.dev/research/ai-search-index/
This Index builds on the one-off State of AI Search snapshot. For which query intents trigger an AI answer at all, see the AI Overview Trigger Index; for how each engine formats its citations, see LLM Citations.
cloro reports whether ChatGPT, Perplexity, Gemini, Copilot, and Google's AI surfaces cite you, how often, and from where — live. Start with 500 free credits.