cloro
Monitoring

GEO Metrics: Mention Rate, Citation Rate, and What to Measure

Ricardo Batista
Founder, cloro
6 min read
GEOAI VisibilityMeasurement
On this page

Generative engine optimization inherited its dashboards before it settled its definitions, so most GEO reporting mixes five different measurements into one “visibility” score. This page defines the GEO metrics separately, gives the traps we have hit measuring each, and says which one answers which question. For how much data each needs before a trend is real, see AI visibility sample size.

What metrics matter for generative engine optimization?

Five GEO metrics, forming a causal chain rather than a menu:

  1. Retrieval. Was your page among the sources the engine pulled while answering? This is the top of the funnel and the most commonly skipped measurement, because it needs per-answer source data rather than answer text. A page never retrieved cannot be cited, whatever its shape.
  2. Citation. Given retrieval, did the answer actually reference your page? Retrieved-but-never-cited and never-retrieved look identical in the answer text and take opposite fixes: the first is a content-shape problem, the second is discovery and authority (what actually moves each).
  3. Mention rate. What share of answers name your brand in the prose, per prompt, per engine. This is the headline brand metric, and it can be high with zero citations when the model knows the brand from training data.
  4. Position. Where in the answer the mention lands. First-mentioned brands are the default recommendation; a mention in the seventh paragraph of a comparison is a different asset.
  5. Share of voice. Your mentions divided by total mentions across a fixed competitor panel, over the same answers. The denominator only means something while the panel is frozen; every panel change rescales the history.

The discipline that makes all five usable: record the denominator next to every rate, and name the engine on every row. A number without its denominator and engine is a vanity score.

They measure different objects on different denominators: mention rate is brand-level, the share of answers whose text names your brand, while citation rate is page-level and per retrieval, citation_count ÷ retrieval_count, how often an engine that pulled your page went on to quote it. Mention rate reads the answer’s prose, whereas citation rate reads its source list.

The trap we hit ourselves: citation rate is not a percentage. One answer can cite the same URL twice, so the ratio exceeds 1.0 routinely, and reading it as “share of retrievals that produced a citation” inverted a priority order in our own reporting once. Treat it as an average with no ceiling.

The two metrics also fail independently, which is why you need both. A brand mentioned constantly but never cited is living off training data, and that presence erodes as engines lean harder on live retrieval. A page cited regularly under a brand nobody names wins the link without winning the recommendation. The first case needs retrievable pages; the second needs the brand named in the pages that get retrieved.

How should I measure whether AI engines cite my brand?

Treat it as the citation stage of the GEO metrics chain and measure it as a rate over scheduled runs of a fixed prompt set, with the answer-level source list as your raw data, and store the raw data rather than the score:

  1. Fix a prompt set and a competitor panel. Both frozen inside a measurement window; date every change.
  2. Run every prompt on every engine daily and capture the full answer with its sources array, not just the text.
  3. Compute per engine, per week: mention rate, citation count per page, and share of voice against the frozen panel.
  4. Keep the answers. Scores computed by your own code from stored raw answers can be audited, recomputed under a new definition, and reconciled against the search metrics you already report. A vendor score with no raw data underneath cannot.

That last point is the structural argument for collecting via an API that returns the raw answer rather than a dashboard that returns a number: your definitions, your denominators, your audit trail.

Is AI visibility comparable across ChatGPT, Gemini and Perplexity?

Only per engine, never as one blended GEO metrics score. The engines convert retrieval into citation at rates that differ by an order of magnitude: measured on our own domain across the same prompt corpus, the retrieval-to-citation ratio was 0.11 on ChatGPT, 1.25 on Google AI Overview, and 2.58 on Gemini. The same page on the same prompt was cited 9 times by Gemini and 0 times by ChatGPT in one of our reads. A pooled “visibility” average over those engines describes no engine that exists.

Stat: 23x spread in retrieval-to-citation conversion across engines, 0.11 on ChatGPT to 2.58 on Gemini for the same domain (source: cloro measurement, 2026-08-08)

Cross-engine comparison is still possible, with two constraints: hold the prompt set identical across engines, and report each engine as its own row. “We are strong on Gemini and invisible on ChatGPT” is an actionable finding with two different fixes; “our visibility is 43” is not.

How quickly do AI search answers change for the same question?

Fast enough that any single answer is already stale as evidence. We ran the same 20 prompts through the same model (claude-sonnet-5, web search on) twice, minutes apart: the second run repeated only about 40% of the first run’s cited domains, with per-prompt overlap ranging from 0.06 to 1.00. Nothing about the world changed between the runs; the variation is the engine’s own sampling.

Day to day the drift compounds: retrieval pools refresh, models rotate behind the interface, and per-engine grounding behavior shifts without notice (measured per engine). The practical consequences: alert on multi-day aggregates rather than single answers, expect week-over-week reads to be the finest sound grain for mid-sized prompt sets, and treat a screenshot of one answer, good or bad, as an anecdote. The sample-size math for separating that noise from a real trend is worked through in AI visibility sample size.

A worked example: five metrics from one week of answers

Numbers make the definitions concrete. Take one prompt (“best rank tracking API”), tracked on two engines, one answer per engine per day for a week: 14 answers. Suppose the stored answers contain:

ChatGPT (7 answers)Gemini (7 answers)
Answers naming your brand25
Answers retrieving your page46
Citations of your page19
Competitor mentions (panel total)1113

The five GEO metrics fall straight out. Retrieval: 4/7 on ChatGPT, 6/7 on Gemini; discovery is working on both. Citation rate: 0.25 on ChatGPT (1 citation over 4 retrievals) against 1.50 on Gemini (9 over 6, exceeding 1.0 because answers cited the page more than once); same page, engines a factor of six apart. Mention rate: 29% on ChatGPT, 71% on Gemini. Share of voice: 2/(2+11) = 15% on ChatGPT, 5/(5+13) = 28% on Gemini, panel-relative on the frozen competitor list. Position comes from where in each answer the mention lands.

Two readings this table of GEO metrics forbids: a pooled “50% mention rate” (true, and describes neither engine), and any month-over-month claim (14 answers per cell resolves nothing, per the sample-size math). What it does support: the ChatGPT gap is a citation problem, not a discovery problem, because the page is retrieved and not quoted, and that diagnosis picks the fix.

Dashboards compress this table into one score by necessity; Semrush’s AI share-of-voice guide shows the packaged version of the same computation. Owning the raw answers means you can always unfold the score back into this table when a number needs defending.

What does it cost to measure GEO metrics properly?

Cheap enough that measurement quality is never the budget’s fault. The worked example above (one prompt, two engines, daily for a week) costs about $0.03: 63 credits at cloro’s published pricing of $0.40 per 1,000 credits on the $100 Hobby tier. Scaling to a serious setup of 100 prompts across 6 engines daily runs about 78,000 credits a month, roughly $31, and the Hobby tier’s 250,000 credits carry it with room for a second market. The number worth internalizing: the data behind a defensible GEO metrics program costs less per month than an hour of the analyst reading it.

That inverts where teams usually economize. Cutting prompts to save money saves almost none and costs statistical power; the expensive failure is the opposite one, paying for a dashboard seat whose number cannot be decomposed when a stakeholder challenges it.

How do I verify my AI visibility platform reports the answers users actually get?

Spot-check it against the surface it claims to measure: take a tracked prompt, open the engine as a logged-out user in the tracked country, and compare the live answer’s citations against what the platform recorded for the same day. Three specific things to check, all failure modes we have measured or documented:

  1. Consumer surface, or model API? Some trackers query a vendor’s developer API (Perplexity’s Sonar and its successor APIs, OpenAI’s API) instead of the consumer interface. The two are different products with different grounding, and the consumer answer is the one your buyers see. Ask the vendor, whether that is Peec, Profound, Otterly or another of the trackers we compare, which surface each engine’s data comes from.
  2. Field completeness. Engines serve some responses in formats that silently drop metadata: OpenAI’s mobile-web format returns ChatGPT answers with no model name and a shortened source list. A tracker unaware of this reports partial answers as complete ones. Ask how the vendor detects and labels degraded responses.
  3. Geography and session. Answers vary by country and by logged-in state. A tracker running from one region with one account state measures that region and that state; check both match your market.

A platform that passes the spot-check on your five most important prompts, per engine, has earned the trust the dashboard asks for. One that cannot explain a mismatch has answered the question too.

Ricardo Batista

About the author

Founder, cloro

Ricardo is one of the founders and engineers behind its SERP and AI-search scraping infrastructure. Before cloro he scaled a financial comparison site to $7M ARR and ran the full-country operations of a unicorn to $65M ARR, then went back to building. He writes about search engine scraping, generative-engine optimization, and turning live search and AI-answer data into something teams can act on.

Frequently asked questions

What metrics matter for generative engine optimization?

Five, in a causal chain: retrieval (was your page in the sources an engine pulled), citation (did the answer quote it), mention rate (what share of answers name your brand), position (where in the answer you appear), and share of voice (your mentions against a fixed competitor panel). Read them per engine, never pooled, and pick the metric that matches the failure you are diagnosing: never-retrieved is a discovery problem, retrieved-but-uncited is a shape problem.

What is the difference between mention rate and citation rate in AI search?

Mention rate is brand-level: the share of answers whose text names your brand, whether or not any link points at you. Citation rate is page-level and, computed properly, is citations per retrieval: how often an engine that pulled your page went on to quote it. It routinely exceeds 1.0 when one answer cites the same URL twice, so it is not a percentage. A brand can have a high mention rate on zero citations (the model knows you from training) and a page can be cited without the answer naming your brand.

Is AI visibility comparable across ChatGPT, Gemini and Perplexity?

Not as one number. In cloro's measurement the same domain converts retrieval into citation at 0.11 on ChatGPT, 1.25 on Google AI Overview and 2.58 on Gemini, so a pooled average describes no engine at all. Compare within an engine over time, or across engines on the same fixed prompt set with the engine named on every row.