Per-market SERP statistics are everywhere, and almost all of them are built the same way: take whatever queries you happen to collect in each country, and report the difference. We published one ourselves. This study asks what is left of those differences when every market is asked the same 150 questions.
Most of the difference does not survive. Austria and Germany sat 28 points apart on AI Overview presence in our own census. On a fixed basket, in the same language, they are 0.7 points apart. The gap was our query mix, not Google's behaviour.
Something real does survive. Across six markets the controlled spread is 18.7 points rather than the 34.9 the census reported, so roughly half the apparent difference between countries is an artefact of what each country was asked. The half that remains is worth knowing about, and it puts the markets in a different order.
Published August 20, 2026 · 1,500 scored Google SERPs · 150 queries · 6 markets · 2 interface languages · single 30-minute window
28 points → 0.7
The Austria against Germany gap in AI Overview presence, before and after holding the questions and the language still. Two markets that looked profoundly different behave the same way.
Half the spread
34.9 points across six markets uncontrolled, 18.7 controlled. A real market effect remains, at about half the size the uncontrolled number implies.
88 points
The gap between the query type most likely to trigger an AI Overview and the least, pooled across every market. Query mix moves this number close to five times as hard as the market does.
Order inverts
Austria ranks highest of the six markets in the census and lowest here. Uncontrolled per-market figures are not merely imprecise, they can point the wrong way.
This is the contrast the study was built for. Austria and Germany share a language, so running the basket in German in both varies one thing: the country. Everything else is pinned.
Census, uncontrolled
95.9% vs 67.9%
28.0 points apart. Each market runs its own prompts.
This study, controlled
48.0% vs 47.3%
0.7 points apart. Same 150 questions, both in German.
A 28-point difference between two neighbouring markets is the kind of number that ends up in a slide deck. It was not describing Google. The Austrian prompts in the monitoring corpus and the German ones are different questions, and different questions carry different features.
We cannot claim the true difference is exactly zero. At 150 queries per cell the interval on a single rate is about 7.9 points either side, so the honest statement is that a gap anywhere near 28 points would have been impossible to miss, and nothing of that size is there.
Each market measured on the same basket, in its own interface language, against what the census reported for that market. The two columns are not measuring the same population, which is the point: the left column is one fixed question set, the right is whatever customers tracking that market happened to ask.
| Market | Controlled | Census | Difference |
|---|---|---|---|
| United States hl=en | 66.0% | 87.2% | -21.2 |
| Japan hl=ja | 64.0% | 90.7% | -26.7 |
| United Kingdom hl=en | 58.7% | 75.8% | -17.1 |
| France hl=fr | 54.0% | 61.0% | -7.0 |
| Austria hl=de | 48.0% | 95.9% | -47.9 |
| Germany hl=de | 47.3% | 67.9% | -20.6 |
Read the columns, not the rows. Every controlled figure is lower than its census counterpart, and that is a property of our basket rather than a finding about Google. The basket is balanced across six query types on purpose, including navigational and product queries that rarely trigger an AI Overview. The monitoring corpus leans conversational and commercial. So the 66.0% for the United States is not "the US AI Overview rate" and should not be quoted as one. What transfers between the two studies is the comparison, not the level.
The ordering is the part worth carrying away. Austria leads the census and comes last here. France sits at the bottom of the census and in the middle here. If a per-market SERP statistic can inverse the ranking of the markets when the questions are held still, then ranking markets by an uncontrolled figure is not a weak measurement, it is the wrong measurement.
The census could say its per-market table was confounded but not how badly. This is the answer. Every query in the basket carries an archetype, and pooling all ten cells gives the range that a market's question mix slides along.
| Archetype | AI Overview | People Also Ask |
|---|---|---|
| Definitional | 97.2% | 100.0% |
| Informational | 87.2% | 96.0% |
| Service | 70.0% | 94.0% |
| Commercial | 54.8% | 98.0% |
| Product | 20.0% | 98.4% |
| Navigational | 9.2% | 44.4% |
Definitional queries return an AI Overview 97.2% of the time. Navigational queries return one 9.2% of the time. That is an 88-point range inside a single market, against 18.7 points across six markets once the questions are fixed.
So a market whose tracked queries skew definitional will post a high AI Overview rate, and a market whose queries skew navigational will post a low one, and neither number says anything about the country. This is exactly what happened to Israel in the census: a 56.1% People Also Ask rate against 90% or better nearly everywhere, on a query set of 130 consumer-recommendation prompts. Note the People Also Ask column here: it sits between 94% and 100% for five archetypes and falls to 44.4% for navigational queries. A market tracking mostly navigational queries would reproduce the Israel result without Israel being involved.
Language was the other candidate for the France and Austria gap, and the census could not test it: the standard search request carries no language parameter, so hl floats with whatever Google infers. Here it is pinned, and each market runs the basket twice, once in its own language and once in English.
| Market | Own language | hl=en | Difference |
|---|---|---|---|
| France hl=fr | 54.0% | 51.3% | +2.7 |
| Germany hl=de | 47.3% | 50.7% | -3.4 |
| Austria hl=de | 48.0% | 50.7% | -2.7 |
| Japan hl=ja | 64.0% | 73.3% | -9.3 |
France moves 2.7 points up in French, Germany 3.3 points down in German, Austria 2.7 down, Japan 9.3 down in Japanese. Three of the four are inside the interval on a single cell, and the directions disagree, so there is no language effect here worth reporting as one. Only Japan approaches readability, and one market moving in one direction is a thing to test rather than a thing to conclude.
One feature does depend on it sharply. Shopping cards appear on 20.0% of Japanese-language results in Japan and on 0.0% of English-language results in the same market, which is the only clean language effect in the study. Everywhere else shopping cards sit between 16.7% and 21.3% regardless of language. We are not going to explain that from one basket, but anyone measuring commerce surfaces in Japan should pin the language before trusting the number.
AI Overview is not deterministic. Google can return one for a query on one fetch and not on the next, so before reading any market difference it is worth knowing how far a cell moves when nothing about it changed. The entire basket was run a second time about twenty minutes later.
| Cell | Run 1 | Run 2 | Movement |
|---|---|---|---|
| jp/en | 73.3% | 74.0% | +0.7 |
| us/en | 66.0% | 68.7% | +2.7 |
| jp/ja | 64.0% | 60.0% | -4.0 |
| gb/en | 58.7% | 57.3% | -1.4 |
| fr/fr | 54.0% | 52.0% | -2.0 |
| fr/en | 51.3% | 50.7% | -0.6 |
| at/en | 50.7% | 48.0% | -2.7 |
| de/en | 50.7% | 50.7% | 0.0 |
| at/de | 48.0% | 48.0% | 0.0 |
| de/de | 47.3% | 49.3% | +2.0 |
Cells move 1.6 points on average and 4.0 points at worst. That is the floor: any market difference smaller than about four points is re-measurement noise, and the 18.7-point spread that survives control is comfortably above it. The residual market effect is real.
The Austria and Germany result replicates, and it changes sign. Austria was 0.7 points above Germany in the first run and 1.3 points below it in the second. A difference that small and that unstable in direction is the signature of no difference at all, which is exactly what a 28-point census gap should not look like.
The levels belong to the basket, not to Google. Every rate here is the rate on 150 queries chosen to spread evenly across six archetypes, which is not how anybody's real query mix looks. Use the contrasts.
Six markets is not the world. The census covers 41, and this panel covers the six that bracket its range plus the pair that shares a language. A seventh market could behave differently and we would not know.
One basket is one intent profile. The archetype table shows how much that matters, so a differently balanced basket would produce different absolute rates in every cell. The comparisons should hold; the numbers would move.
And this is a single window. Google changes: our own census recorded US AI Overview presence falling seven points on a fixed query cohort between June and July 2026. A controlled market panel is a snapshot, and the value of committing the basket and the runner is that the snapshot can be retaken rather than argued about.
150 queries, balanced across six archetypes (definitional, informational, service, commercial, product, navigational) at 25 each. The basket is market-neutral by construction: no national institutions, no single-market brands, no currencies or units, because a query that means something different in Tokyo than in Lyon would reintroduce the confound this study exists to remove.
Ten cells: six markets (United States, United Kingdom, France, Germany, Austria, Japan) crossed with interface language, English everywhere plus each market's own where it differs. 150 queries in each cell, 1,500 scored Google SERPs, all collected inside a single 30-minute window so that time is not a variable. Every call succeeded; nothing is imputed.
Both gl and hl are pinned explicitly on every request. Targeting was verified before the run rather than assumed: the same query returns klm.fr in France, idealo.com in Germany and expedia.co.jp in Japan, sharing 2 of 10 domains with the United States. AI Overview presence is the AI Overview object being returned at all, and it is only populated when the request asks for it, which is the failure mode that would otherwise read as 0% in every market.
Rates carry Wilson 95% intervals, about 7.9 points either side at 150 per cell. Differences smaller than roughly ten points are not readable at this sample size and are not reported as findings anywhere on this page.
The basket, the runner and the scoring are committed in the cloro repository as corpus/query-sets/market-panel.v1.json, run_market_panel.py and score_market_panel.py, so the run can be repeated against a later Google rather than taken on trust.
Related studies
The SERP Feature Census is the uncontrolled companion to this page: 65,945 scored SERPs across 41 markets, far broader, with the query mix varying by market. AI Overviews Around the World covers trigger rates and citation depth across six AI engines and carries the same caveat about market comparisons. This page is the narrow, controlled one: one basket, six markets, Google only.
cloro. (August 20, 2026). SERP Features by Country: The Same 150 Questions in Six Markets. cloro Research. https://cloro.dev/research/serp-features-by-country/
More studies from cloro's monitoring corpus are in the research index.
cloro returns AI Overviews, People Also Ask, ads and the organic list as structured JSON, with gl and hl pinned per request. Start with 500 free credits a month.