Cluster 4 · Measurement

4.2 Share of Voice in AI

Share of Voice in AI measures how often your brand appears in AI answers relative to competitors. It is one of the most useful GEO metrics — and one of the most frequently miscalculated. Its value stands or falls with clean definitions.

Two metrics to keep apart

Mention rate is the share of valid answers in which your brand appears: answers containing your brand, divided by all valid answers. That is a coverage figure — it says how often you are present, not how you compare to others.

Share of mentions is only then a real share: your mentions, divided by the sum of mentions of all measured brands. For that you must decide in advance which brands you measure and how you count when one answer contains several brands.

Report both, and say which one you are reporting. Anyone who says “Share of Voice” without a formula does not know what they are reading. And keep a third distinction next to it that neither metric captures: being mentioned is not being recommended. Measure recommendation separately.

Branded vs non-branded prompts

The crucial distinction in the prompt set: do you ask branded questions (containing your brand name) or non-branded questions (general consumer questions)?

Branded prompts — ‘What does KBC offer as a mortgage?’, ‘Is Coolblue a reliable web shop?’ — mainly measure whether the information is correct and how you are presented: the question is already about you.

Non-branded prompts — ‘Which Belgian bank would you recommend for a first-time buyer?’, ‘Best project management tool for small teams’ — here your brand is not handed to the model. This is the category where real GEO performance becomes visible.

The gap between branded and non-branded presence is often large. Brands that dominate on branded prompts can all but disappear on non-branded questions. That difference is where GEO work has real impact.

Comparative prompts form a third category — ‘KBC vs Belfius for first-time buyers’ — and strategically often the most important one: they are decision moments. But mind the measurement error that lies in wait here: that both brands appear in the prompt does not guarantee that both are treated in the answer. A model can ignore one, correct the premise, or frame the comparison differently — and precisely that is worth measuring. Treat comparative prompts as a separate measurement cluster, in which alongside presence you measure positioning above all: who gets the recommendation, who gets the reservations?

Raw count first, depth separately

A tempting mistake: filtering out incidental mentions because they are “not real presence”. Don’t — a mention is a mention, and whoever filters raw figures on a judgement makes the measurement irreproducible. The right form: count everything in the raw mention rate, and classify the depth of each mention separately — from side item in a list to elaborated recommendation. If you want a stricter metric besides, define a qualified mention rate in advance and report both. That keeps visible what the difference between raw and qualified actually is for you, instead of a filter making the difference invisible.

How many prompts do you need?

There is no universal minimum, and distrust anyone who names one. Small samples give wide uncertainty: twenty prompts at a measured presence of fifty per cent leave an uncertainty margin of roughly twenty percentage points either way. What you need depends on how precisely you want to be able to conclude — decide that in advance, and report the uncertainty next to the figure. How Groundbase fixes sample size, repetition and weighting per sector is part of its own measurement methodology.

Not everything on one heap

Adding up answers from different AI channels into one figure gives every channel, language and prompt type an arbitrary weight. Report separately first — per channel, per prompt category, per language — and only then build a composite score, with weights fixed in advance. The insights are almost always in the segmentation, not in the total.

SoV over time

One measurement is a snapshot. Repeat with the same prompt set, monthly or quarterly. Patterns say more than absolute figures — and never read a shift in isolation from the question of whether the channel’s own answer behaviour changed: a declining line can also be display policy rather than a loss of position.

verification
documented
The definitions are standard statistics; no channel publishes its own visibility figures per brand.
own measurement
The gap between branded and non-branded presence is a recurring pattern in our measurements.
not verifiable
Why a specific brand ends up in a specific answer — the measurement counts outcomes, not causes.