Visibilité dans les LLM

Measuring AI Visibility: What Tools Won't Tell You

GEO measurement is possible, but never exact. Discover the hidden biases in measurement tools and Universem’s method for reliably tracking AI visibility.

GEO · AI Visibility

Measuring AI visibility: what the tools won’t tell you

Corentin DonneauxCorentin Donneaux·JUNE 26, 2026·9 MIN

Can you really know if a brand gets cited by ChatGPT, Gemini or Perplexity? Yes — but not the way you do in SEO. Between probabilistic answers, invisible biases and overly flattering scores, here is what the measurement tools won’t tell you, and the method to measure anyway

Can you really measure a brand’s visibility in AI?

You can measure your visibility in AI, but never the way you do in SEO, and never exactly.

For twenty years, search sold us a comfortable certainty: a keyword, a position, clean tracking over time. Same query, same conditions, same result. We could sleep easy.

Except generative AI never read the manual. Ask ChatGPT your question. Ask it again. And again. Three answers, three different results. Not a bug: an AI is probabilistic where a search engine stays deterministic.

A brand that shows up once might just be luck. The same brand showing up six times out of ten is a signal. That’s the entire shift.

We no longer count a position, we count a frequency.

Definition

GEO — Generative Engine Optimization
A discipline that optimizes and measures a brand’s presence in AI-generated answers (ChatGPT, Gemini, Perplexity…), where SEO optimizes ranking in a search engine’s links.

GEO is the art of making serious decisions on data nobody can prove exists. Understanding this uncertainty is the bare minimum before trusting any AI visibility report.

Why the same question never gives the same answer

In short

An AI generates its answer word by word, drawing from a probability distribution. The answer you read isn’t the answer: it’s an answer among dozens.

An AI generates its answer word by word, picking the most likely term at each step. It’s essentially auto-complete on steroids: it doesn’t recite a stored answer, it recomposes one every time.

Asked “who has the best car insurance?”, the model doesn’t see a single winner. It sees a probability distribution: 27% chance of starting with AG, 22% with Ethias, 19% with AXA, and so on. Every generation is a draw from that distribution.

Given the same question, the model doesn't pick a winner
Given the same question, the model doesn’t pick a winner: it draws from a probability distribution.

Hence the conclusion that throws everyone off: the answer you read isn’t the answer. It’s an answer, one among dozens of possibilities. Measuring an AI on a single pass is like judging a die on a single roll.

How do you know what people actually ask AI?

In short

Nobody knows the real prompts people type: it’s a black box. We don’t guess them, we reconstruct them from real data.

Nobody knows the real prompts people type. In SEO, Google Ads gave us the volume. On the AI side, that number doesn’t exist: what people actually type into an AI remains a black box.

No data doesn’t mean doing nothing. At Universem, we generate the prompts, properly, in four steps:

The Universem method for credible prompts

  1. 1

    Start from keywords

    A keyword always hides a need: the anchor point in real data.

  2. 2

    Layer in personas and People Also Ask

    Google already turns queries into questions; we cross-reference them with target profiles.

  3. 3

    Scrape Reddit when the project allows it

    To capture real, unfiltered questions from actual people.

  4. 4

    Run it all through a trained model

    The result is credible prompts, grounded in real data. Imperfect, we know, but far better than a prompt made up on the spot.

Discover our full GEO approach

What the measurement tools don’t see

The moment a tool scrapes an answer, it stacks a series of invisible layers, and each one distorts the result a little more. Most tools never tell you about this. Three are worth stopping for.

1. The model and its version change the results

ChatGPT, Gemini, Claude, Perplexity, Copilot: each has carved out its own markets. Measuring “the AI” without specifying which model makes no sense at all. And it gets worse within a single model: a GPT-4o and a 5.5 give genuinely different answers.

2. Logged in, logged out, API: the simulated user isn’t the real one

Most tools simulate a logged-out user. But a model doesn’t answer the same way to a logged-in account as it does to an anonymous visitor, or to a free user versus a paid one. Add geolocation, and local tracking on a poorly localized model becomes meaningless.

11 people, one prompt: depending on device and status
11 people, one prompt: depending on the device and whether the user is logged in, the same brand gets mentioned or disappears.

3. History, personalization, grounding and conversation

History personalizes answers: my Claude knows my pipelines run on BigQuery. Then there’s grounding: the model runs its own web searches, which vary each time. And finally the conversation itself: we rephrase, we correct.

Tools measure a one-shot. People, on the other hand, have a conversation.

Mentions, citations, share of voice: what these metrics really tell you

In short

A score on its own means nothing. The two metrics that matter are the mention (the brand is named) and the citation (your source is referenced).

An AI visibility score, on its own, means nothing. GEO arrives with its jargon, and a nice 82 out of 100 is enough to reassure a leadership team. Except that number, as long as nobody knows how it’s calculated, is decorative at best.

The real foundation has two parts: mentions and citations. Mentions carry the most weight; but when no brand is named, your source citation becomes the only field to play on. And when both are present, the mention takes priority.

Criterion Mention Citation
Definition The brand is named Your source is cited as a reference
Impact Strongest: the brand gets verified The only field to play on when no brand is named
Priority Takes priority when both are present Decisive on prompts with no named brand

Definition

Share of mentions
The number of mentions of your domain divided by the total mentions in a given category. The danger isn’t the formula, it’s how it’s presented.
Share of mentions, broken down by category
Share of mentions, broken down by category: the same domain can dominate one segment and disappear from another.

Measuring despite the bias: reach and stability

In short

No data is worse than partial data. The method turns noise into signal through two levers: reach and stability.

So do we give up? No. We measure, but with a method that turns noise into signal, and it rests on two levers.

First, reach: at least ten prompts per intent. Then stability: each prompt run three to ten times. A single run gives a randomly drawn yes/no; ten runs give a clean rate. Reach times stability, always, before drawing any conclusion.

≥ 10

Prompts per intent

3–10×

Runs per prompt

1,000

Prompts to surface the signal

The two anti-noise levers: covering the question (reach)
The two anti-noise levers: covering the question (reach) and making each measurement reliable (stability).

Two non-negotiable principles

  1. 1

    Honesty about the limits

    Everything you just read gets told to the client; transparency about uncertainty is part of the deliverable, not the fine print.

  2. 2

    Transparency on the method

    The client is included in generating the prompts, and every measurement is documented prompt by prompt, with dates.

This isn’t an argument for giving up. It’s an argument for clarity.

Universem · SEO & GEO

Google is gradually opening up Search Console to GEO, ads are arriving on ChatGPT: the ecosystem is going to mature. In the meantime, we’re moving forward half-blind. But we’re moving forward, curious and critical.

Frequently asked questions

Can you measure your AI visibility?+

Yes, but in frequency, never exactly: an AI is probabilistic. You count how many times a brand appears across N runs, not a fixed position.

Why doesn’t ChatGPT give the same answer twice?+

Because it generates its answer word by word, drawing from a probability distribution. The answer you read is one possibility among dozens.

How many prompts do you need for a reliable measurement?+

At least 10 prompts per intent (reach), each run 3 to 10 times (stability). Below that, you’re mostly measuring noise.