For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.
Attensira Logo
Attensira

Guide

The GEO measurement guide

How to measure AI search visibility honestly — what the numbers can prove, what they cannot, and how to tell a defensible report from a flattering one.

Karl-Gustav KallasmaaKarl-Gustav Kallasmaa, Founder & CEOLast updated

What this guide answers

Each of these stands on its own, without the paragraph before it.

  • An AI visibility number is a sample from a non-deterministic system, so it is meaningless without the sample size, the prompt set, the models and the window that produced it.
  • A cell nobody measured must read "not measured", never zero — reporting an unqueried model as zero invents a failure that was never observed.
  • Mention rate, citation rate and share of voice answer three different questions, and only one of them is a share of anything; mention rates across brands do not sum to one.
  • The unit of measurement is the run, not the answer's enthusiasm: an answer that names you nine times counts once, exactly like an answer that names you once.
  • Nobody in this category can attribute a model mention to a visit, a signup or revenue, and any vendor claiming otherwise is selling you a correlation with a causal label on it.

Chapters

Why this guide exists

Every company selling AI search visibility, including us, is selling a number. The number is easy to produce and very hard to produce honestly. It comes out of a system that does not give the same answer twice, sampled at whatever depth the vendor could afford, over a prompt set the customer picked in an afternoon, and it is then rendered as a percentage on a dashboard with a green arrow next to it.

That gap — between how a visibility number is made and how it is presented — is where the category's credibility is being spent. This guide is an attempt to spend less of it. It sets out what can be measured about AI search visibility, what cannot, how to size a sample so the result means something, how the three metrics people conflate actually differ, and what a report has to contain before you should believe it.

It is written to be useful whether or not you ever use our product. Where our own methodology is the example, it is because it is the one we can document line by line, not because the advice depends on it. Where our product cannot do something, this guide says so; the limitations page is the shortest version of that list and it is the page we would rather you read first.

The premise everything else rests on

An AI assistant is not an index you can query for your rank. It is a model that produces text, and the text is different every time.

This is not a quirk to be engineered away. OpenAI documents it as the default behaviour of its API: chat completions are non-deterministic, and model outputs may differ from request to request.[^openai-nondeterministic] Even when you fix the seed and hold every other parameter constant, the documentation warns that determinism may still be impacted by configuration changes made on OpenAI's side.[^openai-seed-caveat] Google is equally explicit about its own surfaces: AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.[^google-varies] And meeting every documented requirement does not entitle a page to appear at all — Google states plainly that compliance does not mean it will crawl, index or serve your content.[^google-no-guarantee]

Take those three sentences seriously and most of the category's marketing collapses. There is no rank to check. There is no position you hold. There is a probability distribution over answers, and every number you will ever see about AI visibility is an estimate of some property of that distribution, computed from a finite sample.

That is not a reason to give up on measurement. It is the reason measurement has to be done properly. Epidemiology, polling and quality control all estimate properties of noisy populations from finite samples, and they do it credibly by being disciplined about four things: the sampling frame, the sample size, the definition of the event being counted, and the treatment of missing data. Those four things are the spine of this guide.

What the five chapters cover

[What can and cannot be measured](/guides/geo-measurement-guide/what-can-and-cannot-be-measured) separates the observable from the unobservable. You can observe what a model said when asked. You cannot observe what any real person was told, how often the question was asked, whether the answer was believed, or whether it caused anything. The chapter draws the line and explains why several popular metrics sit on the wrong side of it — including sentiment, which we do not offer, and attribution, which nobody can offer honestly.[^attensira-no-sentiment][^attensira-no-attribution]

[Sampling and sample size](/guides/geo-measurement-guide/sampling-and-sample-size) is the arithmetic. One run is an anecdote. The chapter works through why, what a confidence interval on a rate actually looks like at small sample sizes,[^nist-wilson] and how to decide how many runs you need before a movement is worth reacting to. It also explains the difference between sampling variance inside the model and drift in the world, and why running your samples concurrently rather than spread across a day is what separates the two.[^attensira-sampling-depth]

[Mention rate, citation rate and share of voice](/guides/geo-measurement-guide/mention-citation-and-share-of-voice) takes apart the three numbers that get used interchangeably and are not interchangeable. Being named is not being linked. Neither is a share of anything: mention rates across a competitive set do not sum to one, because an answer can name five brands or none.[^attensira-mention-not-pie] The chapter also covers the unit-of-observation choice that quietly decides what your metric means — we count runs, so an answer naming you nine times counts once.[^attensira-run-unit]

[Building a prompt set that means something](/guides/geo-measurement-guide/building-a-prompt-set) is the sampling-frame chapter, and it is the one most teams skip. Your prompt set is your population definition. Choose prompts you already win and you have built an instrument that cannot detect failure. The chapter covers question shape, branded versus unbranded prompts, coverage of the buying journey, and how to change a prompt set over time without destroying your ability to compare windows.

[What a defensible report looks like](/guides/geo-measurement-guide/a-defensible-report) assembles the rest into an artefact you can put in front of a sceptical executive: the fields it must carry, the claims it must not make, and a catalogue of the specific ways these numbers are used to mislead — several of which are available to us and which we have therefore designed against.

Five rules that survive being quoted alone

If you take nothing else from the guide, take these.

1. A rate without its denominator is not a measurement. A number reported as a bare percentage hides whether it came from three observations or three hundred. Those two claims deserve wildly different confidence, and any reporting format that renders them identically is destroying the information you most need. Report the pair — the rate and the sample size behind it — everywhere, including in screenshots.

2. Not measured is not zero. If a model was never queried, the correct value is null, rendered as "not measured". Rendering it as zero invents a failure that was never observed.[^attensira-null-vs-zero] This is where plan limits become data-quality problems: a customer on a plan that does not query a given surface should see a coverage gap, not an apparent collapse in visibility on that surface.

3. Only report movement that clears a significance test. Two rates from small samples will differ. That is what small samples do. Before a change goes on a chart with an arrow, it has to clear a test — we use a two-proportion z-test at 95% confidence between the two windows, and when it does not clear, the delta is returned as not real rather than as a number.[^attensira-ztest] "Nothing we can prove" is a legitimate and common answer.

4. State the surface, the model, the country and the window, or the number means nothing. These are not metadata. They are part of the measurement. A mention rate on a consumer app and a mention rate through the same vendor's API are two different quantities, and averaging them produces a third quantity that describes nothing.

5. Never attribute. No tool in this category can connect a model mention to a visit, a signup or revenue.[^attensira-no-attribution] Google's own reporting folds AI feature appearances into overall search traffic rather than separating them out,[^gsc-combined] and assistants that answer without sending a click leave nothing to attribute at all. Treat visibility as a leading indicator, hold it next to your funnel, and label the relationship correlation, because that is what it is.

The measurement stack, top to bottom

It helps to see where each decision sits, because most arguments about AI visibility numbers are actually arguments about different layers.

Every layer above the one you are arguing about determines whether the argument is worth having. A dispute about whether a rate moved from one window to the next is unresolvable if the prompt set changed underneath it, and no amount of statistical sophistication at the estimate layer repairs a sampling frame that was chosen for its optics.

Three axes of variation, and what each one costs you

Answers differ along three axes at once, and conflating them is the most common analytical error in this category.

By model. Different assistants are different systems with different retrieval behaviour, different training data and different linking habits. Google says as much about its own two surfaces: AI Mode and AI Overviews may use different models and techniques, so the responses and links they show will vary.[^google-varies] Across vendors the gap is larger still. A brand can be well established in one assistant's answers and absent from another's for reasons that have nothing to do with the brand. Aggregating across models produces a number whose movement you cannot interpret, because you cannot tell which surface moved.

By country. The same question in the same assistant, asked from a different locale, is a different query — different retrieval, different local sources, sometimes a different language. Country is therefore part of the cell definition, not a filter applied afterwards. This is also where cost multiplies fastest: every country you add multiplies your daily run count by the size of your prompt set times your model count, which is why disciplined programmes track fewer countries deeply rather than many shallowly.

By day, and within a day. Two things move here and they are not the same thing. Sampling variance is the model giving a different answer to an identical request, which OpenAI documents as the default behaviour of its API.[^openai-nondeterministic] Drift is the world changing underneath the question — a new page indexed, a competitor's launch, a model update. The way to tell them apart is to control when the samples are taken. Runs issued concurrently within one cell measure variance in the model; runs spread across days measure variance plus drift, mixed, with no way to separate them afterwards.[^attensira-sampling-depth] If you only ever sample once a day, every number you have is variance and drift added together, and you should stop describing week-to-week wiggles as trends.

There is a fourth axis worth naming even though it is rarely reported: the surface. A vendor's consumer app and its API are different products with different retrieval stacks, and answers from the two should never be averaged into one figure. When they are, the resulting rate describes a system that does not exist.

Two kinds of evidence: the survey and the census

Almost everything in AI visibility measurement is a survey — you ask a sample of questions and generalise. There is one exception worth building a programme around, and it is your own server logs.

When an AI crawler fetches a page, your server records it. That is not a sample; it is a complete record of every request that reached you, with timestamps, paths and user agents. It answers a narrow question exactly, where sampling answers a broad question approximately. The narrow question — did the crawler that feeds this assistant actually retrieve this page, and when — happens to be the prerequisite for most of the broad ones.

The two kinds of evidence fail differently, which is what makes them worth holding together. Logs cannot tell you whether a fetched page was ever used in an answer; a user agent string can be forged, and a fetch is not a citation. Prompt sampling cannot tell you why you are absent; a zero mention rate is compatible with never being crawled, being crawled and not retrieved, or being retrieved and not selected. Put the census next to the survey and the diagnosis narrows: absent from answers and absent from logs is an access problem, absent from answers while present in logs is a content or authority problem, and those two conclusions lead to completely different work.

The practical rule: before you spend a quarter improving content for an assistant, check that the assistant's crawler has actually been to the pages you are improving. Our crawler logs feature exists for that check specifically, and any log pipeline you already run will answer the same question if you keep the user agents.

A worked hypothetical, clearly labelled

The following numbers are invented for illustration. They are not measurements of anything, and no figure in this section describes a real brand, product or study.

Suppose a fictional company tracks 40 prompts across two model surfaces in one country. On a plan issuing one run per prompt, per model, per day, a week of data is 40 prompts, times two surfaces, times seven days: 560 runs. Suppose the brand is named in 84 of them. The point estimate is 0.15.

The point estimate is the least interesting number here. Sliced by model, each surface carries 280 runs; sliced by model and by week, the smallest cell a dashboard will happily render is 40 runs. At 40 observations, a rate near 0.15 carries an interval roughly from 0.07 to 0.29 by the Wilson method the NIST handbook recommends.[^nist-wilson] So a week-over-week move from 0.15 to 0.20 in that cell is entirely consistent with nothing having happened. If it is drawn as a rising line without an interval, the chart is asserting something the data cannot support.

Now suppose the same company adds a third model surface mid-week. The naive aggregate rate will move, because the denominator changed and the new surface behaves differently. Nothing about the brand changed. This is why the window, the frame and the prompt set have to be stated with the number, and why a report that silently changes any of them has broken its own time series.

What we do not claim

Being straight about the limits of this category is more useful than another percentage, so here is our own list, in public.

We cannot tell you a mention produced a visit, a citation produced a signup, or that any of it produced revenue.[^attensira-no-attribution] We have no visibility score and no sentiment metric, deliberately — both are composites that cannot be falsified from published inputs.[^attensira-no-sentiment] We do not know what any individual person was told by an assistant; we know what the assistant said to us, when we asked, in the configuration we asked from. And on plans that issue a single daily run per cell, most week-to-week movements will not clear a significance test, which is the test working rather than the product failing.

Those limits are not a reason to skip measurement. Traditional search analytics sits on comparable foundations — sampled, aggregated, and increasingly consolidated by the platforms themselves — and it is still the most useful instrument most marketing teams own. The point is to know which claims your instrument supports.

The vocabulary this guide uses precisely

These words get used loosely elsewhere. Here they mean one thing each, and the chapters depend on the distinctions.

Run. One issue of one prompt to one model in one country, producing one answer. The run is the unit of observation throughout: it is what appears in a numerator and a denominator, and its outcome is binary. An answer that names you nine times contributes exactly one run to the numerator, the same as an answer that names you once.[^attensira-run-unit] Choosing any other unit — mentions per answer, words about you, position in the paragraph — turns a countable event into a judgement call.

Cell. One combination of prompt, model and country. Cells are where sample sizes actually live. A dashboard's headline number is an aggregate over cells, and the moment a reader filters, they are looking at a cell whose sample size is a fraction of the headline's. Reports that show the denominator at the headline and hide it after filtering are the reason people over-read filtered views.

Depth. How many runs are issued per cell per period. Depth is the lever that buys statistical power, and it is expensive in a way breadth is not: doubling depth doubles cost for the same coverage, whereas doubling the prompt set doubles cost and coverage together. Deciding between them is a real trade-off with no universally correct answer, and it should be made explicitly rather than inherited from a plan default.

Window. The period a rate is computed over. Longer windows have larger denominators and therefore tighter intervals, and they respond more slowly to real change. There is no neutral window length; a seven-day window and a twenty-eight-day window are different instruments, and switching between them mid-report is a way to manufacture a movement.

Successful run. A run that produced an answer. Failed runs — timeouts, outages, blocked requests — belong in a debugging log, not in a denominator. If they enter the denominator, a provider outage reads as a drop in your visibility, which is a fabricated finding with a mundane cause.

How to use this guide

Read it in order if you are designing a measurement programme from scratch. If you already have a dashboard and want to know whether to trust it, start with what a defensible report looks like and work backwards through whichever chapter it sends you to.

If you are evaluating vendors, the fastest test is to ask three questions and watch what happens: what is the sample size behind this number, what does an unmeasured cell render as, and what is the significance test before a change is shown as a change. A vendor who can answer all three in writing has done the work. A vendor who answers with a composite score has not.

Related reading on this site: the AI search visibility guide covers the other half of the problem — how pages come to be found and cited in the first place — and the crawler logs feature page covers the one signal in this category that is a direct server-side observation rather than a sample.

Sources

Every factual claim above is drawn from a source fetched on 3 September 2026 and listed in this page's provenance record: OpenAI's advanced usage documentation, Google Search Central's guidance on AI features, the NIST/SEMATECH e-Handbook of Statistical Methods, and Attensira's own measurement documentation. Where a figure appears without such a source, it is a labelled hypothetical and is marked as one in the text.

Questions people ask

Can you actually measure AI search visibility, or is it guesswork?
You can measure it, but only as a sample. You choose a set of prompts, ask them of named models on a stated cadence, and count how often your brand is named and how often your domain is linked. That is a real measurement of a real population — the answers those models gave to those prompts in that window. It is not a measurement of "how visible you are in AI", because no such population exists to sample from. The distinction matters: the first claim is defensible, the second is not.
How many times do I need to run a prompt before the number means anything?
More than once, and the honest answer depends on how big a change you want to detect. A single run tells you what one answer said on one day. Detecting a modest shift in a rate takes a sample in the low hundreds of runs — the NIST/SEMATECH handbook's worked example needs about 102 observations to detect a change from 0.10 to 0.20 with conventional error rates, and that is a doubling. Smaller effects need much larger samples. Chapter two works through the arithmetic.
Why does my tool show a different number to my competitor's tool?
Because they are measuring different things. Different prompt sets, different models, different countries, different sampling depth, different definitions of "mention", and different windows all produce different numbers from the same underlying reality. None of them is wrong; they are answers to different questions. This is why a number without its methodology is not comparable to anything, including its own value last month.
If a model was never queried, should the report show zero?
No. Zero means "we asked and you were never named". Not measured means "we did not ask". Collapsing the second into the first invents a failure that was never observed, and it is the single most common way an AI visibility report misleads — usually by making a plan's coverage gap look like a brand's problem.
Can any tool prove that an AI mention produced revenue?
No, and be suspicious of anyone who says otherwise. Assistants send traffic without a reliable referrer, Google folds AI feature appearances into overall Search Console web traffic rather than breaking them out, and a person who reads an answer and types your name into a browser three days later leaves no trace linking the two. What you can honestly do is track mention and citation rates as a leading indicator and hold them next to your funnel — as correlation, labelled as correlation.
Is a visibility score a useful number?
Usually not. A single composite score blends measured and unmeasured cells, different sample sizes and arbitrary weights into one figure that cannot be falsified or reproduced. If a score cannot be recomputed from published inputs, it is a ranking dressed as a measurement. Prefer a small number of defined rates, each with its denominator visible.

Sources

Every factual statement above, with the page it came from and the date that page was read.

  1. OpenAI documents that chat completions are non-deterministic by default and that model outputs may differ from request to request.

    developers.openai.com · retrieved

    Chat Completions are non-deterministic by default (which means model outputs may differ from request to request).
  2. OpenAI documents that even with a fixed seed, determinism may be impacted by changes it makes to model configurations on its own side.

    developers.openai.com · retrieved

    Sometimes, determinism may be impacted due to necessary changes OpenAI makes to model configurations on our end.
  3. Google documents that AI Mode and AI Overviews may use different models and techniques, so the responses and links they show will vary.

    developers.google.com · retrieved

    AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.
  4. Google documents that meeting every requirement and policy does not mean Google will crawl, index or serve a page.

    developers.google.com · retrieved

    Just because a page meets all requirements, best practices, and complies with the policies, doesn't mean that Google will crawl, index, or serve its content.
  5. Google documents that appearances in AI features are folded into overall Search Console search traffic under the Web search type, rather than reported as a separate channel.

    developers.google.com · retrieved

  6. Attensira's documentation instructs consumers never to render an unmeasured cell as zero, because reporting null as zero invents a failure that was never observed.

    docs.attensira.com · retrieved

    Never render an unmeasured cell as 0%. Reporting `null` as zero invents a failure that was never observed.
  7. Attensira counts one run as one observation regardless of how many times an answer names the brand.

    docs.attensira.com · retrieved

    An answer that names you nine times is one run in the numerator, exactly like an answer that names you once.
  8. Attensira reports a change between windows only when it clears a two-proportion z-test at 95% confidence, and otherwise returns the delta as not real.

    docs.attensira.com · retrieved

  9. Attensira documents that mention rate does not sum to 100% across brands, because several brands can appear in one answer and none need appear at all.

    docs.attensira.com · retrieved

    treating it as a slice of a pie will lead you to wrong conclusions
  10. Attensira's own limitations page states the product cannot tell you that a mention produced a visit, that a citation produced a signup, or that any of it produced revenue.

    docs.attensira.com · retrieved

    Attensira cannot tell you that a mention produced a visit, that a citation produced a signup, or that any of it produced revenue.
  11. Attensira states there is no visibility score and no sentiment metric in the product.

    docs.attensira.com · retrieved

    There is no visibility score and no sentiment metric in Attensira.
  12. Attensira defines sampling depth as the number of runs issued per prompt, model and country per day, set by plan, with one run per day on Starter and three concurrent runs on Growth and Business.

    docs.attensira.com · retrieved · changes often, check the source

  13. The NIST/SEMATECH e-Handbook gives the Wilson score interval for a proportion and notes Agresti and Coull recommended it for virtually all combinations of n and p.

    itl.nist.gov · retrieved

    its worth does not strongly depend upon the value of n and/or p