For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.
Attensira Logo
Attensira

The GEO measurement guide · chapter 1

What can and cannot be measured in AI search

What can an AI visibility tool actually observe, and what is it guessing at?

Karl-Gustav KallasmaaKarl-Gustav Kallasmaa, Founder & CEOLast updated

The question this chapter answers

An AI visibility tool shows you a number. Before you act on it, you need to know which of two things it is: an observation, or an inference. Observations are recordings of something that happened. Inferences are models of something that did not get recorded. Both can be useful. Only one of them is evidence.

The category blurs the line constantly, usually not out of malice but because the inferred numbers are the ones customers ask for. "How many people saw us in ChatGPT last month" is the question every executive wants answered, and it is unanswerable, so somebody builds an estimate and puts it on a card next to the real numbers in the same typeface. This chapter draws the line explicitly.

What is genuinely observable

What a model said when you asked it. This is the bedrock. You issue a prompt to a named model, from a named country, on a named surface, and you record the answer. It is a real event with a timestamp. Everything defensible in AI visibility measurement is built from a pile of these.

Whether the answer named your brand. Given the answer text and a matching rule you can write down, this is a determinable fact. The matching rule is where the honesty lives — exact string, case handling, whether a possessive or plural counts, whether a partial match counts — and a tool that will not tell you its rule is asking you to trust a number whose definition it is hiding.

Whether the answer linked your domain. Also determinable, and separate from being named. A model can recommend you warmly and link somebody else's review of you, and that is a specific, actionable fact you can only see if the two are counted separately.

Which sources the answer cited. The set of domains an answer linked is recorded in the answer. This is often the most useful observation in the whole category, because it tells you which pages the assistant treats as authoritative in your space — including pages you could plausibly get onto.

Which crawlers fetched which of your pages, and when. This one is not a sample at all. It is your own server log, and it is a complete record rather than an estimate. Its value depends on knowing which user agent does what: OpenAI documents a distinct bot for search and a distinct bot for foundation-model training,[^openai-searchbot] and Perplexity states its search crawler is not used to crawl content for foundation models.[^perplexity-bots] Treating all AI crawlers as one thing throws away the only unambiguous signal you own.

What is not observable, at any price

Impressions. Nobody outside the assistant vendors knows how often an answer containing your brand was shown to a human. There is no public API for it. Google does not break it out: appearances in AI features are folded into overall Search Console search traffic under the Web search type rather than reported separately.[^gsc-combined] So an "AI impressions" number in any third-party product is derived — from query volume estimates, from traffic models, from assumptions. It may be a reasonable estimate. It is not a measurement, and it should never be charted next to measured rates without a visual distinction.

Attribution. No tool in this category can tell you that a mention produced a visit, that a citation produced a signup, or that any of it produced revenue.[^attensira-no-attribution] That sentence is from our own documentation and it applies to every vendor equally. The mechanism is simply absent: a person reads an answer, forms an impression, and arrives at your site three days later by typing your name. There is no identifier that survives that journey. Claims to the contrary are correlations wearing a causal label.

What any individual was told. You measure what the assistant said to your measurement infrastructure, in your configuration, at your time. Real users have different histories, different accounts, different personalisation, different devices and different phrasing. Your sample is drawn from a neighbouring population, not the one you care about. That is a defensible research design — it is what a poll does — provided you say so.

Ranking. There is no position to hold. Assistants do not publish an ordered list of brands, and the order in which names appear in a paragraph is a property of a sentence, not a rank. Any product presenting a "position" is presenting either the order of citations within one answer, which is real but narrow, or a constructed index, which is not a rank at all.

Why you were not mentioned. Absence has many causes that look identical in the data: never crawled, crawled by the wrong bot, crawled and not retrieved, retrieved and not selected, selected and paraphrased without a name. The measurement records the absence, not the reason. Diagnosis requires the second kind of evidence — logs — and even then it narrows the cause rather than determining it.

The awkward middle: things that are observable but not comparable

Some quantities are real observations that still cannot be put on the same axis.

Consumer surfaces and API surfaces. Reading an assistant's consumer app and calling the same vendor's API are two different measurements of two different systems. Coverage differs as well: we document that Google AI has no API reader at all, so what coverage exists comes from consumer reading only and is thinner and less regular than the API-backed channels.[^attensira-surfaces] Averaging across surfaces produces a number describing a system that does not exist.

Different models. A rate aggregated across assistants moves when the mix of runs moves, which happens whenever a model is added, removed or has an outage. The aggregate is legitimate as a summary; it is not legitimate as a time series unless the mix is held constant, and almost nobody holds the mix constant.

Partial days. Today's numbers are incomplete until the day closes, and the incompleteness is not uniform across cells.[^attensira-partial-day] A partial day is not a small full day.

A brand name that is also a word. When the brand name is an everyday word, the mention rate contains noise that averaging will not remove, because the error is systematic rather than random.[^attensira-common-noun] The number is still real; it is measuring something slightly different from what you think.

The distinction that does the most work: zero versus null

Of everything in this chapter, one distinction is worth enforcing in your reporting stack before anything else.

A measured zero means the runs were issued and your brand never appeared. That is a finding, and usually an uncomfortable one: the models know your category and do not know you.

A null means no runs were issued. Perhaps the model was never configured for tracking, in which case it is marked as untracked and no zero applies to it.[^attensira-tracked-false] Perhaps the country was never selected. Perhaps the plan does not cover that surface.

These render identically if you are careless, and the consequences of collapsing them are asymmetric. A null shown as zero manufactures a failure that never happened, and the resulting work — a quarter spent chasing visibility on a surface nobody ever queried — is entirely wasted. Our documentation's instruction on this is blunt: never render an unmeasured cell as 0%, because reporting null as zero invents a failure that was never observed.

Enforce it structurally rather than in a style guide. Make the rate type a pair of a nullable value and a sample size, so that a component physically cannot draw a bar without knowing whether a bar is warranted. Style guides get forgotten; types do not.

A three-question audit of any dashboard

You can classify most of a product's metrics in about ten minutes with three questions, asked of each number on the screen.

Could I have recorded this myself, in principle, with a script? If yes, it is an observation. Mention rate, citation rate, cited domains and crawler hits all pass. Impressions, sentiment scores and attributed revenue all fail, because no script you could write has access to the underlying event.

What is the denominator, and does the interface show it? An observation with a hidden denominator is being presented as more certain than it is. A number that has no denominator at all — a score, an index, a grade — is a construction, and the question becomes whether its construction rule is published.

What does this render when the answer is unknown? Ask the product to show you a model it does not track. If it draws a zero, everything else it shows you is suspect, because the same carelessness will be somewhere you cannot see it.

Numbers that survive all three go in the report. Numbers that fail the first can still be useful, but they belong in a clearly separated section with the word "estimated" attached, not interleaved with measurements.

What to do with an unobservable question

Executives will keep asking for impressions and attribution. The answer is not to refuse and leave them with nothing; it is to reframe to the nearest observable question and be explicit about the substitution.

Each row trades a satisfying answer for a defensible one. That trade is the entire content of this chapter, and taking it consistently is what makes the rest of your reporting worth reading. The next chapter, sampling and sample size, deals with the observable half — and with how much of it you need before a number earns a place on a chart.

Questions people ask

Can I see how many people were shown an answer that mentioned my brand?
No. Assistants do not publish impression counts for brands named in answers, and Google folds appearances in AI features into overall Search Console web search traffic rather than reporting them as a separate channel. Any "AI impressions" figure you are shown is a model of a number, not the number, and you should ask what inputs it was built from before quoting it to anyone.
Is a zero mention rate the same as being invisible?
No. It means you were not named in the runs that were issued, for the prompts in your set, on the models and countries you configured, in that window. Change any of those four and the number can change without anything about your brand changing. A zero is a real observation about a narrow question, not a verdict on your presence in AI generally.
Why do tools not report sentiment about my brand in AI answers?
Some do. We do not, because sentiment is a model judging a model, and the second model's errors are not random — they correlate with phrasing, language and category in ways that bias the aggregate rather than averaging out. A number you cannot reproduce from a stated rule is a number you cannot defend in a meeting, and it is worse than no number when it is wrong in a consistent direction.
My brand name is a common word. Does that break the measurement?
It degrades it, and the degradation does not average away. A brand called "Notion" or "Arc" will be matched in answers that were never about the company, producing a mention rate that is partly noise from a systematic error rather than a random one. The workarounds are all partial — requiring a domain match, requiring the name near category terms — and each trades a false positive rate for a false negative rate. Know which trade your tool made.
Does a crawler visit prove I will appear in answers?
No, and the two are not even the same pipeline. Operators run separate user agents for separate jobs — OpenAI documents one bot for search and a different one for training foundation models, and Perplexity states its search crawler is not used to crawl for foundation models. A fetch by the search crawler is a necessary precondition for being surfaced, not a guarantee of it; Google says outright that compliance does not mean it will crawl, index or serve a page.
Why is today's data incomplete?
Because runs are issued across the day and the day is not over. The incompleteness is also uneven rather than uniform, so a partial day is not a smaller version of a full day and should not be compared to one. Read closed windows, and treat any chart whose last point is today as having one point you cannot use.

Sources

Every factual statement above, with the page it came from and the date that page was read.

  1. Attensira's limitations page states the product cannot tell you that a mention produced a visit, that a citation produced a signup, or that any of it produced revenue.

    docs.attensira.com · retrieved

    Attensira cannot tell you that a mention produced a visit, that a citation produced a signup, or that any of it produced revenue.
  2. Attensira states there is no visibility score and no sentiment metric in the product.

    docs.attensira.com · retrieved

    There is no visibility score and no sentiment metric in Attensira.
  3. Attensira documents that a brand name which is an everyday word produces a mention rate partly composed of noise that averaging cannot remove, because the error is systematic rather than random.

    docs.attensira.com · retrieved

    no amount of averaging removes it, because the error is systematic rather than random
  4. Attensira documents that the current day's numbers are incomplete until the day closes, and that the shape of that incompleteness is not uniform.

    docs.attensira.com · retrieved

    Today's numbers are therefore incomplete until the day closes, and the shape of that incompleteness is not uniform.
  5. Attensira marks a model that was never configured for tracking as tracked false, meaning it was never queried and no zero applies to it.

    docs.attensira.com · retrieved

  6. Google documents that sites appearing in AI features such as AI Overviews and AI Mode are included in overall Search Console search traffic under the Web search type.

    developers.google.com · retrieved

  7. Google documents that meeting every requirement, best practice and policy does not mean Google will crawl, index or serve a page's content.

    developers.google.com · retrieved

    Just because a page meets all requirements, best practices, and complies with the policies, doesn't mean that Google will crawl, index, or serve its content.
  8. OpenAI documents OAI-SearchBot as the user agent for search, distinct from GPTBot, which it documents as used to make its generative AI foundation models more useful and safe.

    developers.openai.com · retrieved

    OAI-SearchBot is for search.
  9. Perplexity documents PerplexityBot as designed to surface and link websites in search results on Perplexity and states it is not used to crawl content for AI foundation models.

    docs.perplexity.ai · retrieved

    It is not used to crawl content for AI foundation models.
  10. Attensira documents that Google AI has no API reader and that coverage for it comes from consumer reading only, with thinner and less regular data than API-backed channels.

    docs.attensira.com · retrieved

    you should expect thinner, less regular data there than on the API-backed channels