The GEO measurement guide · chapter 4
Building a prompt set that means something
How do I choose the questions to track so the resulting numbers actually mean something?
Karl-Gustav Kallasmaa, Founder & CEOLast updated Your prompt set is your sampling frame
In survey research, the sampling frame is the list of things you could possibly sample from. Get it wrong and nothing downstream can save you: a perfectly executed poll of a badly chosen population produces a precise measurement of the wrong thing.
The prompt set is the sampling frame for AI visibility. Every rate you compute is a rate over these questions. Every trend is a trend in how models answer these questions. Every comparison to a competitor is a comparison on these questions. When somebody says "our AI visibility is 0.22", the sentence is incomplete without the phrase "across this set of questions", and the honest version of the sentence names the set.
This makes prompt selection the highest-leverage decision in the whole programme — and the one most often made in twenty minutes by whoever had the dashboard open. It deserves better, because it is the only decision that cannot be fixed retroactively. Sampling depth can be increased next quarter. A prompt set chosen to flatter has already contaminated every number you have.
The failure mode, named plainly
The natural way to pick prompts is to think of questions where you would expect to do well. This is not dishonesty; it is how memory works. You recall the category terms you own, the comparisons you win, the phrasing your own marketing uses.
The result is an instrument that cannot detect failure. It will report high rates forever, it will move only with noise, and the first time it tells you something uncomfortable will be after a competitor has already taken the questions you never tracked.
Two counterweights work.
Source the questions from outside the marketing team. Sales call recordings, support tickets, the free-text field in your onboarding form, community threads, and the questions prospects ask in the first ten minutes of a demo. These are questions people actually asked, phrased the way they actually asked them, which is the definition you want.
Require losses in the set. Before finalising, name the prompts you expect to lose. If there are none, the set is wrong. A useful rule of thumb is that a healthy set contains a meaningful minority of questions where a competitor is the obvious answer today — those are the ones that will tell you whether anything you do is working.
What a good prompt looks like
The mechanics matter more than they sound like they should.
Write it as a complete question, in the words buyers use rather than your own category language.[^attensira-prompt-natural] "What's the best expense tool for a 20-person agency?" is a prompt. "expense tool agency" is a keyword fragment, and keyword-shaped input produces a thin, unnatural answer that tells you nothing about how buyers are actually being advised.[^attensira-prompt-keyword] This is the single most common defect in imported prompt sets, because most teams build theirs from an existing keyword export.
Keep it unbranded unless you are deliberately measuring branded questions. An unbranded prompt is the honest test of whether a model reaches for you unprompted.[^attensira-prompt-unbranded] Branded prompts measure something different and legitimate — whether the model can describe you correctly when told who you are — but the two must be grouped separately, because averaging them produces a rate that is high for reasons you cannot decompose.
Carry the qualifiers a real buyer would carry. Company size, industry, constraint, budget, integration. "Best CRM" is a question nobody asks; "best CRM for a two-person startup that needs to sync with Gmail" is. Qualifiers also make the answer set narrower and the measurement more sensitive, because a generic question invites a generic list that names everyone.
Avoid ambiguity around your own name. If the brand name is an everyday word, the mention rate contains systematic noise that averaging will not remove.[^attensira-common-noun] Prompts that make the ambiguity worse — questions in the semantic neighbourhood of the word itself — should be dropped or grouped apart, because they will pollute the aggregate in one direction.
Covering the journey, not just the bottom of it
A set consisting entirely of "best X for Y" questions measures one moment in the buying process. It is the most commercially interesting moment, and it is also the one where the competition is fiercest and movement is slowest.
A set worth reporting on covers at least four shapes.
Problem questions. "How do I stop my invoices going unpaid?" The buyer does not know the category exists. Being named here is early influence, and these are often the easiest questions to enter.
Category questions. "What is expense management software?" Definitional, high volume, and usually answered from general knowledge rather than from retrieval, which makes them a poor place to spend content effort but a useful baseline.
Comparison and selection questions. "Best expense tool for a 20-person agency." The commercial core. Slow-moving and worth every run you can spend on it.
Evaluation questions. "Is [category leader] worth it for a small team?" and "what are the alternatives to [competitor]?" These are where a challenger brand most often first appears in an answer, and a set with none of them is blind to its own best opportunity.
Group them explicitly, because the four shapes have different base rates and different volatility. Blending them into one aggregate produces a number whose movement depends on which group happened to move, and a report that cannot say which group moved cannot recommend anything.
Sizing the set against the budget
Prompts are not free, and the arithmetic is multiplicative rather than additive. Each combination of prompt, model and country is a cell, and each cell is measured at most once per calendar day.[^attensira-daily-cap] Every model you add multiplies your prompt count by one more daily reading per prompt, per country.[^attensira-model-multiplier] Where prompts are sold as simultaneous slots rather than monthly volume — 50 on the entry tier, 150 and 350 on the higher ones, with a slot returning to the pool the moment a prompt is removed[^attensira-prompt-slots] — the budget is a standing allocation you can rebalance rather than a quota you burn.
That structure has a practical implication most teams miss: a prompt you are not learning from is costing you a prompt you could learn from. Prune. A prompt that has sat at a stable rate for two quarters and never informed a decision is occupying a slot that a question from last week's sales calls would use better.
The trade-off against depth is the one from the sampling chapter, and it resolves differently depending on what you are asking. For a headline trend, breadth wins, because the aggregate denominator is the sum across cells. For "did this project work", depth on the affected cells wins, because a before-and-after on a thin cell will never clear a significance test.
Changing the set without destroying your history
Prompt sets have to change. Products move, markets move, and the questions people ask in a category shift quickly. But every change breaks comparability, because the denominator on either side of the change is composed of different questions.
Treat it as a versioned migration.
Version and date the set. Record which prompts were live in which window. Without this, a year-old chart is uninterpretable and, worse, interpretable incorrectly.
Prefer additions to replacements. Adding a prompt changes the aggregate but leaves every existing cell's series intact. Replacing one destroys a series and starts a new one wearing its name.
Report the stable core separately. Maintain a subset of prompts you commit to never changing, and publish the trend on that subset alongside the full-set number. The stable core is what makes a multi-quarter claim possible at all.
Annotate the chart. A prompt-set change is an event, and it belongs on the timeline next to model updates and site launches. Anyone reading a movement across that line needs to know the instrument changed.
One more source of silent change is worth naming: the models themselves. Google documents that AI Mode and AI Overviews may use different models and techniques, so the responses and links they show will vary.[^google-varies] Your instrument can be perfectly stable while the thing it measures is replaced underneath you. That is not a fixable problem — it is a property of the domain — but it is a reason to annotate what you know about and to distrust any trend claim that spans a known model release without acknowledging it.
A checklist before you commit the set
- Every prompt is a complete question in a buyer's words, not a keyword fragment.
- Branded and unbranded prompts are in separate groups and never averaged.
- The four journey shapes are all represented and separately grouped.
- Someone outside the marketing team supplied at least a third of the questions.
- You can name the prompts you expect to lose.
- The country and model matrix is written down, and its size is multiplied out
into a daily run count you have actually looked at.
- A stable core is designated and committed to for at least a year.
- The whole set is versioned, dated and stored somewhere a reader of next year's
report can find it.
If the set survives that list, the numbers built on it can be defended. The final chapter, what a defensible report looks like, assembles them into something you can hand to a sceptic.
Questions people ask
- How many prompts should I track?
- Fewer than you want to, and chosen more carefully than you want to. Every prompt you add multiplies against your model count and country count in daily runs, so breadth is bought with depth. A tight set of questions your buyers genuinely ask, tracked deeply enough to detect change, beats a long list tracked once a day that can never clear a significance test.
- Should I include my brand name in the prompt?
- Only if you specifically want to measure branded questions, and then keep those prompts in a separate group. An unbranded prompt is the honest test of whether a model reaches for you unprompted; a branded one measures whether the model can describe you when told who you are. Both are useful. Averaging them together produces a number that is neither.
- Can I just paste in my keyword list?
- You can, and the results will mislead you. Keyword-shaped input produces a thin, unnatural answer that tells you nothing about how buyers are actually being advised. Assistants respond to questions; a search engine responds to fragments. Convert each keyword into the sentence a person would actually type or say.
- What happens to my history if I change the prompt set?
- Any comparison that spans the change is broken, because the denominator is composed of different questions on either side. Treat a prompt-set change like a schema migration: version it, date it, and either report the two eras separately or restrict trend comparisons to the prompts common to both.
- How do I stop myself from picking prompts I already win?
- Build the set from evidence outside your own head — sales-call recordings, support tickets, the questions in your onboarding form, community threads — and have someone who is not accountable for the numbers approve the final list. If you cannot name a prompt in your set that you expect to lose, you have built an instrument that cannot detect failure.
- Do I need a separate prompt set per country?
- You need separate measurement per country, which is not the same as separate prompts. The same question in a different locale is a different query with different retrieval and often a different language, so it forms its own cells. Start by translating and localising the same questions, and only diverge the sets where buyers in that market genuinely ask something different.
Sources
Every factual statement above, with the page it came from and the date that page was read.
Attensira's prompt guidance instructs users to write prompts in the words their buyers use rather than their own category language, and to phrase each as a complete question rather than a keyword fragment.
docs.attensira.com · retrieved
“Write prompts in the words your buyers use, not your own category language.”
Attensira's prompt guidance instructs users to avoid their brand name in the prompt unless they specifically want to measure branded questions, describing an unbranded prompt as the honest test of whether a model reaches for the brand unprompted.
docs.attensira.com · retrieved
“an unbranded prompt is the honest test of whether a model reaches for you unprompted”
Attensira documents that keyword-shaped input produces a thin, unnatural answer that tells the user nothing about how buyers are actually being advised.
docs.attensira.com · retrieved
“produces a thin, unnatural answer that tells you nothing about how buyers are actually being advised”
Attensira sells prompts as simultaneous slots rather than monthly allowances, with 50 slots on Starter, 150 on Growth and 350 on Business, and a slot returning to the pool when a prompt is removed.
docs.attensira.com · retrieved · changes often, check the source
Attensira runs each combination of prompt, model and country at most once per workspace-local calendar day.
docs.attensira.com · retrieved
“Each prompt × model × country combination runs at most once per workspace-local calendar day.”
Attensira documents that every model added multiplies the prompt count by one more daily reading per prompt, per country.
docs.attensira.com · retrieved
“every model you add multiplies your prompt count by one more daily reading per prompt, per country”
Attensira documents that a brand name which is an everyday word produces a mention rate partly composed of noise that averaging cannot remove, because the error is systematic rather than random.
docs.attensira.com · retrieved
“no amount of averaging removes it, because the error is systematic rather than random”
Google documents that AI Mode and AI Overviews may use different models and techniques, so the responses and links they show will vary.
developers.google.com · retrieved
“AI Mode and AI Overviews may use different models and techniques, so the set of responses and links they show will vary.”