Prompt tracking vs keyword tracking: two ways to watch demand
A keyword report tells you what people typed and what happened. A prompt panel tells you what a model says when asked. What each approach can observe, what it costs, and where each one lies.
Prompt tracking
Choosing a set of questions a buyer might ask an assistant, putting them to models on a schedule, and recording what the answers say. The sample is generated by you rather than observed from traffic.
Checked 2026-09-04Keyword tracking
Reading the queries that led real users to your pages from a search engine, together with how often you were shown, how often you were clicked, and where you appeared.
Checked 2026-09-04Which one should you choose?
Keyword tracking observes behaviour that already happened to other people. Prompt tracking manufactures observations by asking a model yourself. The first is stronger evidence and covers a shrinking share of the buying journey; the second is weaker evidence and covers the part that is growing. Running only one of them leaves a real hole, and which hole depends on where your buyers ask their questions.
Choose Prompt tracking when
Choose prompt tracking when buyers arrive already repeating an assistant's summary of your category, when you rank well for terms whose clicks are falling, or when you need to know how a model describes you rather than whether it sent you anyone.
Choose Keyword tracking when
Choose keyword tracking when clicks are still the outcome you are paid for, when you need query-level data you did not have to invent, and when you want a free, already-running record of real impressions and positions rather than a sample you designed and must defend.
When neither is the right answer
Neither approach measures whether any of it produced revenue. A search engine can tell you a click happened; nothing here joins an answer a model gave a stranger to a person who later signed up. Connect either number to commercial outcomes with your own analytics, and say out loud that you are inferring the link.
What is specific to this comparison
- Only one of these two approaches requires you to invent the input: a keyword list is discovered from traffic that already happened, while a prompt set has to be written by someone before a single reading exists.
- The capacity models are opposite in kind — prompt slots are a stock that returns when a prompt is removed, whereas keyword reporting is bounded per request by a row limit rather than by how many terms you watch.
- The word position means a different object on each side of this pairing, which is why a chart that plots citation position against Search Console average position is meaningless rather than merely imprecise.
- Statistical gating only appears on the prompt side, because model non-determinism creates run-to-run variance that a record of real impressions simply does not have.
Prompt tracking vs Keyword tracking, criterion by criterion
The short answer
Keyword tracking reads a record of things that already happened to real people: which queries showed your pages, how often, and where. Prompt tracking creates its own record by asking models questions on a schedule and reading what comes back. The first is observation. The second is sampling. Confusing the two is how teams end up presenting a designed experiment as if it were traffic data.
The input problem, which decides everything else
The deepest difference is not the metrics. It is where the list of things to watch comes from.
A keyword list is discovered. The search engine knows which queries led to your site because your site was the destination, and it hands the list back to you. The Search Console performance report gives clicks, impressions, CTR and average position, grouped by dimensions including queries, pages, countries, devices, search appearance and dates. You did not have to guess what people typed. You found out.
A prompt list is written. An assistant that answers a question without sending anyone anywhere has no destination to report, so no equivalent list exists. Somebody has to sit down and write the questions a buyer would ask, phrased as Attensira's documentation puts it — a question a person actually types into an assistant, written the way they would type it. That is a design decision, and it is by far the largest source of bias in the method. A prompt set that quietly favours the terms you already win produces a flattering panel that measures your own assumptions.
The practical consequence: keyword data is wrong in ways the platform can explain, and prompt data is wrong in ways only you can explain. Write down why each prompt is in the set, and review the set as often as you review the numbers.
Cadence, sampling and why the two feel different
Keyword reporting accumulates continuously. Every impression is an event, the platform records it, and the report is a summary of a census rather than a draw from a distribution.
Prompt tracking has a heartbeat. In Attensira, each prompt, model and country combination runs at most once per workspace-local calendar day, with one reading per day on the entry plan and three on the higher ones as of 4 September 2026. That cadence has three consequences worth internalising. Today's figures are partial until the day closes. A day with collection problems produces a smaller sample rather than a distorted rate. And a single day is a thin basis for any claim at all.
It also explains the statistical machinery that has no counterpart on the keyword side. Models are not deterministic — ask the same question twice and you can get two answers — so a movement between two samples might be nothing. Attensira tests each delta with a two-proportion z-test at 95% significance and reports no proven change when the movement does not clear the bar. That is not the tool being cautious for effect; it is the only defensible thing to do with a small sample from a noisy generator.
Nobody applies a significance test to an impression count, because an impression either happened or it did not.
The word "position" means two different things
This is the most common analytical error when the two datasets sit on one dashboard, and it is worth stating plainly.
Google defines average position as the average position of the topmost result from your site. It is a rank among results, on a page, for a query.
A citation position is the mean index of a URL within a model's own citation list, computed per source domain. It answers "when this domain is cited, how far down the list does it sit". It is not a rank against competitors, and it does not become one by being drawn on the same axis as a Search Console line.
Two numbers sharing a word are still two numbers. Label them differently on every chart, and resist the temptation to build a blended "visibility" figure out of them, because the blend inherits the weaknesses of both and the interpretability of neither.
Reproducibility cuts against both, differently
There is a comforting assumption that keyword data is solid and prompt data is flaky. The first half is less true than people think.
Google's own documentation warns that even if a query appears in your list, you might not see your site in results if you run the same query in Google Search. Personalisation, location and timing make a manual check a poor audit of a reported number. So a keyword report is a reliable record of aggregate events and an unreliable guide to what any individual saw.
Prompt tracking is flaky in a more visible way and better instrumented for it. Because the variance is expected, the honest implementations store the evidence: the verbatim answer behind every reading, so a rate can be checked against what was actually said rather than trusted on faith.
The fair summary is that one side hides its variance behind large numbers and the other side exposes it. Exposure is not the same as being worse.
Cost behaves differently too
Keyword reporting is effectively free and already running; the constraint is retrieval, not collection. The Search Analytics API accepts a rowLimit between 1 and 25,000 with a default of 1,000, so at scale the work is pagination rather than budget.
Prompt tracking costs something every time it runs, because every reading is an inference call. That is why capacity shows up as slots — a slot is occupied for as long as a prompt is tracked and returns to the pool the moment you remove it. It is a stock, not a monthly allowance, which changes the discipline: the question is never "have I used up my quota" but "is this prompt still earning its place".
That constraint is healthy. A prompt set small enough to explain is worth more than a large one nobody curates.
What neither one can do
Neither approach tells you what happened next. A search engine can tell you a click occurred; it cannot tell you the click became a customer. Prompt tracking cannot even get that far, because there is no identity join between an answer a model gave someone and a person who later arrived at your site.
Anyone selling a straight line from either number to revenue is inferring it. Draw the line yourself if you must, in your own analytics, and label it as an inference.
How each one fails quietly
Both approaches have a failure mode that produces plausible numbers rather than obviously broken ones, which is the dangerous kind.
Keyword reporting fails quietly through survivorship. It can only show you queries where you were shown at all, so a topic you have no page for is not a low line on the chart — it is absent from the chart. Teams read a clean report and conclude they have covered their category, when what they have covered is the part of the category they already had pages for. The report cannot tell you about the demand you never touched.
Prompt tracking fails quietly through set drift. The set was written once, against a market that has since renamed itself, and every reading afterwards is an accurate measurement of an outdated question. Nothing in the data looks wrong. The rates are stable, the sample sizes are healthy, and the panel is measuring a conversation nobody is having any more. The only defence is a scheduled review of the set itself, treated as seriously as a review of the numbers.
Notice that the two failures are mirror images. One under-reports demand you have not addressed; the other over-reports the relevance of demand you chose to watch. Running both narrows the blind spot, because a topic missing from the keyword report and absent from the prompt set is much harder to overlook than one missing from either alone.
Reading them together without blending them
The useful practice is not a combined score. It is a small set of paired questions you ask of both datasets at once.
Do the two disagree about which topics matter? If your highest-impression queries and your most-answered prompts describe different subjects, one of your two audiences is invisible to you. That is usually the finding worth acting on.
Where do you rank well and get named rarely? Strong positions plus a low rate of appearing in answers points at pages that satisfy a crawler and give a model nothing quotable. The fix is on the page, not in the reporting.
Where do you get named often and rank poorly? The model knows about you from somewhere other than your own site. That is an off-site finding, and no amount of on-page work moves it.
Where has one moved and the other not? Movement in one dataset with stability in the other is either a real divergence or an artefact of the sample. The delta gate answers that on the prompt side; on the keyword side, widen the window until the shape stops changing.
Four questions, two datasets, no blended index. The gap between the two records is the information, and averaging them away is the one operation guaranteed to destroy it.
The one-line version
Track keywords to know what already happened to people who came to you. Track prompts to know what is being said about you where nobody clicks. Keep the two datasets, and the two definitions of position, apart.
Where Attensira fits, and where it does not
Worth a look if your site lives in a Git repository and the gap is between seeing a problem and shipping the fix. Attensira runs a chosen prompt set across assistants on a daily cadence, stores the verbatim answer as the evidence behind every rate, and opens a pull request with the change for your review. If you only need the reporting, cheaper options exist.
See how Attensira compares to bothQuestions people ask
Sources
Every claim on this page, with the page it came from and the date that page was read. Prices and feature lists change; these are what the source said on the date shown, not timeless facts.
- Attensira defines a prompt as a question a person actually types into an assistant, written the way they would type it, rather than a keyword.a question a person actually types into an assistant, written the way they would type ithttps://docs.attensira.com/tracking/prompts — read 2026-09-04
- Each prompt, model and country combination runs at most once per workspace-local calendar day.each prompt × model × country combination runs at most once per workspace-local calendar dayhttps://docs.attensira.com/tracking/prompts — read 2026-09-04
- Sampling depth is set by plan — one reading per day on Starter, three per day on Growth and Business.https://docs.attensira.com/tracking/prompts — read 2026-09-04
- A prompt slot is occupied for as long as a prompt is tracked and returns to the pool the moment the prompt is removed, so capacity is a stock rather than a monthly allowance.https://docs.attensira.com/tracking/prompts — read 2026-09-04
- Models are not deterministic, so the same prompt can return different answers; Attensira tests every reported delta with a two-proportion z-test at 95% significance and reports no proven change when it does not clear the bar.https://docs.attensira.com/reference/limitations — read 2026-09-04
- Citation position is the mean index of a URL within the model's own citation list, computed per source domain, and is not a ranking of a brand against competitors.https://docs.attensira.com/reference/limitations — read 2026-09-04
- There is no identity join between an answer a model gave someone and a person who later arrived at the site, so a mention cannot be attributed to a visit.https://docs.attensira.com/reference/limitations — read 2026-09-04
- The Search Console performance report provides clicks, impressions, CTR and average position, grouped by dimensions including queries, pages, countries, devices, search appearance and dates.https://support.google.com/webmasters/answer/7576553 — read 2026-09-04
- Google defines average position as the average position of the topmost result from your site.The average position of the topmost result from your site.https://support.google.com/webmasters/answer/7576553 — read 2026-09-04
- The Search Analytics API accepts a rowLimit in the valid range 1 to 25,000, defaulting to 1,000 rows per request.https://developers.google.com/webmaster-tools/v1/searchanalytics/query — read 2026-09-04
- Google warns that even if a query appears in your list, you might not see your site in results if you run the same query yourself.Even if a query appears in your list, you might not see your site in results if you run the same query in Google Search.https://support.google.com/webmasters/answer/7576553 — read 2026-09-04