For the complete documentation index, see llms.txt. Every page on this site is also served as Markdown: append `.md` to any URL, or send `Accept: text/markdown`.
Attensira Logo
Attensira

The AI search visibility guide · chapter 4

Why most of your citations are not on your domain

Why does an AI assistant describe my company using someone else's pages instead of mine?

Karl-Gustav KallasmaaKarl-Gustav Kallasmaa, Founder & CEOLast updated

Why does an assistant describe my company using someone else's pages?

Because for the questions buyers actually ask, someone else's page is the better answer. "Which AI visibility tool should a five-person team use" is a comparison question. A vendor's own page cannot answer it - it can only describe one option, and it has an obvious interest in the conclusion. A page that weighs four tools against stated criteria answers it directly. Retrieval picks the page that answers the question, and for most commercial questions that is not yours.

What the numbers say, with the caveat attached

The most-cited figure here is vendor-published, so read it with the discount that deserves. AirOps analysed more than 500 commercial-intent queries across GPT-5, Claude Sonnet 4.5 and Perplexity Sonar, capturing 21,311 brand mentions, and reported that 85% of brand mentions came from external domains while 13.2% came from the brand's own domain[^airops-share] - which the same report expresses as brands being 6.5 times more likely to be mentioned through third-party sources than their own.[^airops-multiple]

What that figure is: one vendor's sample of commercial software queries on three platforms at one point in time, with the method stated.

What it is not: an operator disclosure, a law of nature, or a number that transfers to your category without checking. A regulated industry where the authoritative source is the vendor's documentation will look different. Measure your own before you plan against someone else's.

What survives the caveats is the direction, and the direction is not subtle. If even roughly one mention in seven originates on your own domain, then a programme that only edits your own domain is working on a minority of the surface that decides what assistants say about you.

Why the mechanism produces this

Three documented behaviours combine into it, and none of them are anti-brand sentiment.

Fan-out multiplies the questions. Google says its AI features may issue multiple related searches across subtopics to build one response.[^google-fan-out] Your site answers the sub-questions about your product. Nobody's site answers "how do these four compare" except a third party's.

More links, more diverse. Google says the process identifies more supporting pages and shows a wider, more diverse set of links than classic search.[^google-diverse-links] More slots, and diversity is explicitly a goal - which structurally means not filling them all from one domain.

Credibility signals reward corroboration. The GEO study found that citing credible sources was among the strongest interventions measured.[^geo-cite-sources] A system rewarded for grounded, attributable text prefers a claim that appears in a source with no stake in it. A vendor asserting its own superiority is the weakest available evidence for that claim, and models are increasingly good at noticing the difference.

The work this implies

The uncomfortable part: most of it is not content marketing, and most of it cannot be automated end to end.

Directory and listing accuracy

Category directories, marketplaces, review platforms and "best tools" roundups get cited constantly because they are structurally the answer shape for comparison questions. Most companies have entries on a dozen of these, half stale, several describing a product two versions old. Auditing and correcting those is boring, cheap, and higher-return than another blog post - not because a directory is authoritative, but because it is read, and what it says about you becomes what the model says about you.

Start with the ones already being cited on your prompt set. You do not have to guess which they are: they are in the source lists of the answers you are already sampling.

Reviews you did not write

Review platforms answer the sub-question fan-out generates for almost every commercial query: what do users complain about. If your review presence is three reviews from 2024, the model will use whatever is there, including one unhappy customer with an axe. The fix is asking real customers, at the right moment, and accepting the ones that are not glowing. Manufacturing reviews is both a policy violation on most platforms and a bet that detection will not improve.

Comparison pages that include you

Someone will write the comparison of your category. Better it be accurate. That means being easy to compare with - published pricing, published limits, published docs - so the person writing the comparison does not have to guess, and guess wrong.

This is also the honest argument for writing comparisons yourself, including ones where you lose a criterion. A comparison that concludes you win everything is transparently marketing, and both readers and models discount it. One that says "they are better if you need X" is usable, and usable is what gets cited.

Being fetchable when a buyer asks directly

There is one third-party path that runs back through your own site. Anthropic documents that when a person asks Claude a question, it may fetch sites using its user agent.[^anthropic-user-fetch] A buyer in an evaluation types your name; the assistant goes and reads you. That page had better be readable, current and specific - and not blocked, which is chapter one.

Build the source list before you build anything else

The whole chapter reduces to one artefact, and it takes an afternoon.

  1. Take your prompt set - the ten to thirty questions from chapter two, in a

buyer's words, including the ones where your brand is not named.

  1. Run each one several times on each platform you care about, and record

every URL cited. Not just the ones that mention you. All of them.

  1. Count domains. You will typically find a short head - a handful of domains

appearing across most questions - and a long tail of one-offs. The head is your working list.

  1. Read what the head says about you. For each of those domains: are you

listed at all, is the description current, is the pricing right, are the limits right, is a competitor's entry more complete than yours?

  1. Sort by frequency times wrongness. A domain cited on fifteen of your

twenty questions that describes a feature you retired is the top of the backlog. A perfect entry on a site cited once is not worth an email.

The output is a list of specific corrections at specific URLs, each with an owner. That is a different object from a content calendar, and it is the one that moves what assistants say about you.

Two notes from running this. First, the same handful of domains usually dominates an entire category, which is why the work is finite rather than endless - it is a dozen entries, not a campaign. Second, a surprising share of the errors are your own fault: pricing you changed without telling the directory, a positioning line three quarters out of date, a docs URL that now redirects. Those are the cheapest fixes available and nobody schedules them.

Where the line is

Every off-domain tactic has a version that ages badly, and the pattern is the same one link building went through.

The test that separates the columns: would you be comfortable if the assistant cited the tactic itself? "This brand's entry is accurate and its customers review it well" is fine. "This brand pays for placement in the list it tops" is not the sentence you want in the model's context window.

There is also a practical argument against the right-hand column. The pages in that column are optimised for volume, which means they are exactly the pages that lose trust first when platforms tighten their source selection. You would be buying an asset whose entire value depends on nobody looking closely.

What we do not claim here

We build a product in this space, so the boundary matters. We do not publish customer results for third-party work, because we do not yet have results we can show. We are not claiming a causal link between any specific off-domain action and a citation, because no public dataset establishes one at the level of a single brand. What is defensible is the correlation the vendor data shows, the mechanism Google describes, and the observation that the pages cited instead of yours are a readable specification for what to fix.

The one thing you own that behaves like third-party content

There is an exception to "your domain is a minority of the surface", and it is the most under-used asset most software companies have: documentation.

Docs behave more like third-party content than like marketing content, for three reasons. They answer narrow factual sub-questions, which is what fan-out generates. They are written to be correct rather than persuasive, which is the register the retrieval research rewards. And they are usually structured as discrete question-shaped units already, which extracts cleanly.

The failure modes are equally specific. Docs behind a login cannot be read at all. Docs rendered entirely client-side often cannot be read either. Docs that are a year out of date are worse than absent, because they will be quoted confidently. And docs split across a marketing site and a separate docs domain frequently have one of the two blocked by a CDN rule nobody audited, because the docs domain was set up by a different team.

If you sell to technical buyers, auditing your docs for fetchability and currency is likely to move more than a quarter of blog posts. It is also the only part of this chapter you can fix unilaterally, this week, without asking anyone outside the company.

What to take from this chapter

Pull the source URLs from your own sampled answers, sort them by how often they appear, and look at what the top ten say about you. That list is the off-domain backlog, ordered by impact, and it costs an afternoon to build. Almost nobody does it - most teams work from a keyword list instead, which describes what people search for rather than what the model reads.

Questions people ask

Why does the assistant cite a review site instead of my product page?
Because for a question like "which tool should I use", a page comparing several tools answers the question more completely than any single vendor's page, and a vendor page is a self-description. Your page is one input; a page that weighs you against alternatives is the answer shape.
What share of brand mentions comes from a brand's own domain?
The best public figure is vendor-published, not operator-published. AirOps analysed over 500 commercial-intent queries and 21,311 brand mentions across three platforms and reported 85% of mentions from external domains and 13.2% from the brand's own domain. Treat the direction as informative and the precise number as one sample of one query set.
Can I pay to be added to a "best tools" listicle?
You often can, and the listings that accept payment are usually the ones assistants have the least reason to trust over time. The durable version is being genuinely comparable - accurate entries on directories that verify, and reviews from real customers you did not script.
How do I fix an assistant describing my product incorrectly?
Find the pages it cites when it makes the error, and correct the error at those sources - a stale directory entry, an outdated comparison, a two-year-old review. Publishing a correction on your own site does not overwrite a claim the model is reading from somewhere else.
Is this just link building with a new name?
The mechanism differs. A link passes authority to a URL; a mention gives the model a sentence about you, and the sentence can be right or wrong regardless of whether it links. The work that matters is being described accurately in places that get read, which overlaps with digital PR and barely at all with buying links.

Sources

Every factual statement above, with the page it came from and the date that page was read.

  1. An AirOps analysis of over 500 commercial-intent queries capturing 21,311 brand mentions across GPT-5, Claude Sonnet 4.5 and Perplexity Sonar reported that 85% of brand mentions came from external domains while 13.2% came from the brand's own domain.

    airops.com · retrieved · changes often, check the source

    85% of brand mentions came from external domains, while only 13.2% of mentions came directly from the brands domain
  2. The same AirOps analysis reported that brands are 6.5 times more likely to be mentioned through third-party sources than through their own domains.

    airops.com · retrieved · changes often, check the source

    Brands are 6.5x more likely to be mentioned through third-party sources than their own domains.
  3. Google states that AI Overviews and AI Mode may use a query fan-out technique, issuing multiple related searches across subtopics and data sources to develop a response.

    developers.google.com · retrieved

    may use a "query fan-out" technique - issuing multiple related searches across subtopics and data sources - to develop a response
  4. The GEO study found that adding citations from credible sources was among the three methods producing a 30-40% relative improvement on its position-adjusted word count visibility metric.

    arxiv.org · retrieved

    our top-performing methods, Cite Sources, Quotation Addition, and Statistics Addition, achieved a relative improvement of 30-40% on the Position-Adjusted Word Count metric
  5. Anthropic states that when individuals ask questions to Claude, it may access websites using a Claude-User agent, and that disabling it prevents retrieval of a site's content in response to a user query.

    support.claude.com · retrieved

    When individuals ask questions to Claude, it may access websites using a Claude-User agent.