AI Search

How do AI search tools decide which sources to cite?

No one outside these companies knows the full selection logic, and anyone claiming otherwise is guessing. What is observable is that a source needs to be retrievable at the moment of asking, needs to state its claims in passages that survive being quoted alone, and needs to be corroborated elsewhere so the claim is not resting on a single page. Those three are the parts you can influence.

What we can observe

Most AI answers are produced by retrieving documents and then generating text from them. That retrieval step is the gate, and it behaves in ways that are at least partly visible from the outside.

FactorWhy it appears to matterCan you influence it
RetrievabilityA page that cannot be fetched or parsed cannot be citedYes, directly
Passage clarityA model quotes a passage, not a page. The passage has to stand aloneYes, directly
CorroborationA claim appearing in several independent places is safer to repeatSlowly, and not on your own site
Topical associationBeing consistently connected to a subject across the webSlowly
FreshnessRecent sources appear more often on questions where recency mattersYes
Underlying rankingFor AI features inside search, the results being summarized come from the indexYes, through ordinary SEO

Passage-level quoting changes how you write

This is the practical difference from writing for a ranked results page. A model pulls one paragraph, discards the surrounding context, and presents it as an answer. Two failure modes follow.

A paragraph that depends on the one above it becomes unusable once it is separated. A paragraph that is technically true but ambiguous without context becomes actively misleading, and gets attributed to you.

Writing that survives this states its claim first, keeps the qualification in the same sentence, and does not save the point for the end.

The corroboration problem

A claim that appears only on your own website is a claim from one interested party. The same claim appearing in an industry publication, a directory listing, a conference program and a customer's case study is something else entirely.

This is the slow half of the work and the half most agencies skip, because it involves activity that is not on your website: getting listed accurately, getting mentioned, getting described the same way in each place.

What is genuinely opaque

  • The weighting between these factors, which is not published and changes
  • How much any given model relies on training data against live retrieval for a particular question
  • Why the same question asked twice returns different sources
  • How personalization and prior conversation affect what gets cited

That variability is real. It is also why any figure describing your share of AI citations should be treated as a sample rather than a measurement, and reported that way.

What to do about it

  1. Confirm your pages can be fetched by the crawlers these systems use, which is a different list from traditional search crawlers.
  2. Rewrite your most important claims as standalone passages. One idea, stated fully, in one place.
  3. Add structured data so what the page covers does not need inferring.
  4. Get the facts about your business consistent everywhere they appear off your site.
  5. Sample regularly and watch the direction rather than the number.
Worth being honest about

Everything above is inference from observation. These systems are not documented at the level people writing about them imply, and the ones who sound most certain are usually the ones who have looked least.

Want to see who gets named in your category?

A baseline sample is quick, and the results tend to start a useful conversation.

Book a Strategy Call