No one outside these companies knows the full selection logic, and anyone claiming otherwise is guessing. What is observable is that a source needs to be retrievable at the moment of asking, needs to state its claims in passages that survive being quoted alone, and needs to be corroborated elsewhere so the claim is not resting on a single page. Those three are the parts you can influence.
What we can observe
Most AI answers are produced by retrieving documents and then generating text from them. That retrieval step is the gate, and it behaves in ways that are at least partly visible from the outside.
| Factor | Why it appears to matter | Can you influence it |
|---|---|---|
| Retrievability | A page that cannot be fetched or parsed cannot be cited | Yes, directly |
| Passage clarity | A model quotes a passage, not a page. The passage has to stand alone | Yes, directly |
| Corroboration | A claim appearing in several independent places is safer to repeat | Slowly, and not on your own site |
| Topical association | Being consistently connected to a subject across the web | Slowly |
| Freshness | Recent sources appear more often on questions where recency matters | Yes |
| Underlying ranking | For AI features inside search, the results being summarized come from the index | Yes, through ordinary SEO |
Passage-level quoting changes how you write
This is the practical difference from writing for a ranked results page. A model pulls one paragraph, discards the surrounding context, and presents it as an answer. Two failure modes follow.
A paragraph that depends on the one above it becomes unusable once it is separated. A paragraph that is technically true but ambiguous without context becomes actively misleading, and gets attributed to you.
Writing that survives this states its claim first, keeps the qualification in the same sentence, and does not save the point for the end.
The corroboration problem
A claim that appears only on your own website is a claim from one interested party. The same claim appearing in an industry publication, a directory listing, a conference program and a customer's case study is something else entirely.
This is the slow half of the work and the half most agencies skip, because it involves activity that is not on your website: getting listed accurately, getting mentioned, getting described the same way in each place.
What is genuinely opaque
- The weighting between these factors, which is not published and changes
- How much any given model relies on training data against live retrieval for a particular question
- Why the same question asked twice returns different sources
- How personalization and prior conversation affect what gets cited
That variability is real. It is also why any figure describing your share of AI citations should be treated as a sample rather than a measurement, and reported that way.
What to do about it
- Confirm your pages can be fetched by the crawlers these systems use, which is a different list from traditional search crawlers.
- Rewrite your most important claims as standalone passages. One idea, stated fully, in one place.
- Add structured data so what the page covers does not need inferring.
- Get the facts about your business consistent everywhere they appear off your site.
- Sample regularly and watch the direction rather than the number.
Everything above is inference from observation. These systems are not documented at the level people writing about them imply, and the ones who sound most certain are usually the ones who have looked least.