I. The question nobody asks
Everyone working on AI visibility looks at the same thing. Am I being cited? Which sources does the model choose? How do I steer that selection?
It’s a logical question. Citations are measurable. Tools count them. Dashboards show them. You can work on them.
To be fair: bias is discussed. But almost exclusively as an ethical issue. Gender, race, political slant — there’s an entire literature on that. What hardly exists is bias as a marketing discipline. As something a brand can influence, on a timescale it can actually plan for.
A word choice worth pausing on. Bias works as a hook, but it isn’t technically the right word for what I mean in this piece. Bias implies distortion — that the model sees things crookedly. What I’m describing isn’t crookedness. It’s prior knowledge. ML researchers speak of model priors or brand salience — what the model has already learned about brands and how they relate to categories. From here on I’ll use priors or simply prior knowledge. The title keeps the hook. The rest of the piece tries to be more precise.
Because there’s a layer beneath the citation question. A layer that’s active before the model consults a source. Before it decides what information is relevant. Before anything is retrieved at all.
That layer co-determines who gets a seat at the citation table.
Before the model starts choosing, it has already chosen.
II. Two mechanisms, one answer
An AI answer emerges from two mechanisms that run through each other.
The first is retrieval. At the moment of the question, the model fetches sources, picks which ones are relevant, and weaves them into an answer. This is what citation tools measure. Recent, controllable, influenceable within weeks.
The second is what the model “knows” before the question arrives. Built up during training, from hundreds of billions of text fragments in which brands appear, get described, get compared. The model learns more than facts from that. It learns weight. Which brands keep coming back for certain topics. Which sources are credible. Which names belong, by default, to which category.
The two mechanisms reinforce each other. A brand that’s strongly represented in training data is more easily picked up as a reference during retrieval — the model “thinks” of that brand faster when it searches. A brand that was absent during training has to work harder to be picked up.
How much harder depends on the system. With modern retrieval-first platforms like Perplexity, ChatGPT with search, or Gemini, strong current visibility compensates for a large part of what’s missing in training. The difference is more pronounced with models that lean less on live sources, and with questions for which no verifiable answer exists.
Citation is what the model does today. Prior knowledge is what it learned yesterday.
III. Which mechanism weighs more? It depends on the question.
The balance between the two mechanisms isn’t fixed. It shifts depending on what the user asks — and not primarily on where they sit in the funnel.
The real axis is subjectivity. The more factual the question, the heavier retrieval weighs. The more subjective the question, the heavier priors weigh.
A factual question — “what is the current savings rate at that one bank?” — calls for verifiable information. The model goes actively looking. Whoever is cited right now ends up in the answer. Priors play a part in which sources even get checked, but the heavy lifting is retrieval. This type of question often sits at the bottom of the funnel, but not always. “What are the best heat pumps in Belgium?” looks like an orientation question, but is heavily steered today by current review sites — because the answer is verifiable.
A subjective question is different. “Which bank is good for a starter?” No fixed answer. The model leans on what it already knows. Which brands it sees as strong candidates. Which names come up naturally for this kind of question. Citation is often garnish here rather than basis — a few links supporting an answer the model was going to give anyway.
This touches on something fundamental for anyone looking at AI visibility through a funnel lens. At the top of the funnel, questions tend to be subjective — orientation, discovery, the first shortlist. At the bottom, questions tend to be factual — concrete comparisons, prices, conditions. The funnel correlates with the subjective–factual axis, but doesn’t map onto it cleanly.
What that means in practice: brands trying to win top-of-funnel won’t get there with citation optimisation alone. The phase where shortlists are formed leans heavily on what the model already knew when the user asked the question. And the phase where shortlists are formed is exactly the phase paid media is trying to reach and your website needs to feed.
IV. How prior knowledge forms
Prior knowledge is no mystery. It follows patterns that can be described in three observations.
Frequency wins. A brand that appears more often in training data is more readily used as a reference — even when it isn’t objectively better. AI models partly estimate relevance from repetition. This isn’t a judgement on quality. It’s a consequence of how statistical learning works: what shows up more often gets associated more strongly.
Authority weighs. Not all sources count equally. A mention on Wikipedia, in established media, in encyclopaedic platforms or in government sources is processed differently by the model than a mention on a blog or forum. Not because the model explicitly chooses, but because such sources keep returning during training as sources of truth, and therefore carry more weight in what the model stores as authoritative.
Silence isn’t neutrality. A brand absent from training data isn’t “neutralised”. It’s invisible. The model has no anchor to attach your name to, and falls back on safer brands as a reference. Whoever isn’t in the model’s head doesn’t get the chance to come out of it when a question is asked.
V. Why this is harder than citation
The uncomfortable truth: prior knowledge can’t be fixed with a strong week of content.
Training data is a snapshot that’s already closed by the time the model is released. The next training is months away. What you publish today counts for the next generation — not for the model assessing your brand tomorrow.
That’s no excuse to do nothing. But it does mean the timescale is different. Whoever starts now is building for what the model knows six to eighteen months from now, not what it says this week. Whoever waits for measurable results is waiting structurally too long.
The content you publish today does its real work when a different model is being trained.
VI. What you can actually do
Volume over perfection. For priors, repetition counts more than polish. A brand publishing ten articles on one topic builds a stronger association than a brand with one perfect article. Not flat repetition — but consistently returning to your core themes, from different angles, over a long time.
Letting authority work for you. Mentions on sources the model treats as authoritative carry disproportionate weight in building that prior knowledge. One well-placed mention on a platform that scores high in training data can match years of your own content. That’s why PR no longer works as a separate discipline alongside SEO — it’s become an ingredient of AI visibility.
Consistency over time. A brand that writes consistently about one theme for three years gets a deeper anchor than a brand covering eight themes at once for two. Specialisation beats breadth, even when breadth looks more interesting on paper. The model mostly learns what you are by watching what you repeat.
VII. The sum of two mechanisms
Citation and prior knowledge aren’t alternatives. They’re the two axes on which your AI visibility exists — and they move on different timescales.
Whoever focuses only on citation builds visibility on sand. It works in the short term, but neglects the underlying picture the model is forming of you. One model update can take that foundation away.
Whoever focuses only on prior knowledge acts as if the world stands still. Training cycles are slow, and meanwhile something happens every day in retrieval that makes your brand visible or invisible right now.
The work is in doing both at once, with the awareness that they operate on different timescales. Citation for today’s visibility. Prior knowledge for tomorrow’s. And the awareness that the phase where most orientations happen — subjective questions, first shortlists, the moment before the search behaviour begins — is largely determined by that second mechanism.
At the bottom of the funnel, who gets cited matters. At the top, what the model already knew matters.


