Question.Marketing

How AI actually decides

How AI actually decides who to recommend

Most AI assistants answer in two stages: they retrieve web sources, then generate an answer grounded in what they retrieved. That makes visibility a retrieval problem. But the strongest current evidence suggests something less convenient, that models choose which brands to name from what they already learned in training, and retrieve sources afterwards to support a decision already made.

Both things are happening. Understanding which one you're up against determines what work is worth doing.

Last updated 31 July 2026

Stage one: retrieval, and why it differs by platform

The underlying pattern is retrieval-augmented generation. Before answering, the system fetches relevant sources and grounds its response in them. Your content has to be reachable, parseable and quotable to enter that pool at all.

Where it gets practical is that each platform retrieves from a different place.

PlatformRetrieves fromWhat that means for you
Google AI Overviews / AI ModeGoogle's index, core Search ranking systems, Google Business Profile and Merchant Center for local and commercial queriesYour Google SEO carries over most directly here. Business Profile completeness matters for local.
ChatGPTBing, primarily. Seer Interactive found 87% of SearchGPT citations matched Bing's top ten, against 56% correlation with Google'sBing indexing is not optional. Check Bing Webmaster Tools, most people never have.
PerplexityIts own index, reported at over 200 billion URLs, with heavy weighting toward community sources including RedditReddit and forum presence has genuine retrieval value here. Uncomfortable but true.
GeminiGoogle's index and infrastructureOverlaps with AI Overviews, but the conversational surface behaves differently
ClaudeWeb search at query timeFastest-growing B2B referrer, Goodie found 14 of 16 tracked brands grew their Claude share, typical gain 7.14 points
CopilotBing indexFollows ChatGPT's pattern; also wired into Merchant Center for commerce

The consequence people miss: optimising for Google covers roughly half the surface. A study across 55,936 queries and six LLM search engines found around 37% of the domains those engines cite never appear in traditional search results at all.

Stage two: the part that changes the strategy

Here is the finding that reorders the discipline, and it comes from the most methodologically serious work in the category.

Seer Interactive ran six independent behavioural tests across 362,388 AI responses in 2026. Their hypothesis: the model decides which brands to recommend first, drawing on trained knowledge, then searches for sources that support those choices.

Their formulation is the memorable one: the citation is the bibliography, not the brainstorm.

The corroborating correlation data from Ahrefs, across 75,000 brands:

0.664

Branded web mentions → correlation with AI Overview visibility

Ahrefs, 75,000 brands

0.218

Backlinks → correlation

Ahrefs, 75,000 brands

Branded mentions are roughly three times as predictive. The three strongest signals in that study were all off-site brand signals rather than link metrics.

Why this matters more than any tactic

If the recommendation decision precedes the citation search, then:

  • A brand the model hasn't learned cannot be optimised into a recommendation. Page structure, schema and formatting affect which sources get quoted. They don't change which brands get named.
  • The highest-leverage work is becoming known. Entity clarity, genuine notability, presence in the sources these systems read, review consistency, category authority. Old-fashioned reputation work.
  • Most of what's sold as AEO is bibliography optimisation. Useful, cheaper than it's priced, and insufficient on its own.

The honest caveat

This is a hypothesis with strong supporting evidence, not a settled fact. It's one research team's interpretation of behavioural tests on systems whose owners publish nothing about their internals. It is the best available explanation of the observed data. It could be wrong, or partially right, or right for some platforms and not others.

We're telling you that because the alternative, presenting a hypothesis as mechanism, is how this industry got its reputation.

What the controlled research actually found

The Princeton study

The foundational academic work is GEO: Generative Engine Optimization, Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, published at KDD 2024, with authors from Princeton, IIT Delhi, Georgia Tech and the Allen Institute for AI.

They built GEO-bench: roughly 10,000 queries across eight domains, each paired with the sources a generative engine would draw on. They tested nine content modifications and measured the effect on two metrics they introduced, Position-Adjusted Word Count and Subjective Impression.

What worked:

TacticEffect
Adding statisticsStrongest single result: +41% Position-Adjusted Word Count, +37% Subjective Impression, validated on Perplexity
Citing sourcesLargest equalising effect, a page at position five gained 115.1% relative visibility. Cite other people's work and you become more citable yourself
Adding quotationsConsistent gains, strongest in combination with the above
Fluency optimisationPositive
Technical terminologyPositive in technical domains

What failed: keyword stuffing, unsupported authority claims, and simplification all failed to produce gains, in some cases actively hurt.

The buried finding: effects varied substantially by domain. There is no universal setting. What works in law doesn't work in retail.

Where everyone misquotes it

You will see "Princeton proved GEO increases visibility 40%" everywhere. Three corrections:

  1. 40% is a maximum under favourable conditions, not an average. The three strongest methods produced 30–40% relative improvement against an unoptimised baseline. Low-ranked sources benefited disproportionately; already-visible sources gained far less.
  2. The architecture is dated. The study used Google's top five results for retrieval and GPT-3.5-turbo for synthesis. That is not what any 2026 system does. The findings are directionally valuable and mechanically superseded.
  3. It measured visibility within a generated answer, not business outcomes. No traffic, no conversions, no revenue.

It remains the best controlled experiment the field has. Anyone citing the 40% without those caveats hasn't read past the abstract.

The 23-factor synthesis

The most useful practitioner work is Cyrus Shepard's May 2026 analysis for Zyppy: 23 AI citation factors distilled from 54 studies, patents, experiments and case studies, each scored on repeatability, strength of evidence and official platform support.

Highest scoring: URL accessibility, at 9.5. If crawlers can't reach it, nothing else matters. Check your robots.txt for GPTBot, ClaudeBot, PerplexityBot, Google-Extended and Bingbot, a surprising number of sites are accidentally excluding the systems they're paying to appear in.

Contested: structured data at 5.6. Nearly every study examining it found positive correlation, but LLMs don't ingest schema as training data, so the mechanism is unexplained. Do it. Don't let anyone tell you they know why.

Lowest scoring: llms.txt, at 2.0. No credible evidence of measurable effect in any of the 54 sources. Google has said it won't use it.

Shepard's own headline conclusion: you don't need a new playbook. The overlap with traditional SEO signals is substantial, relevance, trust, topical authority, extractability.

The structural factors that do hold up

Answer position. Zyppy's 2025 analysis found 44.2% of LLM citations come from the first 30% of a page. Systems apply per-URL retrieval caps. Your key claim in paragraph four may never be reached. Lead with the answer, ideally a direct definition in the first 40 to 60 words.

Information gain. Content that restates what's already on the web provides nothing a model can't already generate. Content carrying proprietary data, original research, first-party evidence or genuine expertise becomes worth retrieving. This is the single most reliable content principle in the field and it is the opposite of volume.

Freshness. Ahrefs analysed nearly 17 million citations and found AI-cited content averages 1,064 days old against 1,432 for Google's organic top ten, around 25.7% fresher. Update dates on cornerstone content, genuinely.

Extractability. Clean heading hierarchy, direct definitions, stable terminology, FAQ structure where it's warranted. Pages with three or more schema types have been reported at 13% higher citation likelihood, though treat that figure as indicative.

Corroboration. These systems favour claims supported across multiple independent sources. Consistency of your core facts across your site, directories, review platforms, Wikipedia where legitimate and third-party coverage compounds.

What nobody knows

Worth stating plainly, because it's rare:

  • The weightings. No platform publishes how it ranks retrieval candidates. Every factor list in this industry, including this one, is inferred from correlation and experiment.
  • Why schema correlates. Consistent positive association, unexplained mechanism.
  • How much is parametric versus retrieved. The Seer hypothesis is compelling and unconfirmed.
  • Whether findings transfer between platforms. The Princeton study found domain-level variation within one architecture. Cross-platform transfer is largely untested.
  • How stable any of it is. These systems change without notice or changelog. A tactic validated in March may be irrelevant by September.

Which is why the only defensible approach is to measure your own position, with a stated margin of error, repeatedly.

How we measure →

FAQ

What's the single most important factor?

Whether crawlers can reach your content, it scores 9.5 of 10 on the best available evidence synthesis and everything else depends on it. After that, brand presence: branded mentions correlate roughly three times more strongly with AI visibility than backlinks do.

Does schema markup help AI citations?

Almost every study finds a positive correlation, but the mechanism is unclear, since LLMs don't consume schema as training data. It scores 5.6 of 10 and is flagged as contested. Worth doing; be sceptical of confident explanations.

Does llms.txt help?

No published evidence that it does. Lowest-scoring of 23 factors examined across 54 studies. Google has said it won't use it. It does no harm.

Why does my Google ranking not translate to ChatGPT?

Because ChatGPT retrieves through Bing rather than Google, 87% of its citations matched Bing's top ten against 56% correlation with Google's. Different index, different results. Check your Bing indexing.

Should I optimise differently for each platform?

The foundations are shared and worth doing once. Platform-specific work matters at the margins: Bing indexing for ChatGPT and Copilot, Google Business Profile for AI Overviews on local queries, community presence for Perplexity. The Princeton research found effects vary by domain, so there's no universal configuration.