Methodology
How we measure AI visibility, and why the margin of error is the whole story
Ask an AI the same question twice and you will often get two different answers. That makes any single check an anecdote rather than a measurement, a single run carries a margin of error of roughly ±9 to ±13 percentage points. We sample around 1,500 answers per source per cycle to get that down to about ±1 to ±2, and we publish the interval alongside every figure.
If your current AI visibility report doesn't state a margin of error, it isn't telling you whether anything changed.
Last updated 2026-07-31
The problem: these systems are non-deterministic
Traditional rank tracking works because Google's results are broadly stable. Check position three on Tuesday, check again on Wednesday, you'll usually see position three.
AI answers don't behave like that. Ask ChatGPT "who are the best suppliers of X" ten times and you may get ten overlapping but different lists. Brands appear and vanish. Order shifts. The same prompt, the same day, the same model.
So "we appear in ChatGPT for this prompt" is not a fact. It's the outcome of one sample from a probability distribution.
Think of it as a coin. Flip it ten times and you might get seven heads, that tells you almost nothing about the coin. Flip it a thousand times and you'll land close to the truth. Brand appearance in an AI answer works the same way, except it's usually a rare event, and rare events need even more samples to pin down.
What that does to a single check
Take a brand whose true appearance rate is 5%, it genuinely shows up in one answer in twenty.
Check each prompt once, and your reported figure carries a margin of error of about ±9 points. The true value of 5% could read as anything from 0% to 14%.
Now imagine your agency reports that you moved from 4% to 7% this month. Inside a ±9 point interval, that "improvement" is indistinguishable from nothing happening at all. You could equally have declined.
This is the central failure of AI visibility reporting. A tool that checks each prompt once a day is showing you seven coin flips a week and drawing a trend line through them. The line will move. It will move whether or not anything real happened.
What we do instead
We sample every tracked prompt around 50 times per measurement cycle, across 30 prompts. That's roughly 1,500 answers per source, per cycle.
At that depth, the margin of error on a 5% appearance rate drops to about ±1.1 points, a true range of 3.9% to 6.1%.
That is a number you can put in a board pack. More importantly, it's tight enough to detect a genuine one-to-two point shift week on week, which is what real progress actually looks like.
| Approach | Samples per prompt, per week | Margin of error | Can it detect a 2-point change? |
|---|---|---|---|
| Typical reporting tool (one check daily) | 7 | ±9 to ±13 points | No |
| Directional monitoring | ~15 | ±15 points | No |
| Our standard | ~50 | ±1 to ±2 points | Yes |
The statistics, stated plainly
- Wilson confidence intervals for appearance rates. Appropriate for proportions with rare events and small samples, where the more familiar normal approximation breaks down and produces nonsense like negative lower bounds.
- 95% confidence level on all reported intervals, unless stated otherwise.
- Species-accumulation modelling for competitor discovery, borrowed from ecology, where the same problem exists: how many samples before you've found everything that's out there?
Every figure we report is reproducible from the underlying data. A screenshot isn't.
Interactive: margin of error vs. sample count (Wilson interval, 95% confidence, 5% true appearance rate)
Margin of error
±6.6 pts
Range
1.6% – 14.9%
Reported rate
5.0%
Finding your real competitive set
There's a second measurement problem that nobody talks about.
Every AI answer names other brands. The more you sample, the more distinct competitors you discover, and the list never fully stops growing.
What the data shows consistently:
- Core competitors, the ones named in 10 or more answers, nearly all surface within about 25 runs. These are the brands actually competing with you for the recommendation. Roughly nine of them, typically.
- Occasional mentions, named two to nine times, worth knowing, noisier, around fifteen of them.
- One-offs, named once, a long tail of twenty-plus brands the model occasionally drops in. Chasing these is a waste of your money.
Two questions hide inside "who are our competitors in AI answers?" Have we found the ones that matter? Yes, and quickly. Have we found every brand the model might ever name? No, and we never will. Depth gets you the first. The second is a rabbit hole and we'll tell you so.
This matters commercially because a single-check tool reports whatever competitor happened to appear in one answer. You can spend a quarter chasing a rival that the model names one time in fifty.
What we report, and what we refuse to report
What you get, every cycle:
- Share of voice per platform, with a stated confidence interval
- Whether movement since last cycle is statistically distinguishable from noise, stated explicitly, in words
- Your core competitive set, with appearance rates and intervals
- Prompt-level detail: where you appear, where you don't, who's there instead
- Citation sources: which pages and domains the engines are actually pulling from
- AI crawler activity from server logs, GPTBot, ClaudeBot, PerplexityBot, Google-Extended
- Interventions logged with dates, so cause and effect can be assessed later rather than asserted now
What we won't report:
- A single composite "AI visibility score" with no methodology behind it. These are unfalsifiable by design.
- Movement presented as significant when it sits inside the margin of error.
- Percentage changes without the absolute numbers underneath. Going from one mention to two is a 100% increase and means nothing.
- Competitor comparisons drawn from insufficient samples.
Where our measurement stops
Stating the limits, because a methodology page that claims no limits is marketing.
We measure appearance, not causation. If your share of voice rises after we publish twelve pages, that is correlation with a plausible mechanism. It is not proof the pages caused it. We log interventions with dates so the relationship can be assessed honestly, and we say "consistent with" rather than "caused by."
Prompt selection shapes the result. We measure 30 prompts. Your buyers ask an effectively infinite number. We select for commercial intent and document the set, so you can challenge it, and you should.
Platforms change without notice. Any of these systems can alter its retrieval behaviour overnight, with no changelog. A step change in your numbers may be a platform update rather than anything either of us did. We flag suspected platform effects rather than claiming credit.
Geography and personalisation. Results vary by location and by user context in ways we can only partly control. We measure against a defined country and language and state which.
Mentions versus citations differ. ChatGPT mentions brands roughly 3.2 times more often than it links them, per Seer Interactive. A mention without a link still shapes a buying decision but sends no traffic. We report them separately, because conflating them inflates the numbers.
The question to ask any agency
"What's the margin of error on that figure?"
If they don't know, they're showing you a screenshot. If they say there isn't one, they don't understand the systems they're charging you to measure. If they give you a number, ask how many samples it's based on.
It's one question and it will sort the market for you.
FAQ
Why do AI answers change every time I ask?
These models are non-deterministic, they sample from a probability distribution rather than returning a fixed result. Retrieval also varies as the live web changes. The same prompt can produce genuinely different answers minutes apart.
What's a good margin of error for AI visibility reporting?
±1 to ±2 points at typical share-of-voice levels lets you detect real week-on-week movement. ±9 to ±13, which is what a single daily check produces, does not. Anything above roughly ±5 should be treated as directional only.
How many prompts should be tracked?
We track 30 per brand per cycle at depth. More prompts with fewer runs each is usually the wrong trade, you get a broader picture you can't trust. Depth on commercially relevant prompts beats breadth on everything.
Can you prove your work caused my improvement?
Not with certainty, and we won't claim to. AI platforms change without notice, and we can't run a controlled experiment on your live business. What we can do is log every intervention with a date, report movement with confidence intervals, and be explicit about which relationships are causal claims and which are correlations. Anyone offering you proof of causation in this environment is overselling.
Do you measure Gemini, Claude and Perplexity as well?
Yes. Our core platform tracking covers ChatGPT, Google AI Overviews and Google AI Mode at full depth, and we run a supplementary panel across Gemini, Claude, Perplexity and Copilot. That matters increasingly: Gemini reached 750 million monthly active users and more than doubled its chatbot traffic share, while Claude is the fastest-growing B2B referrer in the data we've seen.
What if the numbers don't move?
We publish that too. Our public ledger includes interventions that produced no measurable change. Link →