Ask the same assistant the same question twice in an hour and you can get different businesses named. This is the single most misunderstood thing about measuring AI visibility, and it changes how any measurement should be read.
Why it happens
Three reasons, and only one of them is about you.
Sampling. These systems generate text probabilistically. Two runs of the same prompt take slightly different paths, and when three businesses are roughly equally supported by the sources, which one gets named first is close to a coin toss.
Live retrieval. Most answer engines fetch sources at query time rather than answering from memory. What is available, and what ranks in the underlying search, differs between runs.
Personalisation and context. In a logged-in session the assistant may have seen you discuss your own business. This is the biggest single source of false confidence in this field, and it is entirely avoidable.
The practical consequence
A single answer is an anecdote. If you asked once, you learned almost nothing.
Two points close together are the same point. In our own measurement across 608 answers, we treat a gap of about two percentage points between adjacent brands as noise rather than as a difference.
Direction over weeks beats magnitude on a day. A brand moving from 4% to 9% over two months is a real change. The same brand appearing in today's answer and not tomorrow's is not.
How to measure so the variance does not fool you
Log out. Use a temporary or logged-out session every time. In your normal account the assistant has seen you talk about your business and will bring it up, which is memory rather than visibility, and it is why people conclude things are fine when they are not.
Repeat. Each prompt several times across several days, not several times in one sitting. Same-session repeats share context and correlate.
Use enough prompts. Ten to thirty. Fewer and one prompt's variance dominates. Many more and nobody reads the actual answers, which is where the useful information is.
Measure fortnightly, not weekly. Weekly checking shows you variance and is a reliable way to talk yourself out of an approach that is working.
What variance does not explain
It does not explain a competitor appearing in eight of your ten prompts while you appear in none. That is not a coin toss, and treating a consistent absence as bad luck is the most expensive mistake available here.
The test is simple. If your name shows up sometimes, you are in the answer set and the variance is deciding the order. If it never shows up across thirty runs, you are not in the set at all and no amount of repetition will surface you.
What it means for reporting
Any AI visibility report should be read as a distribution rather than a scoreboard. The useful lines are the ones that persist: which competitors are consistently named where you are not, and whether your own figure is trending.
This is also why a report showing a movement you cannot reproduce in one sitting is not automatically wrong. Ask how the measurement runs, which engines, how often, and logged in or out. There are legitimate reasons for a gap, and one illegitimate one.
Method
Figures here come from 608 answers collected between 24 May and 24 August 2026 across ChatGPT, Google AI Overviews, Google AI Mode, Perplexity and Gemini, on a rolling schedule rather than in one sitting, which is exactly what makes variance visible rather than invisible.
How much variance is normal
Useful to have a rough scale, because the most common mistake is treating ordinary noise as a result.
Within one engine, same question, days apart. Expect the named set to change by roughly one brand in three checks. A brand that appears two times in three is genuinely present; one that appears once in ten is not.
Between engines, same question, same day. Expect substantial disagreement. These systems draw on different source mixes, and two engines agreeing on all three names is less common than them agreeing on one.
Across a month, same question and conditions. This is where a real change becomes visible, because it is long enough for the noise to average out and short enough that the underlying sources have not shifted much.
What is not variance. A brand that never appears in twenty checks is absent, not unlucky. A brand that appears in every check has a genuine position. The middle is where judgment is needed, and the answer is always more checks rather than more interpretation.
The practical rule: never act on a single check, in either direction. Good news from one run is as unreliable as bad news.
Why we check continuously
The figures above are why the research on this page exists. We run this measurement continuously across every account rather than sampling it, which is the only way to reach a number like 82,000 measured answers and the only way to see a change in the week it happens rather than the month after.
A worked example of reading variance correctly
A tracked set of twenty questions, checked weekly across five engines for eight weeks. What the raw data looked like and what it actually meant.
Question 4 appeared in weeks 1, 3, 4, 6, 7 and 8. Six of eight. That is a genuine position with normal noise around it. No action needed.
Question 11 appeared in week 2 only. One of eight. That is not a position, it is a single lucky draw, and reading it as a win would have been the most expensive mistake available in the whole set.
Question 7 appeared in weeks 6, 7 and 8 and never before. Three consecutive, after five consecutive absences, immediately following a register correction in week 5. That is a real change and it is attributable.
Question 15 appeared on Perplexity every week and never on ChatGPT. Not variance at all. Perplexity was citing a source ChatGPT was not reading, and the fix was finding the equivalent source on the other engine.
The rule that falls out. Three consecutive appearances after a run of absences is a signal. One appearance anywhere is not. Consistent presence on one engine and absence on another is a source question, not a variance question.
See which questions name you
A free AI visibility audit on your own buying questions, across all five answer engines.
Book your audit
Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn