Ten to thirty. Fewer and one prompt's day-to-day variance dominates your whole picture. More and nobody reads the actual answers, which is where the useful information is.
That second failure is the common one and it does not look like a failure. A sixty-prompt tracking set produces a tidy average that goes up and down, and an average is exactly the thing that hides which specific competitor is taking which specific question.
Why the number is smaller than vendors suggest
Tools price by prompt volume, so more prompts is the upgrade path. That is a reasonable business model and a poor guide to what you should measure.
The value in AI visibility tracking is not the score. It is reading the answer, seeing that a named competitor appears in the same four questions every time, and noticing that all four are about pricing. That observation is invisible in an average and obvious in a set of ten.
How to choose the ten
Start from questions, not keywords. People ask assistants long, specific things with their situation attached. "Best commercial plumber Leeds" is a search query. "Who can fix a commercial hot water system on a Sunday in Leeds" is a prompt.
Cover the buying moment, not the research phase. Someone asking what AEO means is not close to buying. Someone asking how to choose an agency is. Weight the set toward the second.
Include the ones you would lose. The temptation is to track questions you might win. The informative ones are the questions where a competitor is strong, because that is where the answer tells you something you did not know.
Include two or three with your brand name. Not for visibility, but to see what the models say about you when they do mention you. They will cheerfully repeat a three-year-old price or a service you dropped.
Fix the set. Changing prompts mid-period changes the result without changing anything real. New prompts start a new period.
Two columns, not one
For every prompt, record two things separately: was your business named in the answer, and did your domain appear as a cited source.
These have different causes and need opposite work. A single blended score shows a middling number and points at the wrong fix, because the two halves move for different reasons.
How often
Fortnightly for the measurement. Weekly shows you variance rather than progress and is a reliable way to talk yourself out of an approach that is working.
Several runs per prompt, spread across days, not several in one sitting. Same-session repeats share context and correlate, which makes them look more consistent than they are.
Always logged out. In your normal account the assistant has seen you discuss your own business and will bring it up. That is memory, not visibility, and it is the biggest single source of false confidence in this field.
When thirty is right
Larger sets earn their place when you have genuinely distinct segments: several locations, several products with different buyers, or several countries. Then it is three sets of ten rather than one set of thirty, reported separately.
Reported as one average, thirty prompts across three segments tells you nothing about any of them.
The set we would build for a small business
Ten prompts:
- Four on choosing: how to choose a provider, what to look for, what it costs, red flags.
- Three on the specific job your best customers arrive with.
- Two on your locality or segment, phrased as someone would say it out loud.
- One on your brand name, to see what is being said.
That is enough to see a pattern, few enough to read every answer, and it takes about half an hour two weeks to run by hand.
What to do when the ten disagree with each other
The most common surprise once someone starts tracking properly is not a low number. It is inconsistency: named on four questions, absent on six, with no obvious pattern.
That inconsistency is information, and reading it correctly saves months.
Named on narrow questions, absent on broad ones. Normal, and the healthiest shape to have. Broad questions have more competitors and older, deeper sources behind them. Keep working the narrow ones and the broad ones follow.
Named on broad questions, absent on narrow ones. Unusual and worth investigating. It generally means the models know your brand but not what you specifically do, which is a positioning problem showing up in the measurement rather than a visibility one.
Named on one engine, absent on the others. Extremely common and mostly about which sources each engine leans on. Check what was cited on the engine where you appear, then look for the equivalent source on the ones where you do not.
Named one week, absent the next, on the same question. Variance rather than a change. This is exactly why a small tracked set checked repeatedly beats a large set checked once.
What a tracked set looks like at scale
The figures above are why the research on this page exists. We run this measurement continuously across every account rather than sampling it, which is the only way to reach a number like 82,000 measured answers and the only way to see a change in the week it happens rather than the month after.
A worked example of building the ten
A commercial insurance broker, choosing their tracked set from scratch.
Where the ten came from. Not keyword research. The last fifty inbound inquiries, read for the question underneath them. Six distinct questions covered forty of the fifty.
The six they kept. "Best insurance broker for a small business", "who can arrange professional indemnity for a consultant", "how much does public liability cost for a trade business", "insurance broker that handles claims for me", "do I need cyber insurance for a small company", "broker for a business with a previous claim".
The four they added. Two competitor-shaped ("alternatives to [named competitor]"), and two they wanted to win rather than already won, on lines they were trying to grow.
What the set showed in week one. Named on one of ten. Cited on four. Two competitors named on seven.
Why ten was the right number. All ten answers were read in full, by a person, every month. That is the actual constraint. A sixty-question set produces a percentage nobody has read behind, and the percentage is the least informative part of the whole exercise.
See which questions name you
A free AI visibility audit on your own buying questions, across all five answer engines.
Book your audit
Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn