An AEO tool tells you which AI answers name your business and which name someone else. That is the whole category, and it is genuinely useful. What tools do not do is the part that changes the answer, and the gap between those two things is where most small business budgets get wasted.
What every tool in this category actually does
Strip the marketing and they all do the same three things.
Run prompts on a schedule. You give it questions your buyers ask. It asks them repeatedly across ChatGPT, Perplexity, Google AI Overviews and usually Gemini, and records what came back.
Count mentions. Whether your brand appeared, whether a competitor did, sometimes how favorably.
Chart it over time. So you can tell a trend from a bad Tuesday.
The differences between tools are which engines they cover, how many prompts you get, whether they separate being cited from being named, and how much the reporting helps you decide what to do. That last one varies enormously and is worth testing before you commit.
The one feature that matters most
Ask whether the tool separates cited from mentioned.
Being cited means your page was used as a source. Being mentioned means your business was named in the answer. These have different causes and need close to opposite work, and a business can be read constantly and named almost never.
A tool that reports one blended visibility figure cannot tell you which situation you are in, which means it cannot tell you whether to write more or to go and get mentioned elsewhere. That is most of the decision.
What no tool does
Claim your listings. The highest-value work in this field is usually completing the citation sources your category runs on, and every one requires a human with your business details and, often, a verification phone call.
Write anything. Some bolt on generation. What comes out is generic, which is the specific failure mode in a field where the point is to say something a model finds worth quoting.
Get you mentioned by other people. The sources that decide who gets named are mostly not yours. In the category we measure, Reddit was cited 225 times across 608 answers, ahead of every agency website in the field. No subscription gets you into a Reddit thread.
Decide what to do first. The output is a list of prompts you are missing from. The useful question is which three matter, and answering that is judgment rather than data.
When a tool is the right purchase
To find out if this matters at all. One month of monitoring tells you whether your buyers' questions produce answers that name anyone in your category. If nobody is named, including competitors, the category is not there yet. That is a legitimate finding and cheap to obtain.
When you have someone who will act on it. A tool with nobody behind it is a subscription to bad news.
When you are already doing the work. If you are claiming listings and writing pages, monitoring tells you whether it is working, which is worth a lot.
When it is the wrong purchase
When you are buying it instead of doing something. This is common and understandable: monitoring feels like progress, produces a dashboard, and the dashboard is evidence of activity. Six months later the chart is flat and it has cost you a few thousand and the time you spent looking at it.
If you do not have someone who will act on what it says, spend the money on a one-off audit instead. It will tell you what to fix once, and then you can decide.
The free version
Before buying anything, do this by hand. Ten prompts, three engines, two columns: named, and cited. Repeat weekly for a month in a logged-out session.
It takes about twenty minutes a week and gives you most of what a tool gives you at this scale. The point at which manual checking becomes the reason you stop is the point a tool earns its subscription. For a small business tracking ten prompts, that point is further away than the vendors suggest.
How to compare two tools in twenty minutes
Feature lists in this category are close to identical and mostly unhelpful. Three checks separate them.
1. Ask both to show you one prompt, in full. Not a score. The actual answer text the engine returned, with the sources it cited. A tool that stores the full answer lets you see why you were left out. A tool that stores only a yes or no gives you a number and no diagnosis.
2. Check whether cited and named are separate fields. Being used as a source and being recommended are different outcomes with different fixes. A tool that blends them into one visibility score cannot tell you which problem you have, which is the only question worth asking at the start.
3. Count the engines and the frequency. Five engines checked weekly beats two engines checked daily, because these systems disagree with each other far more than they disagree with themselves. Ask specifically whether checks run logged out.
Anything beyond those three is a preference. Those three decide whether the subscription produces a diagnosis or a mood.
What happens after the diagnosis
“We went from nobody finding us on AI search to being recommended in almost half of all relevant queries. I didn't think that was possible without months of work and a huge budget.
PROCERT is a building certification firm that had grown entirely on referrals. Within a couple of weeks of the work going live, someone who had never met them inquired directly through the website. Fur Magic went from absent to the most-cited brand in their category in three weeks, ahead of long-established competitors.
Both published with the starting number, the end number and the period: PROCERT and Fur Magic.
A worked example of choosing
A small business trialling two tools side by side for one month.
Tool A. Reported a single visibility score of 12%, tracked across two engines, refreshed daily. Clean interface, useful trend line.
Tool B. Reported named on 3 of 25 and cited on 9 of 25, tracked across five engines, refreshed weekly, and stored the full answer text with its citations.
What Tool A could not answer. Which of the two problems they had. A 12% blended score is consistent with being read constantly and named never, and with being named occasionally and read never, and those need opposite work.
What Tool B showed in ten minutes. Reading the stored answers, their own domain appeared in the citations of nine answers that then recommended a competitor. That is the diagnosis, and it changed the entire plan.
On refresh frequency. Daily on two engines sounded better and was worse. These systems disagree with each other far more than they disagree with themselves, so engine coverage beats refresh rate.
What they kept. Tool B, at a slightly higher price, for the stored answers alone.
See which questions name you
A free AI visibility audit on your own buying questions, across all five answer engines.
Book your audit
Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn