The category is three years old, so nobody has a long track record and the usual signals do not work. Years in business tells you nothing. Client logos tell you who could afford them. There is not yet an award worth weighing.
What is left is the method, and the method is checkable if you know what to ask.
These are the eight things worth walking away from, one thing that looks like a red flag and is not, and the four questions that sort the field fastest.
The measurement problem underneath all eight
We measured 608 AI answers to the questions buyers ask when hiring in this category, across five engines over three months, then traced every answer back to the sources cited to produce it.
“We've always grown through referrals - builders who know us pass our name on. That works, but it only reaches people who already know someone in the industry. Within a couple of weeks of the content going live, we had someone contact us directly through the website. That's a channel we didn't have before - and the enquiries have kept coming.
Two facts do most of the work below. The sources that decide who gets named are largely not anybody's own website. And the category is concentrated enough that a plan aimed at the wrong half of the problem produces nothing at all, rather than producing a smaller version of the result.
Almost every red flag in this list is a symptom of one of those two being ignored.
1. They guarantee you will rank in ChatGPT
There is no ranking. An answer names two or three businesses and everyone else is absent, rather than sitting further down a page nobody scrolls. Anyone promising a position is describing Google, or has not looked closely at what they are selling.
Ask instead: what will you track, and where do I appear on those questions today?
What a credible answer sounds like: here are the twenty questions we will track, here is your baseline on each across five engines, here is what we expect to move first and roughly when. Concrete, checkable, and free of the word "guarantee".
2. The reporting is a traffic chart
AI search frequently produces no click at all. Somebody hears your name in an answer, remembers it, and either acts later or searches for you directly a week afterwards. A dashboard counting sessions shows this as nothing happening.
This is the single most common failure in the category and it is not always dishonest. It is often just an SEO reporting stack pointed at a problem it was not built for.
Ask instead: show me a prompt-level report from a real client, redacted if you need to.
An agency that cannot produce one is measuring something else and calling it AI visibility.
3. They treat cited and recommended as one number
Being cited means your page was used as a source. Being recommended means your business was named in the answer. These are different outcomes with different causes, and a business can be read constantly and named almost never.
Across the answers we measured, the set of brands supplying the reading and the set getting the recommendation overlap far less than anyone expects.
If an agency reports one blended visibility figure and cannot break it down, they cannot tell you which problem you have. The two need close to opposite work: one is fixed on your own pages, the other is fixed almost entirely off them.
4. Everything they propose happens on your website
Your own pages influence whether you get retrieved. Other people's pages influence whether you get recommended.
Look at the table above. The most-cited source in the whole category is a community site. Not a vendor, not a publisher, not anybody's carefully built resource hub.
A plan consisting entirely of publishing on your own domain addresses the half of the problem that is easiest to scope and invoice, and leaves the half that actually decides the answer untouched.
Ask instead: which sources outside my site will you work on, and in what order?
5. The case studies have no numbers, or impossible ones
Thin case studies are forgivable in a young category. Vague ones are not. "Significant uplift in AI visibility" means nothing and is chosen precisely because it means nothing.
Be equally wary of the opposite. Zero to ninety percent in three weeks is not a normal outcome. It usually means something trivial was measured, or the client was in a category nobody had worked yet.
What a real number looks like: a starting point, an end point, a period, and the question set it was measured over. Our own case studies are written that way because anything less is not checkable.
6. They will not say whether the tooling is theirs
Several agencies market a proprietary platform that turns out to be a third-party product they license.
Licensing good software is often the right call and completely fine. Not being straight about it is the problem, because it is a small dishonesty about something trivially easy to check, which tells you what larger claims will look like.
Ask instead: is the platform yours or licensed, and either way, can I log into it?
The second half matters more than the first. A tool you cannot see is a report you have to take on trust.
7. Twelve-month term, monthly reporting
The mismatch is the tell. If the first meaningful report lands at day 30 and you cannot leave until day 365, you have eleven months without leverage.
Month to month exists in this market, and so does a fixed-scope initial engagement that proves the approach before anyone commits further. A twelve-month lock is therefore a choice rather than a necessity, which makes it negotiable.
Ask instead: what is the shortest commitment you offer, and what happens at the first review point?
8. Nobody can tell you what month one contains
A good first month is measurement, a source audit and the start of claiming. It is unglamorous and it is the highest-value work in the whole engagement, because everything afterwards is aimed by it.
If month one is a strategy document, ask what happens in month two. If the answer arrives in abstractions rather than in specific sources and specific questions, they have not done this before.
Not a red flag: a short track record
Everyone has one. The field is three years old and an agency claiming a decade of AI search experience is describing something that did not exist.
Judge the method and the measurement, not the founding date. The right question is not how long they have been doing this. It is whether they can show you a prompt-level report and explain, without being prompted, why Reddit outranks every vendor site in their own category.
The four questions that sort it fastest
- Show me a prompt-level report from a real client.
- What is the difference between being cited and being recommended, and which is my problem?
- Which citation sources matter in my category, and why those?
- How will you show me movement in the first month?
Each is answerable in a sentence or two by someone who does this work, and hard to fake by someone who does not.
What it looks like when none of the flags are there
PROCERT is a building certification firm that had grown entirely on referrals. Within a couple of weeks of the work going live, someone who had never met them inquired directly through the website. That channel did not exist before.
Fur Magic went from absent to the most-cited brand in their category in three weeks, ahead of long-established competitors with far more online presence than they had.
Both are written up with the starting number, the end number and the period, so you can apply red flag five to us. PROCERT and Fur Magic.
A worked example of the flags in one meeting
A real shape these conversations take, composited from several.
The agency opens with a case study showing "a 340% increase in AI visibility" with no starting number, no period and no question set. That is flag five: an impossible-sounding figure with nothing behind it to check.
Asked for a prompt-level report, they offer a dashboard showing organic sessions and impressions. That is flag two.
Asked whether they separate cited from named, they say they track "overall AI visibility as a single score". That is flag three, and it is the one that matters most, because it means they cannot tell you which of two opposite jobs you need.
Asked what happens off your website, the proposal contains twenty articles a month and nothing else. That is flag four.
Asked about the term, it is twelve months with the first report at day thirty. That is flag seven.
None of those individually proves anything. Five of them in one hour is a pattern, and the useful thing is that all five surfaced from four questions you can memorise.
Run the checklist on us
We will show you a prompt-level report, explain cited versus recommended before you ask, name the sources that carry weight in your category, and give you a baseline in week one. Month to month.
See which questions name you
A free AI visibility audit on your own buying questions, across all five answer engines.
Book your audit
Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn