Most advice on hiring in this category is a list of warnings. Warnings tell you who to eliminate. They do not tell you who to choose.
These are the six things worth actively looking for, and the one question that reveals most about whether an agency has actually done this work before.
The data that should shape what you look for
We measured 608 AI answers to the questions buyers ask when hiring in this category, across five engines over three months, and traced each answer back to the sources cited to produce it.
“We've always grown through referrals - builders who know us pass our name on. That works, but it only reaches people who already know someone in the industry. Within a couple of weeks of the content going live, we had someone contact us directly through the website. That's a channel we didn't have before - and the enquiries have kept coming.
Two conclusions follow, and every item below is a way of testing for them.
The category is concentrated: five brands take four namings in ten, and half the field sits below 1.5% visibility. Competence alone does not get you in.
And the sources that decide the answer are mostly not your website. That single fact should reshape what you expect a proposal to contain.
1. They measure two numbers, not one
Being cited as a source and being named in the answer are different outcomes with different causes.
An agency that tracks them separately can tell you which problem you have. One reporting a blended visibility score cannot, and the two need close to opposite work. If you are cited but not named, the fix lives off your website almost entirely. If you are neither, the fix starts on it.
Across the answers we measured, the brands supplying the reading and the brands getting the recommendation overlap far less than anyone expects. Plenty of companies are read constantly and named almost never, and a single blended figure would have pointed every one of them at the wrong fix.
This is the clearest signal that someone has looked at their own data rather than repeating the category's vocabulary.
2. They talk about sources you do not own
Ask which sources matter in your category. A good answer is specific and mostly external: an insurer directory, a licensing register, a trade body, a particular community, a review aggregate that carries unusual weight in your vertical.
We hold citation-source data across 27 industries, scored by how much weight each source carries when an engine builds an answer. The high-weight sources are different in every single one. In home services it is review aggregates. In healthcare it is insurer directories and licensing bodies. In finance it is regulator registers and comparison sites.
An agency whose plan is entirely on-site is addressing the half of the problem that is easiest to scope and invoice.
Ask them to name three sources in your category that carry more weight than your own website, and why. A specific, slightly surprising answer means they have looked. A generic one, Google Business Profile and "the major directories", means they have not.
3. They publish something checkable
Original data, a stated method, a resource that gets quoted by other people.
Not because publishing proves competence, but because it is the mechanism they are selling. Getting named inside other people's content is exactly what a citable resource does, and an agency that has not managed it for itself is describing a technique rather than running one.
The right question is not whether their blog is busy. It is whether anything they have published is the kind of thing another writer would need to cite.
4. Their comparison content is sourced
If they write about competitors, ask where the facts come from. The right answer is the competitor's own published material with a date attached, not a third-party estimate.
This matters well beyond honesty. An agency careless about a checkable claim regarding a named company will be careless about a claim regarding yours, and it is your name on the page.
It is worth knowing how loose this gets. Published pricing in this category is often inconsistent even on a single agency's own site, which is a good reason to confirm any number directly rather than trusting a comparison article, including ours.
5. Month one is boring
A good first month is measurement, a source audit, and the start of claiming. Unglamorous, verifiable, and the highest-value work in the engagement because everything afterwards is aimed by it.
If month one is a strategy document, ask what happens in month two. Strategy is real work and it is a small fraction of what needs doing.
The specific thing to listen for is a baseline. Without one, nothing later can be shown to have moved, and you will be asked to accept a story instead of a measurement.
6. You can see the same screen they see
Ask whether you get a login, not a report.
A monthly deck is a curated view assembled by the person whose work it describes. A live portal is the same view they work from. The difference matters most in the months where progress is uneven, which is most of them.
What good pricing looks like
Not cheap, and not opaque. A written scope, a monthly figure, and terms that let you leave.
Published retainers in this category run from around $2,000 a month at the entry end to $20,000 and beyond for enterprise programs. The figure matters less than what sits inside it: whether third-party source work is included, whether reporting is per prompt, and how long you are committed for.
The signal that outranks the rest
Ask what they would do in your first ninety days, specifically. Not the methodology. The actual actions.
Someone who does this work answers in concrete nouns: these sources, these prompts, these three pages, this outreach, in this order, because of this. Someone who only describes it answers in abstractions.
The difference is audible within a minute, and it is the most reliable test available to a buyer who is not an expert.
What it looks like when it is working
PROCERT had grown entirely on referrals, which works but only reaches people who already know someone in the industry. Within a couple of weeks of the work going live, a stranger inquired directly through the website, and the inquiries kept coming.
Fur Magic went from absent to the most-cited brand in their category in three weeks, ahead of long-established competitors with far more online presence.
Both are published with the starting number, the end number and the period. PROCERT and Fur Magic.
Comparing two good agencies, worked through
The hard case is not spotting a bad agency. It is choosing between two that both answer well.
Ask both for their ninety-day plan in writing, then compare on three axes.
Specificity of sources. Agency A: "claim and correct your top-priority citation sources." Agency B: "your state contractor license record first, because it lists a superseded business name, then the two review platforms your category runs on, then your association listing." B has done the work before the meeting.
Where the effort sits. Agency A: 80% content, 20% everything else. Agency B: 40% claiming and correction in the first six weeks, then content, with one external placement a month throughout. B's split matches where the citations actually come from.
What they promise to show you. Agency A: a monthly report. Agency B: a login, with a baseline in week one. B removes the question of what was left out.
If both are still level after that, ask each what they would do if the named figure has not moved by month four. The answer that includes a specific change of approach beats the answer that asks for more time.
Ask us for the ninety days
We will give you the specific version: the sources, the prompts, the pages and the order, with your own baseline attached rather than a generic plan.
See your ninety-day plan
A free AI visibility audit on your own buying questions, across all five answer engines.
Book your audit
Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn