Most guides to hiring an AEO agency are written as if every agency is doing roughly the same thing at a different price. They are not, and the differences are not the ones the pitch decks talk about.
Here is the process we would use, in the order we would use it, including the free test that tells you what you are actually buying before you speak to anyone.
First, understand what you are up against
We measured 608 AI answers to the questions buyers ask when they are hiring in this category, across five engines over three months. Sixty-six different agencies were named at least once.
That sounds like a crowded, unwinnable field. It is the opposite.
“We've always grown through referrals - builders who know us pass our name on. That works, but it only reaches people who already know someone in the industry. Within a couple of weeks of the content going live, we had someone contact us directly through the website. That's a channel we didn't have before - and the enquiries have kept coming.
Sixty-six brands share 1,302 namings between them. The top five take four namings in every ten. Twenty-six brands, well over a third of everyone who shows up at all, appear in fewer than one answer in a hundred.
A category this concentrated is a category where the gap between invisible and named is much smaller than it looks. It is also one where being average gets you nothing.
What actually decides it
We traced every one of those 608 answers back to the sources cited to produce it. This is the part that should change how you evaluate an agency.
| Source | Answers citing it |
|---|---|
| reddit.com | 225 |
| aeoagency.org | 148 |
| youtube.com | 145 |
| linkedin.com | 134 |
| aeoengine.ai | 130 |
| thatware.co | 125 |
| techradar.com | 83 |
| en.wikipedia.org | 38 |
The most-cited source in the entire category is a community site. Not a vendor. Not a publisher. Reddit is cited 52% more often than the next domain and more often than any agency's own website, including the agencies that get named most.
Sit with that, because it inverts the usual pitch. Almost every proposal you will receive is a plan to publish more content on your own website. Your own website is not what decides the answer.
Being read and being recommended are two different outcomes with two different causes. Publishing on your own site gets you read. Being present in the sources engines trust gets you recommended. Most agencies only sell you the first one.
Step 1: run the free test first
Twenty minutes, before you contact anyone.
Write down ten questions your buyers would genuinely ask. Not keywords. Questions, in the words a customer would use. "Who is the best commercial electrician in Denver." "What software should a mortgage broker use for loan processing." "Which AEO agency should a small ecommerce brand hire."
Run each one in a logged-out or temporary chat, across ChatGPT, Perplexity and Google AI Overviews. Record two columns: were you named, and did your domain appear in the citations.
Three possible outcomes, and each one means something different.
Nobody in your category is named, including competitors. The category is not there yet. Recheck in six months. You are early, which is an advantage, not a problem.
You are cited but not named. Your content is being read and something else is getting the recommendation. This is the most valuable case to fix and the hardest to fix alone, because the work happens off your website.
Neither. Your pages are not being retrieved at all. This is the most fixable situation and the fastest to show movement.
Bring the result to every meeting. It changes the conversation completely: you will know which problem you have, and you will be able to tell within five minutes whether the person opposite recognizes the distinction.
Step 2: shortlist three, not eight
Three is enough to calibrate and few enough to compare properly.
What varies most between agencies is not price. It is whether the reporting measures being named or measures traffic. Three conversations makes that obvious fast.
Places to build a shortlist: ask an assistant, which is both useful and a fair test of the category. Ask peers in your industry. Look at who is publishing something substantive rather than who is advertising.
Step 3: ask these four questions
- Show me a prompt-level report from a real client. Not a traffic chart. A list of specific buying questions, which engine each was asked on, and whether the client was named. If they cannot produce one, they are not measuring the thing you are buying.
- What is the difference between being cited and being recommended, and which is my problem? The answer tells you whether they understand the mechanism or are selling SEO with new vocabulary.
- Which citation sources matter in my category, and why those? The right answer is specific and slightly surprising, because the real answer usually is. We hold citation-source data across 27 industries and it is different in every one.
- How will you show me movement in the first month? Not results in the first month. Movement, from a measured baseline, so you are never guessing.
Step 4: compare scope, not headline price
Ask each for a written scope and a monthly figure, then compare what actually arrives.
A cheaper package with no third-party outreach is not cheaper if outreach is your bottleneck, and by the data above, it usually is. A more expensive one with a twelve-month lock is more expensive than it looks, because you have eleven months without leverage.
Step 5: check the term
Month to month exists in this market. So does a fixed-scope initial engagement ending in a decision point. Both are available, which makes a twelve-month lock a choice rather than a necessity, and therefore negotiable.
What good looks like
The honest test of any agency is what happened to the businesses that hired them. Ours:
PROCERT is a building certification firm that had grown entirely on referrals. Within a couple of weeks of the content going live they took their first direct website inquiry, from someone who had never met them. That is a channel that did not exist before.
Fur Magic went from absent to the most-cited brand in their category in three weeks, ahead of long-established competitors with far more online presence.
Read the full numbers on PROCERT and Fur Magic.
What not to weight too heavily
Years in business. The field is three years old. Nobody has a decade of experience in it and anyone claiming to is describing a different job.
Volume promises. Throughput is an input, not an outcome. "Twenty articles a month" tells you what you will receive, not what will change.
A tool on its own. A subscription that tells you that you are invisible, weekly, is a subscription to bad news. Someone has to do the work.
Where we fit
We do one thing: measure the buying questions your customers actually ask, across ChatGPT, Google AI Overviews, Google AI Mode, Perplexity and Gemini, then do the work that changes who gets named in the reply.
That means a measured baseline in week one, per-prompt reporting in a portal you can open yourself, and work on the third-party sources engines read before they answer, not just on your own site. We work month to month.
A worked example of the whole process
Take a mid-sized commercial cleaning company in Charlotte, twelve staff, most work won through bids and referrals.
The twenty-minute test. They run ten questions logged out: "best commercial cleaner for a Charlotte office", "how much does move-out cleaning cost in Charlotte", "commercial cleaning company that does medical facilities". They are named in none. Their domain appears in the citations for two. That is the cited-but-not-named case, which is the most valuable one to fix and the one nobody diagnoses.
What that changes in the meetings. They now know that publishing more service pages will not help, because their pages are already being read. Two of the three agencies they meet propose exactly that, and the test result is what lets them tell.
The four questions. One agency produces a redacted prompt-level report immediately. One offers a traffic dashboard. One says they can guarantee a ChatGPT ranking, which does not exist.
What they buy. A fixed-scope initial engagement rather than a twelve-month retainer, with the first month spent on their state licensing board record, their trade association listing, and the three review platforms their category runs on. None of that is content work.
Where it lands. The citation figure moves in week five. The named figure moves in month three, on the narrow questions first, exactly as the timeline predicts.
The point of the example is that the twenty minutes at the start changed every decision afterwards.
Start with the test
Run the twenty-minute version yourself this week. If you would rather see the full picture, we will run your real buying questions across all five engines and show you which ones name you, which name a competitor, and which sources decided it.
See which questions name you
A free AI visibility audit on your own buying questions, across all five answer engines.
Book your audit
Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn