Most of the risk in doing this yourself is not doing it badly. It is doing it evenly, stopping early, or drawing the wrong conclusion from a flat result. Here are the five failures we see most, in the order they usually happen.
1. Doing all forty things instead of the four
An audit produces a long list. Four items on it move the number and the rest is housekeeping, and from inside your own business the difference is not obvious.
The result is three months of genuine effort spread across work that was never going to matter, followed by a reasonable conclusion that this does not work.
How to tell early: if your plan has more than about eight items on it and none of them is marked as most important, you are about to do this.
2. Claiming the obvious listings and missing the ones that count
Everyone claims Google Business Profile and a couple of general directories. In most categories the high-weight sources are specific and less obvious: an insurer directory in healthcare, a licensing register in trades, an established legal directory, a trade body listing.
Missing those while completing the general ones produces exactly the pattern people describe as "we did the listings and nothing happened".
How to tell early: ask an assistant which sources it would use to check whether a business in your category is reputable. It will tell you, and the list is usually not the one you claimed.
3. Publishing more when the problem is not your content
This is the expensive one. If the models are already reading your pages and naming someone else, more pages does nothing at all. Publishing harder does not move that number, because the constraint is not how much you have written.
The work that fixes it is off your website entirely, which is why it does not occur to anyone doing the work from inside their own site.
How to tell early: two columns, not one. Track named and cited separately from the first day.
4. Stopping at week twelve
The most common failure and the least discussed. Six hours a month is not a lot, but it is six fragmented hours competing with running a business, and it loses most months.
The damage is not the pause. It is that the work goes stale, the next restart begins from a worse position, and after two restarts most people conclude the approach failed rather than that the scheduling did.
How to tell early: if the work lives in the gaps rather than in a recurring block with something else deliberately dropped, this is coming.
5. Measuring in a logged-in session
Small, embarrassing, and extremely common. If you ask ChatGPT about your category in your normal account, it has seen you talk about your own business and will helpfully bring it up. People conclude they are doing fine.
How to tell early: run the same prompt in a temporary chat. If the answer changes, every measurement you have taken is worthless.
The risks that are overstated
Getting penalized. There is no penalty mechanism in an answer engine. The worst outcome of a badly optimized page is that it does not get retrieved.
Damaging your SEO. Everything that helps here helps there. Clearer pages, consistent listings, real third-party mentions. There is no trade-off to manage.
Being too late. The category is three years old and the distribution is still moving. Late is a real risk in a contested category and mostly imaginary in the rest.
When the risk is worth taking anyway
If you have a named owner with real hours, do it yourself. The five failures above are all avoidable and none of them requires expertise, only sequence: measure with two columns, claim the category-specific sources first, give it six weeks before judging, and put the work somewhere it will not be the thing that slips.
How to de-risk each of the five
None of these require an agency to avoid. They require knowing they exist.
Against doing forty things instead of four: write down your ten buyer questions and your current position on each before you change anything. Every task then has to justify itself against a specific question you are trying to win. Most of the forty cannot.
Against claiming the wrong listings: find the sources by working backwards from the answers rather than forwards from a checklist. Ask an assistant your buyer question, look at what it cited, and claim those. It takes twenty minutes and it beats any generic directory list.
Against publishing when content is not the problem: track cited and named separately from day one. If you are cited and not named, more pages will not help and you have just saved yourself six months.
Against stopping at week twelve: put the monthly re-measurement in the calendar as a recurring event with a fixed short scope. One hour, same ten questions, same conditions. The habit survives what the project does not.
Against measuring logged in: use a temporary or incognito chat every single time, without exception. Your own account has seen you discuss your business and will bring it up, which reads as success and is not.
What the ceiling looks like
“We've always grown through referrals - builders who know us pass our name on. That works, but it only reaches people who already know someone in the industry. Within a couple of weeks of the content going live, we had someone contact us directly through the website. That's a channel we didn't have before - and the enquiries have kept coming.
PROCERT is a building certification firm that had grown entirely on referrals. Within a couple of weeks of the work going live, someone who had never met them inquired directly through the website. Fur Magic went from absent to the most-cited brand in their category in three weeks, ahead of long-established competitors.
Both published with the starting number, the end number and the period: PROCERT and Fur Magic.
A worked example of the fourth risk
Stopping at week twelve is the one that costs most, and it is worth seeing what it looks like.
Weeks one to six went well. Baseline measured. Five sources found and claimed. Two pages published. The citation figure moved from three of twenty to eight of twenty, which is a genuinely good result and visible in the data.
Week seven. A product recall. Everything else stopped, correctly.
Weeks eight to twenty. Nothing. No re-measurement, no further claiming, no external work.
Week twenty-one. Someone asks how the AI search thing is going. Nobody knows, because the last measurement was thirteen weeks ago and the spreadsheet is in someone's downloads folder.
What was actually lost. Not the claiming, which holds. The momentum on the external work, which had not started, and which is the only lever that moves the named figure. The citation gain sat there doing nothing for three months.
What would have prevented it. A recurring calendar entry, one hour, same ten questions, same conditions. Not a project plan. A habit small enough to survive a bad quarter.
See which questions name you
A free AI visibility audit on your own buying questions, across all five answer engines.
Book your audit
Ashur Homa
Built and scaled a digital brand to $100M+ in sales with zero ad spend. Has helped businesses generate millions through AI go-to-market strategy. Leads growth at Omni Eclipse.
Connect on LinkedIn