A number on a report is worth exactly as much as the path back from it. Ours have one: every figure on an AEOSearch report can be walked back to a stored answer you can open and read, with the model that gave it and the date it was given. The methodology page states the rules; this post walks the path in the order the report is built, so you can see what each step does, what it costs, and where it stops.

Step one: the questions, which are yours

The report starts from your brand, its aliases and domains, up to four competitors with theirs, your category and your buyers. From that, a Claude model drafts the buyer questions, 20, 50 or 100 of them, spread across four kinds: best-of questions, comparisons, problem-led questions where nobody has named a brand yet, and category explainers. Then it stops. The set is on the page for you to edit, cut or add to, and no engine is asked anything until you approve it. We built that pause on purpose, and it costs us a sale now and then from a buyer who wanted a button: the questions are the instrument, you know your market better than a model drafting from a form does, and a report is only as good as what it asked.

The drafting and the later diagnosis are the $2 line on the configurator. They're the one part of the price that doesn't scale with the engines, and the line says what it is.

Step two: the asking

Each approved question goes to each engine you switched on, on the model you chose, with live web search or grounding on, three times, or once if you took the single pass. Every answer is stored word for word, with the model that gave it, the timestamp, and whether grounding was on. Grounding matters more than anything else in this step: an answer given without search measures what a model remembers from training, which is the past, and the report is about what your buyer is told today. Search on, or you measured the past is the argument in full. Three passes matter because an AI answer is a draw, not a lookup; the same engine asked the same question three times can name you twice and skip you once, and three is the smallest count that can show that disagreement rather than hide it.

A call that fails is stored as a failure, not as an answer. The run keeps re-asking it inside a thirty-minute window, and what still hasn't answered when the window closes is excluded from every metric, disclosed on the report as a failure, and taken off a one-time bill on its own. An outage on our side or a vendor's is never printed as your invisibility.

Step three: the reading

A machine reads every stored answer once, against a fixed schema, and pulls out the brands it mentioned, in the order they appeared, with a sentiment and the quote that carried each one, and the domains it cited. That's a Claude model with structured output, one call per answer, and the few cents it costs ride on the per-question rate. Matching is done against the names and aliases you listed, case-insensitively; a product name you didn't list as an alias isn't credited to you, and a brand name that is also an ordinary word is disambiguated by your category, so a fruit doesn't count for a phone maker. We keep it that literal on purpose. A matcher that guessed generously would flatter every report a little, and a report that flatters isn't a measurement.

Step four: the counting

The metrics are arithmetic over those extractions, and they're the same arithmetic on every report. Mention rate: the share of answers that named you. Share of AI voice: your mentions against everyone's. Two scoreboards for recommendation, because they answer different questions: how often you were recommended among the answers that named you, and how often among all answers. Citation share: how often your pages were the ones the engine linked. Volatility: how much the three passes disagreed. And the unprompted split, which is the number an owner should read first: a question that names you or a tracked competitor leads the engine, so the unprompted figures are computed only over questions that name no tracked brand at all, and the report says which are which.

From the same rows the report derives the source map, every cited domain with a count and a label saying whether it's yours, a competitor's or a third party's; the recognition ladder, your mention rate by how specific the question was, from branded down to category explainer, with the deepest rung you still appear on marked as where recognition ends; and a buyer-phase split read straight off the question type. For each question you lost, we fetch the pages the engine cited over yours and a Claude model classifies why: content, format, authority, freshness, or schema and entity. That's the prioritized fix list in the base report. The 90-day action plan is a separate call over the finished run, and its own line.

What this method cannot show

Stated plainly, because a method that hides its edges isn't one. Three samples is a small number, and a single question flipping on a single engine can be the market or can be the machine changing its mind; the volatility figure is how we show that rather than bury it, and a monthly delta on 20 questions is noisier than any of us would like. Matching is only as complete as your alias list. A grounded answer is today's answer; next month's is a new measurement, not a correction. We can't see which version of an engine any particular buyer is served, so the model named on the report is the most honest label available and not the whole truth. And nobody, us included, can attribute one action to one result inside an engine that publishes neither its retrieval nor its ranking; a report can show that a number moved after you did something, and it cannot prove the something moved it.

The incentives, stated as well: we profit from the conclusion that this is worth measuring, every step above is also a sales argument, and the way we've chosen to handle that is to put every claim next to the stored row that makes it true. Open any question on your report and the answer is there, in full, with the model beside it. If we have described the method wrongly, the rows will show it, and we'd rather be caught by an answer you can read than believed on a paragraph.

Read one yourself

Build a report in the configurator, approve the questions, and the engines start within a minute. When it lands, don't read the scorecard first. Open the question table, pick the question you most wanted to be named on, and read the three answers the engine actually gave, with the model and the date on each. That's the product. The scorecard is just the count. If you'd like the fix list turned into work, we do it at our agency rates and the next report grades us on the same rows; either way, the proof is something you can read, not something you're asked to believe.

Written by AEOSearch. Vendor details and research findings reflect the article’s publication date.

Read Our Method ↗