Put one uncomfortable question to any AI-visibility number before you spend a dollar on it: how was the answer generated? Not which model, not how many prompts. Was the engine's live web search on or off when it answered?

It sounds like plumbing. It is the whole measurement. A large language model answering from its training data is reciting a snapshot of the web as it existed months or years ago, compressed and blended into a single memory. A model answering with live search on is doing what it does for a real buyer: retrieving current pages, reading them, and composing an answer with citations. These are two different machines that happen to share a chat box, and only one of them is the machine your customers talk to.

Not-grounded numbers flatter the incumbent

Training data over-represents the past, and the past belongs to whoever was big in it. A brand with fifteen years of accumulated web presence can look strong in a not-grounded answer even if its live visibility is collapsing; a two-year-old challenger doing everything right can look invisible. If your measurement is not grounded, you may be tracking history instead of the market, and paying a subscription for the privilege.

Grounded answers move with the web, and that is the entire reason to measure them. Publish the page, get into the roundup, fix the schema, and a grounded answer can reflect it on the next run. That is what makes grounded measurement worth money: it responds to the work you do. A not-grounded score may barely respond, because the training data will not update for months, if it ever does.

This is why every question AEOSearch puts to an assistant's API goes with live web search available, and why each answer we store now records whether the model actually searched. The model decides whether to use it, the way its app does, so the report says, engine by engine, whether its answers searched, with the count where only some did. The record exists precisely so that a number is never presented as something it is not, and the same discipline applies to our own failures.

The second rule: sample, because the machine changes its mind

Grounding is necessary and not sufficient. AI answers are probabilistic: the same grounded question, asked twice, can name different brands in a different order. A single sample is an anecdote wearing a percentage.

The honest fix is boring and mechanical, which is why we like it: ask the same question more than once per engine, store every answer, and report the disagreement instead of averaging it away. We surface that as volatility. On an AEOSearch report you choose one pass or three per question before you pay; the free scorecard asks once, so volatility is only measured where three passes ran. High volatility on a question is not noise to hide; it is the best news in the report. It tells you the engine has not settled on an answer, which usually means the category is still winnable, and winnable categories are where early money earns the most.

Sampling also keeps you from fooling yourself in the happy direction. One run where the engine names you first feels like victory; three runs where you appear once feels like what it is, a coin toss you are starting to influence.

Questions to ask any vendor, including us

Was web search on? Per answer, provably? Can I read the raw answers my score came from? How many samples per question? What happens when the samples disagree? If any answer is a shrug, the number on the dashboard is a vibe with a decimal point, and you should treat it the way you would treat a sales forecast from someone who will not show their pipeline.

Our answers, for the record: search available on every API call, and each answer now records whether the model used it; yes, every number traces to a stored answer and the paid report shows you one verbatim per question; one pass or three per engine, your choice before you pay; disagreement is reported as volatility, never averaged away; and a call that fails is stored as a failure, left out of every rate and re-asked, and a paid report waits for it: if it never answers, its share is refunded or credited. Measurement you would bet budget on has to survive its own methodology page, and ours is public. Let's measure you properly: the scorecard is free.

Written by AEOSearch. Vendor details and research findings reflect the article’s publication date.

Read Our Method ↗