Blog · 2026-08-22 · 6 min read
Search on, or you measured the past
Ask a model with web search off and you learn what it memorized in training, years ago. Ask with search on and you learn what a buyer sees today. Confusing the two is the original sin of AEO measurement, and most of the category is still committing it.
Put one uncomfortable question to any AI-visibility number before you spend a dollar on it: how was the answer generated? Not which model, not how many prompts. Was the engine's live web search on or off when it answered?
It sounds like plumbing. It is the whole measurement. A large language model answering from its training data is reciting a snapshot of the web as it existed months or years ago, compressed and blended into a single memory. A model answering with live search on is doing what it does for a real buyer: retrieving current pages, reading them, and composing an answer with citations. These are two different machines that happen to share a chat box, and only one of them is the machine your customers talk to.
Ungrounded numbers flatter the incumbent
Training data over-represents the past, and the past belongs to whoever was big in it. A brand with fifteen years of accumulated web presence looks strong in an ungrounded answer even if its live visibility is collapsing; a two-year-old challenger doing everything right can look invisible. If your measurement is ungrounded you are not tracking the market. You are tracking history, and paying a subscription for the privilege.
Grounded answers move with the web, and that is the entire reason to measure them. Publish the page, get into the roundup, fix the schema, and a grounded answer can reflect it on the next run. That is what makes grounded measurement worth money: it responds to the work you do. An ungrounded score barely can, because the training data will not update for months, if it ever does.
This is why every AEOSearch question runs with live web search on, and why every stored answer carries a grounded flag. When an engine cannot ground, the report says so. The flag exists precisely so that a number is never presented as something it is not, and the same discipline applies to our own failures.
The second rule: sample, because the machine changes its mind
Grounding is necessary and not sufficient. AI answers are probabilistic: the same grounded question, asked twice, can name different brands in a different order. A single sample is an anecdote wearing a percentage.
The honest fix is boring and mechanical, which is why we like it: ask every question three times per engine, store every answer, and report the disagreement instead of averaging it away. We surface that as volatility. High volatility on a question is not noise to hide; it is the best news in the report. It tells you the engine has not settled on an answer, which usually means the category is still winnable, and winnable categories are where early money earns the most.
Sampling also keeps you from fooling yourself in the happy direction. One run where the engine names you first feels like victory; three runs where you appear once feels like what it is, a coin toss you are starting to influence.
Questions to ask any vendor, including us
Was web search on? Per answer, provably? Can I read the raw answers my score came from? How many samples per question? What happens when the samples disagree? If any answer is a shrug, the number on the dashboard is a vibe with a decimal point, and you should treat it the way you would treat a sales forecast from someone who will not show their pipeline.
Our answers, for the record: always on, flagged per answer; yes, every number traces to a stored answer and the report shows you one verbatim per question; three per engine; disagreement is reported as volatility, never averaged away. Measurement you would bet budget on has to survive its own methodology page, and ours is public. Let's measure you properly: the scorecard is free.
— The AEOSearch team