AEOSearch

Blog · 2026-08-22 · 5 min read

Which model answered? That decides the score

ChatGPT is not one thing. It is a family of models with measurably different search habits, and a visibility score that will not name the model behind it is not a measurement. It is a vibe with a decimal point.

Here is the detail most AI-visibility dashboards would rather you did not ask about: "ChatGPT visibility" is not a measurement of one system. OpenAI serves several models under that one name, and when we probed them on grounded buying questions they behaved like different instruments entirely. gpt-4o-mini issued exactly one web search per question, every single time. The frontier model issued three to six, and chose the number itself. Different retrieval means different pages read, which means different brands named: same engine label, different answer, different score.

So a score labelled "ChatGPT" with no model named is unfalsifiable. You cannot reproduce it, you cannot compare this month to last month, and you cannot know whether the change in your number was the market moving or the vendor quietly swapping models under the label. We have written up what that swap costs in dollars; this post is about what it costs in truth.

What we pin, and where it is allowed to move

Our answer models are pinned per tier, in code, and named on the methodology page. The paid tiers answer on the frontier models, gpt-5.6-sol, claude-sonnet-5 and gemini-3.6-flash, because measuring "what ChatGPT says" with a model no ChatGPT user is actually served is a validity problem, and validity is the product you are buying. The free scorecard answers ChatGPT on its fast model and says so in writing; Gemini answers on 3.6 Flash for everyone, free included.

Changing an answer model here is deliberately expensive. It is a code change gated by a test that fails if any pinned model drifts, and it is a copy change, because the pricing page and the methodology page name what each tier answers on. A model an operator could swap in a settings screen would eventually make a sentence we sell false, silently, and we removed that possibility on purpose. It is the kind of decision that costs us flexibility and buys you a number you can bet on.

And wherever the provider reports it, the stored answer carries the model string it actually answered on, read back from the raw response rather than forward from our intentions. When a report's methodology section says what was measured, that is where it reads it from.

Comparability is the whole game

A visibility score has exactly one job: to be comparable. To last month, to your competitor, to the version of you that did the work. Every uncontrolled variable erodes that. Grounding is one, and we run search-on with the flag stored per answer. Sampling is another, three per question with the disagreement reported as volatility. The model is the third, and it is the one this category talks about least, because naming it is inconvenient.

The test is one question: "which model produced this number, and where is that written down?" If the answer is a shrug, the trend line on the dashboard is comparing unknowns to unknowns, and a trend line of unknowns is how marketing budgets get spent on ghosts. Ours is written on the methodology page and inside every stored answer. Start with the free scorecard and you will see the model named before you see the score.

— The AEOSearch team

Start with the free scorecard.

Your buyers are already asking the engines. Get your grounded baseline.

$0 to start · no card required · every number traces to a stored answer