Stakes first. When your buyer asks ChatGPT which agency to hire, a version answers: the one their plan serves them, on that day, with whatever search behaviour that version has. No report can measure "ChatGPT" in the abstract, any more than a lab can measure "water" without saying which tap. It measures a model, and the honest thing to do with that fact is to let you choose the model, name it on every answer we store, and price it on its own line. So that's what the page does. Each of the five engines on the configurator opens into a list of the versions we sell, each one with its real name, a one-line note on what it buys over the rung below, and a label saying whether its price is a measured cost or the vendor's list price.

The versions are not interchangeable

Technical information, and the part that surprised us. GPT-6 Astra, the newest OpenAI flagship, issues five searches an answer; GPT-5.6 Sol, the model every paid tier answers on, issued three, four, five and six across four probes; GPT-4o mini, what the free scorecard runs, issues exactly one. Those are different instruments. A model that searches once reports what it found on the first page it opened; a model that searches five times has read around, and we'd expect the two to name and cite differently, which is the whole reason the version is yours to choose. The cost tracks the searching, which is why the Sol line reads $91 on a 20-question report and the 4o mini line reads $4.

Claude runs the other way. Claude Opus 5 is the bigger model and the cheaper one on our page, because it reads less retrieved context per answer than Claude Sonnet 5, which bills over a hundred thousand tokens a call on the current search tool; the Claude line at 20 questions reads $78 on Sonnet 5 and $36 on Opus 5. Claude Fable 5.1, the newest and largest, sits between them at $68. Gemini's older 2.5 Flash costs more than the current 3.6 Flash, because Google bills grounding by model family and the older family pays the higher fee. Grok 4.6 reads a hundred thousand tokens of search results an answer and costs accordingly; Grok 4.3, what the paid tiers run, costs a quarter of that. Perplexity's Sonar costs under a cent an answer; its Pro and reasoning rungs sit on the page marked not yet available, because our adapter hasn't answered on them and we don't sell a model we've never called.

Which one to pick

Our default, and the shape the page opens on, is the version the public is actually served on each engine as far as we can tell: Sol, Sonnet 5, Gemini 3.6 Flash, Sonar and Grok 4.3, the same rungs every paid tier answers on. That's a validity choice before it's a cost one. Measuring what ChatGPT says about you with a model no ChatGPT user is served would produce a clean number about nothing, and the report would have to say so in its methodology, which it does wherever the model can be recovered.

Two reasons to leave the default. Take the newest flagship, Astra or Fable 5.1 or Grok 4.6, when your buyers are the kind who pay for the newest thing and your category is the kind they argue about: that's where a five-search answer differs most from a one-search one, and where the extra line buys a reading you can't get cheaper. Take the fast rung, 4o mini or Haiku 4.5, when what you want is an affordable first look at a property you've never measured; the free scorecard already runs 4o mini, and a paid report on it at 20 questions is an honest baseline that says on its face which model produced it. What we'd steer you away from is mixing the two in one report and reading the result as one engine. It's two.

An admission: we don't know which version any given consumer is served. The vendors route by plan, by region and by day, and they don't publish the routing. So the model name on our report is the most honest thing available, not the whole truth; it says which instrument produced the number, and it lets the next report use the same one, which is the only way a change between two reports can mean anything. Versions also retire. Our catalogue treats a model id it has never seen as news rather than as a fact, and a report from March stays readable in September because the model that answered is on every row of it.

What the name buys you

Three things. Every stored answer carries the model that gave it, so a number on the scorecard walks back to a version, not to a brand. The methodology section names the model per engine, so a client or a colleague reading the report six months on knows what they're looking at. And the monthly delta only compares like with like: the same questions on the same versions, so the movement it shows is the market moving and not the instrument. Change a version and the report says so and starts its count again, because a delta between two different instruments is a coincidence dressed as a trend.

Open the configurator, switch on an engine, and open its list. Pick the rung your buyers are served, or the one you want to know about, and the price on the rail follows the choice with the same share on every line. If you'd rather we picked, book a call and we'll choose the versions for your market and say why. Either way the report lands with the model named on every answer, and that name is what makes the next report worth reading beside it.

Written by AEOSearch. Vendor details and research findings reflect the article’s publication date.

Read Our Method ↗