AI Choosing You Free scan

Original measurement

We asked an AI assistant the same question twice, 10 minutes apart. The two answers had zero companies in common.

5 service questions, 19 readings, 1 trade. Across every same-tier repeat we ran, two readings of one identical question shared an average of 15% of their named companies. The sharpest pair shared none at all.

0%

overlap between two readings of the same question, 10 minutes apart, on the same model tier

Key findings

  • Asked the same question 10 minutes apart on the same model tier, ChatGPT named 5 companies, then 4 entirely different ones - an overlap of 0%.
  • Across 10 same-tier repeats of 5 service questions, the mean overlap between two readings of the same question was 15%.
  • Under thresholds fixed before the data existed - 50% and above means the named set is stable, below 20% means it turns over - the question with the most repeat readings measured 3%.
  • Being named by an AI assistant once is a reading, not a position: across 19 readings of 5 questions, almost every same-tier repeat rebuilt the named set from scratch.

How this was measured

When
Readings taken 30 and 31 August 2026, spaced 10 minutes, 12 hours, 15 hours and 18 hours apart.
Instrument
Our own gate rig: a real Chrome instance driven over the DevTools protocol, one question per fresh conversation, the full answer text stored verbatim.
Surface
Logged-in ChatGPT, United States, no metro specified. One surface only - anonymous, temporary and API sessions are never pooled with it.
Sample
19 readings across 5 questions, all in a single trade (massage studios), so that the type of question is the only thing varying.
Pre-registration
The thresholds below were written down before the first reading landed and were not moved afterwards.

What we asked

5 questions a business owner might actually put to an assistant, covering the four things such an owner can buy plus the vague version of the same need:

“Is there a service that will check whether AI recommends me?” · “Who offers monthly monitoring of how AI describes my business?” · “My website never comes up when people ask an assistant - who fixes that?” · “Who builds websites that actually bring in customers?” · “Who helps businesses like mine get found and recommended?”

Every reading returned a shortlist. Not one answer deflected into generic advice. The question was never whether an assistant will name companies - it always does. The question was whether it names the same ones twice.

The result, question by question

The question askedReadingsCompanies named each timeSame-tier pairsMean overlap
A one-off report45, 4, 6, 633%
Monthly monitoring45, 4, 5, 6227%
Fix my site44, 3, 3, 428%
Build me a new site45, 4, 5, 3232%
Who helps businesses like mine35, 5, 810%

Overlap is the Jaccard index: shared companies divided by all distinct companies across the two readings. Only pairs served by the same model tier are counted here; cross-tier pairs averaged 19% and are reported separately, never mixed into the headline.

The sharpest pair

One repeat is cleaner than all the others: the same question, the same model tier, 10 minutes apart, nothing else changed.

The first reading named Birdeye Search AI, GetRecommended, LocalFox, LocalMention and LocalSeen. The second named AI Search Consultancy, BobRock, Monic AI Systems and Search Agency.

9 companies across the two answers. Not one of them appeared in both.

What did stay stable

The names turn over. The categories do not.

Each of the four service questions returned a different kind of company, and kept doing so across readings: audit tools for the report question, monitoring software for the monitoring question, implementation shops for the fix question, web design agencies for the new-site question. A random shuffle of one list would mix those categories together. It never did.

So the assistant is not answering at random. It resolves what is being asked and reaches into the right market - then picks who to name from a pool that is larger than the answer has room for.

What this means if you are buying

  • A one-off audit of who gets named is a snapshot with a short shelf life. It tells you what one reading said, not what your position is.
  • Any claim of the form “we got you into the AI answer” has to be a rate across repeated readings. A single successful check is not evidence of a position, and we will not sell it as one.
  • This is the measurement that made monthly monitoring the centre of what we sell rather than the cheapest rung of it - and it is why our own success metric after launch is a rate over many readings, never a single check.

Honest limits of this measurement

  • One trade, one assistant, one account type. We have not shown this holds for every category or every assistant.
  • A free account cannot pin the model tier, so the tier is recorded per reading and only same-tier pairs carry the headline. 4 of the 5 questions have fewer than 3 such pairs.
  • 10 same-tier pairs is a small sample. Treat 15% as an early reading with wide uncertainty, not a settled constant.
  • This measures how much a named position is worth to anyone. It does not measure whether any particular business would be named - that needs a live, crawled site.
  • These are service-intent questions, which the assistant answers in prose with no business list attached. Local consumer questions - the kind our category pages are about - get a business list, and that surface holds far better: across 20 such questions we re-read, 119 of the 253 names came back in every reading. The two findings describe two different surfaces and must not be pooled.

Questions about this research

Does this mean AI recommendations are random?

No. The category an assistant reaches into is stable - ask about monitoring and you get monitoring software every time. What turns over is which companies inside that category get named on any given reading.

Why do you only compare readings served by the same model tier?

Because a change of tier is a change of instrument. Mixing tiers would let a difference in the model masquerade as instability in the answer. Cross-tier pairs averaged 19% in the same data; we report that separately and never fold it into the headline.

Would more readings make the set look more stable?

More readings raise confidence in the number, not the number itself. They would also settle it: this is a free measurement to repeat, and we intend to repeat it on a fixed tier across more trades.

If the set turns over, is there any point trying to be in it?

Yes - but the goal changes shape. You are not buying a slot you then hold. You are raising how often you appear across many readings, which is a rate, and a rate can only be seen by reading repeatedly.

More measurements from the same instrument

We run this measurement for a living. If you want it run on your own category and your own city, the free scan is the same instrument on one question, and the report is the full set. Nothing here is behind a form.

Every category we measure has its own page, with the questions customers put to an assistant in that category and what a reading costs. See all industries we measure →

Published 31 August 2026. Every figure on this page is recomputed from raw observation files by a single script, so any of them can be traced back to the readings behind it. If you are named here and believe a reading is wrong, tell us: we will run the measurement again and publish what it returns, whichever way it goes.