AI Choosing You Free scan

Measurement log

Every run behind this site

We publish what we measure, including the runs that returned nothing useful. This is the record of every measurement behind this site: what we asked, on what instrument, how many readings, and what came back. It is a log, not a blog - there is an entry when there is a run. 290 readings across 8 runs so far.

Does an assistant name local businesses to their customers?

27 August 2026 – 30 August 2026 · raw_observations.jsonl

The first question worth asking: when an ordinary person asks an assistant for a local service, does it answer with names, or with advice? We asked across trades and metros, one question per fresh conversation, full answer text stored verbatim.

50readings
26distinct questions
2model tiers recorded

It answers with names, every time, in a short business list. This run is the calibration set every later comparison is measured against, and it is where the pool figures on the category pages come from.

The same question, ten minutes apart

30 August 2026 – 31 August 2026 · service_intent_raw.jsonl

Five questions a business owner might put to an assistant when looking for help - the four things such an owner can buy, plus the vague version of the same need. Asked repeatedly, on one model tier, so repeats could be compared.

19readings
15distinct questions
2model tiers recorded

Every reading returned a shortlist; not one deflected into generic advice. But two readings of one identical question shared an average of 15% of their named companies, and the sharpest pair shared none at all. Being named once is a reading, not a position.

Read the finding →

A fifth pass on the same five questions

31 August 2026 · stability_raw.jsonl

A further repeat of the service-intent set, recorded to a separate file so the earlier readings could not be quietly extended. Same surface, same tier discipline.

10readings
5distinct questions
2model tiers recorded

Pooled with the run above, this is what the stability figures rest on. Both files carry the same declared surface, so they are compared with each other and never with the consumer-question readings.

Read the finding →

What an assistant tells a business owner looking for help

29 August 2026 · raw_buyer_probe.jsonl

Ten questions in the owner's own voice - my dental practice does not show up when people ask ChatGPT, who can help my roofing company, who does AI visibility audits for small local businesses.

10readings
10distinct questions
2model tiers recorded

Nine of the ten produced a named shortlist of vendors, with prices and a best-for column. This is a buying surface, not an advice surface - which is the single most commercially useful thing we have measured.

The same ten questions, framed in Austin

30 August 2026 · raw_buyer_us.jsonl

The probe above inherited its operator's location. We ran it again stating a US metro, to find out whether the answer set is national or local.

10readings
10distinct questions
2model tiers recorded

Almost nothing carried over. The names came back Austin-specific, down to firms with the metro in their brand line.

And again, in Miami

30 August 2026 · raw_buyer_miami.jsonl

A third locale, to test whether the Austin result was localisation or noise.

10readings
10distinct questions
1model tier

Roughly 3% of names overlapped between Austin and Miami, and not one company appeared in all three runs. Inside a single metro the answers concentrate hard on one or two firms. This is a per-metro market as an assistant sees it.

Ten trades, three cities, repeated readings

31 August 2026 – 2 September 2026 · vertical_raw.jsonl

Ten verticals, a consumer question each, in Austin, Miami and Columbus, asked more than once so stability could be computed per trade rather than site-wide.

63readings
30distinct questions
2model tiers recorded

Complete. Every reading used web search, so a page can reach these answers at all. The pools differ by more than six times between trades, which is why the remaining 36 categories are worth running. And the sharpest result: asked in three metros, not one of 19 comparable city pairs shared a single business.

Thirty-six more trades, one city, repeated readings

1 September 2026 – 2 September 2026 · vertical_wave2_raw.jsonl

The other thirty-six categories on this site, each asked the consumer question its own customers would ask, in Austin, more than once. One metro rather than three: the previous run had already shown that the answer is rebuilt per city, so a second city would have re-measured a settled question instead of covering new trades.

118readings
36distinct questions
2model tiers recorded

Stopped by decision, not by completion: the account began returning challenges rather than answers, and the readings still missing were third readings of trades that already had two, which is what a stability figure needs. What came back: every single reading used web search, so a page can reach these answers in all forty-six categories we sell to, not just the ten measured before. The pool an assistant draws on ran from two businesses to ten, and of the four hundred and ninety-six pairs of trades, four hundred and ninety-four shared no business at all; the two that did were single clinics that genuinely sell both services. Because the rate limits spread these readings out, they also answered something the earlier run could not: the same stability holds across eight hours that we had only measured across forty minutes. One trade returned ten different businesses across its readings and not a single one of them held.

Read the finding →

How to read this

Every sample size on this page is recomputed from the raw readings when the page is built, so none of them is a number somebody typed. Where a run turned into a published finding there is a link to it; where a run returned nothing useful, it is still here.

The instrument is the same throughout: a real browser driven over the DevTools protocol, one question per fresh conversation, the full answer stored verbatim. Readings are never pooled across surfaces, and a comparison never crosses model tiers.

If you want this run on your own category and city, the free scan is the same instrument on one question.