Does an assistant name local businesses to their customers?
27 August 2026 – 30 August 2026 · raw_observations.jsonl
The first question worth asking: when an ordinary person asks an assistant for a local service, does it answer with names, or with advice? We asked across trades and metros, one question per fresh conversation, full answer text stored verbatim.
50readings
26distinct questions
2model tiers recorded
It answers with names, every time, in a short business list. This run is the calibration set every later comparison is measured against, and it is where the pool figures on the category pages come from.
The same question, ten minutes apart
30 August 2026 – 31 August 2026 · service_intent_raw.jsonl
Five questions a business owner might put to an assistant when looking for help - the four things such an owner can buy, plus the vague version of the same need. Asked repeatedly, on one model tier, so repeats could be compared.
19readings
15distinct questions
2model tiers recorded
Every reading returned a shortlist; not one deflected into generic advice. But two readings of one identical question shared an average of 15% of their named companies, and the sharpest pair shared none at all. Being named once is a reading, not a position.
Read the finding →
A fifth pass on the same five questions
31 August 2026 · stability_raw.jsonl
A further repeat of the service-intent set, recorded to a separate file so the earlier readings could not be quietly extended. Same surface, same tier discipline.
10readings
5distinct questions
2model tiers recorded
Pooled with the run above, this is what the stability figures rest on. Both files carry the same declared surface, so they are compared with each other and never with the consumer-question readings.
Read the finding →
What an assistant tells a business owner looking for help
29 August 2026 · raw_buyer_probe.jsonl
Ten questions in the owner's own voice - my dental practice does not show up when people ask ChatGPT, who can help my roofing company, who does AI visibility audits for small local businesses.
10readings
10distinct questions
2model tiers recorded
Nine of the ten produced a named shortlist of vendors, with prices and a best-for column. This is a buying surface, not an advice surface - which is the single most commercially useful thing we have measured.
The same ten questions, framed in Austin
30 August 2026 · raw_buyer_us.jsonl
The probe above inherited its operator's location. We ran it again stating a US metro, to find out whether the answer set is national or local.
10readings
10distinct questions
2model tiers recorded
Almost nothing carried over. The names came back Austin-specific, down to firms with the metro in their brand line.
And again, in Miami
30 August 2026 · raw_buyer_miami.jsonl
A third locale, to test whether the Austin result was localisation or noise.
10readings
10distinct questions
1model tier
Roughly 3% of names overlapped between Austin and Miami, and not one company appeared in all three runs. Inside a single metro the answers concentrate hard on one or two firms. This is a per-metro market as an assistant sees it.
Ten trades, three cities, repeated readings
31 August 2026 – 2 September 2026 · vertical_raw.jsonl
Ten verticals, a consumer question each, in Austin, Miami and Columbus, asked more than once so stability could be computed per trade rather than site-wide.
63readings
30distinct questions
2model tiers recorded
Complete. Every reading used web search, so a page can reach these answers at all. The pools differ by more than six times between trades, which is why the remaining 36 categories are worth running. And the sharpest result: asked in three metros, not one of 19 comparable city pairs shared a single business.
Thirty-six more trades, one city, repeated readings
1 September 2026 – 2 September 2026 · vertical_wave2_raw.jsonl
The other thirty-six categories on this site, each asked the consumer question its own customers would ask, in Austin, more than once. One metro rather than three: the previous run had already shown that the answer is rebuilt per city, so a second city would have re-measured a settled question instead of covering new trades.
118readings
36distinct questions
2model tiers recorded
Stopped by decision, not by completion: the account began returning challenges rather than answers, and the readings still missing were third readings of trades that already had two, which is what a stability figure needs. What came back: every single reading used web search, so a page can reach these answers in all forty-six categories we sell to, not just the ten measured before. The pool an assistant draws on ran from two businesses to ten, and of the four hundred and ninety-six pairs of trades, four hundred and ninety-four shared no business at all; the two that did were single clinics that genuinely sell both services. Because the rate limits spread these readings out, they also answered something the earlier run could not: the same stability holds across eight hours that we had only measured across forty minutes. One trade returned ten different businesses across its readings and not a single one of them held.
Read the finding →