KesterleyWhat AI says about you

MethodologyVersion 0.2August 2026Public and versioned

How we measure

Every number we report is produced exactly as described here. When the methodology changes, the version changes, and results across versions are never compared silently. The full text is also available as plain text for machine readers.

New in 0.2

Designs are pre-registered and hashed before measurement, every output states its minimum detectable change, conservative matching for common-word brand names is disclosed, and the core measured trio is ChatGPT, Gemini and Claude. Metric definitions are unchanged.

AI availability, defined

AI availability is how easily a brand gets chosen when an AI assistant answers a buyer's question, and how easily an AI agent can read, verify and buy from that brand. It has two halves, reported separately and never averaged: answer availability, the brand's share of the answers across the buying situations of its category, in the assistants' memory and with live search; and agent availability, whether an agent can find, read, trust and act on the brand's pages. It extends the mental and physical availability of the Ehrenberg-Bass research tradition to the AI layer. The full definition, with the metrics and what it is not.

The measurement problem we start from

Independent research on AI brand answers is blunt. Fewer than 1 in 100 identical prompts return the same brand list. Brand identity explains around one percent of the variance of a single response. Different models agree on a category leader less than half the time. Most commercial AI visibility numbers ignore all of this. We treat it as the starting point.

The machine is the respondent

Classical brand research needs consumer surveys, because human memory can only be sampled indirectly. Machines remove that constraint. We interrogate the assistant itself, around a thousand times per category, across its buying situations. It is a census of the machine's memory, not a survey of people.

Two modes, reported separately: memory, with browsing off, which shows what the model already knows. And search, with browsing on, which shows what it finds right now. A brand can be strong in one and absent in the other. That gap is a finding in itself.

Sampling frame: buying situations, not random prompts

For each category we derive 8 to 16 buying situations, the moments in which the need arises, phrased the way real users ask assistants. This applies the category entry point framework from the Ehrenberg-Bass research tradition to machines. Each situation gets several phrasing variants, because phrasing drives more answer variance than repetition. The design is situation by phrasing by model by repeat, around a thousand recorded answers per category.

Pre-registration

For published reports, the complete design (brands, prompts, models, repeats, rules) is frozen before the first question is asked and its SHA-256 hash is published with the results. Anyone can re-hash the design and verify that nothing was selected after seeing the data.

What we report

Consideration set inclusion
How often a brand appears per buying situation, with bootstrap 95% confidence intervals. Never "rank": rank order in AI answers is noise, membership in the answer set is the stable construct.
Share of answers
The brand's share of all brand mentions across situations, weighted by situation importance. The machine analog of mental market share.
Cross model agreement
Reported, not averaged away. Platform disagreement is information.
Entity integrity
Wrong or confused brand claims in answers, each with the recorded evidence.
Citation sources
Which domains actually feed grounded answers in the category. Observed from citations, not guessed.
Reliability of our numbers
Design size, split-half stability, minimum detectable change. When a difference is within noise, the report says so.

The machine facing surface

This is the agent half of AI availability. Ten checks on the brand's public surface: crawler access per bot class, entity records and consistency, product structured data and its freshness, readability without JavaScript or login, feeds and merchant program enrollment where relevant. Content behind registration or social walls is largely invisible to assistants. Walls are findings, not measurement obstacles: what is walled off is, to a machine, barely stocked, and we say so with evidence.

Vocabulary

AI availability
How easily a brand gets chosen when an AI assistant answers a buyer's question, and how easily an AI agent can read, verify and buy from it. Two halves: answer availability and agent availability. Full definition.
Answer availability
The mental half: the brand's share of recorded answers across the buying situations of its category, in memory mode and with live search.
Agent availability
The physical half: whether an AI agent can find, read, trust and act on the brand's pages. Ten checks, four levels.
Buying situation
A moment in which the need for the category arises, phrased the way a buyer asks an assistant. Derived from the category entry point framework; 8 to 16 per category.
Memory and search
Two modes of the same question. Memory: browsing off, what the model already knows. Search: browsing on, what it finds now. Reported separately; the gap is a finding.
Share of answers
The share of recorded answers that include the brand, per situation with a 95% confidence interval, and across situations weighted by importance.
Situations covered
The number of buying situations in which the brand appears at all, above the noise floor. Network size when weighted by importance.
AI mental market share
AI penetration multiplied by network size, as a share of the category: the machine analog of mental market share.
Entity integrity
Whether what assistants say about the brand is true. Each wrong or confused claim is recorded with the answer as evidence.
Citation sources
The domains grounded answers actually cite in a category, counted, not guessed.
Agent-readiness levels
1 Falling behind, 2 The basics only, 3 AI-readable, 4 Agent-ready. From checks passed out of ten.
Minimum detectable change
The smallest movement a design of this size can tell apart from noise. Stated with every report so nobody reads a wobble as a trend.

Category playbooks and engineering configurations are internal documentation. Raw responses are archived per run; every report is reproducible from its archive and pinned to a methodology version.