# Kesterley AI Availability Methodology

**Version 0.2, August 2026. Definition v1.0 of AI availability added September 2026 (documentation; measurement rules unchanged). This document is public.** Every number we report is produced exactly as described here. When the methodology changes, the version changes, and results across versions are not compared silently.

**Changes from v0.1:** (1) designs are pre-registered before measurement: the full design (brands, prompts, models, repeats, rules) is frozen and its SHA-256 hash is published with the results, so prompt or brand selection cannot follow the results; (2) every output now states its minimum detectable change alongside confidence intervals; (3) conservative matching for brand names that are ordinary words is now part of extraction and disclosed; (4) the core measured trio is OpenAI, Google Gemini and Anthropic Claude (Perplexity available on request). Metrics definitions are unchanged.

## Why this exists

AI assistants increasingly stand between buyers and brands. A discipline has grown around this ("GEO", "AEO", "AI visibility"), and its measurement layer is largely broken: single responses carry almost no brand signal (brand identity explains ~1% of response variance; arXiv 2607.13304), fewer than 1 in 100 identical prompts return the same brand list (SparkToro/Gumshoe, n=2,961 runs), models agree on a category leader less than half the time (arXiv 2606.23057), and most tools sample arbitrary prompts with undisclosed methodology.

We measure AI availability the way marketing science measures mental availability: with a defined sampling frame, disclosed design, confidence intervals, and honest reliability reporting. This methodology extends the availability framework of the Ehrenberg-Bass school (mental and physical availability; category entry points, Romaniuk & Sharp) to machine intermediaries. It is our own application of that public framework, not an endorsement by its authors.

## Definitions

- **AI availability**: how easily a brand gets chosen when an AI assistant answers a buyer's question, and how easily an AI agent can read, verify and buy from that brand. Two halves, reported separately and never averaged. Full definition: kesterley.com/ai-availability (also as ai-availability.md).
- **Answer availability** (the mental half): the probability that an AI assistant surfaces a brand across the buying situations of its category, in memory mode and with live search.
- **Agent availability** (the physical half): the degree to which an AI agent can find, parse, evaluate and transact with the brand.
- **Prompt Entry Points (PEPs)**: the machine analog of Category Entry Points: the buying situations of a category, expressed the way real users express them to assistants. PEPs are derived per category along the classic CEP dimensions (why / when / where / while doing what / with whom / how it feels), not invented ad hoc, and each PEP is weighted by its estimated real-world importance.

## What exactly we analyze, and why no questionnaires

Classical mental availability is measured with consumer surveys, because a human memory can only be sampled indirectly. AI availability removes that constraint: **the machine is the respondent.** We interrogate the assistant itself, thousands of times, across the category's buying situations: a census of the machine's memory, not a survey of consumers. Three measurement objects, all digital:

1. **The assistants' answers** (mental side). Queried in two modes and reported separately: **parametric** (no browsing: what the model already "knows"; the machine analog of memory) and **retrieval-augmented** (search/grounding on: what it finds right now; the machine analog of the shelf). A brand can be strong in one and absent in the other; that difference is itself a core finding.
2. **The citation mix of grounded answers.** Retrieval-mode assistants (Perplexity, grounded Gemini) disclose which sources fed each answer. Aggregated per category, this yields the earned-media map of the machine: which domains actually feed AI answers here, and therefore where presence is worth building. No guessing about "what AI reads"; we observe it.
3. **The brand's public machine-facing surface** (physical side): crawler access per bot class, entity records (knowledge graphs, Wikidata/Wikipedia, schema.org Organization/`sameAs`), product structured data and feed freshness, page parseability without JavaScript or login, merchant-program enrollment.

**Walls are findings, not obstacles.** Content behind registration, paywalls or social-app walls is largely invisible to assistants and their crawlers, so we don't need access to it to measure availability; its absence from answers IS the measurement. Where walled platforms leak into models (e.g. licensed corpora such as Reddit), the leak shows up in answers and citations, and we measure it there. A brand whose substance lives behind walls is, to a machine, barely stocked. We say so, with evidence.

**Out of scope, stated plainly:** closed in-platform assistants we cannot query programmatically (e.g. a retailer's internal shopping AI), private social engagement data, and any claim about sales causality.

## Measurement design (mental side)

**Sampling frame.** For a category we define 10–16 PEPs. Each PEP gets 4–8 phrasing variants (paraphrases), because query phrasing explains far more answer variance than repeated sampling of one phrasing (arXiv 2607.13304). Design: `PEP × paraphrase × model × repeat`.

**Measured surfaces.** We query the assistants consumers actually use, via their official APIs: OpenAI, Google Gemini and Anthropic Claude as the core trio, Perplexity on request. API responses are an imperfect proxy for each vendor's consumer app (system prompts and retrieval differ); we disclose this rather than hide it. We never scrape consumer interfaces.

**Pre-registration.** For published category reports, the complete design is frozen before the first API call and hashed (SHA-256). The manifest with the hash is published alongside the results. Anyone can re-hash the design and verify nothing was selected after seeing the data.

**Repeats.** Each cell is sampled multiple times at the provider's default temperature. We report stability, not pretend it away.

**Extraction.** Brand mentions are extracted from answers against a maintained alias list (with fuzzy matching); an inexpensive LLM is used only to assist extraction and never as a measured surface. Brand names that are ordinary words ("On", "Monday", "Remote") are matched conservatively: capitalized form only, distinctive aliases, sentence starts excluded. This under-counts slightly and we disclose it.

### Metrics

1. **AI penetration**: share of PEPs where the brand appears in answers at all (above a noise floor).
2. **PEP network size**: how many distinct PEPs the brand is attached to, weighted by PEP importance.
3. **AI mental market share**: the brand's share of all brand–PEP link mass in the category (the direct analog of Romaniuk's mental market share = mental penetration × network size).
4. **Consideration-set inclusion rate**: how often the brand appears per PEP, with **bootstrap 95% confidence intervals**. We report inclusion, never "rank": rank order in LLM answers is statistically unstable; set membership is the construct that persists (55–77% appearance stability for leading brands vs. near-random ordering).
5. **Cross-model agreement**: reported per brand, because platform disagreement is real information, not noise to average away.
6. **Entity integrity rate**: share of brand-attributed claims in answers that are wrong or belong to a competitor (the machine analog of distinctive-asset uniqueness; weak-entity brands show ~6× higher false-attribute rates).
7. **Reliability disclosure**: every report states the design size, test–retest stability of its own numbers, and the minimum detectable change. If a difference is within noise, we say so.

## Measurement design (physical side)

A 10-point operability review, scored from evidence, not opinion: AI-crawler access policy (per bot class: training / search / user-fetch), Cloudflare or WAF bot posture, entity consistency (name disambiguation, schema.org Organization + `sameAs`, knowledge-graph presence), Product/Offer structured data completeness **and freshness** (stale price/stock is the documented failure mode of agentic checkout), feed availability per surface spec, merchant-program enrollment where relevant, llms.txt (noted for completeness; we tell clients plainly that no measurable citation effect has been demonstrated), retrievability of key claims (are the pages agents cite actually fetchable and parseable), review/rating machine-readability, and transactability path.

## Output

Reports contain no rankings and no causal claims; every recommendation carries its evidence grade. Recommendations are limited to levers with evidence behind them: earned-media breadth, entity maturity, structured-data quality, retrievability. We query official developer interfaces only and never scrape consumer interfaces.

One report per brand: where you stand (the three mental metrics vs. named competitors, with CIs), what agents can and cannot do with you (operability score with evidence), what is wrong in the machine's head about you (entity integrity findings), and a prioritized fix list where every recommendation carries its evidence grade (**strong / moderate / weak**) and an owner. Findings are written so a marketing team can act on them the same week; where the client prefers, we execute the fixes ourselves as a separate engagement.

## Operational notes

Category playbooks (how buying situations are derived and weighted per category), engineering configurations and cost models are internal documentation. What is public is everything needed to judge and verify the measurement: the definitions, the metrics, the design of each published report (frozen and hashed before measurement), and the reliability of the results. Raw responses are archived per run; every report is reproducible from its archive and pinned to a methodology version.
