Sample deliverable
AI visibility baseline: the prompt set and the rules that make it repeatable
A baseline is only useful if the next measurement is done the same way. This sample shows the prompt set, the rule for what counts as a mention, how the evidence is stored, and the limits that come with measuring a system which answers a little differently on every run.
Example, not a client's data
Invented setup: a business with one main service, two competitors it keeps meeting in conversation, and customers in a single country and language. Nobody has looked at what assistants say about the category yet, so this measurement is the starting point before any promotion work is planned. Placeholders stand in for every real name.
The prompt set
Prompts are agreed with you before anything is run, and they stay fixed afterwards. Square brackets mark the placeholders that carry your service, city, industry and competitor names.
| Prompt | Engine | Locale | What counts as a mention |
|---|---|---|---|
| best [service] provider in [city] | ChatGPT | en, [country] | Named among the options, with or without a link |
| who should a small business hire for [service] | Google AI Overviews | en, [country] | Named in the overview text, or cited as a source under it |
| [brand] compared with [competitor] | Perplexity | en, [country] | Named with at least one cited page from the brand's own site |
| is [brand] any good for [service] | Claude | en, [country] | Described, and the description matches the facts on the brand's site |
| how much does [service] cost in [country] | Perplexity | en, [country] | Cited as a source for the range, not merely named in passing |
| alternatives to [competitor] | ChatGPT | en, [country] | Named in the list of alternatives rather than in a closing caveat |
| [service] for [industry] companies | Gemini | en, [country] | Named, or a page from the brand's site cited |
| who can fix [problem] for an online store | Microsoft Copilot | en, [country] | Named among the providers, not only inside a general explanation |
Rules for repeating the measurement
Everything below exists so the second measurement differs from the first only where your visibility differs, and not because the method moved.
- The same prompts, word for word. A rewritten prompt makes the second measurement a different measurement.
- The same locale and interface language, since answers differ by market even when the question does not.
- The same session state: signed out where the engine allows it, no memory, no earlier conversation in the window.
- Several runs of each prompt, because answers drift between runs and a single run is an anecdote.
- The same engine list. An engine added later starts its own baseline instead of joining the old one.
- The model label recorded wherever the engine shows it, so a shift after a model update can be told apart from a shift caused by the work.
- The interval agreed on the brief and then kept, so two measurements are comparable rather than opportunistic.
What is recorded for every run
Evidence is handed over as files, so the measurement can be checked or repeated by someone else, including a contractor who is not us.
| Field | What is recorded | Why it is needed |
|---|---|---|
| Prompt ID | The stable label of the prompt, matching the set above | Rows from two measurements can be lined up without guesswork |
| Engine and model | The engine, plus the model label where the interface shows it | Separates a model change from a change in your visibility |
| Locale | The country and interface language used for the run | The same question is answered differently by market |
| Run number | Which run of that prompt this is | Drift between runs stays visible instead of being averaged away |
| Answer text | The full answer, saved as text | Wording can be searched and quoted later without reopening images |
| Cited sources | Every source the engine listed, in the order it listed them | Shows who represents you in the answer when you are not named |
| Screenshot | The answer as an image, named by prompt, engine and run, with the date of the run | Makes the record checkable by a person, not only by us |
What a baseline does not show
- Traffic. A mention is not a visit, and referrals from assistants are read separately in analytics.
- Cause. The measurement records what the answer said, not why the engine chose it. Reasons are argued from the sources and the site, and labeled as reasoning rather than fact.
- A ranking. There is no position to hold inside an answer, so how often you are named is not a place in a list.
- A promise. Nobody controls whether an engine names a brand, and a guarantee of that is a warning sign rather than an offer.
- The whole market. Only the prompts in the set are measured, and a question nobody agreed to test is simply absent from the data.
Synthetic example, not a client's data
- The prompt set here is short and generic. A real set is built from the questions your customers actually ask and covers the buying journey, not just the obvious query.
- Nothing is counted in this sample. A real baseline states how often each engine named you across the runs, and the same counting rule is applied next time.
- The engines listed are an example of a default set. The real list follows where your customers ask, and can include engines that matter only in your market.
- Answers drift between runs and change after model updates, so a baseline is a snapshot of the moment rather than a constant.
A real baseline arrives as the agreed prompt set, the per engine record with screenshots and cited sources, and a walkthrough where the answers are read together and the gaps become an order of work. Scope is agreed on the brief.
AI Visibility AuditReady for the audit?
Send the site or the account and get a free preliminary check first: the biggest issues, the tier that fits and the fixed price. You decide after that.
Order the auditNo commitment. Response within one business day.