Skip to main content
CyberLab.Team

Sample deliverable

AI visibility baseline: the prompt set and the rules that make it repeatable

A baseline is only useful if the next measurement is done the same way. This sample shows the prompt set, the rule for what counts as a mention, how the evidence is stored, and the limits that come with measuring a system which answers a little differently on every run.

Example, not a client's data

Invented setup: a business with one main service, two competitors it keeps meeting in conversation, and customers in a single country and language. Nobody has looked at what assistants say about the category yet, so this measurement is the starting point before any promotion work is planned. Placeholders stand in for every real name.

The prompt set

Prompts are agreed with you before anything is run, and they stay fixed afterwards. Square brackets mark the placeholders that carry your service, city, industry and competitor names.

PromptEngineLocaleWhat counts as a mention
best [service] provider in [city]ChatGPTen, [country]Named among the options, with or without a link
who should a small business hire for [service]Google AI Overviewsen, [country]Named in the overview text, or cited as a source under it
[brand] compared with [competitor]Perplexityen, [country]Named with at least one cited page from the brand's own site
is [brand] any good for [service]Claudeen, [country]Described, and the description matches the facts on the brand's site
how much does [service] cost in [country]Perplexityen, [country]Cited as a source for the range, not merely named in passing
alternatives to [competitor]ChatGPTen, [country]Named in the list of alternatives rather than in a closing caveat
[service] for [industry] companiesGeminien, [country]Named, or a page from the brand's site cited
who can fix [problem] for an online storeMicrosoft Copiloten, [country]Named among the providers, not only inside a general explanation

Rules for repeating the measurement

Everything below exists so the second measurement differs from the first only where your visibility differs, and not because the method moved.

  • The same prompts, word for word. A rewritten prompt makes the second measurement a different measurement.
  • The same locale and interface language, since answers differ by market even when the question does not.
  • The same session state: signed out where the engine allows it, no memory, no earlier conversation in the window.
  • Several runs of each prompt, because answers drift between runs and a single run is an anecdote.
  • The same engine list. An engine added later starts its own baseline instead of joining the old one.
  • The model label recorded wherever the engine shows it, so a shift after a model update can be told apart from a shift caused by the work.
  • The interval agreed on the brief and then kept, so two measurements are comparable rather than opportunistic.

What is recorded for every run

Evidence is handed over as files, so the measurement can be checked or repeated by someone else, including a contractor who is not us.

FieldWhat is recordedWhy it is needed
Prompt IDThe stable label of the prompt, matching the set aboveRows from two measurements can be lined up without guesswork
Engine and modelThe engine, plus the model label where the interface shows itSeparates a model change from a change in your visibility
LocaleThe country and interface language used for the runThe same question is answered differently by market
Run numberWhich run of that prompt this isDrift between runs stays visible instead of being averaged away
Answer textThe full answer, saved as textWording can be searched and quoted later without reopening images
Cited sourcesEvery source the engine listed, in the order it listed themShows who represents you in the answer when you are not named
ScreenshotThe answer as an image, named by prompt, engine and run, with the date of the runMakes the record checkable by a person, not only by us

What a baseline does not show

  • Traffic. A mention is not a visit, and referrals from assistants are read separately in analytics.
  • Cause. The measurement records what the answer said, not why the engine chose it. Reasons are argued from the sources and the site, and labeled as reasoning rather than fact.
  • A ranking. There is no position to hold inside an answer, so how often you are named is not a place in a list.
  • A promise. Nobody controls whether an engine names a brand, and a guarantee of that is a warning sign rather than an offer.
  • The whole market. Only the prompts in the set are measured, and a question nobody agreed to test is simply absent from the data.

Synthetic example, not a client's data

  • The prompt set here is short and generic. A real set is built from the questions your customers actually ask and covers the buying journey, not just the obvious query.
  • Nothing is counted in this sample. A real baseline states how often each engine named you across the runs, and the same counting rule is applied next time.
  • The engines listed are an example of a default set. The real list follows where your customers ask, and can include engines that matter only in your market.
  • Answers drift between runs and change after model updates, so a baseline is a snapshot of the moment rather than a constant.

A real baseline arrives as the agreed prompt set, the per engine record with screenshots and cited sources, and a walkthrough where the answers are read together and the gaps become an order of work. Scope is agreed on the brief.

AI Visibility Audit

Ready for the audit?

Send the site or the account and get a free preliminary check first: the biggest issues, the tier that fits and the fixed price. You decide after that.

Order the audit

No commitment. Response within one business day.