Measure AI Search Visibility
Record prompts, dates, models, regions, repetitions, answers, and citations. One screenshot cannot establish visibility.
AI answer visibility changes by prompt wording, system, date, region, account state, model, and browsing mode. Benchmark must record those variables and repeat runs.
Question
Define decision before prompts. Example: "Which sources do AI answer systems cite when a marketer asks how to build a GitHub-based content workflow?"
Do not begin with preferred brand or ranking outcome.
Prompt set
Build prompt groups by intent:
- Definition
- How-to
- Tool discovery
- Comparison
- Troubleshooting
- Source request
Freeze prompt set before first scored run. Publish exact text.
Systems and settings
Record:
- Product and model label shown to user
- Browsing or search mode
- Account tier
- Region and interface language
- Run timestamp
- New or continuing conversation
Repetitions
Run each prompt at least three times per system in fresh conversations. More repetitions improve stability estimate but cost more.
Save complete answer, visible citations, cited URLs, and access failures. Screenshots help audit UI but structured text remains main result.
Scoring
Separate measures:
| Measure | Definition |
|---|---|
| Mention rate | Runs naming entity or source |
| Citation rate | Runs linking source |
| Supported citation | Citation supports nearby answer claim |
| Source diversity | Unique cited domains across runs |
| Stability | Similarity of results across repeated runs |
No single overall visibility score until weights have external reason.
Publication
Release prompt file, raw answer records where terms allow, scoring code, dated result table, limitations, and conflicts.
Failure modes
- One favorable answer treated as benchmark
- Prompt contains target brand and inflates mention rate
- Search mode differs between systems
- Citation counted without checking claim support
- Result published without model/date
- Benchmark headline survives after protocol changes
Results expire. Schedule rerun based on material system change, not arbitrary freshness badge.