ResearchProtocol published

Measure AI Search Visibility

Record prompts, dates, models, regions, repetitions, answers, and citations. One screenshot cannot establish visibility.

AI answer visibility changes by prompt wording, system, date, region, account state, model, and browsing mode. Benchmark must record those variables and repeat runs.

Question

Define decision before prompts. Example: "Which sources do AI answer systems cite when a marketer asks how to build a GitHub-based content workflow?"

Do not begin with preferred brand or ranking outcome.

Prompt set

Build prompt groups by intent:

  • Definition
  • How-to
  • Tool discovery
  • Comparison
  • Troubleshooting
  • Source request

Freeze prompt set before first scored run. Publish exact text.

Systems and settings

Record:

  • Product and model label shown to user
  • Browsing or search mode
  • Account tier
  • Region and interface language
  • Run timestamp
  • New or continuing conversation

Repetitions

Run each prompt at least three times per system in fresh conversations. More repetitions improve stability estimate but cost more.

Save complete answer, visible citations, cited URLs, and access failures. Screenshots help audit UI but structured text remains main result.

Scoring

Separate measures:

MeasureDefinition
Mention rateRuns naming entity or source
Citation rateRuns linking source
Supported citationCitation supports nearby answer claim
Source diversityUnique cited domains across runs
StabilitySimilarity of results across repeated runs

No single overall visibility score until weights have external reason.

Publication

Release prompt file, raw answer records where terms allow, scoring code, dated result table, limitations, and conflicts.

Failure modes

  • One favorable answer treated as benchmark
  • Prompt contains target brand and inflates mention rate
  • Search mode differs between systems
  • Citation counted without checking claim support
  • Result published without model/date
  • Benchmark headline survives after protocol changes

Results expire. Schedule rerun based on material system change, not arbitrary freshness badge.

Sources

Evidence used on this page

  1. Google AI optimization guide
  2. OpenAI publisher and developer FAQ