ResearchProtocol published

Measure AI Search Visibility

Record prompts, dates, models, regions, repetitions, answers, and citations. One screenshot cannot establish visibility.

Written by
Marketing Wiki Editors
Reviewed by
Marketing Wiki Editors
Published
Updated
Evidence checked
Sources
2
Direct answer

A reproducible benchmark workflow for mentions, citations, source diversity, answer stability, and factual support across AI answer systems.

AI answer visibility changes by prompt wording, system, date, region, account state, model, and browsing mode. Benchmark must record those variables and repeat runs.

Question#

Define decision before prompts. Example: "Which sources do AI answer systems cite when a marketer asks how to build a GitHub-based content workflow?"

Do not begin with preferred brand or ranking outcome.

Prompt set#

Build prompt groups by intent:

  • Definition
  • How-to
  • Tool discovery
  • Comparison
  • Troubleshooting
  • Source request

Freeze prompt set before first scored run. Publish exact text.

Systems and settings#

Record:

  • Product and model label shown to user
  • Browsing or search mode
  • Account tier
  • Region and interface language
  • Run timestamp
  • New or continuing conversation

Repetitions#

Run each prompt at least three times per system in fresh conversations. More repetitions improve stability estimate but cost more.

Save complete answer, visible citations, cited URLs, and access failures. Screenshots help audit UI but structured text remains main result.

Scoring#

Separate measures:

Scroll table →
MeasureDefinition
Mention rateRuns naming entity or source
Citation rateRuns linking source
Supported citationCitation supports nearby answer claim
Source diversityUnique cited domains across runs
StabilitySimilarity of results across repeated runs

No single overall visibility score until weights have external reason.

Publication#

Release prompt file, raw answer records where terms allow, scoring code, dated result table, limitations, and conflicts.

Failure modes#

  • One favorable answer treated as benchmark
  • Prompt contains target brand and inflates mention rate
  • Search mode differs between systems
  • Citation counted without checking claim support
  • Result published without model/date
  • Benchmark headline survives after protocol changes

Results expire. Schedule rerun based on material system change, not arbitrary freshness badge.

Evidence

Sources behind this page

Claims remain tied to dated source review. Method and corrections stay public.

  1. S-01Google AI optimization guidedevelopers.google.com
  2. S-02OpenAI publisher and developer FAQhelp.openai.com