Benchmark protocolProtocol published

AI Search Visibility Benchmark

Protocol published before results. Prompt set, system settings, repetitions, raw answers, and scoring stay visible.

This research asks which sources AI answer systems surface for common marketing AI questions and how stable those results remain across repeated runs.

Questions

  1. Which domains receive mentions?
  2. Which pages receive clickable citations?
  3. Do citations support nearby claims?
  4. How much do answers change between fresh runs?
  5. Does explicit source request change citation quality?

Cohort

Initial cohort covers question-led prompts about agent architecture, repository workflows, content verification, and AI search measurement. Vendor-ranking prompts remain excluded from first run.

Protocol

  • Freeze prompt text before collection
  • Use fresh conversation for each run
  • Record product, displayed model, browsing mode, region, language, and timestamp
  • Repeat each prompt at least three times per system
  • Store answer text and visible citations
  • Review whether citation supports associated claim

Metrics

mention_rate counts runs naming source or entity. citation_rate counts runs with clickable source. support_rate counts citations that substantively support nearby answer. stability compares results across repetitions.

These metrics remain separate. No weighted overall score yet.

Result status

No result dataset published. Current page describes protocol only. Publishing protocol first reduces incentive to change method after seeing favorable results.

Planned artifacts

research/prompts/v1.jsonl
research/runs/<system>/<date>.jsonl
research/scoring/v1.md
research/results/v1.csv

Limitations

Interfaces, model routing, retrieval indexes, and answer policies change. Result is dated observation, not durable guarantee of citation or rank.

Sources

Evidence used on this page

  1. Google AI optimization guide
  2. OpenAI publisher and developer FAQ