{"schema_version":"2.0","record_type":"workflow","canonical_url":"https://marketingwiki.ai/workflows/measure-ai-search-visibility","id":"workflow-measure-ai-search-visibility","slug":"measure-ai-search-visibility","title":"Measure AI Search Visibility","description":"A reproducible benchmark workflow for mentions, citations, source diversity, answer stability, and factual support across AI answer systems.","dek":"Record prompts, dates, models, regions, repetitions, answers, and citations. One screenshot cannot establish visibility.","label":"Research","status":"Protocol published","stages":["Prompts","Systems","Runs","Score","Publish"],"topics":["AI search","benchmarking","citations"],"publishedAt":"2026-08-10","updatedAt":"2026-08-10","lastVerifiedAt":"2026-08-10","readingMinutes":2,"author":"Marketing Wiki Editors","reviewer":"Marketing Wiki Editors","sources":[{"title":"Google AI optimization guide","url":"https://developers.google.com/search/docs/fundamentals/ai-optimization-guide"},{"title":"OpenAI publisher and developer FAQ","url":"https://help.openai.com/en/articles/12627856-publishers-and-developers-faq"}],"wordCount":314,"body":"AI answer visibility changes by prompt wording, system, date, region, account state, model, and browsing mode. Benchmark must record those variables and repeat runs.\n\n## Question\n\nDefine decision before prompts. Example: \"Which sources do AI answer systems cite when a marketer asks how to build a GitHub-based content workflow?\"\n\nDo not begin with preferred brand or ranking outcome.\n\n## Prompt set\n\nBuild prompt groups by intent:\n\n- Definition\n- How-to\n- Tool discovery\n- Comparison\n- Troubleshooting\n- Source request\n\nFreeze prompt set before first scored run. Publish exact text.\n\n## Systems and settings\n\nRecord:\n\n- Product and model label shown to user\n- Browsing or search mode\n- Account tier\n- Region and interface language\n- Run timestamp\n- New or continuing conversation\n\n## Repetitions\n\nRun each prompt at least three times per system in fresh conversations. More repetitions improve stability estimate but cost more.\n\nSave complete answer, visible citations, cited URLs, and access failures. Screenshots help audit UI but structured text remains main result.\n\n## Scoring\n\nSeparate measures:\n\n| Measure | Definition |\n| --- | --- |\n| Mention rate | Runs naming entity or source |\n| Citation rate | Runs linking source |\n| Supported citation | Citation supports nearby answer claim |\n| Source diversity | Unique cited domains across runs |\n| Stability | Similarity of results across repeated runs |\n\nNo single overall visibility score until weights have external reason.\n\n## Publication\n\nRelease prompt file, raw answer records where terms allow, scoring code, dated result table, limitations, and conflicts.\n\n## Failure modes\n\n- One favorable answer treated as benchmark\n- Prompt contains target brand and inflates mention rate\n- Search mode differs between systems\n- Citation counted without checking claim support\n- Result published without model/date\n- Benchmark headline survives after protocol changes\n\nResults expire. Schedule rerun based on material system change, not arbitrary freshness badge."}