{"schema_version":"2.0","record_type":"article","canonical_url":"https://marketingwiki.ai/articles/what-makes-marketing-content-usable-by-ai-agents","id":"what-makes-marketing-content-usable-by-ai-agents","slug":"what-makes-marketing-content-usable-by-ai-agents","title":"What Makes Marketing Content Usable by AI Agents?","description":"Architecture guide for publishing marketing content through canonical HTML, JSON, JSONL, feeds, sitemaps, MCP, and agent skills without false discovery claims.","dek":"Give agents stable identity, structured records, change signals, provenance, and permission-aware runtime access. Keep every surface tied to one source of truth.","category":"Agent Infrastructure","topics":["AI agents","structured content","MCP","provenance"],"publishedAt":"2026-09-01","updatedAt":"2026-09-14","lastVerifiedAt":"2026-09-14","readingMinutes":10,"author":"Marketing Wiki Research Automation","reviewer":null,"featured":false,"sources":[{"title":"RFC 6596: The Canonical Link Relation","url":"https://www.rfc-editor.org/rfc/rfc6596.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"RFC 8259: The JavaScript Object Notation Data Interchange Format","url":"https://www.rfc-editor.org/rfc/rfc8259.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"JSON Lines format documentation","url":"https://jsonlines.org/?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Schema.org Article type","url":"https://schema.org/Article?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Google general structured data guidelines","url":"https://developers.google.com/search/docs/appearance/structured-data/sd-policies?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"RFC 4287: The Atom Syndication Format","url":"https://www.rfc-editor.org/rfc/rfc4287.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Sitemaps XML format","url":"https://www.sitemaps.org/protocol.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Model Context Protocol 2026-07-28 resources specification","url":"https://modelcontextprotocol.io/specification/2026-07-28/server/resources?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Model Context Protocol 2026-07-28 tools specification","url":"https://modelcontextprotocol.io/specification/2026-07-28/server/tools?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Model Context Protocol 2026-07-28 authorization specification","url":"https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Agent Skills specification","url":"https://agentskills.io/specification?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"RFC 9309: Robots Exclusion Protocol","url":"https://www.rfc-editor.org/rfc/rfc9309.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"OpenAI publisher and developer FAQ","url":"https://help.openai.com/en/articles/12627856-publishers-and-developers-faq?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"},{"title":"Anthropic web crawler guidance","url":"https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents"}],"wordCount":1908,"body":"Marketing content becomes usable by AI agents when it has one stable source of truth and a small set of machine interfaces generated from that source. Canonical HTML gives people and retrieval systems a durable page to read and cite. JSON or JSONL makes records easier to parse in bulk. Structured data identifies entities already visible on the page. Feeds and sitemaps report publication state. MCP resources, tools, and agent skills support direct use by clients that have been configured to use them.\n\n> **Editorial disclosure:** Prepared by Marketing Wiki Research Automation under standing direct-publication authorization and not independently reviewed. Product capabilities are vendor-documented unless labeled otherwise; sources were refreshed on September 1, 2026.\n\nNo surface makes content universally discoverable, cited, installed, or used for training. Agent usability is an interface-design problem with four parts: identity, data, change, and permission.\n\nThis guide covers direct agent consumption architecture. It does not cover search ranking or AI-answer visibility measurement.\n\nFor search eligibility, crawler policy, and visibility measurement, read [SEO vs. AEO vs. GEO](/articles/seo-vs-aeo-vs-geo). The [Model Context Protocol](/index/model-context-protocol) and [Agent Skills](/index/agent-skills) index entries provide shorter definitions of two interfaces used below.\n\n## Surface decision table\n\nPublish each surface only when it has a named consumer and maintenance owner. Generate derived surfaces from the same content record so titles, dates, evidence, and canonical URLs do not drift.\n\n| Surface | Purpose | Main consumer | Maintenance rule | Does not guarantee |\n| --- | --- | --- | --- | --- |\n| Canonical HTML | Human-readable source, stable citation target, full context | People, browsers, web retrieval systems | Render substantive text at one durable URL; keep corrections and visible provenance on-page | Discovery, citation, or correct interpretation |\n| JSON catalog or API | Typed records for filtering and application use | Scripts, data tools, integrated agents | Version schema; keep stable IDs; link every record to canonical HTML | That a general-purpose agent knows endpoint exists |\n| JSONL export | Record-at-a-time bulk processing | Dataset loaders, command-line tools, batch agents | Emit one valid JSON value per line; validate every line; document media type and schema | Standardized discovery or automatic training |\n| Article JSON-LD | Describe page type, author, dates, and citations already present in HTML | Structured-data consumers | Match visible content; remove invented or stale properties | Search feature, agent use, or citation |\n| Atom or RSS feed | Subscription to additions and material updates | Feed readers, monitors, ingestion jobs | Keep stable entry IDs; change update date only for meaningful edits | Complete catalog semantics or immediate processing |\n| XML sitemap | Enumerate public canonical URLs and truthful modification dates | Search crawlers and URL auditors | Include canonical pages; set `lastmod` from content changes, not deploy time | Crawling, indexing, ranking, or agent retrieval |\n| MCP resource | Bounded runtime read of a record or collection | Configured MCP clients | Return stable URI, MIME type, content, and modification metadata; enforce resource permissions | Automatic connection from every MCP-capable client |\n| MCP tool | Filter records or perform an authorized operation | Configured agents with tool access | Publish input/output schemas; validate inputs; rate-limit; require confirmation for sensitive actions | Safe execution without access control and approval |\n| Agent skill | Teach an installed agent when and how to query, verify, and cite records | Skill-compatible agent runtimes | Keep `SKILL.md` short; declare compatibility; test referenced scripts and URLs | Automatic discovery, installation, or cross-client support |\n| `robots.txt` policy | Express crawler preferences by user agent | Compliant crawlers | Set rules per hostname and crawler role; test alongside CDN and bot controls | Authentication, confidentiality, licensing, or guaranteed crawler behavior |\n\n## Make canonical HTML authoritative\n\nFor public editorial content, canonical HTML should remain authoritative. It carries full context, visible corrections, source links, authorship, and accessibility for people. It also gives every machine representation one citation target.\n\n[RFC 6596](https://www.rfc-editor.org/rfc/rfc6596.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) defines `rel=\"canonical\"` as the preferred IRI among duplicate or superset representations. It does not certify accuracy or make one representation authoritative for every use. Use one self-referential canonical URL on the HTML page, then repeat that URL as `canonical_url` in derived records.\n\nDo not maintain article prose independently in HTML, JSON, feeds, and MCP handlers. Store content and metadata once, then render each interface. Independent copies create a predictable failure: an agent reads a newer title from JSON, an older body from HTML, and a deployment timestamp from the sitemap.\n\n## Use JSON for records and JSONL for streams\n\nJSON works well for a bounded response such as one article or a small catalog. A response can include collection metadata, schema version, and an array of records. [RFC 8259](https://www.rfc-editor.org/rfc/rfc8259.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) defines JSON syntax and interoperability requirements, but it does not define the fields in a marketing reference. The publisher must own and version that contract.\n\nJSONL is useful when a consumer should process one record at a time without loading a full array. [JSON Lines documentation](https://jsonlines.org/?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) defines three practical rules: UTF-8, one valid JSON value per line, and a line terminator between values. It also states that `application/jsonl` is not yet a standardized media type. Document chosen response headers instead of implying an IETF standard.\n\nMinimum record:\n\n```json\n{\n  \"schema_version\": \"1.0\",\n  \"id\": \"stable-record-id\",\n  \"canonical_url\": \"https://example.org/articles/stable-record-id\",\n  \"title\": \"Visible article title\",\n  \"summary\": \"Plain-language scope and conclusion.\",\n  \"body_markdown\": \"Full article body...\",\n  \"published_at\": \"2026-08-14\",\n  \"modified_at\": \"2026-08-14\",\n  \"last_verified_at\": \"2026-08-14\",\n  \"language\": \"en\",\n  \"authors\": [{ \"name\": \"Named author\", \"url\": \"https://example.org/about\" }],\n  \"review\": { \"status\": \"reviewer_required\", \"reviewer\": null },\n  \"affiliations\": [],\n  \"sources\": [\n    {\n      \"title\": \"Primary source title\",\n      \"url\": \"https://example.org/source\",\n      \"evidence_type\": \"official\",\n      \"verified_at\": \"2026-08-14\"\n    }\n  ],\n  \"license\": \"https://example.org/data-license\"\n}\n```\n\nFields should say what they mean. `published_at`, `modified_at`, and `last_verified_at` describe different events. `reviewer_required` must not become a reviewer name. Empty evidence should remain empty or unknown, not become an inferred source.\n\n## Treat structured data as description\n\nSchema.org [`Article`](https://schema.org/Article?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) supports properties such as author, citation, publisher, version, and publishing principles. Add only properties that describe visible content and known entities. Google’s [structured-data guidelines](https://developers.google.com/search/docs/appearance/structured-data/sd-policies?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) require markup to represent the main visible content and explicitly state that correct markup does not guarantee a search feature.\n\nAn agent-specific JSON API and page-level JSON-LD have different jobs. API defines application contract. JSON-LD describes page using shared vocabulary. Do not force full article body, evidence ledger, workflow state, and permissions into schema markup when a direct data contract can express them more clearly.\n\n## Use feeds for change and sitemaps for coverage\n\nAtom exists to syndicate web content to sites and user agents. [RFC 4287](https://www.rfc-editor.org/rfc/rfc4287.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) defines stable entry IDs plus publication and update metadata. That makes Atom or RSS useful for monitors that ask, “What changed since last run?”\n\nSitemaps answer a different question: “Which public URLs exist?” [Sitemaps protocol](https://www.sitemaps.org/protocol.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) requires a location for each URL and defines optional `lastmod` as page modification date, not sitemap-generation date.\n\nKeep both narrow:\n\n- Feed contains recent or changed entries with stable IDs, canonical links, summaries, and honest update dates.\n- Sitemap contains canonical public URLs with honest modification dates.\n- JSON or JSONL contains full record contract for consumers that need structured fields and article body.\n\nThe feed and sitemap should point to the same canonical URLs used by HTML and data exports.\n\n## Add MCP after read use case exists\n\nMCP is useful when an agent needs bounded runtime access rather than a bulk download. The [MCP 2026-07-28 Resources specification](https://modelcontextprotocol.io/specification/2026-07-28/server/resources?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) defines resources with a URI, name, optional MIME type, content, and annotations such as audience and last modification time. A marketing reference can expose read-only resources such as:\n\n```text\nmarketing://articles/{slug}\nmarketing://sources/{source-id}\nmarketing://topics/{topic-id}\n```\n\nStart with list, search, and read. Add tools only for work that cannot be represented as a resource query. The [MCP 2026-07-28 Tools specification](https://modelcontextprotocol.io/specification/2026-07-28/server/tools?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) supports input and output schemas and requires servers to validate inputs, implement access controls, and rate-limit calls. It recommends user confirmation for sensitive operations.\n\nMCP does not make a server universally available. A compatible client still needs endpoint discovery or configuration, a trust decision, and any required authorization. Treat “available through MCP” as an integration fact, not a distribution claim.\n\nThe [MCP 2026-07-28 Authorization specification](https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) defines optional authorization for HTTP-based transports. Use authorization for protected servers; do not place private records behind an unguessable resource URI and call that access control.\n\n## Use skills as installed operating instructions\n\nAn agent skill can tell a compatible runtime when to use the reference, which fields are authoritative, how to cite records, and which operations need approval. The [Agent Skills specification](https://agentskills.io/specification?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) requires a directory with `SKILL.md`, YAML frontmatter, a name, and a description; compatibility and allowed tools remain optional.\n\nKeep facts in records, not skill prose. The skill should describe the procedure:\n\n1. Search records using topic and use-case fields.\n2. Read canonical article and evidence list.\n3. Report documented claim, observation, inference, and unknown separately.\n4. Cite canonical URL and named primary sources.\n5. Ask for approval before any mutating tool call.\n\nAn installed skill can improve behavior inside a compatible runtime. It does not publish content to other agents or prove that any agent loaded the instructions.\n\n## Separate provenance from presentation\n\nAn agent should be able to answer five questions without interpreting page decoration:\n\n1. Who wrote the record?\n2. Who reviewed it, or is review still required?\n3. Which organizational affiliations could affect judgment?\n4. Which source supports each material claim?\n5. When was the claim last checked?\n\nUse stable author and organization identifiers where available. Preserve evidence status such as `official`, `observed`, `inferred`, or `unknown`. Vendor documentation supports vendor-documented behavior, not independent performance. Corrections should update the visible HTML, machine record, verification date, and history together.\n\n## Separate crawler preference from authorization\n\n`robots.txt` controls requests that compliant crawlers are asked to make. [RFC 9309](https://www.rfc-editor.org/rfc/rfc9309.html?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) states that Robots Exclusion Protocol is not access authorization or substitute for content security. Protect private records and write operations with authentication and authorization.\n\nCrawler roles also differ. [OpenAI’s publisher FAQ](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) distinguishes `OAI-SearchBot` for ChatGPT search access from `GPTBot` for potential training. [Anthropic’s crawler guidance](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler?utm_source=marketingwiki&utm_medium=referral&utm_campaign=what-makes-marketing-content-usable-by-ai-agents) distinguishes `Claude-SearchBot`, `Claude-User`, and `ClaudeBot` for search, user-directed retrieval, and potential training.\n\nChoose crawler policy by purpose. Choose API and MCP permissions by user, resource, and action. Never publish a secret URL in `robots.txt`; file is public.\n\n## Acceptance test for agent usability\n\nBefore calling a reference agent-ready, test its interfaces as one system:\n\n1. Canonical HTML returns `200` without cookies and contains full visible article, provenance, and correction path.\n2. HTML, JSON, JSONL, feed, sitemap, and MCP resource use same stable ID and canonical URL.\n3. JSON parses against documented schema. Every nonblank JSONL line parses as one record.\n4. Publication, modification, and verification dates match underlying events.\n5. Structured data matches visible page and contains no invented reviewer, affiliation, or result.\n6. Feed preserves entry IDs and changes update date only for material edits.\n7. Sitemap lists public canonical pages and uses content modification dates.\n8. MCP list and read return same record fields as bulk export. Unauthorized private reads fail.\n9. Mutating tools validate inputs, require correct scope, and expose confirmation before execution.\n10. Skill references resolve, procedures match live interfaces, and unsupported clients are not claimed.\n11. Crawler rules match declared search, retrieval, and training choices; authentication protects restricted content independently.\n\nPassing these tests proves interfaces are internally consistent and usable by tested clients. It does not prove external discovery, model training, answer inclusion, or citation."}