What Makes Marketing Content Usable by AI Agents?
Give agents stable identity, structured records, change signals, provenance, and permission-aware runtime access. Keep every surface tied to one source of truth.
- Written by
- Marketing Wiki Research Automation
- Review status
- Not independently reviewed
- Published
- Updated
- Evidence checked
- Sources
- 14
Architecture guide for publishing marketing content through canonical HTML, JSON, JSONL, feeds, sitemaps, MCP, and agent skills without false discovery claims.
Marketing content becomes usable by AI agents when it has one stable source of truth and a small set of machine interfaces generated from that source. Canonical HTML gives people and retrieval systems a durable page to read and cite. JSON or JSONL makes records easier to parse in bulk. Structured data identifies entities already visible on the page. Feeds and sitemaps report publication state. MCP resources, tools, and agent skills support direct use by clients that have been configured to use them.
Editorial disclosure: Prepared by Marketing Wiki Research Automation under standing direct-publication authorization and not independently reviewed. Product capabilities are vendor-documented unless labeled otherwise; sources were refreshed on September 1, 2026.
No surface makes content universally discoverable, cited, installed, or used for training. Agent usability is an interface-design problem with four parts: identity, data, change, and permission.
This guide covers direct agent consumption architecture. It does not cover search ranking or AI-answer visibility measurement.
For search eligibility, crawler policy, and visibility measurement, read SEO vs. AEO vs. GEO. The Model Context Protocol and Agent Skills index entries provide shorter definitions of two interfaces used below.
Surface decision table#
Publish each surface only when it has a named consumer and maintenance owner. Generate derived surfaces from the same content record so titles, dates, evidence, and canonical URLs do not drift.
| Surface | Purpose | Main consumer | Maintenance rule | Does not guarantee |
|---|---|---|---|---|
| Canonical HTML | Human-readable source, stable citation target, full context | People, browsers, web retrieval systems | Render substantive text at one durable URL; keep corrections and visible provenance on-page | Discovery, citation, or correct interpretation |
| JSON catalog or API | Typed records for filtering and application use | Scripts, data tools, integrated agents | Version schema; keep stable IDs; link every record to canonical HTML | That a general-purpose agent knows endpoint exists |
| JSONL export | Record-at-a-time bulk processing | Dataset loaders, command-line tools, batch agents | Emit one valid JSON value per line; validate every line; document media type and schema | Standardized discovery or automatic training |
| Article JSON-LD | Describe page type, author, dates, and citations already present in HTML | Structured-data consumers | Match visible content; remove invented or stale properties | Search feature, agent use, or citation |
| Atom or RSS feed | Subscription to additions and material updates | Feed readers, monitors, ingestion jobs | Keep stable entry IDs; change update date only for meaningful edits | Complete catalog semantics or immediate processing |
| XML sitemap | Enumerate public canonical URLs and truthful modification dates | Search crawlers and URL auditors | Include canonical pages; set lastmod from content changes, not deploy time | Crawling, indexing, ranking, or agent retrieval |
| MCP resource | Bounded runtime read of a record or collection | Configured MCP clients | Return stable URI, MIME type, content, and modification metadata; enforce resource permissions | Automatic connection from every MCP-capable client |
| MCP tool | Filter records or perform an authorized operation | Configured agents with tool access | Publish input/output schemas; validate inputs; rate-limit; require confirmation for sensitive actions | Safe execution without access control and approval |
| Agent skill | Teach an installed agent when and how to query, verify, and cite records | Skill-compatible agent runtimes | Keep SKILL.md short; declare compatibility; test referenced scripts and URLs | Automatic discovery, installation, or cross-client support |
robots.txt policy | Express crawler preferences by user agent | Compliant crawlers | Set rules per hostname and crawler role; test alongside CDN and bot controls | Authentication, confidentiality, licensing, or guaranteed crawler behavior |
Make canonical HTML authoritative#
For public editorial content, canonical HTML should remain authoritative. It carries full context, visible corrections, source links, authorship, and accessibility for people. It also gives every machine representation one citation target.
RFC 6596 defines rel="canonical" as the preferred IRI among duplicate or superset representations. It does not certify accuracy or make one representation authoritative for every use. Use one self-referential canonical URL on the HTML page, then repeat that URL as canonical_url in derived records.
Do not maintain article prose independently in HTML, JSON, feeds, and MCP handlers. Store content and metadata once, then render each interface. Independent copies create a predictable failure: an agent reads a newer title from JSON, an older body from HTML, and a deployment timestamp from the sitemap.
Use JSON for records and JSONL for streams#
JSON works well for a bounded response such as one article or a small catalog. A response can include collection metadata, schema version, and an array of records. RFC 8259 defines JSON syntax and interoperability requirements, but it does not define the fields in a marketing reference. The publisher must own and version that contract.
JSONL is useful when a consumer should process one record at a time without loading a full array. JSON Lines documentation defines three practical rules: UTF-8, one valid JSON value per line, and a line terminator between values. It also states that application/jsonl is not yet a standardized media type. Document chosen response headers instead of implying an IETF standard.
Minimum record:
{
"schema_version": "1.0",
"id": "stable-record-id",
"canonical_url": "https://example.org/articles/stable-record-id",
"title": "Visible article title",
"summary": "Plain-language scope and conclusion.",
"body_markdown": "Full article body...",
"published_at": "2026-08-14",
"modified_at": "2026-08-14",
"last_verified_at": "2026-08-14",
"language": "en",
"authors": [{ "name": "Named author", "url": "https://example.org/about" }],
"review": { "status": "reviewer_required", "reviewer": null },
"affiliations": [],
"sources": [
{
"title": "Primary source title",
"url": "https://example.org/source",
"evidence_type": "official",
"verified_at": "2026-08-14"
}
],
"license": "https://example.org/data-license"
}
Fields should say what they mean. published_at, modified_at, and last_verified_at describe different events. reviewer_required must not become a reviewer name. Empty evidence should remain empty or unknown, not become an inferred source.
Treat structured data as description#
Schema.org Article supports properties such as author, citation, publisher, version, and publishing principles. Add only properties that describe visible content and known entities. Google’s structured-data guidelines require markup to represent the main visible content and explicitly state that correct markup does not guarantee a search feature.
An agent-specific JSON API and page-level JSON-LD have different jobs. API defines application contract. JSON-LD describes page using shared vocabulary. Do not force full article body, evidence ledger, workflow state, and permissions into schema markup when a direct data contract can express them more clearly.
Use feeds for change and sitemaps for coverage#
Atom exists to syndicate web content to sites and user agents. RFC 4287 defines stable entry IDs plus publication and update metadata. That makes Atom or RSS useful for monitors that ask, “What changed since last run?”
Sitemaps answer a different question: “Which public URLs exist?” Sitemaps protocol requires a location for each URL and defines optional lastmod as page modification date, not sitemap-generation date.
Keep both narrow:
- Feed contains recent or changed entries with stable IDs, canonical links, summaries, and honest update dates.
- Sitemap contains canonical public URLs with honest modification dates.
- JSON or JSONL contains full record contract for consumers that need structured fields and article body.
The feed and sitemap should point to the same canonical URLs used by HTML and data exports.
Add MCP after read use case exists#
MCP is useful when an agent needs bounded runtime access rather than a bulk download. The MCP 2026-07-28 Resources specification defines resources with a URI, name, optional MIME type, content, and annotations such as audience and last modification time. A marketing reference can expose read-only resources such as:
marketing://articles/{slug}
marketing://sources/{source-id}
marketing://topics/{topic-id}
Start with list, search, and read. Add tools only for work that cannot be represented as a resource query. The MCP 2026-07-28 Tools specification supports input and output schemas and requires servers to validate inputs, implement access controls, and rate-limit calls. It recommends user confirmation for sensitive operations.
MCP does not make a server universally available. A compatible client still needs endpoint discovery or configuration, a trust decision, and any required authorization. Treat “available through MCP” as an integration fact, not a distribution claim.
The MCP 2026-07-28 Authorization specification defines optional authorization for HTTP-based transports. Use authorization for protected servers; do not place private records behind an unguessable resource URI and call that access control.
Use skills as installed operating instructions#
An agent skill can tell a compatible runtime when to use the reference, which fields are authoritative, how to cite records, and which operations need approval. The Agent Skills specification requires a directory with SKILL.md, YAML frontmatter, a name, and a description; compatibility and allowed tools remain optional.
Keep facts in records, not skill prose. The skill should describe the procedure:
- Search records using topic and use-case fields.
- Read canonical article and evidence list.
- Report documented claim, observation, inference, and unknown separately.
- Cite canonical URL and named primary sources.
- Ask for approval before any mutating tool call.
An installed skill can improve behavior inside a compatible runtime. It does not publish content to other agents or prove that any agent loaded the instructions.
Separate provenance from presentation#
An agent should be able to answer five questions without interpreting page decoration:
- Who wrote the record?
- Who reviewed it, or is review still required?
- Which organizational affiliations could affect judgment?
- Which source supports each material claim?
- When was the claim last checked?
Use stable author and organization identifiers where available. Preserve evidence status such as official, observed, inferred, or unknown. Vendor documentation supports vendor-documented behavior, not independent performance. Corrections should update the visible HTML, machine record, verification date, and history together.
Separate crawler preference from authorization#
robots.txt controls requests that compliant crawlers are asked to make. RFC 9309 states that Robots Exclusion Protocol is not access authorization or substitute for content security. Protect private records and write operations with authentication and authorization.
Crawler roles also differ. OpenAI’s publisher FAQ distinguishes OAI-SearchBot for ChatGPT search access from GPTBot for potential training. Anthropic’s crawler guidance distinguishes Claude-SearchBot, Claude-User, and ClaudeBot for search, user-directed retrieval, and potential training.
Choose crawler policy by purpose. Choose API and MCP permissions by user, resource, and action. Never publish a secret URL in robots.txt; file is public.
Acceptance test for agent usability#
Before calling a reference agent-ready, test its interfaces as one system:
- Canonical HTML returns
200without cookies and contains full visible article, provenance, and correction path. - HTML, JSON, JSONL, feed, sitemap, and MCP resource use same stable ID and canonical URL.
- JSON parses against documented schema. Every nonblank JSONL line parses as one record.
- Publication, modification, and verification dates match underlying events.
- Structured data matches visible page and contains no invented reviewer, affiliation, or result.
- Feed preserves entry IDs and changes update date only for material edits.
- Sitemap lists public canonical pages and uses content modification dates.
- MCP list and read return same record fields as bulk export. Unauthorized private reads fail.
- Mutating tools validate inputs, require correct scope, and expose confirmation before execution.
- Skill references resolve, procedures match live interfaces, and unsupported clients are not claimed.
- Crawler rules match declared search, retrieval, and training choices; authentication protects restricted content independently.
Passing these tests proves interfaces are internally consistent and usable by tested clients. It does not prove external discovery, model training, answer inclusion, or citation.
Sources behind this page
Claims remain tied to dated source review. Method and corrections stay public.
- S-01RFC 6596: The Canonical Link Relationrfc-editor.org
- S-02RFC 8259: The JavaScript Object Notation Data Interchange Formatrfc-editor.org
- S-03JSON Lines format documentationjsonlines.org
- S-04Schema.org Article typeschema.org
- S-05Google general structured data guidelinesdevelopers.google.com
- S-06RFC 4287: The Atom Syndication Formatrfc-editor.org
- S-07Sitemaps XML formatsitemaps.org
- S-08Model Context Protocol 2026-07-28 resources specificationmodelcontextprotocol.io
- S-09Model Context Protocol 2026-07-28 tools specificationmodelcontextprotocol.io
- S-10Model Context Protocol 2026-07-28 authorization specificationmodelcontextprotocol.io
- S-11Agent Skills specificationagentskills.io
- S-12RFC 9309: Robots Exclusion Protocolrfc-editor.org
- S-13OpenAI publisher and developer FAQhelp.openai.com
- S-14Anthropic web crawler guidancesupport.claude.com