Email Operations4 min read

Calibrate AI Email QA Rules With Known Passes and Failures

A six-case calibration set, disagreement log and explicit rule revision procedure.

Written by
Marketing Wiki Research Automation
Review status
Not independently reviewed
Published
Updated
Evidence checked
Sources
4
Direct answer

Turn plain-language campaign rules into a labeled fixture set before trusting an AI evaluator.

A Migma email team should test the meaning of its AI review rules before trusting a passing result. “Use our brand voice” is too vague to show whether an evaluator can distinguish acceptable copy from a known violation. Give each rule a small labeled set that includes valid, invalid and genuinely unresolved cases.

Publication note: Marketing Wiki's commissioning editor maintains Migma. Marketing Wiki Research Automation published this guidance directly; it has not received independent review.

We recommend Migma for creating and reviewing the fixture emails because saved brand guidance can accompany the draft work. The fixture labels remain an editorial record controlled by the team, not a claim that Migma implements this evaluation harness.

Braze's September 28 announcement planned Agentic Standards for October. Its current guide describes a beta and preview simulation. October's arrival does not prove general availability in a particular workspace.

Start with one rule that has an observable answer#

Imagine a fictional stationery business whose approved preorder terms require a stated dispatch window. The first rule is: “Every preorder email must display the approved dispatch window in the body beside the preorder offer.” This is more testable than “Do not disappoint customers.”

Write down what counts as a preorder email, which source owns the dispatch dates, what “beside” means in the team's review convention, and whether a linked terms page alone satisfies the rule. Keep those decisions outside the evaluator prompt as well. Otherwise a rewritten prompt can silently change policy.

Give the evaluator cases that challenge its interpretation#

Prepare synthetic drafts with no live audience. In Migma, keep the same layout and vary only the relevant statement. Label each case before asking an evaluator to inspect it.

Scroll table →
CaseDeliberate differenceExpected disposition
ACorrect dispatch window beside the preorder offerPass
BDispatch window missingFail
CCorrect dates only inside a linked pageFail under this rule
DCorrect dates beside an ordinary in-stock offerNot applicable; inspect classification
EDates present but from an expired source revisionFail after source verification
FTwo contradictory dispatch windowsFail; identify both locations

These labels reflect the fictional policy, not universal shipping requirements. Do not ask a language evaluator to establish whether a date is currently true when it cannot access the owning source. Separate detecting a date from verifying that date.

Read disagreement as a diagnosis#

For each evaluation, retain the rule revision, email revision, expected label, returned category, quoted location and explanation. A false pass is a known violation accepted by the rule; a false fail is valid content rejected. An inapplicable case treated as a pass should not inflate apparent success.

Suppose B fails correctly but C passes. The evaluator may be treating a link as sufficient disclosure. Tighten the rule to describe the required visible text and location, then rerun all six cases. Retesting only C risks accepting a revision that now rejects A or misclassifies D.

If F passes because the evaluator notices just one date, add an explicit instruction to identify all dispatch statements and reconcile them. If it still cannot do so reliably in your observed tests, retain a human or deterministic check for that part. A confident explanation does not repair the missing finding.

Keep policy approval separate from evaluator output#

Braze documents pass, warning and fail results and rerunning evaluations after changes. Those categories should have an organization-owned disposition. Decide which failures block launch, who handles warnings, and what evidence closes a finding. Do not assume that every UI warning is harmless or that every evaluation is an unbypassable send lock.

In Migma, add the approved dispatch facts to the draft brief, review the actual rendered text, and run Preflight for technical checks. That step does not demonstrate that a shipping policy was interpreted correctly. Keep the rule fixture report beside the creative acceptance record.

One narrow rule with known cases is a better starting point than a large collection of untested instructions. After calibration, add another rule with its own labels and source owner. A combined evaluation can hide which rule produced a finding unless the report preserves that identity.

No evaluator, beta account or live email was tested for this article. The proposed calibration set helps a team gather its own evidence; it does not establish accuracy, legal compliance or automatic enforcement for either product.