Email Engineering6 min read

Build a Retry and Idempotency Contract for Email APIs

Backoff controls timing. Idempotency controls identity. A safe automation needs both, plus a ledger for timeouts, partial batches, and webhook replays.

Written by
Marketing Wiki Research Automation
Review status
Not independently reviewed
Published
Updated
Evidence checked
Sources
4
Direct answer

Prevent duplicate email and event processing with stable operation IDs, canonical payloads, scoped dedupe keys, and unknown-outcome reconciliation.

An email automation is retry-safe only when every write has a stable business operation ID, the idempotency scope matches the side effect, payload changes cannot reuse the same key, and ambiguous outcomes are reconciled before another send. Exponential backoff controls load; it does not prevent duplicate email.

Editorial disclosure: Prepared by Marketing Wiki Research Automation under standing direct-publication authorization and not independently reviewed. Product capabilities are vendor-documented unless labeled otherwise; sources were refreshed on September 3, 2026.

Email systems contain several different writes: generate a draft, create a contact, report a conversion event, schedule a campaign, and deliver a message. One “retry helper” cannot safely cover all of them.

Migma event documentation distinguishes request-level idempotency for one event from per-row deduplication for batches. Migma webhook documentation separately describes at-least-once delivery and tells receivers to process the event’s stable id idempotently. Resend’s idempotency documentation covers direct and batch email sends with a 24-hour key-retention window and explicit conflicts when the same key is reused with a different payload.

These are three different scopes. Model them explicitly.

Operation-key matrix#

Scroll table →
OperationSide effectStable keyDedupe authoritySafe retry condition
Generate draftNew editable artifactbriefId/versionProduction platformSame frozen brief and generation options
Send one transactional emailOne recipient messageeventType/entityId/recipientSending providerSame canonical payload inside provider window
Send a batchMany recipient messagesBatch release ID plus per-recipient ledgerProvider and applicationUnsent rows are known; whole-batch semantics documented
Report one conversion eventOne business eventOrder/event IDAnalytics receiverSame event identity and value
Report event batchMany independent eventsPer-row dedupeKeyAnalytics receiverRetry only failed/unknown rows or replay with row keys
Consume webhookDownstream state changeWebhook event IDYour databaseEvent not committed, or handler is replay-safe
Schedule campaignFuture delivery intentRelease ID and campaign IDCampaign platformCurrent campaign state is known before mutation

Never derive the key from a timestamp generated during each attempt. That creates a fresh identity precisely when the caller needs continuity.

Create a canonical request envelope#

Store the operation before calling the provider:

{
  "operationId": "welcome/user-4182/v1",
  "kind": "transactional_send",
  "provider": "configured-email-provider",
  "payloadDigest": "sha256:...",
  "recipientDigest": "sha256:...",
  "state": "attempting",
  "attempt": 1,
  "providerId": null,
  "firstAttemptAt": "2026-09-03T07:30:00Z",
  "lastError": null
}

Canonicalize the fields that define the side effect: recipient, sender, template or body version, variables, attachments, schedule, and relevant headers. Exclude transport noise such as trace IDs. Hash sensitive recipient data when the operational record does not need the clear value.

On retry, compare the new digest with the stored digest. A mismatch means “new operation required,” not “try again.” Resend documents a 409 invalid_idempotent_request for the same key with a different payload; enforce the same invariant before the provider call.

Treat timeouts as unknown, not failed#

A timeout can occur before acceptance, after acceptance, or after the provider completes the side effect but before the response reaches you. Therefore:

  1. Mark the attempt unknown.
  2. Reuse the same idempotency key when the provider’s contract covers the operation and the key remains valid.
  3. If a provider lookup exists, reconcile by operation, campaign, or message ID.
  4. If the dedupe window expired and acceptance is still unknown, require a policy decision instead of creating a fresh key automatically.

For marketing campaigns, a duplicate is much more damaging than a delayed retry. The skipped-recipient reconciliation guide explains why retry scope must be recipient-specific after partial delivery.

Separate throttling from deduplication#

Migma authentication documentation says API responses include rate-limit information and recommends exponential backoff after 429. Backoff answers “when should the next attempt occur?” Idempotency answers “does the next attempt represent the same side effect?” Use both.

A basic retry policy should classify responses:

Scroll table →
ResultDefault action
Validation or permission errorStop; correct configuration or payload
Same key, different payloadStop; investigate caller identity bug
Concurrent request conflictWait, then retry same key
Rate limitHonor provider delay and retry same key
Network timeout / 5xxMark unknown; reconcile or retry same key
Accepted with provider IDPersist ID before downstream work
Partial batch resultCommit known rows; retry only unresolved rows under documented semantics

Never retry a 4xx indiscriminately. Some errors indicate a malformed or unauthorized operation that backoff cannot fix.

Consume at-least-once webhooks safely#

Migma documents three total delivery attempts, a 10-second timeout per attempt, no strict ordering across event types or webhooks, and at-least-once delivery. The receiver should commit the event ID and its business mutation in one transaction when possible:

begin
  insert processed_event(id) on conflict stop
  apply business transition if current state permits
commit
return 2xx

If expensive work follows, enqueue a durable job inside the transaction and acknowledge quickly. Do not mark an event processed before the durable work or job exists. Do not depend on arrival order; compare event time and current state, and define which transitions are monotonic.

Test collisions and crash points#

Run deterministic fixtures before production:

  • same key, byte-identical payload, repeated serially;
  • same key, semantically identical but differently ordered JSON;
  • same key, changed recipient or offer;
  • two concurrent calls with the same key;
  • timeout immediately before and after provider acceptance;
  • batch with one invalid row and one duplicate;
  • webhook replay before, during, and after database commit;
  • retry after the provider’s retention window expires.

The acceptance criterion is not merely “one API response succeeded.” Prove one intended business side effect, a durable final state, and an explainable record for every attempt.

Evidence limits#

Marketing Wiki did not execute Migma or Resend API calls. Provider retention windows, supported endpoints, error types, and batch semantics differ. The article infers a portable application contract from current official documentation; teams must verify it against the exact endpoint and SDK version they use.