Persist one idempotency key and one payload per logical email before the first request. If the response is lost, reuse both. A timeout describes what your client observed; it does not prove the email service rejected the request.
The failure happens between acceptance and acknowledgement
Imagine a checkout commits order 42 and asks an email API to send its receipt. The API saves the message, but the network connection closes before your application receives the response. Your application sees failure; the email service has a valid job.
A retry with a brand-new key creates another logical send. Disabling retries avoids that duplicate but can lose receipts when the first request really never arrived. A stable operation identity lets you recover without guessing which failure occurred.
Idempotency is a contract with a scope and a lifetime. Check the provider’s key retention, account or endpoint scope, concurrent-request behavior, and payload-conflict response. A key is not a promise of global exactly-once delivery forever.
Create the identity with the business event
Write the receipt outbox row in the same database transaction as the completed order. Give it a unique key such as receipt:order-42:v1. If the transaction rolls back, neither the order nor the receipt event should exist.
Persist the exact payload or immutable references that reproduce it. Rendering a template again on each retry can change the text, timestamp, or template version and create a conflict. A deliberate resend is a new business decision with a new identity, not a hidden retry.
Do not put email addresses or access tokens in idempotency keys. Keys often appear in logs. A stable opaque event ID or internal business identifier is enough.
CREATE TABLE receipt_outbox (
event_key text PRIMARY KEY,
payload jsonb NOT NULL,
state text NOT NULL DEFAULT 'pending',
provider_message_id text,
created_at timestamptz NOT NULL DEFAULT now()
);
-- Insert alongside the order transaction, before making network calls.
-- Use worker leases and bounded retry scheduling in production.Classify the result before choosing a retry
Use bounded exponential backoff with jitter for retryable failures. Persist the next attempt time so a process restart does not reset the retry budget or make every worker retry simultaneously. A worker lease prevents two processes from independently treating the same outbox row as unclaimed.
| Result | Action |
|---|---|
| Accepted with a message ID | Save the ID and watch delivery events. Do not submit another message. |
| Network timeout / connection reset | Treat as unresolved. Reconcile or retry the same key and unchanged payload. |
| 429 rate limit | Respect Retry-After and your attempt budget. Preserve the key. |
| Validation or authentication rejection | Correct the underlying problem. Do not repeatedly retry unchanged invalid input. |
| Idempotency conflict | Stop and compare the original payload with the retry. |
| Key expired or unknown retention | Reconcile the original send before creating a new operation. |
Run a lost-response experiment without sending mail
The downloadable Node.js lab accepts a simulated message, deliberately loses the response, and retries with the same key. It also checks a changed-payload conflict and ten matching concurrent calls. It uses an in-memory ledger, so it is a teaching fixture rather than a production idempotency service.
Download the file and run it with Node.js 22 or later. It imports only Node built-ins, uses no credentials, and does not contact a provider.
node idempotency-lab.mjs
# Expected: tests: 4, queued: 1, mode: "demo", sendsRealEmail: falseIdempotent submission is only one boundary
Your application may generate the same business event twice before it reaches the email API. Your webhook endpoint may receive the same notification more than once after delivery. Each boundary needs its own identity and deduplication rule.
Keep an append-only delivery timeline rather than allowing any late event to overwrite the current state. A delayed submitted event should not erase later evidence of server acceptance. Complaints and bounces are distinct events with their own handling, not merely lower or higher numbers in one status ladder.
For an uncertain provider submission, a service may need reconciliation before dispatching again. Do not claim exactly-once delivery simply because the application endpoint accepts idempotency keys.
Keep these failures in your regression suite
For each test, assert both the number of logical messages and the stored state. A successful HTTP response alone cannot prove that a retry policy avoided duplicates.
- The application commits an event and crashes before the first API call.
- The provider accepts a message and the response is lost.
- Two workers claim or retry the same logical event concurrently.
- A retry accidentally renders a different template version.
- A process restarts during backoff or after the provider’s key-retention window.
- Duplicate or out-of-order delivery callbacks arrive after recovery.
Sources and further reading
Written for Stampwing with AI assistance and checked against the linked documentation. Examples are educational; simulated results are labelled. How these resources are maintained.