5.4 KiB
DLQ Recovery Runbook — Email-Processor Dead-Letter Queues
Operational procedure for draining an email-processor dead-letter queue (DLQ)
after a failed async parse. There is no console redrive-to-source for these
queues; recovery is a manual, targeted re-invoke via scripts/reprocess.py.
Scope / mechanism
Both email processors set dead_letter_queue= on the lambda_.Function
construct. This is the legacy per-function Lambda DeadLetterConfig
(asynchronous-invocation DLQ), not an EventInvokeConfig on-failure
Destination — aws lambda get-function-event-invoke-config returns
ResourceNotFoundException for both functions (no destination config exists). A
failed async invocation lands on the DLQ only after Lambda exhausts its
automatic retries.
There is no console redrive-to-source: the SQS console's "redrive to source"
applies only to SQS-to-SQS DLQ relationships, not to a Lambda DeadLetterConfig
target. Recovery is manual, via targeted re-invoke.
Resource inventory (acct 011934824531 seahaven-prod, us-east-1)
| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarm |
|---|---|---|---|---|
| PO | po-email-processor |
CDK-generated; fill from stack resources after the first prod deploy | po-ingest-emails-011934824531 |
po-email-processor-dlq-messages |
| WO | workorder-email-processor |
CDK-generated; fill from stack resources after the first prod deploy | workorder-ingest-emails-011934824531 |
workorder-email-processor-dlq-messages |
DLQ URLs are https://sqs.us-east-1.amazonaws.com/011934824531/<queue-name>.
Both DLQs: 14-day retention, SSE, TLS-enforced, VisibilityTimeout 30s.
(Mgmt-account DLQs were removed with the PLAT-67 stack teardown on 2026-08-05.)
Recovery procedure (no redrive — receive → extract key → targeted re-invoke → verify → purge)
-
Trigger. The
<fn>-dlq-messagesalarm fires (ApproximateNumberOfMessagesVisibleMaximum, 5 min,> 0, eval 1). -
RECEIVE the message (do not purge yet):
aws sqs receive-message \ --queue-url <DLQ_URL> \ --max-number-of-messages 1 \ --visibility-timeout 120 \ --wait-time-seconds 5Capture the
ReceiptHandlefrom the response. -
EXTRACT the S3 key. The
DeadLetterConfigmessageBodyis the original async invocation payload — the S3 event JSON. ReadRecords[0].s3.bucket.nameandRecords[0].s3.object.keyfrom theBody. This is the raw object key that reprocess/S3 emitted (no URL-decoding applied). -
RE-INVOKE (targeted; dry-run first). Confirm the key with a dry-run, then execute:
# PO queue: python scripts/reprocess.py --pipeline po --key '<key>' # dry-run python scripts/reprocess.py --pipeline po --key '<key>' --execute # re-invoke # WO queue: python scripts/reprocess.py --pipeline wo --key '<key>' --executeThis re-invokes the same function with the same raw-key synthetic S3 event (a single object — not
--all). -
VERIFY the write. Confirm the downstream effect landed before proceeding: the DynamoDB item exists / was updated (
purchase-ordersfor PO,WorkOrdersfor WO), and the function's log group shows a clean parse (no new error, no new DLQ message). Do not proceed until verified. -
PURGE the one message. Delete only the processed message by its
ReceiptHandle:aws sqs delete-message --queue-url <DLQ_URL> --receipt-handle '<ReceiptHandle>'Do not
purge-queue— that would drop unexamined breadcrumbs. AWS access is otherwise read-only;delete-messageon a DLQ you are actively draining is the one write this runbook performs.
Recovery windows
- DLQ breadcrumb retention: 14 days (
MessageRetentionPeriod=1209600s, confirmed live on both queues). After 14 days the breadcrumb is gone. - Raw-email S3 retention: 90 days. Both buckets have
RemovalPolicy.RETAIN, but a single Enabled lifecycle rule (Expiration.Days=90,Filter.Prefix='') expires objects bucket-wide. The lifecycle rule OVERRIDES the RETAIN policy — S3 is the real replay floor: a raw email is gone at ~90 days regardless of the table RETAIN policy.
Because 14 days (DLQ) < 90 days (S3), any object referenced by a live DLQ breadcrumb is always still in S3, so DLQ replay within its 14-day window is never blocked by S3 expiry. The 90-day floor binds only for replays reconstructed from other sources (e.g. logs) after the breadcrumb has expired.
Intentional non-DLQ drops (do NOT hunt for these in the DLQ)
Two drop classes never produce a DLQ message because they are fail-closed skips, not errors — the handler returns normally (no raise, no retry, no DLQ message):
- Sender-auth rejections. Logged as a structured
sender_auth_rejectedwarning and skipped; covered by the<fn>-sender-auth-rejectedalarm (log-metric filter), not the DLQ. ai_fallback_rejecteddrops. AI-fallback output that failed the fail-closedvalidate_ai_fallback()gate, emitted as aParseMethod=ai_fallback_rejectedEMF datapoint and dropped without a DynamoDB write; covered by the<fn>-ai-fallback-rejectedalarm, not the DLQ.
If mail is missing but the DLQ is empty, check those two alarms / log filters — the email was intentionally rejected. Re-invoking it via reprocess will just be rejected again; fix the sender-auth config or the upstream email, not the DLQ.