# DLQ Recovery Runbook — Email-Processor Dead-Letter Queues Operational procedure for draining an email-processor dead-letter queue (DLQ) after a failed async parse. There is **no console redrive-to-source** for these queues; recovery is a manual, targeted re-invoke via `scripts/reprocess.py`. ## Scope / mechanism Both email processors set `dead_letter_queue=` on the `lambda_.Function` construct. This is the legacy per-function **Lambda `DeadLetterConfig`** (asynchronous-invocation DLQ), **not** an EventInvokeConfig on-failure Destination — `aws lambda get-function-event-invoke-config` returns `ResourceNotFoundException` for both functions (no destination config exists). A failed async invocation lands on the DLQ only **after** Lambda exhausts its automatic retries. There is **no console redrive-to-source**: the SQS console's "redrive to source" applies only to SQS-to-SQS DLQ relationships, not to a Lambda `DeadLetterConfig` target. Recovery is manual, via targeted re-invoke. ## Resource inventory (acct 011934824531 seahaven-prod, us-east-1) | Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarms | |---|---|---|---|---| | PO | `po-email-processor` | `po-ingest-EmailProcessorDlqA753DED5-Mn33HvhsDPEu` | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` | | WO | `workorder-email-processor` | `WorkorderIngestStack-EmailProcessorDlqA753DED5-qCTHrsoEucas` | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` (depth, visible `> 0`); `workorder-email-processor-dlq-age` (oldest message `>= 86400` s) | DLQ URLs are `https://sqs.us-east-1.amazonaws.com/011934824531/`. Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s. (Mgmt-account DLQs were removed with the PLAT-67 stack teardown on 2026-08-05.) ## Recovery procedure (no redrive — receive → extract key → targeted re-invoke → verify → purge) 1. **Trigger.** The `-dlq-messages` alarm fires (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1). The WO queue also has `workorder-email-processor-dlq-age`, which fires when `ApproximateAgeOfOldestMessage` is `>= 86400` seconds (1 day). That is the ignored-page guard: drain before the 14-day retention window (`1209600` s) expires. 2. **RECEIVE** the message (do not purge yet): ```bash aws sqs receive-message \ --queue-url \ --max-number-of-messages 1 \ --visibility-timeout 120 \ --wait-time-seconds 5 ``` Capture the `ReceiptHandle` from the response. 3. **EXTRACT the S3 key.** The `DeadLetterConfig` message `Body` is the original async invocation payload — the S3 event JSON. Read `Records[0].s3.bucket.name` and `Records[0].s3.object.key` from the `Body`. This is the **raw** object key that reprocess/S3 emitted (no URL-decoding applied). 4. **RE-INVOKE (targeted; dry-run first).** Confirm the key with a dry-run, then execute: ```bash # PO queue: python scripts/reprocess.py --pipeline po --key '' # dry-run python scripts/reprocess.py --pipeline po --key '' --execute # re-invoke # WO queue: python scripts/reprocess.py --pipeline wo --key '' --execute ``` This re-invokes the **same** function with the same raw-key synthetic S3 event (a single object — **not** `--all`). 5. **VERIFY the write.** Confirm the downstream effect landed before proceeding: the DynamoDB item exists / was updated (`purchase-orders` for PO, `work-orders` for WO), and the function's log group shows a clean parse (no new error, no new DLQ message). Do not proceed until verified. 6. **PURGE the one message.** Delete only the processed message by its `ReceiptHandle`: ```bash aws sqs delete-message --queue-url --receipt-handle '' ``` Do **not** `purge-queue` — that would drop unexamined breadcrumbs. AWS access is otherwise read-only; `delete-message` on a DLQ you are actively draining is the one write this runbook performs. ## Recovery windows - **DLQ breadcrumb retention: 14 days** (`MessageRetentionPeriod=1209600s`, confirmed live on both queues). After 14 days the breadcrumb is gone. - **Raw-email S3 retention: 90 days.** Both buckets have `RemovalPolicy.RETAIN`, **but** a single Enabled lifecycle rule (`Expiration.Days=90`, `Filter.Prefix=''`) expires objects bucket-wide. **The lifecycle rule OVERRIDES the RETAIN policy — S3 is the real replay floor:** a raw email is gone at ~90 days regardless of the table RETAIN policy. Because 14 days (DLQ) < 90 days (S3), any object referenced by a live DLQ breadcrumb is always still in S3, so DLQ replay within its 14-day window is never blocked by S3 expiry. The 90-day floor binds only for replays reconstructed from other sources (e.g. logs) after the breadcrumb has expired. ## Intentional non-DLQ drops (do NOT hunt for these in the DLQ) Two drop classes **never** produce a DLQ message because they are fail-closed **skips, not errors** — the handler returns normally (no raise, no retry, no DLQ message): - **Sender-auth rejections.** Logged as a structured `sender_auth_rejected` warning and skipped; covered by the `-sender-auth-rejected` alarm (log-metric filter), **not** the DLQ. - **`ai_fallback_rejected` drops.** AI-fallback output that failed the fail-closed `validate_ai_fallback()` gate, emitted as a `ParseMethod=ai_fallback_rejected` EMF datapoint and dropped without a DynamoDB write; covered by the `-ai-fallback-rejected` alarm, **not** the DLQ. If mail is missing but the DLQ is empty, check those two alarms / log filters — the email was intentionally rejected. Re-invoking it via reprocess will just be rejected again; fix the sender-auth config or the upstream email, not the DLQ.