procurement-ingest/docs/runbook-dlq-recovery.md
2026-08-13 17:39:06 -04:00

122 lines
5.7 KiB
Markdown

# DLQ Recovery Runbook — Email-Processor Dead-Letter Queues
Operational procedure for draining an email-processor dead-letter queue (DLQ)
after a failed async parse. There is **no console redrive-to-source** for these
queues; recovery is a manual, targeted re-invoke via `scripts/reprocess.py`.
## Scope / mechanism
Both email processors set `dead_letter_queue=` on the `lambda_.Function`
construct. This is the legacy per-function **Lambda `DeadLetterConfig`**
(asynchronous-invocation DLQ), **not** an EventInvokeConfig on-failure
Destination — `aws lambda get-function-event-invoke-config` returns
`ResourceNotFoundException` for both functions (no destination config exists). A
failed async invocation lands on the DLQ only **after** Lambda exhausts its
automatic retries.
There is **no console redrive-to-source**: the SQS console's "redrive to source"
applies only to SQS-to-SQS DLQ relationships, not to a Lambda `DeadLetterConfig`
target. Recovery is manual, via targeted re-invoke.
## Resource inventory (acct 011934824531 seahaven-prod, us-east-1)
| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarms |
|---|---|---|---|---|
| PO | `po-email-processor` | `po-ingest-EmailProcessorDlqA753DED5-Mn33HvhsDPEu` | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` |
| WO | `workorder-email-processor` | `WorkorderIngestStack-EmailProcessorDlqA753DED5-qCTHrsoEucas` | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` (depth, visible `> 0`); `workorder-email-processor-dlq-age` (oldest message `>= 86400` s) |
DLQ URLs are `https://sqs.us-east-1.amazonaws.com/011934824531/<queue-name>`.
Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
(Mgmt-account DLQs were removed with the PLAT-67 stack teardown on 2026-08-05.)
## Recovery procedure (no redrive — receive → extract key → targeted re-invoke → verify → purge)
1. **Trigger.** The `<fn>-dlq-messages` alarm fires
(`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1).
The WO queue also has `workorder-email-processor-dlq-age`, which fires when
`ApproximateAgeOfOldestMessage` is `>= 86400` seconds (1 day). That is the
ignored-page guard: drain before the 14-day retention window (`1209600` s)
expires.
2. **RECEIVE** the message (do not purge yet):
```bash
aws sqs receive-message \
--queue-url <DLQ_URL> \
--max-number-of-messages 1 \
--visibility-timeout 120 \
--wait-time-seconds 5
```
Capture the `ReceiptHandle` from the response.
3. **EXTRACT the S3 key.** The `DeadLetterConfig` message `Body` is the original
async invocation payload — the S3 event JSON. Read
`Records[0].s3.bucket.name` and `Records[0].s3.object.key` from the `Body`.
This is the **raw** object key that reprocess/S3 emitted (no URL-decoding
applied).
4. **RE-INVOKE (targeted; dry-run first).** Confirm the key with a dry-run, then
execute:
```bash
# PO queue:
python scripts/reprocess.py --pipeline po --key '<key>' # dry-run
python scripts/reprocess.py --pipeline po --key '<key>' --execute # re-invoke
# WO queue:
python scripts/reprocess.py --pipeline wo --key '<key>' --execute
```
This re-invokes the **same** function with the same raw-key synthetic S3
event (a single object — **not** `--all`).
5. **VERIFY the write.** Confirm the downstream effect landed before proceeding:
the DynamoDB item exists / was updated (`purchase-orders` for PO,
`work-orders` for WO), and the function's log group shows a clean parse (no new
error, no new DLQ message). Do not proceed until verified.
6. **PURGE the one message.** Delete only the processed message by its
`ReceiptHandle`:
```bash
aws sqs delete-message --queue-url <DLQ_URL> --receipt-handle '<ReceiptHandle>'
```
Do **not** `purge-queue` — that would drop unexamined breadcrumbs. AWS access
is otherwise read-only; `delete-message` on a DLQ you are actively draining is
the one write this runbook performs.
## Recovery windows
- **DLQ breadcrumb retention: 14 days** (`MessageRetentionPeriod=1209600s`,
confirmed live on both queues). After 14 days the breadcrumb is gone.
- **Raw-email S3 retention: 90 days.** Both buckets have
`RemovalPolicy.RETAIN`, **but** a single Enabled lifecycle rule
(`Expiration.Days=90`, `Filter.Prefix=''`) expires objects bucket-wide. **The
lifecycle rule OVERRIDES the RETAIN policy — S3 is the real replay floor:** a
raw email is gone at ~90 days regardless of the table RETAIN policy.
Because 14 days (DLQ) < 90 days (S3), any object referenced by a live DLQ
breadcrumb is always still in S3, so DLQ replay within its 14-day window is never
blocked by S3 expiry. The 90-day floor binds only for replays reconstructed from
other sources (e.g. logs) after the breadcrumb has expired.
## Intentional non-DLQ drops (do NOT hunt for these in the DLQ)
Two drop classes **never** produce a DLQ message because they are fail-closed
**skips, not errors** — the handler returns normally (no raise, no retry, no DLQ
message):
- **Sender-auth rejections.** Logged as a structured `sender_auth_rejected`
warning and skipped; covered by the `<fn>-sender-auth-rejected` alarm
(log-metric filter), **not** the DLQ.
- **`ai_fallback_rejected` drops.** AI-fallback output that failed the
fail-closed `validate_ai_fallback()` gate, emitted as a
`ParseMethod=ai_fallback_rejected` EMF datapoint and dropped without a
DynamoDB write; covered by the `<fn>-ai-fallback-rejected` alarm, **not** the
DLQ.
If mail is missing but the DLQ is empty, check those two alarms / log filters —
the email was intentionally rejected. Re-invoking it via reprocess will just be
rejected again; fix the sender-auth config or the upstream email, not the DLQ.