From 40884e1bf0f726298dafee6da130524e44902faf Mon Sep 17 00:00:00 2001 From: Adam Moussa Date: Thu, 13 Aug 2026 16:59:53 -0400 Subject: [PATCH] fix(wo): add DLQ age alarm before 14-day expiry --- README.md | 2 +- docs/runbook-dlq-recovery.md | 12 ++++++++---- terraform/wo_alarms.tf | 18 ++++++++++++++++++ 3 files changed, 27 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index 2946af5..4cb3fec 100644 --- a/README.md +++ b/README.md @@ -175,7 +175,7 @@ The `-sender-auth-rejected` alarm closes the silent-drop gap in INFRA-107: a The `-duration` and `-throttles` alarms for `po-email-processor` and `workorder-email-processor` supersede the orphaned, CLI-created `Lambda-Duration-*` / `Lambda-Throttles-*` alarms (deleted post-deploy). -**DLQ alarms** (`AWS/SQS`): `po-email-processor-dlq-messages` and `workorder-email-processor-dlq-messages` fire when any message is visible on an email-processor DLQ (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) — a message there means an email was dropped after Lambda exhausted its async retries. Recovery from a DLQ message (no console redrive) is documented in the [DLQ recovery runbook](docs/runbook-dlq-recovery.md). +**DLQ alarms** (`AWS/SQS`): `po-email-processor-dlq-messages` and `workorder-email-processor-dlq-messages` fire when any message is visible on an email-processor DLQ (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) — a message there means an email was dropped after Lambda exhausted its async retries. WO also has `workorder-email-processor-dlq-age` (`ApproximateAgeOfOldestMessage` Maximum, 5 min, `>= 86400` s, eval 1) so an undrained breadcrumb pages again well before 14-day retention expiry. Recovery from a DLQ message (no console redrive) is documented in the [DLQ recovery runbook](docs/runbook-dlq-recovery.md). **SHOC emitter alarms** (bespoke — these don't fit the standard-Lambda-alarm helper's shape): `workorder-shoc-emitter-iterator-age` (`IteratorAge` Maximum, 5 min, `>= 600000` ms, eval 3 / datapoints 2) fires when the stream lags ≥ 10 minutes — SHOC is likely down and the shard is blocking on retries, which is exactly the ordered-backpressure design working, but an operator should know. `workorder-shoc-emitter-failures-messages` and `workorder-shoc-emitter-rejected-messages` (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) page the replay runbook: the `-failures` queue receives ESM failure **metadata** for retry-exhausted records, the `-rejected` queue receives the **full parked payloads** of non-retryable 4xx deliveries (see the [SHOC webhook feed](#shoc-webhook-feed-workorder-shoc-emitter) section and `scripts/replay_shoc_webhooks.py`). All three: ALARM-only → `site-alerts`, `NOT_BREACHING`, per the house rules above. diff --git a/docs/runbook-dlq-recovery.md b/docs/runbook-dlq-recovery.md index 2a23d20..36aff64 100644 --- a/docs/runbook-dlq-recovery.md +++ b/docs/runbook-dlq-recovery.md @@ -20,10 +20,10 @@ target. Recovery is manual, via targeted re-invoke. ## Resource inventory (acct 011934824531 seahaven-prod, us-east-1) -| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarm | +| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarms | |---|---|---|---|---| -| PO | `po-email-processor` | _CDK-generated; fill from stack resources after the first prod deploy_ | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` | -| WO | `workorder-email-processor` | _CDK-generated; fill from stack resources after the first prod deploy_ | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` | +| PO | `po-email-processor` | `po-ingest-EmailProcessorDlqA753DED5-Mn33HvhsDPEu` | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` | +| WO | `workorder-email-processor` | `WorkorderIngestStack-EmailProcessorDlqA753DED5-qCTHrsoEucas` | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` (depth, visible `> 0`); `workorder-email-processor-dlq-age` (oldest message `>= 86400` s) | DLQ URLs are `https://sqs.us-east-1.amazonaws.com/011934824531/`. Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s. @@ -33,6 +33,10 @@ Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s. 1. **Trigger.** The `-dlq-messages` alarm fires (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1). + The WO queue also has `workorder-email-processor-dlq-age`, which fires when + `ApproximateAgeOfOldestMessage` is `>= 86400` seconds (1 day). That is the + ignored-page guard: drain before the 14-day retention window (`1209600` s) + expires. 2. **RECEIVE** the message (do not purge yet): @@ -69,7 +73,7 @@ Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s. 5. **VERIFY the write.** Confirm the downstream effect landed before proceeding: the DynamoDB item exists / was updated (`purchase-orders` for PO, - `WorkOrders` for WO), and the function's log group shows a clean parse (no new + `work-orders` for WO), and the function's log group shows a clean parse (no new error, no new DLQ message). Do not proceed until verified. 6. **PURGE the one message.** Delete only the processed message by its diff --git a/terraform/wo_alarms.tf b/terraform/wo_alarms.tf index 1f2c5ed..f6418fa 100644 --- a/terraform/wo_alarms.tf +++ b/terraform/wo_alarms.tf @@ -65,6 +65,24 @@ resource "aws_cloudwatch_metric_alarm" "wo_email_processor_dlq" { } } +resource "aws_cloudwatch_metric_alarm" "wo_email_processor_dlq_age" { + alarm_name = "workorder-email-processor-dlq-age" + alarm_description = "workorder-email-processor DLQ oldest message is >= 1 day old (drain before 14-day retention expiry)" + comparison_operator = "GreaterThanOrEqualToThreshold" + evaluation_periods = 1 + metric_name = "ApproximateAgeOfOldestMessage" + namespace = "AWS/SQS" + period = 300 + statistic = "Maximum" + threshold = 86400 + treat_missing_data = "notBreaching" + alarm_actions = [data.aws_sns_topic.site_alerts.arn] + + dimensions = { + QueueName = aws_sqs_queue.wo_email_processor_dlq.name + } +} + resource "aws_cloudwatch_metric_alarm" "wo_email_processor_duration" { alarm_name = "workorder-email-processor-duration" alarm_description = "workorder-email-processor p95 duration approaching the 60s timeout"