fix(wo): add DLQ age alarm before 14-day expiry

This commit is contained in:
Adam Moussa 2026-08-13 16:59:53 -04:00
parent a73603e16a
commit 40884e1bf0
No known key found for this signature in database
3 changed files with 27 additions and 5 deletions

View file

@ -175,7 +175,7 @@ The `<fn>-sender-auth-rejected` alarm closes the silent-drop gap in INFRA-107: a
The `<fn>-duration` and `<fn>-throttles` alarms for `po-email-processor` and `workorder-email-processor` supersede the orphaned, CLI-created `Lambda-Duration-*` / `Lambda-Throttles-*` alarms (deleted post-deploy).
**DLQ alarms** (`AWS/SQS`): `po-email-processor-dlq-messages` and `workorder-email-processor-dlq-messages` fire when any message is visible on an email-processor DLQ (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) — a message there means an email was dropped after Lambda exhausted its async retries. Recovery from a DLQ message (no console redrive) is documented in the [DLQ recovery runbook](docs/runbook-dlq-recovery.md).
**DLQ alarms** (`AWS/SQS`): `po-email-processor-dlq-messages` and `workorder-email-processor-dlq-messages` fire when any message is visible on an email-processor DLQ (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) — a message there means an email was dropped after Lambda exhausted its async retries. WO also has `workorder-email-processor-dlq-age` (`ApproximateAgeOfOldestMessage` Maximum, 5 min, `>= 86400` s, eval 1) so an undrained breadcrumb pages again well before 14-day retention expiry. Recovery from a DLQ message (no console redrive) is documented in the [DLQ recovery runbook](docs/runbook-dlq-recovery.md).
**SHOC emitter alarms** (bespoke — these don't fit the standard-Lambda-alarm helper's shape): `workorder-shoc-emitter-iterator-age` (`IteratorAge` Maximum, 5 min, `>= 600000` ms, eval 3 / datapoints 2) fires when the stream lags ≥ 10 minutes — SHOC is likely down and the shard is blocking on retries, which is exactly the ordered-backpressure design working, but an operator should know. `workorder-shoc-emitter-failures-messages` and `workorder-shoc-emitter-rejected-messages` (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) page the replay runbook: the `-failures` queue receives ESM failure **metadata** for retry-exhausted records, the `-rejected` queue receives the **full parked payloads** of non-retryable 4xx deliveries (see the [SHOC webhook feed](#shoc-webhook-feed-workorder-shoc-emitter) section and `scripts/replay_shoc_webhooks.py`). All three: ALARM-only → `site-alerts`, `NOT_BREACHING`, per the house rules above.

View file

@ -20,10 +20,10 @@ target. Recovery is manual, via targeted re-invoke.
## Resource inventory (acct 011934824531 seahaven-prod, us-east-1)
| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarm |
| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarms |
|---|---|---|---|---|
| PO | `po-email-processor` | _CDK-generated; fill from stack resources after the first prod deploy_ | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` |
| WO | `workorder-email-processor` | _CDK-generated; fill from stack resources after the first prod deploy_ | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` |
| PO | `po-email-processor` | `po-ingest-EmailProcessorDlqA753DED5-Mn33HvhsDPEu` | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` |
| WO | `workorder-email-processor` | `WorkorderIngestStack-EmailProcessorDlqA753DED5-qCTHrsoEucas` | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` (depth, visible `> 0`); `workorder-email-processor-dlq-age` (oldest message `>= 86400` s) |
DLQ URLs are `https://sqs.us-east-1.amazonaws.com/011934824531/<queue-name>`.
Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
@ -33,6 +33,10 @@ Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
1. **Trigger.** The `<fn>-dlq-messages` alarm fires
(`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1).
The WO queue also has `workorder-email-processor-dlq-age`, which fires when
`ApproximateAgeOfOldestMessage` is `>= 86400` seconds (1 day). That is the
ignored-page guard: drain before the 14-day retention window (`1209600` s)
expires.
2. **RECEIVE** the message (do not purge yet):
@ -69,7 +73,7 @@ Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
5. **VERIFY the write.** Confirm the downstream effect landed before proceeding:
the DynamoDB item exists / was updated (`purchase-orders` for PO,
`WorkOrders` for WO), and the function's log group shows a clean parse (no new
`work-orders` for WO), and the function's log group shows a clean parse (no new
error, no new DLQ message). Do not proceed until verified.
6. **PURGE the one message.** Delete only the processed message by its

View file

@ -65,6 +65,24 @@ resource "aws_cloudwatch_metric_alarm" "wo_email_processor_dlq" {
}
}
resource "aws_cloudwatch_metric_alarm" "wo_email_processor_dlq_age" {
alarm_name = "workorder-email-processor-dlq-age"
alarm_description = "workorder-email-processor DLQ oldest message is >= 1 day old (drain before 14-day retention expiry)"
comparison_operator = "GreaterThanOrEqualToThreshold"
evaluation_periods = 1
metric_name = "ApproximateAgeOfOldestMessage"
namespace = "AWS/SQS"
period = 300
statistic = "Maximum"
threshold = 86400
treat_missing_data = "notBreaching"
alarm_actions = [data.aws_sns_topic.site_alerts.arn]
dimensions = {
QueueName = aws_sqs_queue.wo_email_processor_dlq.name
}
}
resource "aws_cloudwatch_metric_alarm" "wo_email_processor_duration" {
alarm_name = "workorder-email-processor-duration"
alarm_description = "workorder-email-processor p95 duration approaching the 60s timeout"