mirror of
https://github.com/Sea-Haven-Industries/procurement-ingest.git
synced 2026-09-30 07:13:13 +00:00
fix(wo): add DLQ age alarm before 14-day expiry
This commit is contained in:
parent
a73603e16a
commit
3f0ff8e82b
3 changed files with 27 additions and 5 deletions
|
|
@ -175,7 +175,7 @@ The `<fn>-sender-auth-rejected` alarm closes the silent-drop gap in INFRA-107: a
|
||||||
|
|
||||||
The `<fn>-duration` and `<fn>-throttles` alarms for `po-email-processor` and `workorder-email-processor` supersede the orphaned, CLI-created `Lambda-Duration-*` / `Lambda-Throttles-*` alarms (deleted post-deploy).
|
The `<fn>-duration` and `<fn>-throttles` alarms for `po-email-processor` and `workorder-email-processor` supersede the orphaned, CLI-created `Lambda-Duration-*` / `Lambda-Throttles-*` alarms (deleted post-deploy).
|
||||||
|
|
||||||
**DLQ alarms** (`AWS/SQS`): `po-email-processor-dlq-messages` and `workorder-email-processor-dlq-messages` fire when any message is visible on an email-processor DLQ (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) — a message there means an email was dropped after Lambda exhausted its async retries. Recovery from a DLQ message (no console redrive) is documented in the [DLQ recovery runbook](docs/runbook-dlq-recovery.md).
|
**DLQ alarms** (`AWS/SQS`): `po-email-processor-dlq-messages` and `workorder-email-processor-dlq-messages` fire when any message is visible on an email-processor DLQ (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) — a message there means an email was dropped after Lambda exhausted its async retries. WO also has `workorder-email-processor-dlq-age` (`ApproximateAgeOfOldestMessage` Maximum, 5 min, `>= 86400` s, eval 1) so an undrained breadcrumb pages again well before 14-day retention expiry. Recovery from a DLQ message (no console redrive) is documented in the [DLQ recovery runbook](docs/runbook-dlq-recovery.md).
|
||||||
|
|
||||||
**SHOC emitter alarms** (bespoke — these don't fit the standard-Lambda-alarm helper's shape): `workorder-shoc-emitter-iterator-age` (`IteratorAge` Maximum, 5 min, `>= 600000` ms, eval 3 / datapoints 2) fires when the stream lags ≥ 10 minutes — SHOC is likely down and the shard is blocking on retries, which is exactly the ordered-backpressure design working, but an operator should know. `workorder-shoc-emitter-failures-messages` and `workorder-shoc-emitter-rejected-messages` (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) page the replay runbook: the `-failures` queue receives ESM failure **metadata** for retry-exhausted records, the `-rejected` queue receives the **full parked payloads** of non-retryable 4xx deliveries (see the [SHOC webhook feed](#shoc-webhook-feed-workorder-shoc-emitter) section and `scripts/replay_shoc_webhooks.py`). All three: ALARM-only → `site-alerts`, `NOT_BREACHING`, per the house rules above.
|
**SHOC emitter alarms** (bespoke — these don't fit the standard-Lambda-alarm helper's shape): `workorder-shoc-emitter-iterator-age` (`IteratorAge` Maximum, 5 min, `>= 600000` ms, eval 3 / datapoints 2) fires when the stream lags ≥ 10 minutes — SHOC is likely down and the shard is blocking on retries, which is exactly the ordered-backpressure design working, but an operator should know. `workorder-shoc-emitter-failures-messages` and `workorder-shoc-emitter-rejected-messages` (`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1) page the replay runbook: the `-failures` queue receives ESM failure **metadata** for retry-exhausted records, the `-rejected` queue receives the **full parked payloads** of non-retryable 4xx deliveries (see the [SHOC webhook feed](#shoc-webhook-feed-workorder-shoc-emitter) section and `scripts/replay_shoc_webhooks.py`). All three: ALARM-only → `site-alerts`, `NOT_BREACHING`, per the house rules above.
|
||||||
|
|
||||||
|
|
|
||||||
|
|
@ -20,10 +20,10 @@ target. Recovery is manual, via targeted re-invoke.
|
||||||
|
|
||||||
## Resource inventory (acct 011934824531 seahaven-prod, us-east-1)
|
## Resource inventory (acct 011934824531 seahaven-prod, us-east-1)
|
||||||
|
|
||||||
| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarm |
|
| Pipeline | Function | DLQ queue name | Raw-email bucket | DLQ alarms |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| PO | `po-email-processor` | _CDK-generated; fill from stack resources after the first prod deploy_ | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` |
|
| PO | `po-email-processor` | `po-ingest-EmailProcessorDlqA753DED5-Mn33HvhsDPEu` | `po-ingest-emails-011934824531` | `po-email-processor-dlq-messages` |
|
||||||
| WO | `workorder-email-processor` | _CDK-generated; fill from stack resources after the first prod deploy_ | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` |
|
| WO | `workorder-email-processor` | `WorkorderIngestStack-EmailProcessorDlqA753DED5-qCTHrsoEucas` | `workorder-ingest-emails-011934824531` | `workorder-email-processor-dlq-messages` (depth, visible `> 0`); `workorder-email-processor-dlq-age` (oldest message `>= 86400` s) |
|
||||||
|
|
||||||
DLQ URLs are `https://sqs.us-east-1.amazonaws.com/011934824531/<queue-name>`.
|
DLQ URLs are `https://sqs.us-east-1.amazonaws.com/011934824531/<queue-name>`.
|
||||||
Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
|
Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
|
||||||
|
|
@ -33,6 +33,10 @@ Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
|
||||||
|
|
||||||
1. **Trigger.** The `<fn>-dlq-messages` alarm fires
|
1. **Trigger.** The `<fn>-dlq-messages` alarm fires
|
||||||
(`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1).
|
(`ApproximateNumberOfMessagesVisible` Maximum, 5 min, `> 0`, eval 1).
|
||||||
|
The WO queue also has `workorder-email-processor-dlq-age`, which fires when
|
||||||
|
`ApproximateAgeOfOldestMessage` is `>= 86400` seconds (1 day). That is the
|
||||||
|
ignored-page guard: drain before the 14-day retention window (`1209600` s)
|
||||||
|
expires.
|
||||||
|
|
||||||
2. **RECEIVE** the message (do not purge yet):
|
2. **RECEIVE** the message (do not purge yet):
|
||||||
|
|
||||||
|
|
@ -69,7 +73,7 @@ Both DLQs: 14-day retention, SSE, TLS-enforced, `VisibilityTimeout` 30s.
|
||||||
|
|
||||||
5. **VERIFY the write.** Confirm the downstream effect landed before proceeding:
|
5. **VERIFY the write.** Confirm the downstream effect landed before proceeding:
|
||||||
the DynamoDB item exists / was updated (`purchase-orders` for PO,
|
the DynamoDB item exists / was updated (`purchase-orders` for PO,
|
||||||
`WorkOrders` for WO), and the function's log group shows a clean parse (no new
|
`work-orders` for WO), and the function's log group shows a clean parse (no new
|
||||||
error, no new DLQ message). Do not proceed until verified.
|
error, no new DLQ message). Do not proceed until verified.
|
||||||
|
|
||||||
6. **PURGE the one message.** Delete only the processed message by its
|
6. **PURGE the one message.** Delete only the processed message by its
|
||||||
|
|
|
||||||
|
|
@ -65,6 +65,24 @@ resource "aws_cloudwatch_metric_alarm" "wo_email_processor_dlq" {
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
resource "aws_cloudwatch_metric_alarm" "wo_email_processor_dlq_age" {
|
||||||
|
alarm_name = "workorder-email-processor-dlq-age"
|
||||||
|
alarm_description = "workorder-email-processor DLQ oldest message is >= 1 day old (drain before 14-day retention expiry)"
|
||||||
|
comparison_operator = "GreaterThanOrEqualToThreshold"
|
||||||
|
evaluation_periods = 1
|
||||||
|
metric_name = "ApproximateAgeOfOldestMessage"
|
||||||
|
namespace = "AWS/SQS"
|
||||||
|
period = 300
|
||||||
|
statistic = "Maximum"
|
||||||
|
threshold = 86400
|
||||||
|
treat_missing_data = "notBreaching"
|
||||||
|
alarm_actions = [data.aws_sns_topic.site_alerts.arn]
|
||||||
|
|
||||||
|
dimensions = {
|
||||||
|
QueueName = aws_sqs_queue.wo_email_processor_dlq.name
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
resource "aws_cloudwatch_metric_alarm" "wo_email_processor_duration" {
|
resource "aws_cloudwatch_metric_alarm" "wo_email_processor_duration" {
|
||||||
alarm_name = "workorder-email-processor-duration"
|
alarm_name = "workorder-email-processor-duration"
|
||||||
alarm_description = "workorder-email-processor p95 duration approaching the 60s timeout"
|
alarm_description = "workorder-email-processor p95 duration approaching the 60s timeout"
|
||||||
|
|
|
||||||
Loading…
Add table
Reference in a new issue