Add CloudWatch alarm coverage (MonitoringConstruct) #61

Closed
amoussa1229 wants to merge 3 commits from infra-cw-alarm-coverage into main
amoussa1229 commented 2026-06-17 17:59:05 +00:00 (Migrated from github.com)

Summary

Adds CloudWatch alarm coverage for the seahaven-slack-bot CDK stack via a new MonitoringConstruct (lib/constructs/monitoring.ts) wired into the stack. Every alarm:

  • sends an SNS alarm action only (no OK action) to the shared site-alerts topic (arn:aws:sns:us-east-1:328440206208:site-alerts, imported once)
  • treats missing data as NOT_BREACHING
  • is named repo-namespaced kebab-case seahaven-<fn>-<signal>

36 alarms synthesize cleanly (npx tsc --noEmit + cdk synth both pass).

Alarm coverage

Target Alarms
Lambda — all 9 functions (slack-processor, app-home, qbo-oauth, qbo-lookup, maps-lookup, wo-po-lookup, po-sync, workorder-sync, notion-sync) Errors (≥1/5min), Throttles (≥1/5min), Duration (p99, eval3/dp2, ~80% of timeout) = 27 alarms
DynamoDB — seahaven-conversations, seahaven-unanswered-questions ThrottledRequests + SystemErrors = 4 alarms
API Gateway v2 — seahaven-slack-webhook 5xx (≥1), 4xx (≥10 sustained), Latency (p99 ≥ 3s) — ApiId dim, v2 metric names = 3 alarms
ECS Fargate — seahaven-socket-mode CPU (≥85%), Memory (≥85%) AWS/ECS (always) + RunningTaskCount (<1, gated) = 3 alarms

Exposes qbo-oauth Lambda and the ECS cluster/service as public readonly handles without changing any logical IDs.

Commits

  1. c28d5b1 — base alarm coverage (Lambda + DynamoDB + API Gateway + ECS CPU/Mem). No cost/config impact.
  2. ea173a0 — NEEDS ADAM SIGN-OFF (cost/config) — enables Container Insights on the seahaven-socket-mode cluster + adds the running-tasks alarm. Isolated so it can be dropped if declined.
  3. 090f2a1 — README monitoring section.

⚠️ NEEDS ADAM SIGN-OFF

1. Duration thresholds (p99, ~80% of each function's timeout)

Function Timeout p99 threshold
seahaven-slack-processor 5 min 240 s
seahaven-app-home 10 s 8 s
seahaven-qbo-oauth 15 s 12 s
seahaven-qbo-lookup 30 s 24 s
seahaven-maps-lookup 30 s 24 s
seahaven-wo-po-lookup 30 s 24 s
seahaven-po-sync 15 min 720 s
seahaven-workorder-sync 5 min 240 s
seahaven-notion-sync 5 min 240 s

2. Container Insights (cost/config) — commit ea173a0

RunningTaskCount is only published when Container Insights is enabled on the cluster, which adds CloudWatch metric + log-ingestion cost. Commit 2 enables containerInsightsV2: ENABLED on the seahaven-socket-mode cluster and turns on the running-tasks alarm. To decline: drop commit ea173a0 — the CPU/Memory alarms (commit 1) use the standard AWS/ECS namespace and are unaffected.

DynamoDB dimension findings

ThrottledRequests and SystemErrors have no valid TableName-only aggregate — CDK's bare metricThrottledRequests / metricSystemErrors helpers are deprecated/invalid. The alarms use the per-operations metric-math helpers (metricThrottledRequestsForOperations / metricSystemErrorsForOperations), which build math expressions across operations at the TableName dimension.

CloudWatch caps an alarm math expression at 10 metrics; DynamoDB defines 14 operations, so the operation set is scoped to the 6 CRUD operations these tables use: GetItem, PutItem, UpdateItem, DeleteItem, Query, BatchWriteItem. (Confirmed via cdk synth: each DDB alarm = 6-metric math, AWS/DynamoDB namespace, TableName dimension.)

Note: neither table currently emits these metrics (verified via cloudwatch list-metrics) — expected, as error/throttle metrics only appear once an error/throttle occurs. NOT_BREACHING keeps the alarms OK until then.

Confluence map (id 1540098) delta

The AWS Architecture Map "seahaven-slack-bot" subgraph should note the new monitoring layer: a MonitoringConstruct emitting 36 CloudWatch alarms → site-alerts SNS topic. If commit 2 is accepted, also note Container Insights enabled on the seahaven-socket-mode cluster. (Documentation update outstanding — to be applied once the PR is merged and the Container Insights decision is made.)

Not deployed / not merged

Per Wave 1 rules: one PR, base main, no merge/deploy.

## Summary Adds CloudWatch alarm coverage for the `seahaven-slack-bot` CDK stack via a new `MonitoringConstruct` (`lib/constructs/monitoring.ts`) wired into the stack. Every alarm: - sends an **SNS alarm action only** (no OK action) to the shared `site-alerts` topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`, imported once) - treats missing data as `NOT_BREACHING` - is named repo-namespaced kebab-case `seahaven-<fn>-<signal>` **36 alarms** synthesize cleanly (`npx tsc --noEmit` + `cdk synth` both pass). ## Alarm coverage | Target | Alarms | |---|---| | **Lambda** — all 9 functions (slack-processor, app-home, qbo-oauth, qbo-lookup, maps-lookup, wo-po-lookup, po-sync, workorder-sync, notion-sync) | Errors (≥1/5min), Throttles (≥1/5min), Duration (p99, eval3/dp2, ~80% of timeout) = 27 alarms | | **DynamoDB** — `seahaven-conversations`, `seahaven-unanswered-questions` | ThrottledRequests + SystemErrors = 4 alarms | | **API Gateway v2** — `seahaven-slack-webhook` | 5xx (≥1), 4xx (≥10 sustained), Latency (p99 ≥ 3s) — `ApiId` dim, v2 metric names = 3 alarms | | **ECS Fargate** — `seahaven-socket-mode` | CPU (≥85%), Memory (≥85%) `AWS/ECS` (always) + RunningTaskCount (<1, gated) = 3 alarms | Exposes `qbo-oauth` Lambda and the ECS cluster/service as `public readonly` handles **without changing any logical IDs**. ## Commits 1. `c28d5b1` — base alarm coverage (Lambda + DynamoDB + API Gateway + ECS CPU/Mem). **No cost/config impact.** 2. `ea173a0` — **NEEDS ADAM SIGN-OFF (cost/config)** — enables Container Insights on the `seahaven-socket-mode` cluster + adds the `running-tasks` alarm. Isolated so it can be **dropped** if declined. 3. `090f2a1` — README monitoring section. ## ⚠️ NEEDS ADAM SIGN-OFF ### 1. Duration thresholds (p99, ~80% of each function's timeout) | Function | Timeout | p99 threshold | |---|---|---| | seahaven-slack-processor | 5 min | 240 s | | seahaven-app-home | 10 s | 8 s | | seahaven-qbo-oauth | 15 s | 12 s | | seahaven-qbo-lookup | 30 s | 24 s | | seahaven-maps-lookup | 30 s | 24 s | | seahaven-wo-po-lookup | 30 s | 24 s | | seahaven-po-sync | 15 min | 720 s | | seahaven-workorder-sync | 5 min | 240 s | | seahaven-notion-sync | 5 min | 240 s | ### 2. Container Insights (cost/config) — commit `ea173a0` `RunningTaskCount` is only published when **Container Insights** is enabled on the cluster, which adds CloudWatch metric + log-ingestion cost. Commit 2 enables `containerInsightsV2: ENABLED` on the `seahaven-socket-mode` cluster and turns on the `running-tasks` alarm. **To decline: drop commit `ea173a0`** — the CPU/Memory alarms (commit 1) use the standard `AWS/ECS` namespace and are unaffected. ## DynamoDB dimension findings `ThrottledRequests` and `SystemErrors` have **no valid `TableName`-only aggregate** — CDK's bare `metricThrottledRequests` / `metricSystemErrors` helpers are deprecated/invalid. The alarms use the per-operations metric-math helpers (`metricThrottledRequestsForOperations` / `metricSystemErrorsForOperations`), which build math expressions across operations at the `TableName` dimension. CloudWatch caps an alarm math expression at **10 metrics**; DynamoDB defines 14 operations, so the operation set is scoped to the **6 CRUD operations** these tables use: GetItem, PutItem, UpdateItem, DeleteItem, Query, BatchWriteItem. (Confirmed via `cdk synth`: each DDB alarm = 6-metric math, `AWS/DynamoDB` namespace, `TableName` dimension.) Note: neither table currently emits these metrics (verified via `cloudwatch list-metrics`) — expected, as error/throttle metrics only appear once an error/throttle occurs. `NOT_BREACHING` keeps the alarms OK until then. ## Confluence map (id 1540098) delta The **AWS Architecture Map** "seahaven-slack-bot" subgraph should note the new monitoring layer: a `MonitoringConstruct` emitting 36 CloudWatch alarms → `site-alerts` SNS topic. If commit 2 is accepted, also note Container Insights enabled on the `seahaven-socket-mode` cluster. *(Documentation update outstanding — to be applied once the PR is merged and the Container Insights decision is made.)* ## Not deployed / not merged Per Wave 1 rules: one PR, base `main`, no merge/deploy.
amoussa1229 commented 2026-06-17 18:18:33 +00:00 (Migrated from github.com)

Closing without merge: seahaven-slack-bot is being deprecated soon, so the alarm suite added here (36 alarms + Container Insights) won't be maintained. Kept as reference on the branch if needed. Per the Wave 1 alarm-coverage rollout decision (2026-06-17).

Closing without merge: seahaven-slack-bot is being deprecated soon, so the alarm suite added here (36 alarms + Container Insights) won't be maintained. Kept as reference on the branch if needed. Per the Wave 1 alarm-coverage rollout decision (2026-06-17).
This repo is archived. You cannot comment on pull requests.
No description provided.