feat(exec-aide): CloudWatch alarm coverage (Lambda + DynamoDB + ECS) #59

Closed
amoussa1229 wants to merge 2 commits from infra-cloudwatch-alarm-coverage into main

2 commits

Author SHA1 Message Date
25277f8dcf feat(exec-aide): RunningTaskCount alarm + Container Insights [NEEDS ADAM SIGN-OFF: cost/config]
Add the exec-aide-listener-no-running-tasks CloudWatch alarm
(RunningTaskCount Minimum < 1, eval 3 / datapoints 2, ALARM-only,
NOT_BREACHING) so we page if the Slack Socket Mode listener has no
running task and the bot goes dark.

RunningTaskCount is only emitted in the ECS/ContainerInsights namespace,
so this commit also enables Container Insights on the exec-aide cluster
(containerInsightsV2: ENABLED). That is a cost/config change (extra
CloudWatch ingestion/storage for the cluster) and is isolated to this
commit pending Adam's sign-off; the rest of the alarm coverage does not
depend on it.

Also refreshes the README CDK Constructs + new Monitoring & Alarms
section (and drops the stale 'reminder' Lambda reference removed in #54).
2026-06-17 13:57:56 -04:00
e5f68f1cd3 feat(exec-aide): CloudWatch alarm coverage for Lambdas, DynamoDB, ECS
Add ALARM-only CloudWatch alarms routed to the shared site-alerts SNS
topic (imported once via Topic.fromTopicArn and injected into both
constructs via props). All alarms use treatMissingData NOT_BREACHING and
have no OK / InsufficientData actions, mirroring the proposal-system
alarm construct.

Lambda (fetch-classify, daily-digest, conversation):
- Errors  (Sum >= 1, eval 1)
- Throttles (Sum >= 1, eval 1)
- Duration (p99, eval 3 / datapoints 2, ~80% of timeout:
  96000ms for the 120s fns, 144000ms for conversation's 180s)
  -- thresholds pending Adam sign-off.

DynamoDB exec-aide table:
- ThrottledRequests and SystemErrors. These metrics are NOT published at
  the bare TableName dimension (CDK's metricThrottledRequests /
  metricSystemErrors are deprecated as invalid); they are keyed by the
  Operation dimension. Used the *ForOperations math helpers scoped to the
  6 operations this single-table app issues (GetItem/PutItem/Query/Scan/
  UpdateItem/DeleteItem) to stay within the 10-metric math-expr cap.

ECS exec-aide-listener Fargate service (AWS/ECS, no Container Insights):
- CPU and Memory utilization (Average > 80%, eval 3 / datapoints 2).
- Service assigned to a const (logical id 'Service' unchanged) so metrics
  can reference it.

The RunningTaskCount alarm (requires Container Insights) is intentionally
deferred to a separate sign-off-gated commit.
2026-06-17 13:56:25 -04:00