feat(exec-aide): RunningTaskCount alarm + Container Insights [NEEDS ADAM SIGN-OFF: cost/config]

Add the exec-aide-listener-no-running-tasks CloudWatch alarm
(RunningTaskCount Minimum < 1, eval 3 / datapoints 2, ALARM-only,
NOT_BREACHING) so we page if the Slack Socket Mode listener has no
running task and the bot goes dark.

RunningTaskCount is only emitted in the ECS/ContainerInsights namespace,
so this commit also enables Container Insights on the exec-aide cluster
(containerInsightsV2: ENABLED). That is a cost/config change (extra
CloudWatch ingestion/storage for the cluster) and is isolated to this
commit pending Adam's sign-off; the rest of the alarm coverage does not
depend on it.

Also refreshes the README CDK Constructs + new Monitoring & Alarms
section (and drops the stale 'reminder' Lambda reference removed in #54).
This commit is contained in:
Adam Moussa 2026-06-17 13:57:56 -04:00
parent e5f68f1cd3
commit 25277f8dcf
2 changed files with 53 additions and 2 deletions

View file

@ -52,8 +52,29 @@ The conversation Lambda has access to these tools:
### CDK Constructs
- `lib/constructs/email-pipeline.ts` — DynamoDB table, all four Lambda functions (fetch-classify, daily-digest, conversation, reminder), EventBridge schedules
- `lib/constructs/socket-mode.ts` — VPC, ECS cluster, Fargate service, ECR repo (image built automatically via `ContainerImage.fromAsset()`)
- `lib/constructs/email-pipeline.ts` — DynamoDB table, the three pipeline Lambda functions (fetch-classify, daily-digest, conversation), EventBridge schedules, and their CloudWatch alarms (Lambda + DynamoDB)
- `lib/constructs/socket-mode.ts` — VPC, ECS cluster, Fargate service, ECR repo (image built automatically via `ContainerImage.fromAsset()`), and the listener's CloudWatch alarms (CPU, memory, running-task count)
## Monitoring & Alarms
All CloudWatch alarms are ALARM-only (no OK / InsufficientData actions), use `treatMissingData: NOT_BREACHING`, and publish to the shared cross-stack `site-alerts` SNS topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`, managed by `seahaven-account-baseline`, encrypted with `alias/seahaven-alarm-topics`). The topic is imported once at the stack level and injected into both constructs via props. Alarm names follow `exec-aide-<resource>-<signal>`.
| Alarm | Resource | Condition |
|---|---|---|
| `exec-aide-<fn>-errors` | each Lambda | Errors Sum ≥ 1 over 5 min |
| `exec-aide-<fn>-throttles` | each Lambda | Throttles Sum ≥ 1 over 5 min |
| `exec-aide-<fn>-duration` | each Lambda | p99 Duration > ~80% of timeout (96000 ms for the 120 s fns, 144000 ms for conversation's 180 s), eval 3 / datapoints 2 |
| `exec-aide-table-throttles` | DynamoDB `exec-aide` | ThrottledRequests Sum ≥ 1 (summed across the operations the app issues) |
| `exec-aide-table-system-errors` | DynamoDB `exec-aide` | SystemErrors Sum ≥ 1 (summed across the operations the app issues) |
| `exec-aide-listener-cpu-high` | Fargate service | CPU Average > 80%, eval 3 / datapoints 2 |
| `exec-aide-listener-memory-high` | Fargate service | Memory Average > 80%, eval 3 / datapoints 2 |
| `exec-aide-listener-no-running-tasks` | Fargate service | RunningTaskCount Minimum < 1, eval 3 / datapoints 2 |
`<fn>` ∈ {`fetch-classify`, `daily-digest`, `conversation`}.
Notes:
- DynamoDB `ThrottledRequests` / `SystemErrors` are not published at the bare `TableName` dimension — AWS keys them by `Operation`. The alarms use CDK's `*ForOperations` math helpers scoped to the six operations this single-table app issues (GetItem, PutItem, Query, Scan, UpdateItem, DeleteItem) to stay within CloudWatch's 10-metric math-expression limit.
- `exec-aide-listener-no-running-tasks` reads `RunningTaskCount` from the `ECS/ContainerInsights` namespace, which requires **Container Insights** to be enabled on the `exec-aide` cluster (a cost/config change).
## Classification Rules

View file

@ -71,6 +71,11 @@ export class SocketModeConstruct extends Construct {
const cluster = new ecs.Cluster(this, 'Cluster', {
clusterName: 'exec-aide',
vpc,
// NEEDS ADAM SIGN-OFF (cost/config): Container Insights is required for
// the RunningTaskCount alarm below (that metric lives in the
// ECS/ContainerInsights namespace; AWS/ECS does not publish it). Enabling
// it incurs additional CloudWatch ingestion/storage cost for this cluster.
containerInsightsV2: ecs.ContainerInsights.ENABLED,
});
const taskDef = new ecs.FargateTaskDefinition(this, 'TaskDef', {
@ -180,5 +185,30 @@ export class SocketModeConstruct extends Construct {
datapointsToAlarm: 2,
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
}).addAlarmAction(alarmAction);
// NEEDS ADAM SIGN-OFF (cost/config): RunningTaskCount. The desiredCount is
// 1; this fires if the listener has no running task (Slack Socket Mode goes
// dark — DMs and @mentions stop being handled). RunningTaskCount is only
// emitted in the ECS/ContainerInsights namespace, which requires Container
// Insights to be enabled on the cluster (done above).
new cloudwatch.Alarm(this, 'ListenerNoRunningTasks', {
alarmName: 'exec-aide-listener-no-running-tasks',
alarmDescription: 'exec-aide-listener Fargate service has no running tasks — Slack Socket Mode is down',
metric: new cloudwatch.Metric({
namespace: 'ECS/ContainerInsights',
metricName: 'RunningTaskCount',
dimensionsMap: {
ClusterName: cluster.clusterName,
ServiceName: service.serviceName,
},
statistic: 'Minimum',
period: cdk.Duration.minutes(5),
}),
threshold: 1,
comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
evaluationPeriods: 3,
datapointsToAlarm: 2,
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
}).addAlarmAction(alarmAction);
}
}