diff --git a/README.md b/README.md index f5bd814..1f74ec4 100644 --- a/README.md +++ b/README.md @@ -82,6 +82,35 @@ EventBridge (daily 02:00 UTC) This bot reads the `purchase-orders` DynamoDB table **read-only** (via the `po-sync` and `wo-po-lookup` Lambdas, both granted `grantReadData`). The table is owned by the `procurement-ingest` repo (`po-ingest` stack), which is the sole authoritative writer. Any change to the `purchase-orders` schema must be coordinated with `procurement-ingest` (owner) and `payments-dashboard` (the other read-only consumer). +## Monitoring & Alarms + +CloudWatch alarm coverage is defined in `lib/constructs/monitoring.ts` (`MonitoringConstruct`). Every alarm sends an **alarm action only** (no OK action) to the shared `site-alerts` SNS topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`) and treats missing data as `NOT_BREACHING`. Alarm names are repo-namespaced kebab-case: `seahaven--`. + +| Target | Alarms | +|---|---| +| **Lambda** (all 9 functions) | `errors` (≥1 / 5 min), `throttles` (≥1 / 5 min), `duration` (p99 ≥ ~80% of each function's timeout, 3 eval / 2 datapoints) | +| **DynamoDB** `seahaven-conversations`, `seahaven-unanswered-questions` | `throttles` (ThrottledRequests), `system-errors` (SystemErrors) | +| **API Gateway v2** `seahaven-slack-webhook` | `5xx` (≥1), `4xx` (≥10 sustained), `latency` (p99 ≥ 3s) | +| **ECS Fargate** `seahaven-socket-mode` | `cpu` (≥85%), `memory` (≥85%), `running-tasks` (< 1)¹ | + +Per-function Duration thresholds (p99, ~80% of timeout): + +| Function | Timeout | Threshold | +|---|---|---| +| `seahaven-slack-processor` | 5 min | 240 s | +| `seahaven-app-home` | 10 s | 8 s | +| `seahaven-qbo-oauth` | 15 s | 12 s | +| `seahaven-qbo-lookup` | 30 s | 24 s | +| `seahaven-maps-lookup` | 30 s | 24 s | +| `seahaven-wo-po-lookup` | 30 s | 24 s | +| `seahaven-po-sync` | 15 min | 720 s | +| `seahaven-workorder-sync` | 5 min | 240 s | +| `seahaven-notion-sync` | 5 min | 240 s | + +**DynamoDB metric note:** `ThrottledRequests` and `SystemErrors` have no valid `TableName`-only aggregate (the bare CDK helpers are deprecated). The alarms use the per-operations metric-math helpers, scoped to the six CRUD operations these tables use (GetItem, PutItem, UpdateItem, DeleteItem, Query, BatchWriteItem) to stay under CloudWatch's 10-metric alarm-math cap. + +¹ **`running-tasks` requires Container Insights** on the `seahaven-socket-mode` cluster (`containerInsightsV2: ENABLED`), which adds CloudWatch metric + log-ingestion cost. The CPU and Memory alarms use the standard `AWS/ECS` namespace and need no Container Insights. + ## Prerequisites - Node.js 22+ @@ -170,6 +199,7 @@ lib/ notion-sync.ts EventBridge daily cron + notion-sync Lambda po-sync.ts EventBridge daily cron + po-sync Lambda workorder-sync.ts EventBridge daily cron + workorder-sync Lambda + monitoring.ts CloudWatch alarm coverage → site-alerts SNS topic services/ socket-mode/ Socket Mode ECS service (Dockerfile, TypeScript, @slack/socket-mode) lambda/