Document CloudWatch alarm coverage in README
This commit is contained in:
parent
ea173a0966
commit
090f2a1f12
1 changed files with 30 additions and 0 deletions
30
README.md
30
README.md
|
|
@ -82,6 +82,35 @@ EventBridge (daily 02:00 UTC)
|
|||
|
||||
This bot reads the `purchase-orders` DynamoDB table **read-only** (via the `po-sync` and `wo-po-lookup` Lambdas, both granted `grantReadData`). The table is owned by the `procurement-ingest` repo (`po-ingest` stack), which is the sole authoritative writer. Any change to the `purchase-orders` schema must be coordinated with `procurement-ingest` (owner) and `payments-dashboard` (the other read-only consumer).
|
||||
|
||||
## Monitoring & Alarms
|
||||
|
||||
CloudWatch alarm coverage is defined in `lib/constructs/monitoring.ts` (`MonitoringConstruct`). Every alarm sends an **alarm action only** (no OK action) to the shared `site-alerts` SNS topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`) and treats missing data as `NOT_BREACHING`. Alarm names are repo-namespaced kebab-case: `seahaven-<fn>-<signal>`.
|
||||
|
||||
| Target | Alarms |
|
||||
|---|---|
|
||||
| **Lambda** (all 9 functions) | `errors` (≥1 / 5 min), `throttles` (≥1 / 5 min), `duration` (p99 ≥ ~80% of each function's timeout, 3 eval / 2 datapoints) |
|
||||
| **DynamoDB** `seahaven-conversations`, `seahaven-unanswered-questions` | `throttles` (ThrottledRequests), `system-errors` (SystemErrors) |
|
||||
| **API Gateway v2** `seahaven-slack-webhook` | `5xx` (≥1), `4xx` (≥10 sustained), `latency` (p99 ≥ 3s) |
|
||||
| **ECS Fargate** `seahaven-socket-mode` | `cpu` (≥85%), `memory` (≥85%), `running-tasks` (< 1)¹ |
|
||||
|
||||
Per-function Duration thresholds (p99, ~80% of timeout):
|
||||
|
||||
| Function | Timeout | Threshold |
|
||||
|---|---|---|
|
||||
| `seahaven-slack-processor` | 5 min | 240 s |
|
||||
| `seahaven-app-home` | 10 s | 8 s |
|
||||
| `seahaven-qbo-oauth` | 15 s | 12 s |
|
||||
| `seahaven-qbo-lookup` | 30 s | 24 s |
|
||||
| `seahaven-maps-lookup` | 30 s | 24 s |
|
||||
| `seahaven-wo-po-lookup` | 30 s | 24 s |
|
||||
| `seahaven-po-sync` | 15 min | 720 s |
|
||||
| `seahaven-workorder-sync` | 5 min | 240 s |
|
||||
| `seahaven-notion-sync` | 5 min | 240 s |
|
||||
|
||||
**DynamoDB metric note:** `ThrottledRequests` and `SystemErrors` have no valid `TableName`-only aggregate (the bare CDK helpers are deprecated). The alarms use the per-operations metric-math helpers, scoped to the six CRUD operations these tables use (GetItem, PutItem, UpdateItem, DeleteItem, Query, BatchWriteItem) to stay under CloudWatch's 10-metric alarm-math cap.
|
||||
|
||||
¹ **`running-tasks` requires Container Insights** on the `seahaven-socket-mode` cluster (`containerInsightsV2: ENABLED`), which adds CloudWatch metric + log-ingestion cost. The CPU and Memory alarms use the standard `AWS/ECS` namespace and need no Container Insights.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
- Node.js 22+
|
||||
|
|
@ -170,6 +199,7 @@ lib/
|
|||
notion-sync.ts EventBridge daily cron + notion-sync Lambda
|
||||
po-sync.ts EventBridge daily cron + po-sync Lambda
|
||||
workorder-sync.ts EventBridge daily cron + workorder-sync Lambda
|
||||
monitoring.ts CloudWatch alarm coverage → site-alerts SNS topic
|
||||
services/
|
||||
socket-mode/ Socket Mode ECS service (Dockerfile, TypeScript, @slack/socket-mode)
|
||||
lambda/
|
||||
|
|
|
|||
Reference in a new issue