From 90c10d38d314c58f8527af01b84d549387cf8a39 Mon Sep 17 00:00:00 2001 From: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> Date: Wed, 17 Jun 2026 14:45:47 -0400 Subject: [PATCH] Add CloudWatch alarm coverage for all functions, table, and HTTP API (#126) * Add CloudWatch alarm coverage for all functions, table, and HTTP API Extend the in-template Lambda-- alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests is not emitted at the TableName-only dimension (only TableName+Operation), so the table-level alarm would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Correct DynamoDB alarm docs and drop sign-off wording README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments. --- README.md | 42 +++++ template.yaml | 435 ++++++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 477 insertions(+) diff --git a/README.md b/README.md index e789c92..00d9aea 100644 --- a/README.md +++ b/README.md @@ -189,6 +189,48 @@ sam build sam deploy ``` +## Monitoring & Alarms + +All CloudWatch alarms are defined in `template.yaml` and notify the shared +`site-alerts` SNS topic (→ AWS Chatbot → Slack). None set `OKActions` — recovery +is not paged. Alarm names follow the in-template convention `Lambda--` +(e.g. `Lambda-Errors-afterhours-ring-scheduler`). + +**Lambda alarms** (all six functions: `afterhours-shift-manager`, +`afterhours-weekly-post`, `afterhours-roster-sync`, `afterhours-ring-scheduler`, +`afterhours-holiday-router`, `afterhours-release-notifier`): + +| Alarm | Metric | Condition | Notes | +|---|---|---|---| +| `Lambda-Errors-` | `Errors` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` | +| `Lambda-Duration-` | `Duration` (Maximum, ms) | `>= ~80% of timeout`, 2 of 3 5-min periods | Thresholds: 24000 ms (30s-timeout fns) / 48000 ms (60s-timeout fns) | +| `Lambda-Throttles-` | `Throttles` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` | + +**DynamoDB alarm** (`afterhours-shifts` table): + +| Alarm | Metric | Condition | Notes | +|---|---|---|---| +| `DDB-ReadThrottle-afterhours-shifts` | `ReadThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` | +| `DDB-WriteThrottle-afterhours-shifts` | `WriteThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` | + +`ReadThrottleEvents` / `WriteThrottleEvents` are the table-level throttle +signals: AWS/DynamoDB emits them at the `TableName` dimension, so these alarms +transition normally. `ThrottledRequests` and `SystemErrors` are intentionally +**not** alarmed: AWS emits them only at `TableName`+`Operation` granularity, so a +`TableName`-only alarm would sit permanently in `INSUFFICIENT_DATA`. + +**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway` +v2 metrics, `ApiId` dimension): + +| Alarm | Metric | Condition | +|---|---|---| +| `ApiGateway-4xx-` | `4xx` (Sum) | `>= 5` over one 5-min period | +| `ApiGateway-5xx-` | `5xx` (Sum) | `>= 1` over one 5-min period | +| `ApiGateway-Latency-` | `Latency` (p99, ms) | `>= 3000` ms, 2 of 3 5-min periods | + +> Duration and API latency thresholds are starting points and may be tuned after +> observing real traffic. + ## Releases & Versioning The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`** — it is diff --git a/template.yaml b/template.yaml index 9ecbc77..68f325a 100644 --- a/template.yaml +++ b/template.yaml @@ -455,6 +455,441 @@ Resources: Action: lambda:InvokeFunction Resource: !GetAtt ReleaseNotifierFunction.Arn + # =========================================================================== + # CloudWatch alarm coverage (Wave 1 PR A). All alarms page the same + # site-alerts SNS topic → AWS Chatbot → Slack as the existing Errors alarms. + # No OKActions by design (the existing Errors alarms have none either). + # Naming follows the in-template convention: Lambda--${Fn}. + # =========================================================================== + + # --- Lambda Errors alarms (clone of HolidayRouterErrorAlarm) --- + # Errors for ring-scheduler + holiday-router already exist above. + + RosterSyncErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${RosterSyncFunction}" + AlarmDescription: "Roster sync Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref RosterSyncFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ReleaseNotifierErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${ReleaseNotifierFunction}" + AlarmDescription: "Release notifier Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref ReleaseNotifierFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # ORPHAN ADOPTION: live alarms named Lambda-Errors-afterhours-shift-manager + # and Lambda-Errors-afterhours-weekly-post already exist OUTSIDE the stack. + # They MUST be deleted immediately before this stack deploys, or CloudFormation + # will fail to create these resources (AlarmName collision). See PR body. + SlackBotErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${SlackBotFunction}" + AlarmDescription: "Slack bot Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref SlackBotFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + WeeklyPostErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${WeeklyPostFunction}" + AlarmDescription: "Weekly post Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref WeeklyPostFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) --- + # Thresholds set to ~80% of each function's timeout, with 2-of-3 evaluation. + # Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default); + # roster-sync/ring-scheduler/holiday-router = 60s. + SlackBotDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${SlackBotFunction}" + AlarmDescription: "Slack bot Lambda duration approaching its 30s timeout (>=24s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref SlackBotFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 24000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + WeeklyPostDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${WeeklyPostFunction}" + AlarmDescription: "Weekly post Lambda duration approaching its 30s timeout (>=24s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref WeeklyPostFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 24000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RosterSyncDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${RosterSyncFunction}" + AlarmDescription: "Roster sync Lambda duration approaching its 60s timeout (>=48s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref RosterSyncFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 48000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RingSchedulerDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${RingSchedulerFunction}" + AlarmDescription: "Ring scheduler Lambda duration approaching its 60s timeout (>=48s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref RingSchedulerFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 48000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + HolidayRouterDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${HolidayRouterFunction}" + AlarmDescription: "Holiday router Lambda duration approaching its 60s timeout (>=48s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref HolidayRouterFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 48000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ReleaseNotifierDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${ReleaseNotifierFunction}" + AlarmDescription: "Release notifier Lambda duration approaching its 30s timeout (>=24s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref ReleaseNotifierFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 24000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- Lambda Throttles alarms (Sum; threshold 1 over one 5-min period) --- + SlackBotThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${SlackBotFunction}" + AlarmDescription: "Slack bot Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref SlackBotFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + WeeklyPostThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${WeeklyPostFunction}" + AlarmDescription: "Weekly post Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref WeeklyPostFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RosterSyncThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${RosterSyncFunction}" + AlarmDescription: "Roster sync Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref RosterSyncFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RingSchedulerThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${RingSchedulerFunction}" + AlarmDescription: "Ring scheduler Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref RingSchedulerFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + HolidayRouterThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${HolidayRouterFunction}" + AlarmDescription: "Holiday router Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref HolidayRouterFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ReleaseNotifierThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${ReleaseNotifierFunction}" + AlarmDescription: "Release notifier Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref ReleaseNotifierFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- DynamoDB alarms (afterhours-shifts table) --- + # ReadThrottleEvents / WriteThrottleEvents are the table-level throttle + # signals: AWS/DynamoDB emits them at the TableName dimension, so these + # alarms transition normally. (ThrottledRequests and SystemErrors are NOT + # emitted at TableName-only granularity — only at TableName+Operation — so + # alarms on them sit permanently in INSUFFICIENT_DATA and never fire.) + ShiftTableReadThrottleAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "DDB-ReadThrottle-${ShiftTable}" + AlarmDescription: "afterhours-shifts table had one or more read throttle events" + Namespace: AWS/DynamoDB + MetricName: ReadThrottleEvents + Dimensions: + - Name: TableName + Value: !Ref ShiftTable + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ShiftTableWriteThrottleAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "DDB-WriteThrottle-${ShiftTable}" + AlarmDescription: "afterhours-shifts table had one or more write throttle events" + Namespace: AWS/DynamoDB + MetricName: WriteThrottleEvents + Dimensions: + - Name: TableName + Value: !Ref ShiftTable + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- API Gateway v2 (HTTP API) alarms on the implicit ServerlessHttpApi --- + # AWS::ApiGatewayV2 metric names: 4xx, 5xx, Latency; dimension ApiId. + ApiGateway4xxAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "ApiGateway-4xx-${ServerlessHttpApi}" + AlarmDescription: "Elevated 4xx responses on the Slack events HTTP API" + Namespace: AWS/ApiGateway + MetricName: 4xx + Dimensions: + - Name: ApiId + Value: !Ref ServerlessHttpApi + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 5 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ApiGateway5xxAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "ApiGateway-5xx-${ServerlessHttpApi}" + AlarmDescription: "5xx responses on the Slack events HTTP API" + Namespace: AWS/ApiGateway + MetricName: 5xx + Dimensions: + - Name: ApiId + Value: !Ref ServerlessHttpApi + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # Latency p99 via ExtendedStatistic. ~3000ms target chosen alongside the + # Lambda Duration thresholds (Slack requires a fast 3s ack). + ApiGatewayLatencyAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "ApiGateway-Latency-${ServerlessHttpApi}" + AlarmDescription: "p99 latency on the Slack events HTTP API exceeded 3s" + Namespace: AWS/ApiGateway + MetricName: Latency + Dimensions: + - Name: ApiId + Value: !Ref ServerlessHttpApi + ExtendedStatistic: p99 + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 3000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + # --- CloudWatch Log Groups (explicit 60-day retention) --- SlackBotLogGroup: Type: AWS::Logs::LogGroup