diff --git a/README.md b/README.md index e789c92..3cf8cdd 100644 --- a/README.md +++ b/README.md @@ -189,6 +189,45 @@ sam build sam deploy ``` +## Monitoring & Alarms + +All CloudWatch alarms are defined in `template.yaml` and notify the shared +`site-alerts` SNS topic (→ AWS Chatbot → Slack). None set `OKActions` — recovery +is not paged. Alarm names follow the in-template convention `Lambda--` +(e.g. `Lambda-Errors-afterhours-ring-scheduler`). + +**Lambda alarms** (all six functions: `afterhours-shift-manager`, +`afterhours-weekly-post`, `afterhours-roster-sync`, `afterhours-ring-scheduler`, +`afterhours-holiday-router`, `afterhours-release-notifier`): + +| Alarm | Metric | Condition | Notes | +|---|---|---|---| +| `Lambda-Errors-` | `Errors` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` | +| `Lambda-Duration-` | `Duration` (Maximum, ms) | `>= ~80% of timeout`, 2 of 3 5-min periods | Thresholds: 24000 ms (30s-timeout fns) / 48000 ms (60s-timeout fns) | +| `Lambda-Throttles-` | `Throttles` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` | + +**DynamoDB alarm** (`afterhours-shifts` table): + +| Alarm | Metric | Condition | Notes | +|---|---|---|---| +| `DynamoDB-ThrottledRequests-afterhours-shifts` | `ThrottledRequests` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` | + +`SystemErrors` is intentionally **not** alarmed: AWS emits it only at +`TableName`+`Operation` granularity, so a `TableName`-only alarm would sit +permanently in `INSUFFICIENT_DATA`. + +**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway` +v2 metrics, `ApiId` dimension): + +| Alarm | Metric | Condition | +|---|---|---| +| `ApiGateway-4xx-` | `4xx` (Sum) | `>= 5` over one 5-min period | +| `ApiGateway-5xx-` | `5xx` (Sum) | `>= 1` over one 5-min period | +| `ApiGateway-Latency-` | `Latency` (p99, ms) | `>= 3000` ms, 2 of 3 5-min periods | + +> Duration and API latency thresholds are starting points and may be tuned after +> observing real traffic. + ## Releases & Versioning The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`** — it is diff --git a/template.yaml b/template.yaml index 9ecbc77..a67ac6f 100644 --- a/template.yaml +++ b/template.yaml @@ -455,6 +455,424 @@ Resources: Action: lambda:InvokeFunction Resource: !GetAtt ReleaseNotifierFunction.Arn + # =========================================================================== + # CloudWatch alarm coverage (Wave 1 PR A). All alarms page the same + # site-alerts SNS topic → AWS Chatbot → Slack as the existing Errors alarms. + # No OKActions by design (the existing Errors alarms have none either). + # Naming follows the in-template convention: Lambda--${Fn}. + # =========================================================================== + + # --- Lambda Errors alarms (clone of HolidayRouterErrorAlarm) --- + # Errors for ring-scheduler + holiday-router already exist above. + + RosterSyncErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${RosterSyncFunction}" + AlarmDescription: "Roster sync Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref RosterSyncFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ReleaseNotifierErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${ReleaseNotifierFunction}" + AlarmDescription: "Release notifier Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref ReleaseNotifierFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # ORPHAN ADOPTION: live alarms named Lambda-Errors-afterhours-shift-manager + # and Lambda-Errors-afterhours-weekly-post already exist OUTSIDE the stack. + # They MUST be deleted immediately before this stack deploys, or CloudFormation + # will fail to create these resources (AlarmName collision). See PR body. + SlackBotErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${SlackBotFunction}" + AlarmDescription: "Slack bot Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref SlackBotFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + WeeklyPostErrorAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Errors-${WeeklyPostFunction}" + AlarmDescription: "Weekly post Lambda reported one or more errors" + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref WeeklyPostFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) --- + # Thresholds set to ~80% of each function's timeout. ALL Duration thresholds + # and the 3/2 evaluation are PENDING ADAM SIGN-OFF (see PR body). + # Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default); + # roster-sync/ring-scheduler/holiday-router = 60s. + SlackBotDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${SlackBotFunction}" + AlarmDescription: "Slack bot Lambda duration approaching its 30s timeout (>=24s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref SlackBotFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 24000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + WeeklyPostDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${WeeklyPostFunction}" + AlarmDescription: "Weekly post Lambda duration approaching its 30s timeout (>=24s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref WeeklyPostFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 24000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RosterSyncDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${RosterSyncFunction}" + AlarmDescription: "Roster sync Lambda duration approaching its 60s timeout (>=48s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref RosterSyncFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 48000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RingSchedulerDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${RingSchedulerFunction}" + AlarmDescription: "Ring scheduler Lambda duration approaching its 60s timeout (>=48s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref RingSchedulerFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 48000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + HolidayRouterDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${HolidayRouterFunction}" + AlarmDescription: "Holiday router Lambda duration approaching its 60s timeout (>=48s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref HolidayRouterFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 48000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ReleaseNotifierDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Duration-${ReleaseNotifierFunction}" + AlarmDescription: "Release notifier Lambda duration approaching its 30s timeout (>=24s)" + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref ReleaseNotifierFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 24000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- Lambda Throttles alarms (Sum; threshold 1 over one 5-min period) --- + SlackBotThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${SlackBotFunction}" + AlarmDescription: "Slack bot Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref SlackBotFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + WeeklyPostThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${WeeklyPostFunction}" + AlarmDescription: "Weekly post Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref WeeklyPostFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RosterSyncThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${RosterSyncFunction}" + AlarmDescription: "Roster sync Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref RosterSyncFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + RingSchedulerThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${RingSchedulerFunction}" + AlarmDescription: "Ring scheduler Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref RingSchedulerFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + HolidayRouterThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${HolidayRouterFunction}" + AlarmDescription: "Holiday router Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref HolidayRouterFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ReleaseNotifierThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "Lambda-Throttles-${ReleaseNotifierFunction}" + AlarmDescription: "Release notifier Lambda was throttled (concurrency limit hit)" + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref ReleaseNotifierFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- DynamoDB alarms (afterhours-shifts table) --- + # ThrottledRequests emits at the TableName dimension → safe to alarm. + # NOTE: SystemErrors is INTENTIONALLY OMITTED. AWS emits AWS/DynamoDB + # SystemErrors only at TableName+Operation granularity, never TableName-only, + # so a TableName-only SystemErrors alarm would sit permanently in + # INSUFFICIENT_DATA. (Verified: no DynamoDB table in this account has ever + # emitted SystemErrors, consistent with documented per-Operation emission.) + ShiftTableThrottleAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "DynamoDB-ThrottledRequests-${ShiftTable}" + AlarmDescription: "afterhours-shifts table had one or more throttled requests" + Namespace: AWS/DynamoDB + MetricName: ThrottledRequests + Dimensions: + - Name: TableName + Value: !Ref ShiftTable + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # --- API Gateway v2 (HTTP API) alarms on the implicit ServerlessHttpApi --- + # AWS::ApiGatewayV2 metric names: 4xx, 5xx, Latency; dimension ApiId. + ApiGateway4xxAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "ApiGateway-4xx-${ServerlessHttpApi}" + AlarmDescription: "Elevated 4xx responses on the Slack events HTTP API" + Namespace: AWS/ApiGateway + MetricName: 4xx + Dimensions: + - Name: ApiId + Value: !Ref ServerlessHttpApi + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 5 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + ApiGateway5xxAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "ApiGateway-5xx-${ServerlessHttpApi}" + AlarmDescription: "5xx responses on the Slack events HTTP API" + Namespace: AWS/ApiGateway + MetricName: 5xx + Dimensions: + - Name: ApiId + Value: !Ref ServerlessHttpApi + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + + # Latency p99 via ExtendedStatistic. ~3000ms target is PENDING ADAM SIGN-OFF + # alongside the Lambda Duration thresholds (Slack requires a fast 3s ack). + ApiGatewayLatencyAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: !Sub "ApiGateway-Latency-${ServerlessHttpApi}" + AlarmDescription: "p99 latency on the Slack events HTTP API exceeded 3s" + Namespace: AWS/ApiGateway + MetricName: Latency + Dimensions: + - Name: ApiId + Value: !Ref ServerlessHttpApi + ExtendedStatistic: p99 + Period: 300 + EvaluationPeriods: 3 + DatapointsToAlarm: 2 + Threshold: 3000 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + # --- CloudWatch Log Groups (explicit 60-day retention) --- SlackBotLogGroup: Type: AWS::Logs::LogGroup