Add CloudWatch alarm coverage for all functions, table, and HTTP API (#126)
Some checks are pending
Deploy / deploy (push) Waiting to run
Deploy / release (push) Blocked by required conditions

* Add CloudWatch alarm coverage for all functions, table, and HTTP API

Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage:

- Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan
  adoption of slack-bot and weekly-post (live alarms of those exact names
  already exist outside the stack and must be deleted before deploy).
- Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions.
- Throttles (Sum, >=1/5min) for all six functions.
- DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors
  omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit
  permanently in INSUFFICIENT_DATA.
- API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the
  implicit ServerlessHttpApi.

All alarms page the shared site-alerts SNS topic, no OKActions,
TreatMissingData notBreaching. README updated with a Monitoring & Alarms section.

Duration and API latency thresholds pending sign-off.

* Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents

ThrottledRequests is not emitted at the TableName-only dimension (only
TableName+Operation), so the table-level alarm would sit permanently in
INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and
WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension.

* Correct DynamoDB alarm docs and drop sign-off wording

README DynamoDB section now lists the alarms actually shipped
(DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead
of the stale ThrottledRequests alarm. Thresholds are owner-approved, so
remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
This commit is contained in:
Adam Moussa 2026-06-17 14:45:47 -04:00 • committed by GitHub
parent 06bcb9b1b7
commit 90c10d38d3
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 477 additions and 0 deletions

View file

@ -189,6 +189,48 @@ sam build
sam deploy sam deploy
``` ```
## Monitoring & Alarms
All CloudWatch alarms are defined in `template.yaml` and notify the shared
`site-alerts` SNS topic (→ AWS Chatbot → Slack). None set `OKActions` — recovery
is not paged. Alarm names follow the in-template convention `Lambda-<Metric>-<fn>`
(e.g. `Lambda-Errors-afterhours-ring-scheduler`).
**Lambda alarms** (all six functions: `afterhours-shift-manager`,
`afterhours-weekly-post`, `afterhours-roster-sync`, `afterhours-ring-scheduler`,
`afterhours-holiday-router`, `afterhours-release-notifier`):
| Alarm | Metric | Condition | Notes |
|---|---|---|---|
| `Lambda-Errors-<fn>` | `Errors` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
| `Lambda-Duration-<fn>` | `Duration` (Maximum, ms) | `>= ~80% of timeout`, 2 of 3 5-min periods | Thresholds: 24000 ms (30s-timeout fns) / 48000 ms (60s-timeout fns) |
| `Lambda-Throttles-<fn>` | `Throttles` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
**DynamoDB alarm** (`afterhours-shifts` table):
| Alarm | Metric | Condition | Notes |
|---|---|---|---|
| `DDB-ReadThrottle-afterhours-shifts` | `ReadThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
| `DDB-WriteThrottle-afterhours-shifts` | `WriteThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
`ReadThrottleEvents` / `WriteThrottleEvents` are the table-level throttle
signals: AWS/DynamoDB emits them at the `TableName` dimension, so these alarms
transition normally. `ThrottledRequests` and `SystemErrors` are intentionally
**not** alarmed: AWS emits them only at `TableName`+`Operation` granularity, so a
`TableName`-only alarm would sit permanently in `INSUFFICIENT_DATA`.
**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway`
v2 metrics, `ApiId` dimension):
| Alarm | Metric | Condition |
|---|---|---|
| `ApiGateway-4xx-<apiId>` | `4xx` (Sum) | `>= 5` over one 5-min period |
| `ApiGateway-5xx-<apiId>` | `5xx` (Sum) | `>= 1` over one 5-min period |
| `ApiGateway-Latency-<apiId>` | `Latency` (p99, ms) | `>= 3000` ms, 2 of 3 5-min periods |
> Duration and API latency thresholds are starting points and may be tuned after
> observing real traffic.
## Releases & Versioning ## Releases & Versioning
The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`** — it is The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`** — it is

View file

@ -455,6 +455,441 @@ Resources:
Action: lambda:InvokeFunction Action: lambda:InvokeFunction
Resource: !GetAtt ReleaseNotifierFunction.Arn Resource: !GetAtt ReleaseNotifierFunction.Arn
# ===========================================================================
# CloudWatch alarm coverage (Wave 1 PR A). All alarms page the same
# site-alerts SNS topic → AWS Chatbot → Slack as the existing Errors alarms.
# No OKActions by design (the existing Errors alarms have none either).
# Naming follows the in-template convention: Lambda-<Metric>-${Fn}.
# ===========================================================================
# --- Lambda Errors alarms (clone of HolidayRouterErrorAlarm) ---
# Errors for ring-scheduler + holiday-router already exist above.
RosterSyncErrorAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Errors-${RosterSyncFunction}"
AlarmDescription: "Roster sync Lambda reported one or more errors"
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref RosterSyncFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
ReleaseNotifierErrorAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Errors-${ReleaseNotifierFunction}"
AlarmDescription: "Release notifier Lambda reported one or more errors"
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref ReleaseNotifierFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# ORPHAN ADOPTION: live alarms named Lambda-Errors-afterhours-shift-manager
# and Lambda-Errors-afterhours-weekly-post already exist OUTSIDE the stack.
# They MUST be deleted immediately before this stack deploys, or CloudFormation
# will fail to create these resources (AlarmName collision). See PR body.
SlackBotErrorAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Errors-${SlackBotFunction}"
AlarmDescription: "Slack bot Lambda reported one or more errors"
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref SlackBotFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
WeeklyPostErrorAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Errors-${WeeklyPostFunction}"
AlarmDescription: "Weekly post Lambda reported one or more errors"
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref WeeklyPostFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) ---
# Thresholds set to ~80% of each function's timeout, with 2-of-3 evaluation.
# Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default);
# roster-sync/ring-scheduler/holiday-router = 60s.
SlackBotDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Duration-${SlackBotFunction}"
AlarmDescription: "Slack bot Lambda duration approaching its 30s timeout (>=24s)"
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref SlackBotFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 3
DatapointsToAlarm: 2
Threshold: 24000
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
WeeklyPostDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Duration-${WeeklyPostFunction}"
AlarmDescription: "Weekly post Lambda duration approaching its 30s timeout (>=24s)"
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref WeeklyPostFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 3
DatapointsToAlarm: 2
Threshold: 24000
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
RosterSyncDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Duration-${RosterSyncFunction}"
AlarmDescription: "Roster sync Lambda duration approaching its 60s timeout (>=48s)"
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref RosterSyncFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 3
DatapointsToAlarm: 2
Threshold: 48000
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
RingSchedulerDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Duration-${RingSchedulerFunction}"
AlarmDescription: "Ring scheduler Lambda duration approaching its 60s timeout (>=48s)"
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref RingSchedulerFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 3
DatapointsToAlarm: 2
Threshold: 48000
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
HolidayRouterDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Duration-${HolidayRouterFunction}"
AlarmDescription: "Holiday router Lambda duration approaching its 60s timeout (>=48s)"
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref HolidayRouterFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 3
DatapointsToAlarm: 2
Threshold: 48000
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
ReleaseNotifierDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Duration-${ReleaseNotifierFunction}"
AlarmDescription: "Release notifier Lambda duration approaching its 30s timeout (>=24s)"
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref ReleaseNotifierFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 3
DatapointsToAlarm: 2
Threshold: 24000
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# --- Lambda Throttles alarms (Sum; threshold 1 over one 5-min period) ---
SlackBotThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Throttles-${SlackBotFunction}"
AlarmDescription: "Slack bot Lambda was throttled (concurrency limit hit)"
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref SlackBotFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
WeeklyPostThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Throttles-${WeeklyPostFunction}"
AlarmDescription: "Weekly post Lambda was throttled (concurrency limit hit)"
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref WeeklyPostFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
RosterSyncThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Throttles-${RosterSyncFunction}"
AlarmDescription: "Roster sync Lambda was throttled (concurrency limit hit)"
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref RosterSyncFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
RingSchedulerThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Throttles-${RingSchedulerFunction}"
AlarmDescription: "Ring scheduler Lambda was throttled (concurrency limit hit)"
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref RingSchedulerFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
HolidayRouterThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Throttles-${HolidayRouterFunction}"
AlarmDescription: "Holiday router Lambda was throttled (concurrency limit hit)"
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref HolidayRouterFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
ReleaseNotifierThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "Lambda-Throttles-${ReleaseNotifierFunction}"
AlarmDescription: "Release notifier Lambda was throttled (concurrency limit hit)"
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref ReleaseNotifierFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# --- DynamoDB alarms (afterhours-shifts table) ---
# ReadThrottleEvents / WriteThrottleEvents are the table-level throttle
# signals: AWS/DynamoDB emits them at the TableName dimension, so these
# alarms transition normally. (ThrottledRequests and SystemErrors are NOT
# emitted at TableName-only granularity — only at TableName+Operation — so
# alarms on them sit permanently in INSUFFICIENT_DATA and never fire.)
ShiftTableReadThrottleAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "DDB-ReadThrottle-${ShiftTable}"
AlarmDescription: "afterhours-shifts table had one or more read throttle events"
Namespace: AWS/DynamoDB
MetricName: ReadThrottleEvents
Dimensions:
- Name: TableName
Value: !Ref ShiftTable
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
ShiftTableWriteThrottleAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "DDB-WriteThrottle-${ShiftTable}"
AlarmDescription: "afterhours-shifts table had one or more write throttle events"
Namespace: AWS/DynamoDB
MetricName: WriteThrottleEvents
Dimensions:
- Name: TableName
Value: !Ref ShiftTable
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# --- API Gateway v2 (HTTP API) alarms on the implicit ServerlessHttpApi ---
# AWS::ApiGatewayV2 metric names: 4xx, 5xx, Latency; dimension ApiId.
ApiGateway4xxAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "ApiGateway-4xx-${ServerlessHttpApi}"
AlarmDescription: "Elevated 4xx responses on the Slack events HTTP API"
Namespace: AWS/ApiGateway
MetricName: 4xx
Dimensions:
- Name: ApiId
Value: !Ref ServerlessHttpApi
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 5
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
ApiGateway5xxAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "ApiGateway-5xx-${ServerlessHttpApi}"
AlarmDescription: "5xx responses on the Slack events HTTP API"
Namespace: AWS/ApiGateway
MetricName: 5xx
Dimensions:
- Name: ApiId
Value: !Ref ServerlessHttpApi
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 1
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# Latency p99 via ExtendedStatistic. ~3000ms target chosen alongside the
# Lambda Duration thresholds (Slack requires a fast 3s ack).
ApiGatewayLatencyAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: !Sub "ApiGateway-Latency-${ServerlessHttpApi}"
AlarmDescription: "p99 latency on the Slack events HTTP API exceeded 3s"
Namespace: AWS/ApiGateway
MetricName: Latency
Dimensions:
- Name: ApiId
Value: !Ref ServerlessHttpApi
ExtendedStatistic: p99
Period: 300
EvaluationPeriods: 3
DatapointsToAlarm: 2
Threshold: 3000
ComparisonOperator: GreaterThanOrEqualToThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# --- CloudWatch Log Groups (explicit 60-day retention) --- # --- CloudWatch Log Groups (explicit 60-day retention) ---
SlackBotLogGroup: SlackBotLogGroup:
Type: AWS::Logs::LogGroup Type: AWS::Logs::LogGroup