Correct DynamoDB alarm docs and drop sign-off wording

README DynamoDB section now lists the alarms actually shipped
(DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead
of the stale ThrottledRequests alarm. Thresholds are owner-approved, so
remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
This commit is contained in:
Adam Moussa 2026-06-17 14:33:57 -04:00
parent 0a299d5525
commit 7aff5952b2
2 changed files with 10 additions and 8 deletions

View file

@ -210,11 +210,14 @@ is not paged. Alarm names follow the in-template convention `Lambda-<Metric>-<fn
| Alarm | Metric | Condition | Notes |
|---|---|---|---|
| `DynamoDB-ThrottledRequests-afterhours-shifts` | `ThrottledRequests` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
| `DDB-ReadThrottle-afterhours-shifts` | `ReadThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
| `DDB-WriteThrottle-afterhours-shifts` | `WriteThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
`SystemErrors` is intentionally **not** alarmed: AWS emits it only at
`TableName`+`Operation` granularity, so a `TableName`-only alarm would sit
permanently in `INSUFFICIENT_DATA`.
`ReadThrottleEvents` / `WriteThrottleEvents` are the table-level throttle
signals: AWS/DynamoDB emits them at the `TableName` dimension, so these alarms
transition normally. `ThrottledRequests` and `SystemErrors` are intentionally
**not** alarmed: AWS emits them only at `TableName`+`Operation` granularity, so a
`TableName`-only alarm would sit permanently in `INSUFFICIENT_DATA`.
**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway`
v2 metrics, `ApiId` dimension):

View file

@ -546,8 +546,7 @@ Resources:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) ---
# Thresholds set to ~80% of each function's timeout. ALL Duration thresholds
# and the 3/2 evaluation are PENDING ADAM SIGN-OFF (see PR body).
# Thresholds set to ~80% of each function's timeout, with 2-of-3 evaluation.
# Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default);
# roster-sync/ring-scheduler/holiday-router = 60s.
SlackBotDurationAlarm:
@ -869,8 +868,8 @@ Resources:
AlarmActions:
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
# Latency p99 via ExtendedStatistic. ~3000ms target is PENDING ADAM SIGN-OFF
# alongside the Lambda Duration thresholds (Slack requires a fast 3s ack).
# Latency p99 via ExtendedStatistic. ~3000ms target chosen alongside the
# Lambda Duration thresholds (Slack requires a fast 3s ack).
ApiGatewayLatencyAlarm:
Type: AWS::CloudWatch::Alarm
Properties: