mirror of
https://github.com/Sea-Haven-Industries/afterhours-shift-manager.git
synced 2026-10-06 13:32:04 +00:00
Correct DynamoDB alarm docs and drop sign-off wording
README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
This commit is contained in:
parent
0a299d5525
commit
7aff5952b2
2 changed files with 10 additions and 8 deletions
11
README.md
11
README.md
|
|
@ -210,11 +210,14 @@ is not paged. Alarm names follow the in-template convention `Lambda-<Metric>-<fn
|
||||||
|
|
||||||
| Alarm | Metric | Condition | Notes |
|
| Alarm | Metric | Condition | Notes |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| `DynamoDB-ThrottledRequests-afterhours-shifts` | `ThrottledRequests` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
| `DDB-ReadThrottle-afterhours-shifts` | `ReadThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
||||||
|
| `DDB-WriteThrottle-afterhours-shifts` | `WriteThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
||||||
|
|
||||||
`SystemErrors` is intentionally **not** alarmed: AWS emits it only at
|
`ReadThrottleEvents` / `WriteThrottleEvents` are the table-level throttle
|
||||||
`TableName`+`Operation` granularity, so a `TableName`-only alarm would sit
|
signals: AWS/DynamoDB emits them at the `TableName` dimension, so these alarms
|
||||||
permanently in `INSUFFICIENT_DATA`.
|
transition normally. `ThrottledRequests` and `SystemErrors` are intentionally
|
||||||
|
**not** alarmed: AWS emits them only at `TableName`+`Operation` granularity, so a
|
||||||
|
`TableName`-only alarm would sit permanently in `INSUFFICIENT_DATA`.
|
||||||
|
|
||||||
**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway`
|
**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway`
|
||||||
v2 metrics, `ApiId` dimension):
|
v2 metrics, `ApiId` dimension):
|
||||||
|
|
|
||||||
|
|
@ -546,8 +546,7 @@ Resources:
|
||||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||||
|
|
||||||
# --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) ---
|
# --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) ---
|
||||||
# Thresholds set to ~80% of each function's timeout. ALL Duration thresholds
|
# Thresholds set to ~80% of each function's timeout, with 2-of-3 evaluation.
|
||||||
# and the 3/2 evaluation are PENDING ADAM SIGN-OFF (see PR body).
|
|
||||||
# Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default);
|
# Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default);
|
||||||
# roster-sync/ring-scheduler/holiday-router = 60s.
|
# roster-sync/ring-scheduler/holiday-router = 60s.
|
||||||
SlackBotDurationAlarm:
|
SlackBotDurationAlarm:
|
||||||
|
|
@ -869,8 +868,8 @@ Resources:
|
||||||
AlarmActions:
|
AlarmActions:
|
||||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||||
|
|
||||||
# Latency p99 via ExtendedStatistic. ~3000ms target is PENDING ADAM SIGN-OFF
|
# Latency p99 via ExtendedStatistic. ~3000ms target chosen alongside the
|
||||||
# alongside the Lambda Duration thresholds (Slack requires a fast 3s ack).
|
# Lambda Duration thresholds (Slack requires a fast 3s ack).
|
||||||
ApiGatewayLatencyAlarm:
|
ApiGatewayLatencyAlarm:
|
||||||
Type: AWS::CloudWatch::Alarm
|
Type: AWS::CloudWatch::Alarm
|
||||||
Properties:
|
Properties:
|
||||||
|
|
|
||||||
Loading…
Add table
Reference in a new issue