mirror of
https://github.com/Sea-Haven-Industries/afterhours-shift-manager.git
synced 2026-09-30 10:13:11 +00:00
Add CloudWatch alarm coverage for all functions, table, and HTTP API
Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off.
This commit is contained in:
parent
06bcb9b1b7
commit
38ab8a1487
2 changed files with 457 additions and 0 deletions
39
README.md
39
README.md
|
|
@ -189,6 +189,45 @@ sam build
|
|||
sam deploy
|
||||
```
|
||||
|
||||
## Monitoring & Alarms
|
||||
|
||||
All CloudWatch alarms are defined in `template.yaml` and notify the shared
|
||||
`site-alerts` SNS topic (→ AWS Chatbot → Slack). None set `OKActions` — recovery
|
||||
is not paged. Alarm names follow the in-template convention `Lambda-<Metric>-<fn>`
|
||||
(e.g. `Lambda-Errors-afterhours-ring-scheduler`).
|
||||
|
||||
**Lambda alarms** (all six functions: `afterhours-shift-manager`,
|
||||
`afterhours-weekly-post`, `afterhours-roster-sync`, `afterhours-ring-scheduler`,
|
||||
`afterhours-holiday-router`, `afterhours-release-notifier`):
|
||||
|
||||
| Alarm | Metric | Condition | Notes |
|
||||
|---|---|---|---|
|
||||
| `Lambda-Errors-<fn>` | `Errors` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
||||
| `Lambda-Duration-<fn>` | `Duration` (Maximum, ms) | `>= ~80% of timeout`, 2 of 3 5-min periods | Thresholds: 24000 ms (30s-timeout fns) / 48000 ms (60s-timeout fns) |
|
||||
| `Lambda-Throttles-<fn>` | `Throttles` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
||||
|
||||
**DynamoDB alarm** (`afterhours-shifts` table):
|
||||
|
||||
| Alarm | Metric | Condition | Notes |
|
||||
|---|---|---|---|
|
||||
| `DynamoDB-ThrottledRequests-afterhours-shifts` | `ThrottledRequests` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
||||
|
||||
`SystemErrors` is intentionally **not** alarmed: AWS emits it only at
|
||||
`TableName`+`Operation` granularity, so a `TableName`-only alarm would sit
|
||||
permanently in `INSUFFICIENT_DATA`.
|
||||
|
||||
**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway`
|
||||
v2 metrics, `ApiId` dimension):
|
||||
|
||||
| Alarm | Metric | Condition |
|
||||
|---|---|---|
|
||||
| `ApiGateway-4xx-<apiId>` | `4xx` (Sum) | `>= 5` over one 5-min period |
|
||||
| `ApiGateway-5xx-<apiId>` | `5xx` (Sum) | `>= 1` over one 5-min period |
|
||||
| `ApiGateway-Latency-<apiId>` | `Latency` (p99, ms) | `>= 3000` ms, 2 of 3 5-min periods |
|
||||
|
||||
> Duration and API latency thresholds are starting points and may be tuned after
|
||||
> observing real traffic.
|
||||
|
||||
## Releases & Versioning
|
||||
|
||||
The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`** — it is
|
||||
|
|
|
|||
418
template.yaml
418
template.yaml
|
|
@ -455,6 +455,424 @@ Resources:
|
|||
Action: lambda:InvokeFunction
|
||||
Resource: !GetAtt ReleaseNotifierFunction.Arn
|
||||
|
||||
# ===========================================================================
|
||||
# CloudWatch alarm coverage (Wave 1 PR A). All alarms page the same
|
||||
# site-alerts SNS topic → AWS Chatbot → Slack as the existing Errors alarms.
|
||||
# No OKActions by design (the existing Errors alarms have none either).
|
||||
# Naming follows the in-template convention: Lambda-<Metric>-${Fn}.
|
||||
# ===========================================================================
|
||||
|
||||
# --- Lambda Errors alarms (clone of HolidayRouterErrorAlarm) ---
|
||||
# Errors for ring-scheduler + holiday-router already exist above.
|
||||
|
||||
RosterSyncErrorAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Errors-${RosterSyncFunction}"
|
||||
AlarmDescription: "Roster sync Lambda reported one or more errors"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Errors
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref RosterSyncFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
ReleaseNotifierErrorAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Errors-${ReleaseNotifierFunction}"
|
||||
AlarmDescription: "Release notifier Lambda reported one or more errors"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Errors
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref ReleaseNotifierFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
# ORPHAN ADOPTION: live alarms named Lambda-Errors-afterhours-shift-manager
|
||||
# and Lambda-Errors-afterhours-weekly-post already exist OUTSIDE the stack.
|
||||
# They MUST be deleted immediately before this stack deploys, or CloudFormation
|
||||
# will fail to create these resources (AlarmName collision). See PR body.
|
||||
SlackBotErrorAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Errors-${SlackBotFunction}"
|
||||
AlarmDescription: "Slack bot Lambda reported one or more errors"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Errors
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref SlackBotFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
WeeklyPostErrorAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Errors-${WeeklyPostFunction}"
|
||||
AlarmDescription: "Weekly post Lambda reported one or more errors"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Errors
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref WeeklyPostFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
# --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) ---
|
||||
# Thresholds set to ~80% of each function's timeout. ALL Duration thresholds
|
||||
# and the 3/2 evaluation are PENDING ADAM SIGN-OFF (see PR body).
|
||||
# Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default);
|
||||
# roster-sync/ring-scheduler/holiday-router = 60s.
|
||||
SlackBotDurationAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Duration-${SlackBotFunction}"
|
||||
AlarmDescription: "Slack bot Lambda duration approaching its 30s timeout (>=24s)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Duration
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref SlackBotFunction
|
||||
Statistic: Maximum
|
||||
Period: 300
|
||||
EvaluationPeriods: 3
|
||||
DatapointsToAlarm: 2
|
||||
Threshold: 24000
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
WeeklyPostDurationAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Duration-${WeeklyPostFunction}"
|
||||
AlarmDescription: "Weekly post Lambda duration approaching its 30s timeout (>=24s)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Duration
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref WeeklyPostFunction
|
||||
Statistic: Maximum
|
||||
Period: 300
|
||||
EvaluationPeriods: 3
|
||||
DatapointsToAlarm: 2
|
||||
Threshold: 24000
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
RosterSyncDurationAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Duration-${RosterSyncFunction}"
|
||||
AlarmDescription: "Roster sync Lambda duration approaching its 60s timeout (>=48s)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Duration
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref RosterSyncFunction
|
||||
Statistic: Maximum
|
||||
Period: 300
|
||||
EvaluationPeriods: 3
|
||||
DatapointsToAlarm: 2
|
||||
Threshold: 48000
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
RingSchedulerDurationAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Duration-${RingSchedulerFunction}"
|
||||
AlarmDescription: "Ring scheduler Lambda duration approaching its 60s timeout (>=48s)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Duration
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref RingSchedulerFunction
|
||||
Statistic: Maximum
|
||||
Period: 300
|
||||
EvaluationPeriods: 3
|
||||
DatapointsToAlarm: 2
|
||||
Threshold: 48000
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
HolidayRouterDurationAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Duration-${HolidayRouterFunction}"
|
||||
AlarmDescription: "Holiday router Lambda duration approaching its 60s timeout (>=48s)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Duration
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref HolidayRouterFunction
|
||||
Statistic: Maximum
|
||||
Period: 300
|
||||
EvaluationPeriods: 3
|
||||
DatapointsToAlarm: 2
|
||||
Threshold: 48000
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
ReleaseNotifierDurationAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Duration-${ReleaseNotifierFunction}"
|
||||
AlarmDescription: "Release notifier Lambda duration approaching its 30s timeout (>=24s)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Duration
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref ReleaseNotifierFunction
|
||||
Statistic: Maximum
|
||||
Period: 300
|
||||
EvaluationPeriods: 3
|
||||
DatapointsToAlarm: 2
|
||||
Threshold: 24000
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
# --- Lambda Throttles alarms (Sum; threshold 1 over one 5-min period) ---
|
||||
SlackBotThrottlesAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Throttles-${SlackBotFunction}"
|
||||
AlarmDescription: "Slack bot Lambda was throttled (concurrency limit hit)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Throttles
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref SlackBotFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
WeeklyPostThrottlesAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Throttles-${WeeklyPostFunction}"
|
||||
AlarmDescription: "Weekly post Lambda was throttled (concurrency limit hit)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Throttles
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref WeeklyPostFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
RosterSyncThrottlesAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Throttles-${RosterSyncFunction}"
|
||||
AlarmDescription: "Roster sync Lambda was throttled (concurrency limit hit)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Throttles
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref RosterSyncFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
RingSchedulerThrottlesAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Throttles-${RingSchedulerFunction}"
|
||||
AlarmDescription: "Ring scheduler Lambda was throttled (concurrency limit hit)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Throttles
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref RingSchedulerFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
HolidayRouterThrottlesAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Throttles-${HolidayRouterFunction}"
|
||||
AlarmDescription: "Holiday router Lambda was throttled (concurrency limit hit)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Throttles
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref HolidayRouterFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
ReleaseNotifierThrottlesAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "Lambda-Throttles-${ReleaseNotifierFunction}"
|
||||
AlarmDescription: "Release notifier Lambda was throttled (concurrency limit hit)"
|
||||
Namespace: AWS/Lambda
|
||||
MetricName: Throttles
|
||||
Dimensions:
|
||||
- Name: FunctionName
|
||||
Value: !Ref ReleaseNotifierFunction
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
# --- DynamoDB alarms (afterhours-shifts table) ---
|
||||
# ThrottledRequests emits at the TableName dimension → safe to alarm.
|
||||
# NOTE: SystemErrors is INTENTIONALLY OMITTED. AWS emits AWS/DynamoDB
|
||||
# SystemErrors only at TableName+Operation granularity, never TableName-only,
|
||||
# so a TableName-only SystemErrors alarm would sit permanently in
|
||||
# INSUFFICIENT_DATA. (Verified: no DynamoDB table in this account has ever
|
||||
# emitted SystemErrors, consistent with documented per-Operation emission.)
|
||||
ShiftTableThrottleAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "DynamoDB-ThrottledRequests-${ShiftTable}"
|
||||
AlarmDescription: "afterhours-shifts table had one or more throttled requests"
|
||||
Namespace: AWS/DynamoDB
|
||||
MetricName: ThrottledRequests
|
||||
Dimensions:
|
||||
- Name: TableName
|
||||
Value: !Ref ShiftTable
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
# --- API Gateway v2 (HTTP API) alarms on the implicit ServerlessHttpApi ---
|
||||
# AWS::ApiGatewayV2 metric names: 4xx, 5xx, Latency; dimension ApiId.
|
||||
ApiGateway4xxAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "ApiGateway-4xx-${ServerlessHttpApi}"
|
||||
AlarmDescription: "Elevated 4xx responses on the Slack events HTTP API"
|
||||
Namespace: AWS/ApiGateway
|
||||
MetricName: 4xx
|
||||
Dimensions:
|
||||
- Name: ApiId
|
||||
Value: !Ref ServerlessHttpApi
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 5
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
ApiGateway5xxAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "ApiGateway-5xx-${ServerlessHttpApi}"
|
||||
AlarmDescription: "5xx responses on the Slack events HTTP API"
|
||||
Namespace: AWS/ApiGateway
|
||||
MetricName: 5xx
|
||||
Dimensions:
|
||||
- Name: ApiId
|
||||
Value: !Ref ServerlessHttpApi
|
||||
Statistic: Sum
|
||||
Period: 300
|
||||
EvaluationPeriods: 1
|
||||
Threshold: 1
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
# Latency p99 via ExtendedStatistic. ~3000ms target is PENDING ADAM SIGN-OFF
|
||||
# alongside the Lambda Duration thresholds (Slack requires a fast 3s ack).
|
||||
ApiGatewayLatencyAlarm:
|
||||
Type: AWS::CloudWatch::Alarm
|
||||
Properties:
|
||||
AlarmName: !Sub "ApiGateway-Latency-${ServerlessHttpApi}"
|
||||
AlarmDescription: "p99 latency on the Slack events HTTP API exceeded 3s"
|
||||
Namespace: AWS/ApiGateway
|
||||
MetricName: Latency
|
||||
Dimensions:
|
||||
- Name: ApiId
|
||||
Value: !Ref ServerlessHttpApi
|
||||
ExtendedStatistic: p99
|
||||
Period: 300
|
||||
EvaluationPeriods: 3
|
||||
DatapointsToAlarm: 2
|
||||
Threshold: 3000
|
||||
ComparisonOperator: GreaterThanOrEqualToThreshold
|
||||
TreatMissingData: notBreaching
|
||||
AlarmActions:
|
||||
- !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts"
|
||||
|
||||
# --- CloudWatch Log Groups (explicit 60-day retention) ---
|
||||
SlackBotLogGroup:
|
||||
Type: AWS::Logs::LogGroup
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue