mirror of
https://github.com/Sea-Haven-Industries/front-integrations.git
synced 2026-09-30 03:43:12 +00:00
Add CloudWatch alarm coverage for front-integrations (#11)
Some checks failed
Deploy / deploy (push) Has been cancelled
Some checks failed
Deploy / deploy (push) Has been cancelled
* Add CloudWatch alarm coverage for front-integrations Both Lambdas and the front-sla-alerts table previously had zero alarm coverage, so failures or runaway runs went unnoticed until someone checked logs. Wire a standard alarm set to the shared site-alerts SNS topic (ALARM-only, TreatMissingData notBreaching) per Wave 1 conventions. - Lambda Errors + Throttles alarms for front-sla-monitor and front-user-sync (Sum, threshold 0). - Lambda Duration alarms (Max, threshold 270000 = 90% of the shared 300s timeout) for both functions. - DynamoDB ThrottledRequests + SystemErrors alarms on front-sla-alerts. Document the alarm set in the README. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests and SystemErrors are not emitted at the TableName-only dimension (only TableName+Operation), so these table-level alarms would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Fix README DynamoDB alarm rows to match shipped alarms Replace stale front-sla-alerts-throttled-requests / -system-errors rows with the alarms actually shipped: front-sla-alerts-read-throttle (ReadThrottleEvents) and front-sla-alerts-write-throttle (WriteThrottleEvents).
This commit is contained in:
parent
98b30ba1ec
commit
7dcd2d7ccc
2 changed files with 175 additions and 0 deletions
15
README.md
15
README.md
|
|
@ -60,6 +60,21 @@ Business time counts weekday hours only (Mon-Fri, Eastern time). Alerts are only
|
||||||
- **DynamoDB:** `front-sla-alerts` — alert history per conversation + monitor state, 7-day TTL
|
- **DynamoDB:** `front-sla-alerts` — alert history per conversation + monitor state, 7-day TTL
|
||||||
- **EventBridge:** SLA check every 15 min during business hours; user sync daily at 6 AM ET weekdays
|
- **EventBridge:** SLA check every 15 min during business hours; user sync daily at 6 AM ET weekdays
|
||||||
|
|
||||||
|
## Monitoring
|
||||||
|
|
||||||
|
CloudWatch alarms publish to the shared `site-alerts` SNS topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`). All alarms are ALARM-only (no OK/recovery notification) and treat missing data as `notBreaching`.
|
||||||
|
|
||||||
|
| Alarm | Metric | Trigger |
|
||||||
|
|---|---|---|
|
||||||
|
| `front-sla-monitor-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window |
|
||||||
|
| `front-sla-monitor-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window |
|
||||||
|
| `front-sla-monitor-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) |
|
||||||
|
| `front-user-sync-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window |
|
||||||
|
| `front-user-sync-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window |
|
||||||
|
| `front-user-sync-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) |
|
||||||
|
| `front-sla-alerts-read-throttle` | DynamoDB `ReadThrottleEvents` (Sum) | Any read throttle on the table |
|
||||||
|
| `front-sla-alerts-write-throttle` | DynamoDB `WriteThrottleEvents` (Sum) | Any write throttle on the table |
|
||||||
|
|
||||||
## Secrets (Secrets Manager)
|
## Secrets (Secrets Manager)
|
||||||
|
|
||||||
| Secret | Purpose |
|
| Secret | Purpose |
|
||||||
|
|
|
||||||
160
template.yaml
160
template.yaml
|
|
@ -115,6 +115,108 @@ Resources:
|
||||||
Description: Check Front conversations for SLA breaches every 15 min during business hours
|
Description: Check Front conversations for SLA breaches every 15 min during business hours
|
||||||
Enabled: true
|
Enabled: true
|
||||||
|
|
||||||
|
# SLA Monitor alarms
|
||||||
|
SlaMonitorErrorsAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-sla-monitor-errors
|
||||||
|
AlarmDescription: front-sla-monitor invocation errors
|
||||||
|
Namespace: AWS/Lambda
|
||||||
|
MetricName: Errors
|
||||||
|
Dimensions:
|
||||||
|
- Name: FunctionName
|
||||||
|
Value: !Ref SlaMonitorFunction
|
||||||
|
Statistic: Sum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 0
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
|
SlaMonitorThrottlesAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-sla-monitor-throttles
|
||||||
|
AlarmDescription: front-sla-monitor invocations throttled
|
||||||
|
Namespace: AWS/Lambda
|
||||||
|
MetricName: Throttles
|
||||||
|
Dimensions:
|
||||||
|
- Name: FunctionName
|
||||||
|
Value: !Ref SlaMonitorFunction
|
||||||
|
Statistic: Sum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 0
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
|
SlaMonitorDurationAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-sla-monitor-duration
|
||||||
|
AlarmDescription: front-sla-monitor approaching its 300s timeout (>90%)
|
||||||
|
Namespace: AWS/Lambda
|
||||||
|
MetricName: Duration
|
||||||
|
Dimensions:
|
||||||
|
- Name: FunctionName
|
||||||
|
Value: !Ref SlaMonitorFunction
|
||||||
|
Statistic: Maximum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 270000
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
|
# DynamoDB alarms — front-sla-alerts table
|
||||||
|
# ReadThrottleEvents / WriteThrottleEvents are emitted at the TableName
|
||||||
|
# dimension, so these alarms transition normally. (ThrottledRequests and
|
||||||
|
# SystemErrors are only emitted at TableName+Operation granularity, never
|
||||||
|
# TableName-only, so alarms on them sit permanently in INSUFFICIENT_DATA
|
||||||
|
# and never fire — these are the correct table-level throttle signals.)
|
||||||
|
AlertsTableReadThrottleAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-sla-alerts-read-throttle
|
||||||
|
AlarmDescription: front-sla-alerts DynamoDB table had one or more read throttle events
|
||||||
|
Namespace: AWS/DynamoDB
|
||||||
|
MetricName: ReadThrottleEvents
|
||||||
|
Dimensions:
|
||||||
|
- Name: TableName
|
||||||
|
Value: !Ref AlertsTable
|
||||||
|
Statistic: Sum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 0
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
|
AlertsTableWriteThrottleAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-sla-alerts-write-throttle
|
||||||
|
AlarmDescription: front-sla-alerts DynamoDB table had one or more write throttle events
|
||||||
|
Namespace: AWS/DynamoDB
|
||||||
|
MetricName: WriteThrottleEvents
|
||||||
|
Dimensions:
|
||||||
|
- Name: TableName
|
||||||
|
Value: !Ref AlertsTable
|
||||||
|
Statistic: Sum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 0
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Google User Sync
|
# Google User Sync
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|
@ -155,6 +257,64 @@ Resources:
|
||||||
Description: Sync Google Workspace user profiles to Front daily at 6 AM ET
|
Description: Sync Google Workspace user profiles to Front daily at 6 AM ET
|
||||||
Enabled: true
|
Enabled: true
|
||||||
|
|
||||||
|
# User Sync alarms
|
||||||
|
UserSyncErrorsAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-user-sync-errors
|
||||||
|
AlarmDescription: front-user-sync invocation errors
|
||||||
|
Namespace: AWS/Lambda
|
||||||
|
MetricName: Errors
|
||||||
|
Dimensions:
|
||||||
|
- Name: FunctionName
|
||||||
|
Value: !Ref UserSyncFunction
|
||||||
|
Statistic: Sum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 0
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
|
UserSyncThrottlesAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-user-sync-throttles
|
||||||
|
AlarmDescription: front-user-sync invocations throttled
|
||||||
|
Namespace: AWS/Lambda
|
||||||
|
MetricName: Throttles
|
||||||
|
Dimensions:
|
||||||
|
- Name: FunctionName
|
||||||
|
Value: !Ref UserSyncFunction
|
||||||
|
Statistic: Sum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 0
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
|
UserSyncDurationAlarm:
|
||||||
|
Type: AWS::CloudWatch::Alarm
|
||||||
|
Properties:
|
||||||
|
AlarmName: front-user-sync-duration
|
||||||
|
AlarmDescription: front-user-sync approaching its 300s timeout (>90%)
|
||||||
|
Namespace: AWS/Lambda
|
||||||
|
MetricName: Duration
|
||||||
|
Dimensions:
|
||||||
|
- Name: FunctionName
|
||||||
|
Value: !Ref UserSyncFunction
|
||||||
|
Statistic: Maximum
|
||||||
|
Period: 300
|
||||||
|
EvaluationPeriods: 1
|
||||||
|
Threshold: 270000
|
||||||
|
ComparisonOperator: GreaterThanThreshold
|
||||||
|
TreatMissingData: notBreaching
|
||||||
|
AlarmActions:
|
||||||
|
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
|
||||||
|
|
||||||
Outputs:
|
Outputs:
|
||||||
SlaMonitorFunctionArn:
|
SlaMonitorFunctionArn:
|
||||||
Description: Front SLA Monitor Lambda ARN
|
Description: Front SLA Monitor Lambda ARN
|
||||||
|
|
|
||||||
Loading…
Add table
Reference in a new issue