Add CloudWatch alarm coverage for front-integrations

Both Lambdas and the front-sla-alerts table previously had zero alarm
coverage, so failures or runaway runs went unnoticed until someone
checked logs. Wire a standard alarm set to the shared site-alerts SNS
topic (ALARM-only, TreatMissingData notBreaching) per Wave 1 conventions.

- Lambda Errors + Throttles alarms for front-sla-monitor and
  front-user-sync (Sum, threshold 0).
- Lambda Duration alarms (Max, threshold 270000 = 90% of the shared
  300s timeout) for both functions.
- DynamoDB ThrottledRequests + SystemErrors alarms on front-sla-alerts.

Document the alarm set in the README.
This commit is contained in:
Adam Moussa 2026-06-17 13:53:30 -04:00
parent 98b30ba1ec
commit efa31783e2
2 changed files with 175 additions and 0 deletions

View file

@ -60,6 +60,21 @@ Business time counts weekday hours only (Mon-Fri, Eastern time). Alerts are only
- **DynamoDB:** `front-sla-alerts` — alert history per conversation + monitor state, 7-day TTL
- **EventBridge:** SLA check every 15 min during business hours; user sync daily at 6 AM ET weekdays
## Monitoring
CloudWatch alarms publish to the shared `site-alerts` SNS topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`). All alarms are ALARM-only (no OK/recovery notification) and treat missing data as `notBreaching`.
| Alarm | Metric | Trigger |
|---|---|---|
| `front-sla-monitor-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window |
| `front-sla-monitor-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window |
| `front-sla-monitor-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) |
| `front-user-sync-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window |
| `front-user-sync-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window |
| `front-user-sync-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) |
| `front-sla-alerts-throttled-requests` | DynamoDB `ThrottledRequests` (Sum) | Any throttled request on the table |
| `front-sla-alerts-system-errors` | DynamoDB `SystemErrors` (Sum) | Any HTTP 500 from the table |
## Secrets (Secrets Manager)
| Secret | Purpose |

View file

@ -115,6 +115,108 @@ Resources:
Description: Check Front conversations for SLA breaches every 15 min during business hours
Enabled: true
# SLA Monitor alarms
SlaMonitorErrorsAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-monitor-errors
AlarmDescription: front-sla-monitor invocation errors
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref SlaMonitorFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
SlaMonitorThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-monitor-throttles
AlarmDescription: front-sla-monitor invocations throttled
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref SlaMonitorFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
SlaMonitorDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-monitor-duration
AlarmDescription: front-sla-monitor approaching its 300s timeout (>90%)
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref SlaMonitorFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 1
Threshold: 270000
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
# DynamoDB alarms — front-sla-alerts table
# ThrottledRequests and SystemErrors are documented at dimensions
# (TableName, Operation); CloudWatch also publishes the TableName-only
# aggregate, which is what we alarm on here. Neither has emitted yet
# (no throttle/5xx events to date), so the dimension was confirmed against
# the AWS DynamoDB metrics reference rather than live data.
AlertsTableThrottleAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-alerts-throttled-requests
AlarmDescription: front-sla-alerts DynamoDB table is throttling requests
Namespace: AWS/DynamoDB
MetricName: ThrottledRequests
Dimensions:
- Name: TableName
Value: !Ref AlertsTable
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
AlertsTableSystemErrorsAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-alerts-system-errors
AlarmDescription: front-sla-alerts DynamoDB table returned HTTP 500 system errors
Namespace: AWS/DynamoDB
MetricName: SystemErrors
Dimensions:
- Name: TableName
Value: !Ref AlertsTable
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
# ---------------------------------------------------------------------------
# Google User Sync
# ---------------------------------------------------------------------------
@ -155,6 +257,64 @@ Resources:
Description: Sync Google Workspace user profiles to Front daily at 6 AM ET
Enabled: true
# User Sync alarms
UserSyncErrorsAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-user-sync-errors
AlarmDescription: front-user-sync invocation errors
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref UserSyncFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
UserSyncThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-user-sync-throttles
AlarmDescription: front-user-sync invocations throttled
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref UserSyncFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
UserSyncDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-user-sync-duration
AlarmDescription: front-user-sync approaching its 300s timeout (>90%)
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref UserSyncFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 1
Threshold: 270000
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
Outputs:
SlaMonitorFunctionArn:
Description: Front SLA Monitor Lambda ARN