Add CloudWatch alarm coverage for front-integrations (#11)
Some checks failed
Deploy / deploy (push) Has been cancelled

* Add CloudWatch alarm coverage for front-integrations

Both Lambdas and the front-sla-alerts table previously had zero alarm
coverage, so failures or runaway runs went unnoticed until someone
checked logs. Wire a standard alarm set to the shared site-alerts SNS
topic (ALARM-only, TreatMissingData notBreaching) per Wave 1 conventions.

- Lambda Errors + Throttles alarms for front-sla-monitor and
  front-user-sync (Sum, threshold 0).
- Lambda Duration alarms (Max, threshold 270000 = 90% of the shared
  300s timeout) for both functions.
- DynamoDB ThrottledRequests + SystemErrors alarms on front-sla-alerts.

Document the alarm set in the README.

* Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents

ThrottledRequests and SystemErrors are not emitted at the TableName-only
dimension (only TableName+Operation), so these table-level alarms would sit
permanently in INSUFFICIENT_DATA and never fire. Replace with
ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the
TableName dimension.

* Fix README DynamoDB alarm rows to match shipped alarms

Replace stale front-sla-alerts-throttled-requests / -system-errors rows
with the alarms actually shipped: front-sla-alerts-read-throttle
(ReadThrottleEvents) and front-sla-alerts-write-throttle
(WriteThrottleEvents).
This commit is contained in:
Adam Moussa 2026-06-17 14:45:58 -04:00 • committed by GitHub
parent 98b30ba1ec
commit 7dcd2d7ccc
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 175 additions and 0 deletions

View file

@ -60,6 +60,21 @@ Business time counts weekday hours only (Mon-Fri, Eastern time). Alerts are only
- **DynamoDB:** `front-sla-alerts` — alert history per conversation + monitor state, 7-day TTL - **DynamoDB:** `front-sla-alerts` — alert history per conversation + monitor state, 7-day TTL
- **EventBridge:** SLA check every 15 min during business hours; user sync daily at 6 AM ET weekdays - **EventBridge:** SLA check every 15 min during business hours; user sync daily at 6 AM ET weekdays
## Monitoring
CloudWatch alarms publish to the shared `site-alerts` SNS topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`). All alarms are ALARM-only (no OK/recovery notification) and treat missing data as `notBreaching`.
| Alarm | Metric | Trigger |
|---|---|---|
| `front-sla-monitor-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window |
| `front-sla-monitor-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window |
| `front-sla-monitor-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) |
| `front-user-sync-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window |
| `front-user-sync-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window |
| `front-user-sync-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) |
| `front-sla-alerts-read-throttle` | DynamoDB `ReadThrottleEvents` (Sum) | Any read throttle on the table |
| `front-sla-alerts-write-throttle` | DynamoDB `WriteThrottleEvents` (Sum) | Any write throttle on the table |
## Secrets (Secrets Manager) ## Secrets (Secrets Manager)
| Secret | Purpose | | Secret | Purpose |

View file

@ -115,6 +115,108 @@ Resources:
Description: Check Front conversations for SLA breaches every 15 min during business hours Description: Check Front conversations for SLA breaches every 15 min during business hours
Enabled: true Enabled: true
# SLA Monitor alarms
SlaMonitorErrorsAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-monitor-errors
AlarmDescription: front-sla-monitor invocation errors
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref SlaMonitorFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
SlaMonitorThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-monitor-throttles
AlarmDescription: front-sla-monitor invocations throttled
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref SlaMonitorFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
SlaMonitorDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-monitor-duration
AlarmDescription: front-sla-monitor approaching its 300s timeout (>90%)
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref SlaMonitorFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 1
Threshold: 270000
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
# DynamoDB alarms — front-sla-alerts table
# ReadThrottleEvents / WriteThrottleEvents are emitted at the TableName
# dimension, so these alarms transition normally. (ThrottledRequests and
# SystemErrors are only emitted at TableName+Operation granularity, never
# TableName-only, so alarms on them sit permanently in INSUFFICIENT_DATA
# and never fire — these are the correct table-level throttle signals.)
AlertsTableReadThrottleAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-alerts-read-throttle
AlarmDescription: front-sla-alerts DynamoDB table had one or more read throttle events
Namespace: AWS/DynamoDB
MetricName: ReadThrottleEvents
Dimensions:
- Name: TableName
Value: !Ref AlertsTable
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
AlertsTableWriteThrottleAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-sla-alerts-write-throttle
AlarmDescription: front-sla-alerts DynamoDB table had one or more write throttle events
Namespace: AWS/DynamoDB
MetricName: WriteThrottleEvents
Dimensions:
- Name: TableName
Value: !Ref AlertsTable
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Google User Sync # Google User Sync
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
@ -155,6 +257,64 @@ Resources:
Description: Sync Google Workspace user profiles to Front daily at 6 AM ET Description: Sync Google Workspace user profiles to Front daily at 6 AM ET
Enabled: true Enabled: true
# User Sync alarms
UserSyncErrorsAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-user-sync-errors
AlarmDescription: front-user-sync invocation errors
Namespace: AWS/Lambda
MetricName: Errors
Dimensions:
- Name: FunctionName
Value: !Ref UserSyncFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
UserSyncThrottlesAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-user-sync-throttles
AlarmDescription: front-user-sync invocations throttled
Namespace: AWS/Lambda
MetricName: Throttles
Dimensions:
- Name: FunctionName
Value: !Ref UserSyncFunction
Statistic: Sum
Period: 300
EvaluationPeriods: 1
Threshold: 0
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
UserSyncDurationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: front-user-sync-duration
AlarmDescription: front-user-sync approaching its 300s timeout (>90%)
Namespace: AWS/Lambda
MetricName: Duration
Dimensions:
- Name: FunctionName
Value: !Ref UserSyncFunction
Statistic: Maximum
Period: 300
EvaluationPeriods: 1
Threshold: 270000
ComparisonOperator: GreaterThanThreshold
TreatMissingData: notBreaching
AlarmActions:
- !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts
Outputs: Outputs:
SlaMonitorFunctionArn: SlaMonitorFunctionArn:
Description: Front SLA Monitor Lambda ARN Description: Front SLA Monitor Lambda ARN