From efa31783e26d64109e3c3db2ffee46544f3426fb Mon Sep 17 00:00:00 2001 From: Adam Moussa Date: Wed, 17 Jun 2026 13:53:30 -0400 Subject: [PATCH] Add CloudWatch alarm coverage for front-integrations Both Lambdas and the front-sla-alerts table previously had zero alarm coverage, so failures or runaway runs went unnoticed until someone checked logs. Wire a standard alarm set to the shared site-alerts SNS topic (ALARM-only, TreatMissingData notBreaching) per Wave 1 conventions. - Lambda Errors + Throttles alarms for front-sla-monitor and front-user-sync (Sum, threshold 0). - Lambda Duration alarms (Max, threshold 270000 = 90% of the shared 300s timeout) for both functions. - DynamoDB ThrottledRequests + SystemErrors alarms on front-sla-alerts. Document the alarm set in the README. --- README.md | 15 +++++ template.yaml | 160 ++++++++++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 175 insertions(+) diff --git a/README.md b/README.md index 85743ba..96802b2 100644 --- a/README.md +++ b/README.md @@ -60,6 +60,21 @@ Business time counts weekday hours only (Mon-Fri, Eastern time). Alerts are only - **DynamoDB:** `front-sla-alerts` — alert history per conversation + monitor state, 7-day TTL - **EventBridge:** SLA check every 15 min during business hours; user sync daily at 6 AM ET weekdays +## Monitoring + +CloudWatch alarms publish to the shared `site-alerts` SNS topic (`arn:aws:sns:us-east-1:328440206208:site-alerts`). All alarms are ALARM-only (no OK/recovery notification) and treat missing data as `notBreaching`. + +| Alarm | Metric | Trigger | +|---|---|---| +| `front-sla-monitor-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window | +| `front-sla-monitor-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window | +| `front-sla-monitor-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) | +| `front-user-sync-errors` | Lambda `Errors` (Sum) | Any invocation error in a 5-min window | +| `front-user-sync-throttles` | Lambda `Throttles` (Sum) | Any throttled invocation in a 5-min window | +| `front-user-sync-duration` | Lambda `Duration` (Max) | Run exceeds 270s (90% of the 300s timeout) | +| `front-sla-alerts-throttled-requests` | DynamoDB `ThrottledRequests` (Sum) | Any throttled request on the table | +| `front-sla-alerts-system-errors` | DynamoDB `SystemErrors` (Sum) | Any HTTP 500 from the table | + ## Secrets (Secrets Manager) | Secret | Purpose | diff --git a/template.yaml b/template.yaml index 09baba5..11810ca 100644 --- a/template.yaml +++ b/template.yaml @@ -115,6 +115,108 @@ Resources: Description: Check Front conversations for SLA breaches every 15 min during business hours Enabled: true + # SLA Monitor alarms + SlaMonitorErrorsAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-sla-monitor-errors + AlarmDescription: front-sla-monitor invocation errors + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref SlaMonitorFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + + SlaMonitorThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-sla-monitor-throttles + AlarmDescription: front-sla-monitor invocations throttled + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref SlaMonitorFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + + SlaMonitorDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-sla-monitor-duration + AlarmDescription: front-sla-monitor approaching its 300s timeout (>90%) + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref SlaMonitorFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 1 + Threshold: 270000 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + + # DynamoDB alarms — front-sla-alerts table + # ThrottledRequests and SystemErrors are documented at dimensions + # (TableName, Operation); CloudWatch also publishes the TableName-only + # aggregate, which is what we alarm on here. Neither has emitted yet + # (no throttle/5xx events to date), so the dimension was confirmed against + # the AWS DynamoDB metrics reference rather than live data. + AlertsTableThrottleAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-sla-alerts-throttled-requests + AlarmDescription: front-sla-alerts DynamoDB table is throttling requests + Namespace: AWS/DynamoDB + MetricName: ThrottledRequests + Dimensions: + - Name: TableName + Value: !Ref AlertsTable + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + + AlertsTableSystemErrorsAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-sla-alerts-system-errors + AlarmDescription: front-sla-alerts DynamoDB table returned HTTP 500 system errors + Namespace: AWS/DynamoDB + MetricName: SystemErrors + Dimensions: + - Name: TableName + Value: !Ref AlertsTable + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + # --------------------------------------------------------------------------- # Google User Sync # --------------------------------------------------------------------------- @@ -155,6 +257,64 @@ Resources: Description: Sync Google Workspace user profiles to Front daily at 6 AM ET Enabled: true + # User Sync alarms + UserSyncErrorsAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-user-sync-errors + AlarmDescription: front-user-sync invocation errors + Namespace: AWS/Lambda + MetricName: Errors + Dimensions: + - Name: FunctionName + Value: !Ref UserSyncFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + + UserSyncThrottlesAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-user-sync-throttles + AlarmDescription: front-user-sync invocations throttled + Namespace: AWS/Lambda + MetricName: Throttles + Dimensions: + - Name: FunctionName + Value: !Ref UserSyncFunction + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 0 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + + UserSyncDurationAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: front-user-sync-duration + AlarmDescription: front-user-sync approaching its 300s timeout (>90%) + Namespace: AWS/Lambda + MetricName: Duration + Dimensions: + - Name: FunctionName + Value: !Ref UserSyncFunction + Statistic: Maximum + Period: 300 + EvaluationPeriods: 1 + Threshold: 270000 + ComparisonOperator: GreaterThanThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts + Outputs: SlaMonitorFunctionArn: Description: Front SLA Monitor Lambda ARN