From f7c44778b53f71c3e1fbdb0048d7150b330b9f73 Mon Sep 17 00:00:00 2001 From: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> Date: Mon, 29 Jun 2026 15:10:02 -0400 Subject: [PATCH] [#142] Fix payroll email: SES domain identity + send-as pin + failure alarm (#143) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The weekly-post pay-summary email to payroll failed with SES AccessDenied every Monday since v1.10.1: the role granted ses:SendEmail on identity/noreply@seahaven.com, but that address is not a verified SES identity — it is covered by the verified domain identity seahaven.com, which is what SES authorizes against. Grant the domain ARN instead. Pin the grant with a ses:FromAddress condition (= noreply@seahaven.com, the existing SES_SENDER) so the domain-wide identity can't be used to send-as any other @seahaven.com mailbox (BEC blast radius). Surfaced by /sh-security-review; matches the existing single-sender intent. Add a CloudWatch metric-filter alarm on the swallowed "Failed to send pay summary" log line -> site-alerts. The email send is wrapped in try/except so a delivery failure never increments the Lambda Errors metric; this is the only signal that surfaces a silent payroll failure. Closes #142 --- CHANGELOG.md | 9 ++++++ src/slack-bot/CHANGELOG.md | 9 ++++++ template.yaml | 56 ++++++++++++++++++++++++++++++++++++-- 3 files changed, 71 insertions(+), 3 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 1359743..a01eb10 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,15 @@ fine and still supported. --- +## v1.13.1 — June 29, 2026 + +**Fixed: the weekly pay summary email to payroll is sending again.** A +permissions change had quietly been blocking the Monday pay-summary email to +payroll since mid-June. The Slack pay report and the schedule post were never +affected — only the email to payroll. That's now fixed, the two missed weeks +were re-sent to payroll, and we've added an alert so a future email failure +can't slip by unnoticed. + ## v1.13.0 — June 27, 2026 **The two-week schedule post now sticks to the bottom of the channel.** It used diff --git a/src/slack-bot/CHANGELOG.md b/src/slack-bot/CHANGELOG.md index 1359743..a01eb10 100644 --- a/src/slack-bot/CHANGELOG.md +++ b/src/slack-bot/CHANGELOG.md @@ -10,6 +10,15 @@ fine and still supported. --- +## v1.13.1 — June 29, 2026 + +**Fixed: the weekly pay summary email to payroll is sending again.** A +permissions change had quietly been blocking the Monday pay-summary email to +payroll since mid-June. The Slack pay report and the schedule post were never +affected — only the email to payroll. That's now fixed, the two missed weeks +were re-sent to payroll, and we've added an alert so a future email failure +can't slip by unnoticed. + ## v1.13.0 — June 27, 2026 **The two-week schedule post now sticks to the bottom of the channel.** It used diff --git a/template.yaml b/template.yaml index e2532d4..9e3145d 100644 --- a/template.yaml +++ b/template.yaml @@ -166,14 +166,26 @@ Resources: Action: - ses:SendEmail Resource: - # Only the single sending identity (noreply@seahaven.com), not - # every identity in the account. - - !Sub "arn:aws:ses:${AWS::Region}:${AWS::AccountId}:identity/noreply@seahaven.com" + # The From address (noreply@seahaven.com) is NOT a standalone + # verified SES identity — it is covered by the verified DOMAIN + # identity seahaven.com, which is the identity SES authorizes + # SendEmail against. Granting identity/noreply@seahaven.com (a + # non-existent identity) caused AccessDenied; the domain ARN is + # the correct least-privilege resource. Scoped to this one domain, + # not identity/* (every identity in the account). + - !Sub "arn:aws:ses:${AWS::Region}:${AWS::AccountId}:identity/seahaven.com" # The sending identity has a default configuration set # (seahaven-email-events); SES authorizes SendEmail against the # config-set resource too, so it must be granted alongside the # identity or the send is denied. Scoped to the known set name. - !Sub "arn:aws:ses:${AWS::Region}:${AWS::AccountId}:configuration-set/seahaven-email-events" + # The domain identity covers every address under seahaven.com, so + # without this the role could send-as any @seahaven.com mailbox. + # Pin the From to the one legitimate sender (matches SES_SENDER) so + # a compromised function can't spoof internal senders (BEC). + Condition: + StringEquals: + ses:FromAddress: noreply@seahaven.com Events: # EST: 7am ET = 12:00 UTC (Nov-Mar) WeeklyPostEST: @@ -551,6 +563,44 @@ Resources: AlarmActions: - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + # --- Payroll email delivery failure (log-metric alarm) --- + # The pay-summary email send is wrapped in try/except so a delivery failure + # never aborts the schedule post — which means it does NOT surface on the + # Lambda Errors metric above (the handler exits "success"). This filter turns + # the "Failed to send pay summary" log line into a metric so a silent payroll + # delivery failure still pages site-alerts. DefaultValue 0 keeps the metric + # reporting on every run so the alarm sits in OK, not INSUFFICIENT_DATA. + PayrollEmailFailureMetricFilter: + Type: AWS::Logs::MetricFilter + Properties: + # !Ref the stack-managed log group (not a literal name) so CloudFormation + # orders this filter after the group exists. + LogGroupName: !Ref WeeklyPostLogGroup + FilterPattern: '"Failed to send pay summary"' + MetricTransformations: + - MetricName: PayrollEmailFailures + MetricNamespace: AfterHours/WeeklyPost + MetricValue: "1" + # DefaultValue 0 means non-matching runs emit 0, so the metric stays + # populated and the alarm rests in OK rather than INSUFFICIENT_DATA. + DefaultValue: 0 + + PayrollEmailFailureAlarm: + Type: AWS::CloudWatch::Alarm + Properties: + AlarmName: PayrollEmailFailure-afterhours-weekly-post + AlarmDescription: "Weekly-post failed to email the pay summary to payroll (SES send error)" + Namespace: AfterHours/WeeklyPost + MetricName: PayrollEmailFailures + Statistic: Sum + Period: 300 + EvaluationPeriods: 1 + Threshold: 1 + ComparisonOperator: GreaterThanOrEqualToThreshold + TreatMissingData: notBreaching + AlarmActions: + - !Sub "arn:aws:sns:${AWS::Region}:${AWS::AccountId}:site-alerts" + # --- Lambda Duration alarms (Maximum, ms; 2-of-3 evaluation) --- # Thresholds set to ~80% of each function's timeout, with 2-of-3 evaluation. # Timeouts: slack-bot/weekly-post/release-notifier = 30s (global default);