PR #248 disabled the SES send by blanking PAYROLL_RECIPIENTS and dropping the
ses:SendEmail grant. This removes the now-dead path: _send_pay_email and
_build_pay_email_html, the SES_SENDER and PAYROLL_RECIPIENTS env, and the
PayrollEmailFailure metric filter and alarm that only fired on that path.
Slack schedule post, pay-summary DM, and checkcomponents enqueue unchanged.
* feat(roster): add Bearer PUT/DELETE roster API
Identity hire needs to write Slack IDs onto roster rows without a stale
daily 3CX sync clearing them, using the existing HTTP client contract.
* fix(roster): strip Secrets Manager token whitespace
A file:// secret commonly includes a trailing newline, so compare_digest
must strip the cached value the same way it strips the Bearer header.
* feat(pay): send after-hours lines to paychex checkcomponents (PLAT-154)
* style(pay): drop trailing blank line in weekly post tests
* fix(pay): skip duplicate checkcomponents send on weekly-post retry
The weekly-post pay-summary email to payroll failed with SES AccessDenied
every Monday since v1.10.1: the role granted ses:SendEmail on
identity/noreply@seahaven.com, but that address is not a verified SES
identity — it is covered by the verified domain identity seahaven.com,
which is what SES authorizes against. Grant the domain ARN instead.
Pin the grant with a ses:FromAddress condition (= noreply@seahaven.com,
the existing SES_SENDER) so the domain-wide identity can't be used to
send-as any other @seahaven.com mailbox (BEC blast radius). Surfaced by
/sh-security-review; matches the existing single-sender intent.
Add a CloudWatch metric-filter alarm on the swallowed "Failed to send
pay summary" log line -> site-alerts. The email send is wrapped in
try/except so a delivery failure never increments the Lambda Errors
metric; this is the only signal that surfaces a silent payroll failure.
Closes#142
* Fix auth and race-condition flaws in shift commands
Four confirmed findings from the 2026-06-17 security sweep:
- register_user let any Slack user overwrite an extension already
bound to a different user (account takeover). Add a DynamoDB
ConditionExpression so a write only succeeds when the extension is
unclaimed or already this user's; raise ExtensionAlreadyRegistered
otherwise and surface a clear Slack message.
- The `rate` subcommand was routed without the is_admin flag, so any
user could set $0 pay rates. Gate _handle_rate on is_admin, matching
the admin-command guard.
- `/oncall pick` used a plain put_item (TOCTOU): two concurrent picks
both won. Use the atomic claim_open_shift conditional claim so the
loser gets an "already picked up" message.
- swap-accept overwrote a shift independently claimed after the swap
was initiated. Add reassign_if_held_by, a conditional write that only
applies the swap while the override is still the requester's (or on
the weekly fallback), and notify the accepter otherwise.
Add tests for the register-ownership guard and the rate admin guard.
Refs: INFRA
* Scope shift-manager Lambda IAM to least privilege
The nightly sweep flagged four over-broad permissions. Scope each to
only what the function actually reads (verified against source):
- WeeklyPost: secrets to slack-bot-token-* only (was the whole
afterhours-shift-manager/* namespace); SES SendEmail to the single
noreply@seahaven.com identity (was identity/*).
- RosterSync and RingScheduler: secrets to 3cx-* only (was the whole
namespace); both read only the 3cx domain/client-id/client-secret.
SlackBotFunction and HolidayRouter wildcards are left unchanged — out
of scope for this sweep.
Refs: INFRA
* fix: re-validate shift holder on swap-accept (sh-security-review RIHB-1)
reassign_if_held_by trusted 'no override row' as 'still the requester's',
but a weekly-held shift also has no override row. An admin clear or weekly
edit between swap-init and accept could move the shift to a third party
with no override, letting the accept steal it (CWE-367, confirmed HIGH).
Re-resolve the current holder at accept and abort if it is no longer the
requester. Adds regression test + seeds the holder in existing accept tests.
* fix: complete IAM least-privilege sweep (sh-security-review)
HolidayRouter secrets scope afterhours-shift-manager/* -> /3cx-* (reads
only 3cx secrets); RingScheduler DynamoDBCrudPolicy -> DynamoDBReadPolicy
(read-only at runtime). SlackBot wildcard left as-is (reads across all
sub-prefixes; verified defensible).
* Add CloudWatch alarm coverage for all functions, table, and HTTP API
Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage:
- Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan
adoption of slack-bot and weekly-post (live alarms of those exact names
already exist outside the stack and must be deleted before deploy).
- Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions.
- Throttles (Sum, >=1/5min) for all six functions.
- DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors
omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit
permanently in INSUFFICIENT_DATA.
- API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the
implicit ServerlessHttpApi.
All alarms page the shared site-alerts SNS topic, no OKActions,
TreatMissingData notBreaching. README updated with a Monitoring & Alarms section.
Duration and API latency thresholds pending sign-off.
* Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents
ThrottledRequests is not emitted at the TableName-only dimension (only
TableName+Operation), so the table-level alarm would sit permanently in
INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and
WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension.
* Correct DynamoDB alarm docs and drop sign-off wording
README DynamoDB section now lists the alarms actually shipped
(DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead
of the stale ThrottledRequests alarm. Thresholds are owner-approved, so
remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
The live afterhours-ring-scheduler Lambda had no error alarm. Add a
CloudWatch Errors alarm mirroring the existing HolidayRouterErrorAlarm:
AWS/Lambda Errors, Sum over one 5-min period, threshold >=1, missing
data notBreaching, paging the site-alerts SNS topic.
The orphaned Lambda-Errors-3cx-ring-group-scheduler alarm (pointing at
a renamed/absent function) is being removed separately.
Holiday day-shifts (08:00-17:00 ET) with N slots and 1.5x pay. A new
afterhours-holiday-router Lambda, fired by per-holiday EventBridge Scheduler
one-offs, repoints IVR 800 (key-0 + no-input/timeout) to holiday queue 802 and
sets 802's membership to the day's assignees (ext 100 fallback when unfilled),
reverting at 17:00. Pickups after a shift starts go through an admin Approve/Deny
flow for both regular and holiday shifts. Pay (weekly post + /oncall pay) shows
holiday rates distinctly.
Adds HOLIDAY and PICKUP_REQUEST DynamoDB record types, scheduler IAM scoped to
holiday-* schedules with conditioned PassRole, and the holiday-router function
with a 60-day log group and error alarm.
* Fix payroll email: grant SES config-set permission + isolate failures
The weekly pay-summary email to payroll has been failing with SES
AccessDenied since 2026-06-08. The sending identity (seahaven.com) gained
a default configuration set (seahaven-email-events), and SES authorizes
SendEmail against the config-set ARN as well as the identity — but the
WeeklyPostFunction role only granted ses:SendEmail on identity/*.
- template.yaml: add the configuration-set ARN (scoped to the known set
name) to the SES policy so sends are authorized again.
- weekly-post/app.py: wrap _send_pay_email in try/except so a delivery
failure can never abort the handler before the Slack schedule post.
Previously the SES error also blocked the two-week schedule post.
- Add a regression test covering the isolation.
Cross-family GPT-4.1 IAM review: APPROVE.
* Bump to v1.10.1 in CHANGELOG and sync App Home copy
The release job moved from release.yaml into deploy.yaml to clear
CodeQL's workflow_run findings, but the OIDC invoke role's trust still
pinned job_workflow_ref to release.yaml. That denied the AssumeRole at
the release job's Configure-AWS step, so the v1.10.0 announcement never
fired. Point the condition at deploy.yaml (the inline release job's
top-level workflow) so the token's job_workflow_ref matches.
* Add changelog-driven releases and App Home tab
Version the bot continuously from CHANGELOG.md (the single source of
truth for both the version and the staff-readable notes) and surface
changes to users in two ways:
- A new afterhours-release-notifier Lambda posts a "What's New" message
to the shift channel on minor/major releases (patches stay silent).
- The bot gains an App Home "About" tab showing what it does, the
command list, and the current version's notes.
release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN
events don't start downstream workflows), checks out the deployed commit,
and tags + publishes a GitHub Release + invokes the notifier. It assumes a
dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the
notifier; the account's cfn role gates role creation on that boundary.
The manual Version Bump workflow is retired. A CI guard enforces that a
CHANGELOG edit is a clean SemVer bump and that the in-package copy matches.
* Harden release workflow and regex against CodeQL findings
Address three code-scanning alerts on the PR:
- Critical (actions/untrusted-checkout): split release.yaml into a
read-only `prepare` job that checks out and runs repo code, and a
privileged `publish` job (contents:write + OIDC) that never checks out
repo code — it tags, releases, and invokes purely through the GitHub
and AWS APIs. Also assert head_branch == main.
- High x2 (py/polynomial-redos): rewrite the italic and link regexes in
markdown_to_mrkdwn with possessive quantifiers and exclusive character
classes so they run in linear time on adversarial input. Adds a
regression test.
* Move release/announce into Deploy workflow to clear CodeQL
The workflow_run-triggered release.yaml kept tripping CodeQL's
privileged-context rules (untrusted-checkout, then cache-poisoning) —
CodeQL distrusts any workflow_run that checks out a ref, regardless of
the main-only guarantee, and there is no autofix.
Fold the release job into deploy.yaml gated on `needs: deploy`. A
push-to-main run is a trusted context, so checking out and running repo
code with write/OIDC is safe there. This still gates on deploy success
and serializes via the deploy concurrency group, and removes the
separate workflow entirely.
Applies seahaven-lambda-execution-boundary to all SAM auto-generated
function execution roles via Globals.Function.PermissionsBoundary.
Required so the github-cfn-execution-role scope-down (INFRA-97) can
safely constrain role creation without blocking Lambda deploys.
No explicit AWS::IAM::Role resources exist in this template.
Refs: INFRA-103
Add AccessLogSettings on the implicit HTTP API stage pointing at a new
/aws/apigateway/afterhours-shift-manager log group with 90-day retention,
plus DefaultRouteSettings throttling (100 rps / 50 burst). Mirrors the
M-18 pattern landed on payments-dashboard.
/oncall swap no longer reassigns immediately. It now writes a pending SWAP
record and DMs the target Accept/Decline buttons; the shift only moves once
they accept.
- schedule.py: create_pending_swap / get_swap / mark_swap_verified /
clear_swap (PK=SWAP, date/shift SK mirroring OVERRIDE, status + timestamps
+ expires_at for TTL). A new request supersedes a prior pending one.
- app.py: _handle_swap creates the pending swap + DMs the target (requires the
target be Slack-linked; rejects self-swap). New module-level
handle_swap_accept / handle_swap_decline + two @app.action registrations.
Accept writes the override, repoints 3CX when it's the active shift, marks
the swap verified, notifies the channel + requester. Decline clears it and
DMs the requester. Lazy expiry: accept is rejected once the shift has started
(_shift_start/_shift_started).
- blocks.py: build_swap_request_blocks (Accept/Decline) + build_swap_resolved_blocks.
- template.yaml: enable DynamoDB TTL on expires_at so abandoned pending swaps
self-clean.
- tests: swap schedule methods, swap blocks, rewritten test_handle_swap
(pending + DM, no immediate override), new test_swap_accept_decline. 174 passed.
- README: swap behavior + SWAP item type + TTL.
The verified SWAP status is what #84 (24h drop guard) will query.
Closes#83
* Add arm64, log retention, and compliance fixes
- Set arm64 architecture globally for all Lambda functions
- Add explicit CloudWatch log groups with 60-day retention
- Add missing WeeklyPostFunctionArn to stack outputs
- Add Dependabot assignees for both ecosystems
- Add samconfig.toml.example for onboarding
* Restructure src/ to per-function layout with shared Layer
Move from flat src/ to per-function directories:
- src/slack-bot/ — Slack Bolt Lambda handler
- src/weekly-post/ — Monday schedule + pay post
- src/roster-sync/ — Daily 3CX roster sync
- src/shared/ — Lambda Layer with schedule, blocks, three_cx_client
Each function has its own requirements.txt and CodeUri. Shared
modules are deployed as a SAM Layer (afterhours-shared) importable
as `from shared.X import Y`.
* Migrate secrets from SSM Parameter Store to Secrets Manager
- Slack bot token and signing secret now read from Secrets Manager
- 3CX credentials (domain, client-id, client-secret) moved to
Secrets Manager under afterhours-shift-manager/3cx-* prefix
- Channel ID is now a non-secret CloudFormation parameter (ShiftChannel)
- Add shared secrets.py helper for Secrets Manager reads
- Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue
* Merge ring-scheduler-3cx as 4th Lambda function
- Add afterhours-ring-scheduler Lambda with 4 EventBridge rules
(daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group
routing updates
- Extract shared ring_scheduler.py module for direct ring group
updates from both the scheduled Lambda and the Slack bot
- Replace cross-Lambda invoke with direct update_ring_group() call
in the Slack bot — eliminates lambda:InvokeFunction dependency
- Use RingGroup API (correct) instead of Queue API (was wrong in
the original ring-scheduler repo)
- Eliminate YAML config fallback — DynamoDB is the sole schedule
source
- Add RingGroupNumber CloudFormation parameter
* Add schedule post live-update and old post deletion (#40, #41)
- Store schedule message timestamp in DynamoDB (SCHEDULE_POST record)
- Delete previous week's schedule post before posting the new one
- Live-update the schedule post via chat_update after any
pick/drop/swap/button-pickup so it always reflects current state
* Disallow past shifts and add day/night labels (#43, #42)
- Reject /oncall pick and /oncall drop for past dates
- Show ephemeral error when stale pickup buttons are clicked
- Hide pickup buttons for dates in the past
- Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to
schedule lines, pickup buttons, and shift change notifications
* Add admin slash commands for shift and roster management (#39)
- /oncall admin override <date> <ext> — assign a shift
- /oncall admin open <date> — mark shift as open
- /oncall admin clear <date> — remove override, revert to weekly
- /oncall admin roster add/remove/rename — manage roster entries
- Admin access gated by admin_users list in DynamoDB CONFIG
- Help message shows admin commands for admin users
* Update README for merged architecture and new features
* Switch from RingGroup API to Queue API at extension 801
The 3CX routing was changed from ring group 800 to queue 801 in a
previous PR on ring-scheduler-3cx. Updates all callers and the SAM
template parameter default accordingly.
* Pass SAM parameter overrides in deploy workflow
* Fix review findings: IAM, routing guards, past-date check, roster safety
- Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query)
- Button pickup: update 3CX for active shift type, not just night
- Pick/drop/swap commands: only update 3CX when shift type is active
- Swap command: add missing past-date guard
- add_roster_entry: reject if extension already exists
- Apply ruff formatting
* Add error handling to ring scheduler 3CX call
* Fix weekend day shift commands and admin 3CX routing
- Add _find_employee_shift() to check both day/night on weekends
- Drop/swap now correctly find and operate on weekend day shifts
- Pick finds first available shift type on weekends
- Admin override/open/clear update 3CX for same-day active shifts
* Fix dependabot directories and admin weekend shift handling
Dependabot now scans per-function requirement directories instead
of the repo root. Admin override/open/clear commands accept an
optional day/night parameter for weekend day shift management.
* Fix weekend day shift active window to 8am-5pm
Before midnight-8am on weekends incorrectly reported the day shift
as active when the previous night shift is still running.
* Show shift type label for both weekend shifts in notifications
Night shift notifications on weekends were missing the type label,
making them ambiguous. Also fix schedule post text fallback to use
this_monday instead of now for the start date.
* Extract determine_shift_type into shared layer
Eliminates duplicated weekend day/night boundary logic between
the ring scheduler and Slack bot Lambdas.
* Fix weekly schedule fallback start date
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
* Include weekend shift type in command confirmations
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
* Apply ruff formatting to app.py
* Only show day/night shift labels on weekends in schedule display
Weekday shifts are always night — the label was redundant clutter.
* Deduplicate 3CX forwarding payload and add shift type to pick command
Extract _update_forwarding helper in ThreeCXClient to share the
payload between queue and ring group methods. Add optional day/night
argument to /oncall pick so users can target a specific weekend shift.
* Consolidate WEEKEND_DAYS and fix weekday pickup button labels
Import WEEKEND_DAYS from shared.schedule instead of redefining in
blocks.py and weekly-post/app.py. Gate pickup button day/night
labels on weekends only, matching all other display surfaces.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
Slash command handlers (drop, pick, swap) were posting notifications to
command["channel_id"] — wherever the command was run. If someone ran
/oncall drop from a DM, the notification went there instead of the
schedule channel. Pickup buttons didn't have this problem because
body["channel"]["id"] is always the channel where the button lives.
Added SHIFT_CHANNEL_PARAM to the SlackBotFunction env vars, read it on
cold start, and route all slash command shift-change notifications to
the configured schedule channel.
Closes#21
- Pay summary sent as DM to designated user instead of channel
- Email shows only per-person totals, no shift breakdown
- Renamed email subject/header to "Bonus Pay Summary"
Sends an HTML pay summary email to configured payroll recipients
alongside the existing Slack post. Uses SES with noreply@seahaven.com
as sender. Recipients configured via PAYROLL_RECIPIENTS env var
(set to adam@seahaven.com for testing, switch to payroll@seahaven.com
for production).
Closes#3
Syncs DynamoDB roster from a configured 3CX group daily at 6am ET
(before the 7am schedule post). Includes ThreeCXClient for XAPI
authentication and group member queries. Preserves existing
slack_user_id links and guards against accidental roster wipes.
Closes#4
Slack Bolt app on Lambda for managing on-call shifts. Employees can
pick up, drop, and swap shifts via /oncall commands. Changes update
3CX ring group 800 routing in real time for same-day shifts.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>