* Fix auth and race-condition flaws in shift commands
Four confirmed findings from the 2026-06-17 security sweep:
- register_user let any Slack user overwrite an extension already
bound to a different user (account takeover). Add a DynamoDB
ConditionExpression so a write only succeeds when the extension is
unclaimed or already this user's; raise ExtensionAlreadyRegistered
otherwise and surface a clear Slack message.
- The `rate` subcommand was routed without the is_admin flag, so any
user could set $0 pay rates. Gate _handle_rate on is_admin, matching
the admin-command guard.
- `/oncall pick` used a plain put_item (TOCTOU): two concurrent picks
both won. Use the atomic claim_open_shift conditional claim so the
loser gets an "already picked up" message.
- swap-accept overwrote a shift independently claimed after the swap
was initiated. Add reassign_if_held_by, a conditional write that only
applies the swap while the override is still the requester's (or on
the weekly fallback), and notify the accepter otherwise.
Add tests for the register-ownership guard and the rate admin guard.
Refs: INFRA
* Scope shift-manager Lambda IAM to least privilege
The nightly sweep flagged four over-broad permissions. Scope each to
only what the function actually reads (verified against source):
- WeeklyPost: secrets to slack-bot-token-* only (was the whole
afterhours-shift-manager/* namespace); SES SendEmail to the single
noreply@seahaven.com identity (was identity/*).
- RosterSync and RingScheduler: secrets to 3cx-* only (was the whole
namespace); both read only the 3cx domain/client-id/client-secret.
SlackBotFunction and HolidayRouter wildcards are left unchanged — out
of scope for this sweep.
Refs: INFRA
* fix: re-validate shift holder on swap-accept (sh-security-review RIHB-1)
reassign_if_held_by trusted 'no override row' as 'still the requester's',
but a weekly-held shift also has no override row. An admin clear or weekly
edit between swap-init and accept could move the shift to a third party
with no override, letting the accept steal it (CWE-367, confirmed HIGH).
Re-resolve the current holder at accept and abort if it is no longer the
requester. Adds regression test + seeds the holder in existing accept tests.
* fix: complete IAM least-privilege sweep (sh-security-review)
HolidayRouter secrets scope afterhours-shift-manager/* -> /3cx-* (reads
only 3cx secrets); RingScheduler DynamoDBCrudPolicy -> DynamoDBReadPolicy
(read-only at runtime). SlackBot wildcard left as-is (reads across all
sub-prefixes; verified defensible).
* Add CloudWatch alarm coverage for all functions, table, and HTTP API
Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage:
- Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan
adoption of slack-bot and weekly-post (live alarms of those exact names
already exist outside the stack and must be deleted before deploy).
- Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions.
- Throttles (Sum, >=1/5min) for all six functions.
- DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors
omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit
permanently in INSUFFICIENT_DATA.
- API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the
implicit ServerlessHttpApi.
All alarms page the shared site-alerts SNS topic, no OKActions,
TreatMissingData notBreaching. README updated with a Monitoring & Alarms section.
Duration and API latency thresholds pending sign-off.
* Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents
ThrottledRequests is not emitted at the TableName-only dimension (only
TableName+Operation), so the table-level alarm would sit permanently in
INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and
WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension.
* Correct DynamoDB alarm docs and drop sign-off wording
README DynamoDB section now lists the alarms actually shipped
(DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead
of the stale ThrottledRequests alarm. Thresholds are owner-approved, so
remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
The live afterhours-ring-scheduler Lambda had no error alarm. Add a
CloudWatch Errors alarm mirroring the existing HolidayRouterErrorAlarm:
AWS/Lambda Errors, Sum over one 5-min period, threshold >=1, missing
data notBreaching, paging the site-alerts SNS topic.
The orphaned Lambda-Errors-3cx-ring-group-scheduler alarm (pointing at
a renamed/absent function) is being removed separately.
get_ivr/set_ivr_routes/extract_ivr_routes targeted a nonexistent IVRs entity
set with an Options[].Route/TimeoutForward shape. The live 3CX IVR is the
Receptionists entity: the no-input/timeout route is the scalar TimeoutForwardDN,
and the key-0 route is a child of the Forwards collection (matched by Input=='0'),
written via a parent deep-PATCH. Routes are now destination numbers. Caught by the
live prod round-trip (get_ivr returned 405) before any holiday ran; rewritten and
re-verified against the live PBX. v1.11.1.
Holiday day-shifts (08:00-17:00 ET) with N slots and 1.5x pay. A new
afterhours-holiday-router Lambda, fired by per-holiday EventBridge Scheduler
one-offs, repoints IVR 800 (key-0 + no-input/timeout) to holiday queue 802 and
sets 802's membership to the day's assignees (ext 100 fallback when unfilled),
reverting at 17:00. Pickups after a shift starts go through an admin Approve/Deny
flow for both regular and holiday shifts. Pay (weekly post + /oncall pay) shows
holiday rates distinctly.
Adds HOLIDAY and PICKUP_REQUEST DynamoDB record types, scheduler IAM scoped to
holiday-* schedules with conditioned PassRole, and the holiday-router function
with a 60-day log group and error alarm.
* Fix payroll email: grant SES config-set permission + isolate failures
The weekly pay-summary email to payroll has been failing with SES
AccessDenied since 2026-06-08. The sending identity (seahaven.com) gained
a default configuration set (seahaven-email-events), and SES authorizes
SendEmail against the config-set ARN as well as the identity — but the
WeeklyPostFunction role only granted ses:SendEmail on identity/*.
- template.yaml: add the configuration-set ARN (scoped to the known set
name) to the SES policy so sends are authorized again.
- weekly-post/app.py: wrap _send_pay_email in try/except so a delivery
failure can never abort the handler before the Slack schedule post.
Previously the SES error also blocked the two-week schedule post.
- Add a regression test covering the isolation.
Cross-family GPT-4.1 IAM review: APPROVE.
* Bump to v1.10.1 in CHANGELOG and sync App Home copy
The release job moved from release.yaml into deploy.yaml to clear
CodeQL's workflow_run findings, but the OIDC invoke role's trust still
pinned job_workflow_ref to release.yaml. That denied the AssumeRole at
the release job's Configure-AWS step, so the v1.10.0 announcement never
fired. Point the condition at deploy.yaml (the inline release job's
top-level workflow) so the token's job_workflow_ref matches.
* Add changelog-driven releases and App Home tab
Version the bot continuously from CHANGELOG.md (the single source of
truth for both the version and the staff-readable notes) and surface
changes to users in two ways:
- A new afterhours-release-notifier Lambda posts a "What's New" message
to the shift channel on minor/major releases (patches stay silent).
- The bot gains an App Home "About" tab showing what it does, the
command list, and the current version's notes.
release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN
events don't start downstream workflows), checks out the deployed commit,
and tags + publishes a GitHub Release + invokes the notifier. It assumes a
dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the
notifier; the account's cfn role gates role creation on that boundary.
The manual Version Bump workflow is retired. A CI guard enforces that a
CHANGELOG edit is a clean SemVer bump and that the in-package copy matches.
* Harden release workflow and regex against CodeQL findings
Address three code-scanning alerts on the PR:
- Critical (actions/untrusted-checkout): split release.yaml into a
read-only `prepare` job that checks out and runs repo code, and a
privileged `publish` job (contents:write + OIDC) that never checks out
repo code — it tags, releases, and invokes purely through the GitHub
and AWS APIs. Also assert head_branch == main.
- High x2 (py/polynomial-redos): rewrite the italic and link regexes in
markdown_to_mrkdwn with possessive quantifiers and exclusive character
classes so they run in linear time on adversarial input. Adds a
regression test.
* Move release/announce into Deploy workflow to clear CodeQL
The workflow_run-triggered release.yaml kept tripping CodeQL's
privileged-context rules (untrusted-checkout, then cache-poisoning) —
CodeQL distrusts any workflow_run that checks out a ref, regardless of
the main-only guarantee, and there is no autofix.
Fold the release job into deploy.yaml gated on `needs: deploy`. A
push-to-main run is a trusted context, so checking out and running repo
code with write/OIDC is safe there. This still gates on deploy success
and serializes via the deploy concurrency group, and removes the
separate workflow entirely.
Applies seahaven-lambda-execution-boundary to all SAM auto-generated
function execution roles via Globals.Function.PermissionsBoundary.
Required so the github-cfn-execution-role scope-down (INFRA-97) can
safely constrain role creation without blocking Lambda deploys.
No explicit AWS::IAM::Role resources exist in this template.
Refs: INFRA-103
Add AccessLogSettings on the implicit HTTP API stage pointing at a new
/aws/apigateway/afterhours-shift-manager log group with 90-day retention,
plus DefaultRouteSettings throttling (100 rps / 50 burst). Mirrors the
M-18 pattern landed on payments-dashboard.
* Add dependency-review caller workflow
Add a pull_request-triggered caller that invokes the org-level
callable-dependency-review workflow to scan dependency changes and
fail on high-severity advisories.
* chore: retrigger checks
* chore: retrigger dep review (post-fix)
Human-readable history of the bot from launch (Apr 3) through v1.9.2, written
for a non-technical reader. Recent entries map to the new semantic version tags
(v1.7.19-v1.9.2); earlier work is grouped by dated milestone since those
releases were daily auto-versioned.
We now apply semantic version tags deliberately (see v1.7.19–v1.9.2), so the
daily auto-bump is no longer wanted.
- Drop the `schedule:` cron (and the now-unneeded DST guard) — the workflow
runs only on `workflow_dispatch`, with patch/minor/major options.
- Remove the Slack notification entirely: the "Update changelog canvas" and
"Post to Slack" steps (and the PR/bullet collection that fed them) are gone,
along with their SLACK_* secret usage.
- Keep the core behavior: compute the next version from the latest tag + chosen
bump and push an annotated tag.
Renames the workflow "Daily Version Bump" -> "Version Bump".
SETUP.md step 2 told operators to store the Slack token/signing secret in
SSM Parameter Store, which (a) violates the secrets-and-config handbook
(API tokens/signing values must live in Secrets Manager) and (b) contradicts
the IaC — the stack reads Secrets Manager and the IAM roles only grant
secretsmanager:GetSecretValue on afterhours-shift-manager/*, so following the
old instructions would break the deploy.
- Section 2 now uses `aws secretsmanager create-secret` for all five secrets
(slack-bot-token, slack-signing-secret, 3cx-domain/client-id/client-secret) —
the 3CX secrets were also previously undocumented.
- The Slack channel ID is not a secret; documented as the `ShiftChannel` deploy
parameter (--parameter-overrides) instead of an SSM SecureString.
Addresses violation 2 of #68.
A user can no longer `/oncall drop` a shift inside the 24h window before it
starts — inside that window coverage must be handed off via a verified swap
(target accepts) or opened by an admin.
- app.py: _within_drop_lock(date, shift_type) (24h before _shift_start);
guard in _handle_drop after the ownership check. Admin `open` is a separate
handler and is unaffected (bypasses the lock).
- Removed the now-unreachable 3CX-repoint-on-drop branch: a same-day shift is
always inside the lock, so a drop never reaches mark_open for today.
- Help text + README note the 24h rule.
- tests: rewritten test_handle_drop (outside/inside-24h per shift type,
weekend day, admin bypass, plus the existing guard-precedence cases) and
direct _shift_start/_shift_started/_within_drop_lock helper tests. 185 passed.
Closes#84
/oncall swap no longer reassigns immediately. It now writes a pending SWAP
record and DMs the target Accept/Decline buttons; the shift only moves once
they accept.
- schedule.py: create_pending_swap / get_swap / mark_swap_verified /
clear_swap (PK=SWAP, date/shift SK mirroring OVERRIDE, status + timestamps
+ expires_at for TTL). A new request supersedes a prior pending one.
- app.py: _handle_swap creates the pending swap + DMs the target (requires the
target be Slack-linked; rejects self-swap). New module-level
handle_swap_accept / handle_swap_decline + two @app.action registrations.
Accept writes the override, repoints 3CX when it's the active shift, marks
the swap verified, notifies the channel + requester. Decline clears it and
DMs the requester. Lazy expiry: accept is rejected once the shift has started
(_shift_start/_shift_started).
- blocks.py: build_swap_request_blocks (Accept/Decline) + build_swap_resolved_blocks.
- template.yaml: enable DynamoDB TTL on expires_at so abandoned pending swaps
self-clean.
- tests: swap schedule methods, swap blocks, rewritten test_handle_swap
(pending + DM, no immediate override), new test_swap_accept_decline. 174 passed.
- README: swap behavior + SWAP item type + TTL.
The verified SWAP status is what #84 (24h drop guard) will query.
Closes#83
* Add pytest suite and wire it into CI
Stands up the first automated tests for the repo (151 tests) and turns on
the CI test step.
- Lift slack-bot handlers out of create_app() closures to module level so
they're unit-testable; create_app is now a thin Bolt-wiring layer. No
behavior change (handler entrypoints and create_app signature unchanged).
- tests/ mirrors src/: shared layer (schedule, blocks, 3CX client,
ring_scheduler, secrets) + all four Lambdas (pay math, drop/swap/pick/
admin/register/rate, pickup button, roster sync, queue scheduler).
- All boundaries mocked: DynamoDB/SES/Secrets via moto, 3CX HTTP via
responses, Slack via fakes, time via freezegun. No real network/AWS.
- pyproject.toml pytest config (pythonpath=src/shared, importlib mode);
per-package conftest loads each app.py under a unique name to avoid the
four-app.py collision. tests/requirements.txt for test-only deps.
- ci.yaml: run-tests: true (reusable workflow auto-installs deps) and lint
the tests dir too.
- README Testing section.
Closes#85
* Add least-privilege permissions block to CI workflow
Resolves the CodeQL actions/missing-workflow-permissions alert: the CI
workflow now restricts GITHUB_TOKEN to contents: read (it only checks out,
lints, and runs tests).
* Stop logging extension numbers in 3CX queue updates
Resolves 3 high CodeQL py/clear-text-logging-sensitive-data alerts: the
queue/ring-group forwarding logs no longer include the routed extension
values (closed/holiday/extension). Non-sensitive context (resource id,
queue number) is retained.