* fix: delete+repost schedule on weekly rollover for bottom placement
The Monday rollover was chat_update-ing in place, which only refreshes
content without moving the message to the bottom. Now it chat_deletes
the old post and chat_postMessages a fresh one so the schedule lands
at the bottom every Monday, independent of in-week activity.
Also added info-level logging to the bump handler silent return paths
so skipped bumps are observable at runtime.
* fix: roll back weekly repost when its ts can't be persisted
The Monday rollover deletes the old post then reposts a fresh one, but only
saved the new ts as its last step. If the save failed (or the Lambda died)
after the post landed, the async retry would read the stale, already-deleted
ts, no-op its delete, and post a second schedule — orphaning the first at the
bottom of the channel.
Wrap the save so a failure after a successful repost best-effort deletes the
fresh message before re-raising, letting the retry start clean. Mirrors the
orphan-avoidance the activity bump already has.
---------
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
* Add Slack admin modals + App Home admin section
Replace the two most error-prone positional admin commands with Block Kit
modals (override and holiday-add) opened from a new App Home admin section,
while keeping the typed subcommands as a fallback. Validation and side
effects are factored into shared helpers so the modal and command paths
can't drift, and every action/view handler re-checks is_admin against
get_admin_users() so a modal opened from Home can't bypass authorization.
Adds Schedule.list_overrides for the upcoming-overrides overview.
Refs: #136
* Update changelog date to July 02, 2026
* Add point-and-click admin actions in Slack for easier overrides and holidays
* [#136] Add admin UI evaluation spike doc (#140)
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
The weekly-post pay-summary email to payroll failed with SES AccessDenied
every Monday since v1.10.1: the role granted ses:SendEmail on
identity/noreply@seahaven.com, but that address is not a verified SES
identity — it is covered by the verified domain identity seahaven.com,
which is what SES authorizes against. Grant the domain ARN instead.
Pin the grant with a ses:FromAddress condition (= noreply@seahaven.com,
the existing SES_SENDER) so the domain-wide identity can't be used to
send-as any other @seahaven.com mailbox (BEC blast radius). Surfaced by
/sh-security-review; matches the existing single-sender intent.
Add a CloudWatch metric-filter alarm on the swallowed "Failed to send
pay summary" log line -> site-alerts. The email send is wrapped in
try/except so a delivery failure never increments the Lambda Errors
metric; this is the only signal that surfaces a silent payroll failure.
Closes#142
* Expand /oncall date parser to accept more formats
Users entering everyday forms like 7/3/26 hit a generic parse
failure because the parser only accepted four-digit years and a
bare m/d. Add two-digit-year and month-name (with optional
ordinal/year) formats, treating explicit-year inputs as fixed and
keeping the year-less roll-forward for bare m/d. Update the help
text and per-command parse hints to match.
Refs: #133
* Let drop pick a shift and flag night rows
Drop now accepts an optional [day|night|holiday] qualifier and, when a
date carries more than one shift the user holds, asks which to drop
instead of silently releasing the holiday or weekend day shift. The
static post also tags day/night rows with distinct glyphs and labels on
weekdays, so the after-hours row is unmistakable.
Refs: #134
* Refine night-row labeling and drop notifications
Weekday rows are night-only, so the moon glyph alone marks the
after-hours shift; the verbose time-range label is kept only on
weekend rows where day and night shifts coexist. Channel shift-change
notifications now label the shift on every day for parity. Thread the
caller's user_id through the regular-drop path instead of re-reading it
off the employee record, and document the user-facing changes.
Refs: #134
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
* Expand /oncall date parser to accept more formats
Users entering everyday forms like 7/3/26 hit a generic parse
failure because the parser only accepted four-digit years and a
bare m/d. Add two-digit-year and month-name (with optional
ordinal/year) formats, treating explicit-year inputs as fixed and
keeping the year-less roll-forward for bare m/d. Update the help
text and per-command parse hints to match.
Refs: #133
* Add dash 4-digit year format and pin year boundary
Dash inputs like 7-3-2026 previously returned None because only the
slash variant had a 4-digit-year format. Add %m-%d-%Y so dash and slash
behave alike, and pin the two-digit-year century boundary with a test.
Refs: #133
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
* Fix auth and race-condition flaws in shift commands
Four confirmed findings from the 2026-06-17 security sweep:
- register_user let any Slack user overwrite an extension already
bound to a different user (account takeover). Add a DynamoDB
ConditionExpression so a write only succeeds when the extension is
unclaimed or already this user's; raise ExtensionAlreadyRegistered
otherwise and surface a clear Slack message.
- The `rate` subcommand was routed without the is_admin flag, so any
user could set $0 pay rates. Gate _handle_rate on is_admin, matching
the admin-command guard.
- `/oncall pick` used a plain put_item (TOCTOU): two concurrent picks
both won. Use the atomic claim_open_shift conditional claim so the
loser gets an "already picked up" message.
- swap-accept overwrote a shift independently claimed after the swap
was initiated. Add reassign_if_held_by, a conditional write that only
applies the swap while the override is still the requester's (or on
the weekly fallback), and notify the accepter otherwise.
Add tests for the register-ownership guard and the rate admin guard.
Refs: INFRA
* Scope shift-manager Lambda IAM to least privilege
The nightly sweep flagged four over-broad permissions. Scope each to
only what the function actually reads (verified against source):
- WeeklyPost: secrets to slack-bot-token-* only (was the whole
afterhours-shift-manager/* namespace); SES SendEmail to the single
noreply@seahaven.com identity (was identity/*).
- RosterSync and RingScheduler: secrets to 3cx-* only (was the whole
namespace); both read only the 3cx domain/client-id/client-secret.
SlackBotFunction and HolidayRouter wildcards are left unchanged — out
of scope for this sweep.
Refs: INFRA
* fix: re-validate shift holder on swap-accept (sh-security-review RIHB-1)
reassign_if_held_by trusted 'no override row' as 'still the requester's',
but a weekly-held shift also has no override row. An admin clear or weekly
edit between swap-init and accept could move the shift to a third party
with no override, letting the accept steal it (CWE-367, confirmed HIGH).
Re-resolve the current holder at accept and abort if it is no longer the
requester. Adds regression test + seeds the holder in existing accept tests.
* fix: complete IAM least-privilege sweep (sh-security-review)
HolidayRouter secrets scope afterhours-shift-manager/* -> /3cx-* (reads
only 3cx secrets); RingScheduler DynamoDBCrudPolicy -> DynamoDBReadPolicy
(read-only at runtime). SlackBot wildcard left as-is (reads across all
sub-prefixes; verified defensible).
* Add CloudWatch alarm coverage for all functions, table, and HTTP API
Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage:
- Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan
adoption of slack-bot and weekly-post (live alarms of those exact names
already exist outside the stack and must be deleted before deploy).
- Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions.
- Throttles (Sum, >=1/5min) for all six functions.
- DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors
omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit
permanently in INSUFFICIENT_DATA.
- API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the
implicit ServerlessHttpApi.
All alarms page the shared site-alerts SNS topic, no OKActions,
TreatMissingData notBreaching. README updated with a Monitoring & Alarms section.
Duration and API latency thresholds pending sign-off.
* Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents
ThrottledRequests is not emitted at the TableName-only dimension (only
TableName+Operation), so the table-level alarm would sit permanently in
INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and
WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension.
* Correct DynamoDB alarm docs and drop sign-off wording
README DynamoDB section now lists the alarms actually shipped
(DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead
of the stale ThrottledRequests alarm. Thresholds are owner-approved, so
remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
The live afterhours-ring-scheduler Lambda had no error alarm. Add a
CloudWatch Errors alarm mirroring the existing HolidayRouterErrorAlarm:
AWS/Lambda Errors, Sum over one 5-min period, threshold >=1, missing
data notBreaching, paging the site-alerts SNS topic.
The orphaned Lambda-Errors-3cx-ring-group-scheduler alarm (pointing at
a renamed/absent function) is being removed separately.
get_ivr/set_ivr_routes/extract_ivr_routes targeted a nonexistent IVRs entity
set with an Options[].Route/TimeoutForward shape. The live 3CX IVR is the
Receptionists entity: the no-input/timeout route is the scalar TimeoutForwardDN,
and the key-0 route is a child of the Forwards collection (matched by Input=='0'),
written via a parent deep-PATCH. Routes are now destination numbers. Caught by the
live prod round-trip (get_ivr returned 405) before any holiday ran; rewritten and
re-verified against the live PBX. v1.11.1.
Holiday day-shifts (08:00-17:00 ET) with N slots and 1.5x pay. A new
afterhours-holiday-router Lambda, fired by per-holiday EventBridge Scheduler
one-offs, repoints IVR 800 (key-0 + no-input/timeout) to holiday queue 802 and
sets 802's membership to the day's assignees (ext 100 fallback when unfilled),
reverting at 17:00. Pickups after a shift starts go through an admin Approve/Deny
flow for both regular and holiday shifts. Pay (weekly post + /oncall pay) shows
holiday rates distinctly.
Adds HOLIDAY and PICKUP_REQUEST DynamoDB record types, scheduler IAM scoped to
holiday-* schedules with conditioned PassRole, and the holiday-router function
with a 60-day log group and error alarm.