mirror of
https://github.com/Sea-Haven-Industries/afterhours-shift-manager.git
synced 2026-09-30 09:03:11 +00:00
* feat(roster): add Bearer PUT/DELETE roster API Identity hire needs to write Slack IDs onto roster rows without a stale daily 3CX sync clearing them, using the existing HTTP client contract. * fix(roster): strip Secrets Manager token whitespace A file:// secret commonly includes a trailing newline, so compare_digest must strip the cached value the same way it strips the Bearer header.
323 lines
19 KiB
Markdown
323 lines
19 KiB
Markdown
# After-Hours Shift Manager
|
|
|
|

|
|

|
|

|
|

|
|
|
|
Slack bot for managing after-hours on-call shifts at Sea Haven Industries. Employees can pick up, drop, and swap shifts directly from Slack. Changes automatically update 3CX queue routing via the integrated ring scheduler.
|
|
|
|
## How It Works
|
|
|
|
A recurring weekly schedule assigns employees to after-hours phone duty. Weekend shifts are split into Day (8am-5pm) and Night (5pm-8am). Any unassigned shift shows as **Available** in Slack with a pickup button. When someone picks up or drops a shift for today, the 3CX queue is updated immediately. Future changes take effect when the ring scheduler runs at 8am daily and 5pm on weekends.
|
|
|
|
The weekly schedule post is updated live when shifts change, and the previous week's post is automatically deleted when the new one goes out.
|
|
|
|
The bot also has an **About** page: open the bot in Slack and click its **Home** tab to see what it does, the full command list, and the latest "What's New" (see [Releases & Versioning](#releases--versioning)).
|
|
|
|
## Slack Commands
|
|
|
|
| Command | Description |
|
|
|---|---|
|
|
| `/oncall` | Show this week's schedule |
|
|
| `/oncall next` | Show next week's schedule |
|
|
| `/oncall pick <date>` | Pick up an available shift — instant if the shift hasn't started; if it's already underway (but not ended) it needs admin approval (see [Late-pickup approval](#late-pickup-approval)) |
|
|
| `/oncall drop <date>` | Drop your shift (marks it available) — blocked within 24h of shift start; swap or ask an admin instead |
|
|
| `/oncall swap <date> @person` | Request a swap — the other person gets an Accept/Decline DM and the shift only moves once they accept |
|
|
| `/oncall register <ext>` | Link your Slack account to your phone extension |
|
|
| `/oncall roster` | Show all employees and their link status |
|
|
| `/oncall pay` | Show last week's bonus pay summary |
|
|
| `/oncall rate` | Show current shift pay rates |
|
|
| `/oncall rate default <amount>` | Set the default per-shift rate |
|
|
| `/oncall rate <ext> <amount>` | Set a per-person shift rate |
|
|
| `/oncall help` | Show help |
|
|
|
|
### Admin Commands
|
|
|
|
Available to users listed in `admin_users` in the CONFIG record:
|
|
|
|
| Command | Description |
|
|
|---|---|
|
|
| `/oncall admin override <date> <ext>` | Assign a shift to an extension |
|
|
| `/oncall admin open <date>` | Mark a shift as open |
|
|
| `/oncall admin clear <date>` | Remove override (revert to weekly) |
|
|
| `/oncall admin roster add <ext> <name>` | Add an employee to the roster |
|
|
| `/oncall admin roster remove <ext>` | Remove an employee |
|
|
| `/oncall admin roster rename <ext> <name>` | Rename an employee |
|
|
| `/oncall admin holiday add <date> <slots> [x<mult>] <label>` | Schedule a holiday day shift (8am-5pm ET) with N slots, optional pay multiplier override (e.g. `x2`), and a label |
|
|
| `/oncall admin holiday remove <date>` | Remove a scheduled holiday and its activate/deactivate schedules |
|
|
| `/oncall admin holiday list` | List today-and-future scheduled holidays |
|
|
|
|
Dates accept: `today`, `tomorrow`, `monday`-`sunday`, `4/5`, `2026-04-05`
|
|
|
|
Example: `/oncall admin holiday add 2026-07-04 2 x2 Independence Day` schedules a
|
|
2-slot holiday paying 2x. Omitting the `x<mult>` token uses the default
|
|
`holiday_multiplier` from CONFIG (1.5). See [Holidays](#holidays) for the full flow.
|
|
|
|
## Architecture
|
|
|
|
- **Runtime**: Python 3.12 on AWS Lambda (arm64)
|
|
- **Data**: DynamoDB single-table (`afterhours-shifts`)
|
|
- **IaC**: AWS SAM (`template.yaml`) with shared Lambda Layer
|
|
- **Slack**: Slack Bolt framework with `/oncall` slash command
|
|
- **3CX Integration**: Queue routing updated directly via 3CX Queue XAPI
|
|
- **Secrets**: AWS Secrets Manager (`afterhours-shift-manager/*`)
|
|
|
|
### Lambda Functions
|
|
|
|
| Function | Trigger | Purpose |
|
|
|---|---|---|
|
|
| `afterhours-shift-manager` | API Gateway (POST /slack/events) | Slack bot — handles `/oncall` commands and interactive buttons |
|
|
| `afterhours-weekly-post` | EventBridge (Monday 7am ET) | Posts weekly schedule to Slack, sends pay report email |
|
|
| `afterhours-roster-sync` | EventBridge (daily 6am ET) | Syncs employee roster from 3CX |
|
|
| `afterhours-roster-api` | API Gateway (PUT /roster, DELETE /roster/{extension}) | Bearer-authenticated roster upsert/delete for the identity processor |
|
|
| `afterhours-ring-scheduler` | EventBridge (daily 8am ET + weekend 5pm ET) | Updates 3CX queue routing based on who's on shift |
|
|
| `afterhours-holiday-router` | EventBridge Scheduler (per-holiday one-off: 8am activate / 5pm deactivate ET) | Repoints the IVR to the holiday queue and sets queue agents for a holiday day shift; reverts at 5pm (see [Holidays](#holidays)) |
|
|
| `afterhours-release-notifier` | Invoked by the Deploy workflow's release job on minor/major releases | Posts a "What's New" announcement to the shift channel |
|
|
|
|
### Project Layout
|
|
|
|
```
|
|
src/
|
|
slack-bot/ Slack Bolt Lambda (handler + app); ships CHANGELOG.md for App Home
|
|
weekly-post/ Monday schedule + pay post
|
|
roster-sync/ Daily 3CX roster sync
|
|
roster-api/ HTTP PUT/DELETE /roster for identity hire/offboard
|
|
ring-scheduler/ 3CX queue routing updates
|
|
holiday-router/ 3CX IVR/queue repoint for holiday day shifts (activate/deactivate)
|
|
release-notifier/ Posts release announcements to Slack
|
|
shared/ Lambda Layer (schedule, blocks, changelog, 3CX client, secrets)
|
|
scripts/ changelog CLI + CI guard + in-package copy sync
|
|
tests/ pytest suite (mirrors src/, one dir per Lambda + shared)
|
|
```
|
|
|
|
### DynamoDB Schema
|
|
|
|
Single table with `PK` / `SK` keys:
|
|
|
|
| PK | SK | Description |
|
|
|---|---|---|
|
|
| `ROSTER` | `<extension>` | Employee: name, extension, slack_user_id |
|
|
| `WEEKLY` | `<DayName>` | Default weekly schedule: extension, name |
|
|
| `OVERRIDE` | `<YYYY-MM-DD>` | Date override from pickup/drop (or `OPEN`) |
|
|
| `SWAP` | `<YYYY-MM-DD>` | Pending/verified swap request: requester, target, status, `expires_at` (TTL) |
|
|
| `HOLIDAY` | `<YYYY-MM-DD>` | Holiday day shift (one per date): `slots` (int), `assignees` (MAP keyed by extension — `{"114": {name, claimed_at}}`), `multiplier` (Decimal, defaults to `CONFIG.holiday_multiplier` = 1.5, overridable per holiday), `label`, `created_at`, `created_by`, `activated` (bool), `schedule_names` (list) |
|
|
| `PICKUP_REQUEST` | `<YYYY-MM-DD>[-DAY]#<ext>` | Pending late-pickup awaiting admin approval: `requester_ext`, `requester_name`, `requester_slack`, `shift_type`, `is_holiday`, `status`, `created_at`, `expires_at` (TTL = shift end) |
|
|
| `SCHEDULE_POST` | `<channel_id>` | Current schedule message timestamp |
|
|
| `PAY` | `<YYYY-MM-DD>` | Weekly pay record (Monday date key) |
|
|
| `CONFIG` | `CONFIG` | Settings: shift_rate, fallback_extension, admin_users, `ring_group`, `holiday_multiplier` (default holiday pay multiplier, 1.5), `holiday_queue` (3CX queue repointed during holidays, default 802), `ivr_number` (3CX IVR repointed during holidays, default 800), `captured_ivr_routes` (original IVR routes saved at holiday activation, restored at deactivation) |
|
|
|
|
Weekend day-shift rows use a `-DAY` suffix on the SK (e.g. `OVERRIDE` / `2026-04-05-DAY`). The table has TTL enabled on `expires_at` so abandoned pending swaps and pickup requests self-clean.
|
|
|
|
Shift priority for any date is **HOLIDAY > OVERRIDE > WEEKLY** — a holiday record wins over a regular override, which wins over the standing weekly schedule.
|
|
|
|
Holiday slot claims are atomic Map updates so concurrent pickers can't oversubscribe:
|
|
|
|
- **Claim** — `SET assignees.#ext` guarded by `attribute_not_exists(assignees.#ext) AND size(assignees) < :slots`.
|
|
- **Release** — `REMOVE assignees.#ext` guarded by `attribute_exists`.
|
|
- **Swap** — `REMOVE #from SET #to` guarded by `attribute_exists(#from) AND attribute_not_exists(#to)`.
|
|
|
|
**Swap flow:** `/oncall swap` writes a `pending` `SWAP` record and DMs the target Accept/Decline buttons; it does **not** reassign the shift. On Accept, the override is written, 3CX is repointed if it's the active shift, and the record is marked `verified`. On Decline (or once the shift has started) the request is dropped and the shift stays with the original owner.
|
|
|
|
### Holidays
|
|
|
|
A **holiday** is a single day-only shift (08:00-17:00 ET), one `HOLIDAY` record per
|
|
date, that can hold multiple people (`slots`). Admins manage holidays with
|
|
`/oncall admin holiday add|remove|list`. A holiday takes priority over a regular
|
|
override and the weekly schedule for that date, and pays at its `multiplier`
|
|
(per-holiday override, else `CONFIG.holiday_multiplier`, default 1.5). Open slots
|
|
show in the schedule with a pickup button; claims, releases, and swaps are atomic
|
|
Map updates on the record (see above) so the slot count can't be oversubscribed.
|
|
|
|
**Holiday-router + Scheduler flow.** When an admin adds a holiday, the slack-bot
|
|
creates two **one-off EventBridge Scheduler** schedules for that date —
|
|
`holiday-activate-<YYYYMMDD>` at 08:00 ET and `holiday-deactivate-<YYYYMMDD>` at
|
|
17:00 ET — whose names are stored on the record's `schedule_names`. Scheduler
|
|
assumes `HolidaySchedulerExecutionRole` to invoke `afterhours-holiday-router`:
|
|
|
|
- **Activate (08:00):** capture both IVR `ivr_number` (800) routes — key-0 **and**
|
|
no-input/timeout — into `CONFIG.captured_ivr_routes` (skipped if they already
|
|
point at the holiday queue, so re-runs don't clobber the originals), set
|
|
`holiday_queue` (802) agents to the holiday's assignees (or `[fallback_extension]`
|
|
= `[100]` when no slots are filled — set **once** here), repoint **both** IVR 800
|
|
routes to queue 802, and mark `activated = True`. Idempotent.
|
|
- **Deactivate (17:00):** restore both IVR routes from `captured_ivr_routes` (only
|
|
routes still pointing at the queue, defensive against manual changes), clear them,
|
|
empty queue 802's agents, and mark `activated = False`. Idempotent.
|
|
|
|
Queue 801 (the daily ring-scheduler queue) is left untouched. If an admin adds a
|
|
holiday whose 08:00-17:00 window is already open, the slack-bot **inline-activates**
|
|
it immediately (invoking the router) rather than waiting for the 08:00 schedule.
|
|
Removing a holiday deletes the record and any outstanding schedules.
|
|
|
|
### Late-pickup approval
|
|
|
|
Picking up a shift behaves differently depending on timing, for **both** regular
|
|
and holiday shifts:
|
|
|
|
- **Before the shift starts** — immediate pickup (the prior behaviour, unchanged).
|
|
- **After the shift has started but before it ends** (08:00 for a day/holiday shift,
|
|
17:00 for a night shift) — the shift is **not** claimed yet. The bot writes a
|
|
`pending` `PICKUP_REQUEST` and DMs **every admin** Approve/Deny buttons (mirroring
|
|
the verified-swap flow). The **first admin to approve wins** (the claim is
|
|
conditional, so a second approval is a safe no-op). On approve, the shift is
|
|
claimed (regular: `set_override`; holiday: atomic `claim_holiday_slot`), the
|
|
requester and channel are notified, and 3CX is fired if the window is live —
|
|
regular shifts call `_update_3cx_routing(picker)` when it's today's active shift;
|
|
holidays refresh queue 802's agents to the current assignees. On deny, the request
|
|
is cleared and the requester is told.
|
|
- **After the shift has ended** — rejected outright; it's too late to pick up.
|
|
|
|
A slot claimed after the shift has started always needs an admin to approve it.
|
|
|
|
### Secrets Manager
|
|
|
|
| Secret | Description |
|
|
|---|---|
|
|
| `afterhours-shift-manager/slack-bot-token` | Slack bot OAuth token (`xoxb-...`) |
|
|
| `afterhours-shift-manager/slack-signing-secret` | Slack app signing secret |
|
|
| `afterhours-shift-manager/3cx-domain` | 3CX FQDN (e.g. `company.3cx.us`) |
|
|
| `afterhours-shift-manager/3cx-client-id` | 3CX OAuth2 client ID |
|
|
| `afterhours-shift-manager/3cx-client-secret` | 3CX OAuth2 client secret |
|
|
| `afterhours-shift-manager/roster-api-token` | Bearer token for PUT/DELETE `/roster`. Duplicate the same value into the seahaven-prod secret `paychex-integrations/afterhours-roster-token`. |
|
|
|
|
### Roster HTTP API
|
|
|
|
Identity hire/offboard in `paychex-integrations` calls this API. It is a separate Lambda on the same implicit HTTP API as Slack (`POST /slack/events` is unchanged).
|
|
|
|
| Method | Path | Body | Success |
|
|
|---|---|---|---|
|
|
| PUT | `/roster` | `{"name","extension","slack_user_id"}` (all required strings) | 200 `{"ok":true}` |
|
|
| DELETE | `/roster/{extension}` | none | 204 empty body, including when the row is already gone |
|
|
|
|
Header: `Authorization: Bearer {token}`. Missing or wrong token is 401. Invalid JSON or fields is 400. A secret-read failure is 503.
|
|
|
|
Set processor `AFTERHOURS_BASE_URL` to the stack output `AfterhoursApiBaseUrl`:
|
|
|
|
```
|
|
https://${ServerlessHttpApi}.execute-api.us-east-1.amazonaws.com
|
|
```
|
|
|
|
That value is the API origin only. Do not append `/mgmt` or `/roster`.
|
|
|
|
Daily `afterhours-roster-sync` still owns the 3CX `DEFAULT` group at 6am ET: rows absent from that group are deleted. Hire is safe because 3CX create (into `DEFAULT`) happens before the roster PUT. An HTTP-only row that is not in that group will be removed on the next sync. Sync preserves `slack_user_id` on existing rows and does not overwrite a just-created API row's Slack id.
|
|
|
|
Token rotation is a maintenance-window action. The Lambda caches the token per execution environment. Update the mgmt secret and the prod copy together, then recycle `afterhours-roster-api`. Updating only one copy, or recycling environments out of order, causes 401s until both sides match.
|
|
|
|
## Documentation
|
|
|
|
The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This project's `afterhours-shift-manager` stack is represented there as a Mermaid subgraph.
|
|
|
|
- **[AWS Architecture Map](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098)** (Confluence, IT space, page 1540098)
|
|
|
|
## Deployment
|
|
|
|
Merges to `main` are automatically deployed via **GitHub Actions** using reusable SAM workflows from the Sea Haven org.
|
|
|
|
For manual deploys:
|
|
|
|
```bash
|
|
sam build
|
|
sam deploy
|
|
```
|
|
|
|
## Monitoring & Alarms
|
|
|
|
All CloudWatch alarms are defined in `template.yaml` and notify the shared
|
|
`site-alerts` SNS topic (→ AWS Chatbot → Slack). None set `OKActions` — recovery
|
|
is not paged. Alarm names follow the in-template convention `Lambda-<Metric>-<fn>`
|
|
(e.g. `Lambda-Errors-afterhours-ring-scheduler`).
|
|
|
|
**Lambda alarms** (all seven functions: `afterhours-shift-manager`,
|
|
`afterhours-weekly-post`, `afterhours-roster-sync`, `afterhours-roster-api`,
|
|
`afterhours-ring-scheduler`, `afterhours-holiday-router`,
|
|
`afterhours-release-notifier`):
|
|
|
|
| Alarm | Metric | Condition | Notes |
|
|
|---|---|---|---|
|
|
| `Lambda-Errors-<fn>` | `Errors` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
| `Lambda-Duration-<fn>` | `Duration` (Maximum, ms) | `>= ~80% of timeout`, 2 of 3 5-min periods | Thresholds: 24000 ms (30s-timeout fns) / 48000 ms (60s-timeout fns) |
|
|
| `Lambda-Throttles-<fn>` | `Throttles` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
|
|
**DynamoDB alarm** (`afterhours-shifts` table):
|
|
|
|
| Alarm | Metric | Condition | Notes |
|
|
|---|---|---|---|
|
|
| `DDB-ReadThrottle-afterhours-shifts` | `ReadThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
| `DDB-WriteThrottle-afterhours-shifts` | `WriteThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
|
|
`ReadThrottleEvents` / `WriteThrottleEvents` are the table-level throttle
|
|
signals: AWS/DynamoDB emits them at the `TableName` dimension, so these alarms
|
|
transition normally. `ThrottledRequests` and `SystemErrors` are intentionally
|
|
**not** alarmed: AWS emits them only at `TableName`+`Operation` granularity, so a
|
|
`TableName`-only alarm would sit permanently in `INSUFFICIENT_DATA`.
|
|
|
|
**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway`
|
|
v2 metrics, `ApiId` dimension):
|
|
|
|
| Alarm | Metric | Condition |
|
|
|---|---|---|
|
|
| `ApiGateway-4xx-<apiId>` | `4xx` (Sum) | `>= 5` over one 5-min period |
|
|
| `ApiGateway-5xx-<apiId>` | `5xx` (Sum) | `>= 1` over one 5-min period |
|
|
| `ApiGateway-Latency-<apiId>` | `Latency` (p99, ms) | `>= 3000` ms, 2 of 3 5-min periods |
|
|
|
|
> Duration and API latency thresholds are starting points and may be tuned after
|
|
> observing real traffic.
|
|
|
|
## Releases & Versioning
|
|
|
|
The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`** — it is
|
|
the single source of truth for both the version number and the human-readable
|
|
notes. There is no separate tagging tool.
|
|
|
|
**To cut a release**, in your feature PR add a new `## vX.Y.Z — Month D, YYYY`
|
|
section at the top of `CHANGELOG.md` (plain language, written for on-call staff),
|
|
bumping per SemVer, then run `python scripts/sync_changelog.py` to update the
|
|
in-package copy. The `Changelog Guard` PR check enforces that the bump is a clean
|
|
single SemVer step above the latest tag and that the two copies match.
|
|
|
|
On the **deploy-then-merge** path, once the merge's Deploy succeeds, the Deploy
|
|
workflow's `release` job (`needs: deploy`) tags the new version, publishes a
|
|
GitHub Release with the notes, and — for **minor and major** bumps only (patches
|
|
stay silent) — invokes `afterhours-release-notifier` to post a "What's New"
|
|
message in the shift channel. The **App Home** tab ("About" page on the bot)
|
|
always shows the current version's notes, read from the CHANGELOG that ships in
|
|
the slack-bot package.
|
|
|
|
> The release job lives inside the Deploy workflow (gated on `needs: deploy`)
|
|
> rather than a separate `workflow_run`-triggered workflow. A push-to-main run is
|
|
> a trusted context, so checking out and running repo code with write/OIDC is safe
|
|
> — whereas `workflow_run` is flagged by CodeQL for untrusted checkout. Gating on
|
|
> `needs: deploy` still guarantees we never announce a version that isn't live.
|
|
|
|
**One-time setup (per environment):** after the first deploy creates the
|
|
`ReleaseNotifyInvokeRole`, copy its ARN from the `ReleaseNotifyInvokeRoleArn` stack
|
|
output into the repo **variable** `RELEASE_NOTIFY_INVOKE_ROLE_ARN` (Settings →
|
|
Secrets and variables → Actions → Variables). Until it's set, releases still tag
|
|
and publish but skip the Slack announcement (with a warning).
|
|
|
|
> **Convention note (deliberate deviation).** The Sea Haven handbook says internal
|
|
> SAM stacks generally need no versioning and that tags are applied manually. This
|
|
> bot is versioned by owner choice (it has staff-facing release notes) and tagged
|
|
> automatically by the Deploy workflow's release job. This is intentional — not drift.
|
|
|
|
## Testing
|
|
|
|
Unit tests use `pytest` with all external boundaries mocked — DynamoDB / SES /
|
|
Secrets Manager via `moto`, 3CX HTTP via `responses`, Slack via fakes, and time
|
|
via `freezegun`. No test touches the network or real AWS.
|
|
|
|
```bash
|
|
python -m venv .venv && source .venv/bin/activate
|
|
pip install -r tests/requirements.txt # test-only deps
|
|
pip install -r src/slack-bot/requirements.txt \
|
|
-r src/weekly-post/requirements.txt \
|
|
-r src/shared/requirements.txt # runtime deps the imports need
|
|
pytest
|
|
```
|
|
|
|
Each Lambda has its own `app.py`, so the per-package `conftest.py` loads each one
|
|
under a unique module name (importlib mode) to avoid collisions. CI runs the same
|
|
suite on every PR via the org `ci-python-sam` workflow (`run-tests: true`).
|
|
|
|
See [SETUP.md](SETUP.md) for full deployment and Slack app creation instructions.
|