mirror of
https://github.com/Sea-Haven-Industries/afterhours-shift-manager.git
synced 2026-09-30 07:53:11 +00:00
* Add CloudWatch alarm coverage for all functions, table, and HTTP API Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests is not emitted at the TableName-only dimension (only TableName+Operation), so the table-level alarm would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Correct DynamoDB alarm docs and drop sign-off wording README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
290 lines
17 KiB
Markdown
290 lines
17 KiB
Markdown
# After-Hours Shift Manager
|
|
|
|

|
|

|
|

|
|

|
|
|
|
Slack bot for managing after-hours on-call shifts at Sea Haven Industries. Employees can pick up, drop, and swap shifts directly from Slack. Changes automatically update 3CX queue routing via the integrated ring scheduler.
|
|
|
|
## How It Works
|
|
|
|
A recurring weekly schedule assigns employees to after-hours phone duty. Weekend shifts are split into Day (8am-5pm) and Night (5pm-8am). Any unassigned shift shows as **Available** in Slack with a pickup button. When someone picks up or drops a shift for today, the 3CX queue is updated immediately. Future changes take effect when the ring scheduler runs at 8am daily and 5pm on weekends.
|
|
|
|
The weekly schedule post is updated live when shifts change, and the previous week's post is automatically deleted when the new one goes out.
|
|
|
|
The bot also has an **About** page: open the bot in Slack and click its **Home** tab to see what it does, the full command list, and the latest "What's New" (see [Releases & Versioning](#releases--versioning)).
|
|
|
|
## Slack Commands
|
|
|
|
| Command | Description |
|
|
|---|---|
|
|
| `/oncall` | Show this week's schedule |
|
|
| `/oncall next` | Show next week's schedule |
|
|
| `/oncall pick <date>` | Pick up an available shift — instant if the shift hasn't started; if it's already underway (but not ended) it needs admin approval (see [Late-pickup approval](#late-pickup-approval)) |
|
|
| `/oncall drop <date>` | Drop your shift (marks it available) — blocked within 24h of shift start; swap or ask an admin instead |
|
|
| `/oncall swap <date> @person` | Request a swap — the other person gets an Accept/Decline DM and the shift only moves once they accept |
|
|
| `/oncall register <ext>` | Link your Slack account to your phone extension |
|
|
| `/oncall roster` | Show all employees and their link status |
|
|
| `/oncall pay` | Show last week's bonus pay summary |
|
|
| `/oncall rate` | Show current shift pay rates |
|
|
| `/oncall rate default <amount>` | Set the default per-shift rate |
|
|
| `/oncall rate <ext> <amount>` | Set a per-person shift rate |
|
|
| `/oncall help` | Show help |
|
|
|
|
### Admin Commands
|
|
|
|
Available to users listed in `admin_users` in the CONFIG record:
|
|
|
|
| Command | Description |
|
|
|---|---|
|
|
| `/oncall admin override <date> <ext>` | Assign a shift to an extension |
|
|
| `/oncall admin open <date>` | Mark a shift as open |
|
|
| `/oncall admin clear <date>` | Remove override (revert to weekly) |
|
|
| `/oncall admin roster add <ext> <name>` | Add an employee to the roster |
|
|
| `/oncall admin roster remove <ext>` | Remove an employee |
|
|
| `/oncall admin roster rename <ext> <name>` | Rename an employee |
|
|
| `/oncall admin holiday add <date> <slots> [x<mult>] <label>` | Schedule a holiday day shift (8am-5pm ET) with N slots, optional pay multiplier override (e.g. `x2`), and a label |
|
|
| `/oncall admin holiday remove <date>` | Remove a scheduled holiday and its activate/deactivate schedules |
|
|
| `/oncall admin holiday list` | List today-and-future scheduled holidays |
|
|
|
|
Dates accept: `today`, `tomorrow`, `monday`-`sunday`, `4/5`, `2026-04-05`
|
|
|
|
Example: `/oncall admin holiday add 2026-07-04 2 x2 Independence Day` schedules a
|
|
2-slot holiday paying 2x. Omitting the `x<mult>` token uses the default
|
|
`holiday_multiplier` from CONFIG (1.5). See [Holidays](#holidays) for the full flow.
|
|
|
|
## Architecture
|
|
|
|
- **Runtime**: Python 3.12 on AWS Lambda (arm64)
|
|
- **Data**: DynamoDB single-table (`afterhours-shifts`)
|
|
- **IaC**: AWS SAM (`template.yaml`) with shared Lambda Layer
|
|
- **Slack**: Slack Bolt framework with `/oncall` slash command
|
|
- **3CX Integration**: Queue routing updated directly via 3CX Queue XAPI
|
|
- **Secrets**: AWS Secrets Manager (`afterhours-shift-manager/*`)
|
|
|
|
### Lambda Functions
|
|
|
|
| Function | Trigger | Purpose |
|
|
|---|---|---|
|
|
| `afterhours-shift-manager` | API Gateway (POST /slack/events) | Slack bot — handles `/oncall` commands and interactive buttons |
|
|
| `afterhours-weekly-post` | EventBridge (Monday 7am ET) | Posts weekly schedule to Slack, sends pay report email |
|
|
| `afterhours-roster-sync` | EventBridge (daily 6am ET) | Syncs employee roster from 3CX |
|
|
| `afterhours-ring-scheduler` | EventBridge (daily 8am ET + weekend 5pm ET) | Updates 3CX queue routing based on who's on shift |
|
|
| `afterhours-holiday-router` | EventBridge Scheduler (per-holiday one-off: 8am activate / 5pm deactivate ET) | Repoints the IVR to the holiday queue and sets queue agents for a holiday day shift; reverts at 5pm (see [Holidays](#holidays)) |
|
|
| `afterhours-release-notifier` | Invoked by the Deploy workflow's release job on minor/major releases | Posts a "What's New" announcement to the shift channel |
|
|
|
|
### Project Layout
|
|
|
|
```
|
|
src/
|
|
slack-bot/ Slack Bolt Lambda (handler + app); ships CHANGELOG.md for App Home
|
|
weekly-post/ Monday schedule + pay post
|
|
roster-sync/ Daily 3CX roster sync
|
|
ring-scheduler/ 3CX queue routing updates
|
|
holiday-router/ 3CX IVR/queue repoint for holiday day shifts (activate/deactivate)
|
|
release-notifier/ Posts release announcements to Slack
|
|
shared/ Lambda Layer (schedule, blocks, changelog, 3CX client, secrets)
|
|
scripts/ changelog CLI + CI guard + in-package copy sync
|
|
tests/ pytest suite (mirrors src/, one dir per Lambda + shared)
|
|
```
|
|
|
|
### DynamoDB Schema
|
|
|
|
Single table with `PK` / `SK` keys:
|
|
|
|
| PK | SK | Description |
|
|
|---|---|---|
|
|
| `ROSTER` | `<extension>` | Employee: name, extension, slack_user_id |
|
|
| `WEEKLY` | `<DayName>` | Default weekly schedule: extension, name |
|
|
| `OVERRIDE` | `<YYYY-MM-DD>` | Date override from pickup/drop (or `OPEN`) |
|
|
| `SWAP` | `<YYYY-MM-DD>` | Pending/verified swap request: requester, target, status, `expires_at` (TTL) |
|
|
| `HOLIDAY` | `<YYYY-MM-DD>` | Holiday day shift (one per date): `slots` (int), `assignees` (MAP keyed by extension — `{"114": {name, claimed_at}}`), `multiplier` (Decimal, defaults to `CONFIG.holiday_multiplier` = 1.5, overridable per holiday), `label`, `created_at`, `created_by`, `activated` (bool), `schedule_names` (list) |
|
|
| `PICKUP_REQUEST` | `<YYYY-MM-DD>[-DAY]#<ext>` | Pending late-pickup awaiting admin approval: `requester_ext`, `requester_name`, `requester_slack`, `shift_type`, `is_holiday`, `status`, `created_at`, `expires_at` (TTL = shift end) |
|
|
| `SCHEDULE_POST` | `<channel_id>` | Current schedule message timestamp |
|
|
| `PAY` | `<YYYY-MM-DD>` | Weekly pay record (Monday date key) |
|
|
| `CONFIG` | `CONFIG` | Settings: shift_rate, fallback_extension, admin_users, `ring_group`, `holiday_multiplier` (default holiday pay multiplier, 1.5), `holiday_queue` (3CX queue repointed during holidays, default 802), `ivr_number` (3CX IVR repointed during holidays, default 800), `captured_ivr_routes` (original IVR routes saved at holiday activation, restored at deactivation) |
|
|
|
|
Weekend day-shift rows use a `-DAY` suffix on the SK (e.g. `OVERRIDE` / `2026-04-05-DAY`). The table has TTL enabled on `expires_at` so abandoned pending swaps and pickup requests self-clean.
|
|
|
|
Shift priority for any date is **HOLIDAY > OVERRIDE > WEEKLY** — a holiday record wins over a regular override, which wins over the standing weekly schedule.
|
|
|
|
Holiday slot claims are atomic Map updates so concurrent pickers can't oversubscribe:
|
|
|
|
- **Claim** — `SET assignees.#ext` guarded by `attribute_not_exists(assignees.#ext) AND size(assignees) < :slots`.
|
|
- **Release** — `REMOVE assignees.#ext` guarded by `attribute_exists`.
|
|
- **Swap** — `REMOVE #from SET #to` guarded by `attribute_exists(#from) AND attribute_not_exists(#to)`.
|
|
|
|
**Swap flow:** `/oncall swap` writes a `pending` `SWAP` record and DMs the target Accept/Decline buttons; it does **not** reassign the shift. On Accept, the override is written, 3CX is repointed if it's the active shift, and the record is marked `verified`. On Decline (or once the shift has started) the request is dropped and the shift stays with the original owner.
|
|
|
|
### Holidays
|
|
|
|
A **holiday** is a single day-only shift (08:00-17:00 ET), one `HOLIDAY` record per
|
|
date, that can hold multiple people (`slots`). Admins manage holidays with
|
|
`/oncall admin holiday add|remove|list`. A holiday takes priority over a regular
|
|
override and the weekly schedule for that date, and pays at its `multiplier`
|
|
(per-holiday override, else `CONFIG.holiday_multiplier`, default 1.5). Open slots
|
|
show in the schedule with a pickup button; claims, releases, and swaps are atomic
|
|
Map updates on the record (see above) so the slot count can't be oversubscribed.
|
|
|
|
**Holiday-router + Scheduler flow.** When an admin adds a holiday, the slack-bot
|
|
creates two **one-off EventBridge Scheduler** schedules for that date —
|
|
`holiday-activate-<YYYYMMDD>` at 08:00 ET and `holiday-deactivate-<YYYYMMDD>` at
|
|
17:00 ET — whose names are stored on the record's `schedule_names`. Scheduler
|
|
assumes `HolidaySchedulerExecutionRole` to invoke `afterhours-holiday-router`:
|
|
|
|
- **Activate (08:00):** capture both IVR `ivr_number` (800) routes — key-0 **and**
|
|
no-input/timeout — into `CONFIG.captured_ivr_routes` (skipped if they already
|
|
point at the holiday queue, so re-runs don't clobber the originals), set
|
|
`holiday_queue` (802) agents to the holiday's assignees (or `[fallback_extension]`
|
|
= `[100]` when no slots are filled — set **once** here), repoint **both** IVR 800
|
|
routes to queue 802, and mark `activated = True`. Idempotent.
|
|
- **Deactivate (17:00):** restore both IVR routes from `captured_ivr_routes` (only
|
|
routes still pointing at the queue, defensive against manual changes), clear them,
|
|
empty queue 802's agents, and mark `activated = False`. Idempotent.
|
|
|
|
Queue 801 (the daily ring-scheduler queue) is left untouched. If an admin adds a
|
|
holiday whose 08:00-17:00 window is already open, the slack-bot **inline-activates**
|
|
it immediately (invoking the router) rather than waiting for the 08:00 schedule.
|
|
Removing a holiday deletes the record and any outstanding schedules.
|
|
|
|
### Late-pickup approval
|
|
|
|
Picking up a shift behaves differently depending on timing, for **both** regular
|
|
and holiday shifts:
|
|
|
|
- **Before the shift starts** — immediate pickup (the prior behaviour, unchanged).
|
|
- **After the shift has started but before it ends** (08:00 for a day/holiday shift,
|
|
17:00 for a night shift) — the shift is **not** claimed yet. The bot writes a
|
|
`pending` `PICKUP_REQUEST` and DMs **every admin** Approve/Deny buttons (mirroring
|
|
the verified-swap flow). The **first admin to approve wins** (the claim is
|
|
conditional, so a second approval is a safe no-op). On approve, the shift is
|
|
claimed (regular: `set_override`; holiday: atomic `claim_holiday_slot`), the
|
|
requester and channel are notified, and 3CX is fired if the window is live —
|
|
regular shifts call `_update_3cx_routing(picker)` when it's today's active shift;
|
|
holidays refresh queue 802's agents to the current assignees. On deny, the request
|
|
is cleared and the requester is told.
|
|
- **After the shift has ended** — rejected outright; it's too late to pick up.
|
|
|
|
A slot claimed after the shift has started always needs an admin to approve it.
|
|
|
|
### Secrets Manager
|
|
|
|
| Secret | Description |
|
|
|---|---|
|
|
| `afterhours-shift-manager/slack-bot-token` | Slack bot OAuth token (`xoxb-...`) |
|
|
| `afterhours-shift-manager/slack-signing-secret` | Slack app signing secret |
|
|
| `afterhours-shift-manager/3cx-domain` | 3CX FQDN (e.g. `company.3cx.us`) |
|
|
| `afterhours-shift-manager/3cx-client-id` | 3CX OAuth2 client ID |
|
|
| `afterhours-shift-manager/3cx-client-secret` | 3CX OAuth2 client secret |
|
|
|
|
## Deployment
|
|
|
|
Merges to `main` are automatically deployed via **GitHub Actions** using reusable SAM workflows from the Sea Haven org.
|
|
|
|
For manual deploys:
|
|
|
|
```bash
|
|
sam build
|
|
sam deploy
|
|
```
|
|
|
|
## Monitoring & Alarms
|
|
|
|
All CloudWatch alarms are defined in `template.yaml` and notify the shared
|
|
`site-alerts` SNS topic (→ AWS Chatbot → Slack). None set `OKActions` — recovery
|
|
is not paged. Alarm names follow the in-template convention `Lambda-<Metric>-<fn>`
|
|
(e.g. `Lambda-Errors-afterhours-ring-scheduler`).
|
|
|
|
**Lambda alarms** (all six functions: `afterhours-shift-manager`,
|
|
`afterhours-weekly-post`, `afterhours-roster-sync`, `afterhours-ring-scheduler`,
|
|
`afterhours-holiday-router`, `afterhours-release-notifier`):
|
|
|
|
| Alarm | Metric | Condition | Notes |
|
|
|---|---|---|---|
|
|
| `Lambda-Errors-<fn>` | `Errors` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
| `Lambda-Duration-<fn>` | `Duration` (Maximum, ms) | `>= ~80% of timeout`, 2 of 3 5-min periods | Thresholds: 24000 ms (30s-timeout fns) / 48000 ms (60s-timeout fns) |
|
|
| `Lambda-Throttles-<fn>` | `Throttles` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
|
|
**DynamoDB alarm** (`afterhours-shifts` table):
|
|
|
|
| Alarm | Metric | Condition | Notes |
|
|
|---|---|---|---|
|
|
| `DDB-ReadThrottle-afterhours-shifts` | `ReadThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
| `DDB-WriteThrottle-afterhours-shifts` | `WriteThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
|
|
|
|
`ReadThrottleEvents` / `WriteThrottleEvents` are the table-level throttle
|
|
signals: AWS/DynamoDB emits them at the `TableName` dimension, so these alarms
|
|
transition normally. `ThrottledRequests` and `SystemErrors` are intentionally
|
|
**not** alarmed: AWS emits them only at `TableName`+`Operation` granularity, so a
|
|
`TableName`-only alarm would sit permanently in `INSUFFICIENT_DATA`.
|
|
|
|
**API Gateway alarms** (implicit HTTP API `ServerlessHttpApi`, `AWS/ApiGateway`
|
|
v2 metrics, `ApiId` dimension):
|
|
|
|
| Alarm | Metric | Condition |
|
|
|---|---|---|
|
|
| `ApiGateway-4xx-<apiId>` | `4xx` (Sum) | `>= 5` over one 5-min period |
|
|
| `ApiGateway-5xx-<apiId>` | `5xx` (Sum) | `>= 1` over one 5-min period |
|
|
| `ApiGateway-Latency-<apiId>` | `Latency` (p99, ms) | `>= 3000` ms, 2 of 3 5-min periods |
|
|
|
|
> Duration and API latency thresholds are starting points and may be tuned after
|
|
> observing real traffic.
|
|
|
|
## Releases & Versioning
|
|
|
|
The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`** — it is
|
|
the single source of truth for both the version number and the human-readable
|
|
notes. There is no separate tagging tool.
|
|
|
|
**To cut a release**, in your feature PR add a new `## vX.Y.Z — Month D, YYYY`
|
|
section at the top of `CHANGELOG.md` (plain language, written for on-call staff),
|
|
bumping per SemVer, then run `python scripts/sync_changelog.py` to update the
|
|
in-package copy. The `Changelog Guard` PR check enforces that the bump is a clean
|
|
single SemVer step above the latest tag and that the two copies match.
|
|
|
|
On the **deploy-then-merge** path, once the merge's Deploy succeeds, the Deploy
|
|
workflow's `release` job (`needs: deploy`) tags the new version, publishes a
|
|
GitHub Release with the notes, and — for **minor and major** bumps only (patches
|
|
stay silent) — invokes `afterhours-release-notifier` to post a "What's New"
|
|
message in the shift channel. The **App Home** tab ("About" page on the bot)
|
|
always shows the current version's notes, read from the CHANGELOG that ships in
|
|
the slack-bot package.
|
|
|
|
> The release job lives inside the Deploy workflow (gated on `needs: deploy`)
|
|
> rather than a separate `workflow_run`-triggered workflow. A push-to-main run is
|
|
> a trusted context, so checking out and running repo code with write/OIDC is safe
|
|
> — whereas `workflow_run` is flagged by CodeQL for untrusted checkout. Gating on
|
|
> `needs: deploy` still guarantees we never announce a version that isn't live.
|
|
|
|
**One-time setup (per environment):** after the first deploy creates the
|
|
`ReleaseNotifyInvokeRole`, copy its ARN from the `ReleaseNotifyInvokeRoleArn` stack
|
|
output into the repo **variable** `RELEASE_NOTIFY_INVOKE_ROLE_ARN` (Settings →
|
|
Secrets and variables → Actions → Variables). Until it's set, releases still tag
|
|
and publish but skip the Slack announcement (with a warning).
|
|
|
|
> **Convention note (deliberate deviation).** The Sea Haven handbook says internal
|
|
> SAM stacks generally need no versioning and that tags are applied manually. This
|
|
> bot is versioned by owner choice (it has staff-facing release notes) and tagged
|
|
> automatically by the Deploy workflow's release job. This is intentional — not drift.
|
|
|
|
## Testing
|
|
|
|
Unit tests use `pytest` with all external boundaries mocked — DynamoDB / SES /
|
|
Secrets Manager via `moto`, 3CX HTTP via `responses`, Slack via fakes, and time
|
|
via `freezegun`. No test touches the network or real AWS.
|
|
|
|
```bash
|
|
python -m venv .venv && source .venv/bin/activate
|
|
pip install -r tests/requirements.txt # test-only deps
|
|
pip install -r src/slack-bot/requirements.txt \
|
|
-r src/weekly-post/requirements.txt \
|
|
-r src/shared/requirements.txt # runtime deps the imports need
|
|
pytest
|
|
```
|
|
|
|
Each Lambda has its own `app.py`, so the per-package `conftest.py` loads each one
|
|
under a unique module name (importlib mode) to avoid collisions. CI runs the same
|
|
suite on every PR via the org `ci-python-sam` workflow (`run-tests: true`).
|
|
|
|
See [SETUP.md](SETUP.md) for full deployment and Slack app creation instructions.
|