afterhours-shift-manager/SETUP.md
Adam Moussa 7c5d44e9cf
feat(api): collapse Slack, portal, and jobs onto Fargate (PLAT-216)
Move HTTP and scheduled work onto one always-on Flask task so after-hours
loses Lambda cold start without changing the Cognito or roster contracts.
2026-09-21 14:42:34 -04:00

218 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# After-Hours Shift Manager — Setup Guide
## 1. Create a Slack App
1. Go to [api.slack.com/apps](https://api.slack.com/apps) → **Create New App** → **From scratch**
2. Name: `After-Hours Shift Manager`, pick your workspace
3. Under **OAuth & Permissions**, add these **Bot Token Scopes**:
- `chat:write` — post messages to channels
- `commands` — register slash commands
- `chat:write.public` — post to channels the bot isn't in
4. Under **Slash Commands**, create a new command:
- Command: `/oncall`
- Request URL: `(fill in after deploy — see step 4)`
- Short Description: `Manage after-hours on-call shifts`
- Usage Hint: `[schedule | pick <date> | drop <date> | swap <date> @person | register <ext> | roster | help]`
5. Under **Interactivity & Shortcuts**, toggle **ON**:
- Request URL: same URL as the slash command (the `/slack/events` endpoint)
6. **Install to Workspace** — approve the permissions
7. Copy the **Bot User OAuth Token** (`xoxb-...`) and **Signing Secret** (under Basic Information)
## 2. Store Secrets in AWS Secrets Manager
The stack reads these from Secrets Manager (the IAM roles grant
`secretsmanager:GetSecretValue` on `afterhours-shift-manager/*`):
```bash
# Slack
aws secretsmanager create-secret \
--name afterhours-shift-manager/slack-bot-token \
--secret-string "xoxb-YOUR-BOT-TOKEN"
aws secretsmanager create-secret \
--name afterhours-shift-manager/slack-signing-secret \
--secret-string "YOUR-SIGNING-SECRET"
# 3CX Queue XAPI (used by roster-sync and ring-scheduler)
aws secretsmanager create-secret \
--name afterhours-shift-manager/3cx-domain \
--secret-string "yourcompany.3cx.us"
aws secretsmanager create-secret \
--name afterhours-shift-manager/3cx-client-id \
--secret-string "YOUR-3CX-CLIENT-ID"
aws secretsmanager create-secret \
--name afterhours-shift-manager/3cx-client-secret \
--secret-string "YOUR-3CX-CLIENT-SECRET"
# Roster HTTP API bearer token (plain string). Duplicate the same value into
# seahaven-prod as paychex-integrations/afterhours-roster-token. Generate the
# token into a temp file, pass --secret-string file://..., then delete the file.
# Never paste the value into chat, Terraform, or a PR.
aws secretsmanager create-secret \
--name afterhours-shift-manager/roster-api-token \
--secret-string file://./roster-api-token.tmp
```
Create `afterhours-shift-manager/roster-api-token` **before** the first deploy that
includes `afterhours-roster-api`, or live PUT/DELETE calls return 503.
Rotation is coordinated: write the new value to both
`afterhours-shift-manager/roster-api-token` and
`paychex-integrations/afterhours-roster-token`, then recycle
`afterhours-roster-api` so cached execution environments pick it up. Updating
only one copy causes 401s. The identity processor `AFTERHOURS_BASE_URL` is the
Terraform output `api_origin` (origin only, no `/roster` suffix). Flip that HCP
variable on `paychex-integrations-prod` at cutover after DynamoDB is copied.
Daily roster-sync still removes DynamoDB rows that are not in the 3CX `DEFAULT`
group. Hire stays safe because 3CX create lands the extension in that group
before the identity processor PUTs `/roster`.
> The Slack **channel ID** is not a secret. It is Terraform variable
> `shift_channel` (default `C0APATP612N`), not stored in Secrets Manager.
## 3. HCP Terraform and GitHub Environment
Workspaces tagged `app:afterhours-shift-manager`:
`afterhours-shift-manager-prod` in `seahaven-prod` (account `011934824531`) and
`afterhours-shift-manager-dev` in `seahaven-dev` (account `710827005802`).
Create the dev workspace before any Slack URL flip. HCP variable `environment`
is `prod` or `dev`.
First apply uses the hcptf-bootstrap window (exact `StringEquals` trust, never
`StringLike`). Creating `afterhours-shift-manager-ecs-task-boundary` is
`iam:CreatePolicy` and needs that window.
1. Create the HCP workspace. Auto-apply off. No project-level variable set.
Working directory `terraform`. File trigger prefix `terraform/**` only.
Speculative plans on. VCS on `main`.
2. From `seahaven-org-baseline`:
`scripts/create-hcptf-bootstrap-roles.sh --account prod --allow-workspace afterhours-shift-manager-prod`
3. Point workspace `TFC_AWS_APPLY_ROLE_ARN` / `TFC_AWS_PLAN_ROLE_ARN` at
`hcptf-bootstrap` / `hcptf-bootstrap-plan`. Set `TFC_AWS_PROVIDER_AUTH=true`.
4. One manual apply with `schedules_enabled=false`. This creates the scoped
`hcptf-*` roles, the Lambda boundary, and the rest of the stack. If 3CX
secrets already exist from seahaven-door-unlock-api, import those three
names instead of creating them:
`terraform import 'aws_secretsmanager_secret.this["afterhours-shift-manager/3cx-domain"]' afterhours-shift-manager/3cx-domain`
(and the client-id / client-secret names). Do not overwrite 3CX values.
5. Retarget `TFC_AWS_*` to `hcptf-afterhours-shift-manager` /
`hcptf-afterhours-shift-manager-plan`. Re-run the create script with no
`--allow-workspace`.
6. Second manual apply as the scoped role. Then seal auto-apply on.
GitHub Environment `prod`: reviewers, branch policy `main` and `v*`, Environment
variable `DEPLOY_ROLE_ARN` = Terraform output `github_deploy_role_arn`.
GitHub Environment `dev`: same `DEPLOY_ROLE_ARN` in the matching account,
branch policy as needed for image CD.
Function zips: Actions → Deploy on push to `main`, or `workflow_dispatch`.
Image CD: Actions → Deploy API (`deploy-api.yaml`). Keep `schedules_enabled=false`
and `ecs_schedules_enabled=false` until the Fargate cutover below.
HCP outputs to copy: `slack_request_url`, `api_origin`, `fargate_origin`,
`holiday_scheduler_role_arn`, `github_deploy_role_arn`, `jobs_queue_arn`.
## 4. Set the Slack Request URL
Reuse the existing Slack app. After the zip deploy, copy `slack_request_url`
from HCP outputs:
- **Slash Commands** → edit `/oncall` → set **Request URL** to that URL
- **Interactivity & Shortcuts** → set **Request URL** to the same URL
Do this in the cutover window, not before DynamoDB is copied.
## 5. Seed the Schedule
```bash
python scripts/seed_schedule.py
```
This populates the DynamoDB table with the employee roster and default weekly schedule.
## 6. Invite the Bot & Register Users
1. Create a channel (e.g. `#after-hours-shifts`) and invite the bot: `/invite @After-Hours Shift Manager`
2. Each employee links their Slack account by running:
```
/oncall register 114
```
(using their own extension number)
## 7. Update the 3CX Scheduler (optional)
To have the 3CX scheduler read overrides from DynamoDB (so Slack-driven changes apply to future dates automatically), deploy the updated 3CX scheduler from the `feature/dynamodb-shift-integration` branch. See that branch's changes for details.
Without this step, the Slack bot still works — it invokes the 3CX scheduler Lambda directly for same-day changes. Future-date overrides would only take effect if the scheduler reads DynamoDB.
## 8. Prod cutover (PLAT-74)
Avoid Monday 06:00-08:00 ET and any holiday 08:00/17:00 ET window. Dry-run the
scripts first (`--execute` is required for writes).
1. Merge this repo's PR (SAM CD is gone). First HCP apply is the bootstrap
window above with `schedules_enabled=false`.
2. Copy DynamoDB `afterhours-shifts` mgmt → prod. Verify item counts:
`python scripts/cutover/copy_dynamodb.py --src-profile mgmt --dst-profile prod`
then `--execute`.
3. Confirm secrets in prod. `copy_secrets.py` writes Slack bot token, Slack
signing secret, and roster-api-token into empty Terraform shells and skips
dest names that already have a value. It never writes 3CX secrets:
`python scripts/cutover/copy_secrets.py --src-profile mgmt --dst-profile prod`
then `--execute`. Strip trailing newlines is built in.
4. GHA `workflow_dispatch` (or the merge deploy; re-run if it raced apply) to
overwrite stubs.
5. Recreate outstanding future `holiday-activate-*` / `holiday-deactivate-*`
in prod against the new router ARN and scheduler role:
`python scripts/cutover/recreate_holiday_schedules.py --src-profile mgmt --dst-profile prod`
6. Instant cut: Slack Request URL → prod `/slack/events`; Paychex HCP variable
`afterhours_base_url` on `paychex-integrations-prod` → prod `api_origin`;
`schedules_enabled=true` via a terraform-only merge; disable mgmt EventBridge.
Smoke: Slack `/oncall`, roster PUT/DELETE, weekly-post SendMessage (or
simulate-principal-policy plus one smoke message), ring-scheduler invoke,
holiday GetSchedule.
7. Seal auto-apply on. Delete the mgmt SAM stack. Remove the mgmt SQS principal
from `paychex-checkcomponents`. Update Confluence AWS Architecture Map and
check PLAT-71 item 4.
Do not dual-run 3CX writers. Do not flip `afterhours_base_url` before DynamoDB
is copied.
## 9. Fargate cutover (PLAT-216)
Dual-run ECS beside API Gateway. Do not dual-write 3CX. Do not flip Slack
without the `afterhours-shift-manager-dev` workspace already serving
`/api/health`.
1. Image deploy via `deploy-api.yaml`. `GET /api/health` reports the real sha.
2. Recreate outstanding `holiday-activate-*` / `holiday-deactivate-*` onto
jobs SQS (same class of work as `recreate_holiday_schedules.py`):
`python scripts/cutover/retarget_holiday_schedules_to_sqs.py --profile prod --queue-arn <jobs_queue_arn>`
then `--execute`.
3. Instant cut: Slack Request URL, Paychex `AFTERHOURS_BASE_URL`, portal
`VITE_SHIFTS_API_BASE` → `https://afterhours.seahaven.com` (dev hostname in
portal-dev). Smoke `/oncall`, roster PUT/DELETE, `GET /api/shifts`, one
holiday GetSchedule, ring job.
4. Set `ecs_schedules_enabled=true` and keep `schedules_enabled=false`. Confirm
no 3CX writers remain on Lambda.
5. After smoke, remove API Gateway, the eight Lambdas, zip packaging, and
Lambda alarms in a follow-up apply. ALB 5xx/latency and unhealthy-host
alarms stay.
## Commands Reference
| Command | Description |
|---|---|
| `/oncall` | Show this week's schedule |
| `/oncall next` | Show next week's schedule |
| `/oncall pick <date>` | Pick up a shift |
| `/oncall drop <date>` | Drop your shift (marks it open) |
| `/oncall swap <date> @person` | Hand your shift to someone else |
| `/oncall register <ext>` | Link your Slack to your extension |
| `/oncall roster` | Show all employees and their link status |
| `/oncall help` | Show help |
Dates can be: `today`, `tomorrow`, `monday`–`sunday`, `4/5`, `2026-04-05`