afterhours-shift-manager/SETUP.md
Adam Moussa 7c5d44e9cf
feat(api): collapse Slack, portal, and jobs onto Fargate (PLAT-216)
Move HTTP and scheduled work onto one always-on Flask task so after-hours
loses Lambda cold start without changing the Cognito or roster contracts.
2026-09-21 14:42:34 -04:00

10 KiB
Raw Blame History

After-Hours Shift Manager — Setup Guide

1. Create a Slack App

  1. Go to api.slack.com/apps → Create New App → From scratch
  2. Name: After-Hours Shift Manager, pick your workspace
  3. Under OAuth & Permissions, add these Bot Token Scopes:
    • chat:write — post messages to channels
    • commands — register slash commands
    • chat:write.public — post to channels the bot isn't in
  4. Under Slash Commands, create a new command:
    • Command: /oncall
    • Request URL: (fill in after deploy — see step 4)
    • Short Description: Manage after-hours on-call shifts
    • Usage Hint: [schedule | pick <date> | drop <date> | swap <date> @person | register <ext> | roster | help]
  5. Under Interactivity & Shortcuts, toggle ON:
    • Request URL: same URL as the slash command (the /slack/events endpoint)
  6. Install to Workspace — approve the permissions
  7. Copy the Bot User OAuth Token (xoxb-...) and Signing Secret (under Basic Information)

2. Store Secrets in AWS Secrets Manager

The stack reads these from Secrets Manager (the IAM roles grant secretsmanager:GetSecretValue on afterhours-shift-manager/*):

# Slack
aws secretsmanager create-secret \
  --name afterhours-shift-manager/slack-bot-token \
  --secret-string "xoxb-YOUR-BOT-TOKEN"

aws secretsmanager create-secret \
  --name afterhours-shift-manager/slack-signing-secret \
  --secret-string "YOUR-SIGNING-SECRET"

# 3CX Queue XAPI (used by roster-sync and ring-scheduler)
aws secretsmanager create-secret \
  --name afterhours-shift-manager/3cx-domain \
  --secret-string "yourcompany.3cx.us"

aws secretsmanager create-secret \
  --name afterhours-shift-manager/3cx-client-id \
  --secret-string "YOUR-3CX-CLIENT-ID"

aws secretsmanager create-secret \
  --name afterhours-shift-manager/3cx-client-secret \
  --secret-string "YOUR-3CX-CLIENT-SECRET"

# Roster HTTP API bearer token (plain string). Duplicate the same value into
# seahaven-prod as paychex-integrations/afterhours-roster-token. Generate the
# token into a temp file, pass --secret-string file://..., then delete the file.
# Never paste the value into chat, Terraform, or a PR.
aws secretsmanager create-secret \
  --name afterhours-shift-manager/roster-api-token \
  --secret-string file://./roster-api-token.tmp

Create afterhours-shift-manager/roster-api-token before the first deploy that includes afterhours-roster-api, or live PUT/DELETE calls return 503.

Rotation is coordinated: write the new value to both afterhours-shift-manager/roster-api-token and paychex-integrations/afterhours-roster-token, then recycle afterhours-roster-api so cached execution environments pick it up. Updating only one copy causes 401s. The identity processor AFTERHOURS_BASE_URL is the Terraform output api_origin (origin only, no /roster suffix). Flip that HCP variable on paychex-integrations-prod at cutover after DynamoDB is copied.

Daily roster-sync still removes DynamoDB rows that are not in the 3CX DEFAULT group. Hire stays safe because 3CX create lands the extension in that group before the identity processor PUTs /roster.

The Slack channel ID is not a secret. It is Terraform variable shift_channel (default C0APATP612N), not stored in Secrets Manager.

3. HCP Terraform and GitHub Environment

Workspaces tagged app:afterhours-shift-manager: afterhours-shift-manager-prod in seahaven-prod (account 011934824531) and afterhours-shift-manager-dev in seahaven-dev (account 710827005802). Create the dev workspace before any Slack URL flip. HCP variable environment is prod or dev.

First apply uses the hcptf-bootstrap window (exact StringEquals trust, never StringLike). Creating afterhours-shift-manager-ecs-task-boundary is iam:CreatePolicy and needs that window.

  1. Create the HCP workspace. Auto-apply off. No project-level variable set. Working directory terraform. File trigger prefix terraform/** only. Speculative plans on. VCS on main.
  2. From seahaven-org-baseline: scripts/create-hcptf-bootstrap-roles.sh --account prod --allow-workspace afterhours-shift-manager-prod
  3. Point workspace TFC_AWS_APPLY_ROLE_ARN / TFC_AWS_PLAN_ROLE_ARN at hcptf-bootstrap / hcptf-bootstrap-plan. Set TFC_AWS_PROVIDER_AUTH=true.
  4. One manual apply with schedules_enabled=false. This creates the scoped hcptf-* roles, the Lambda boundary, and the rest of the stack. If 3CX secrets already exist from seahaven-door-unlock-api, import those three names instead of creating them: terraform import 'aws_secretsmanager_secret.this["afterhours-shift-manager/3cx-domain"]' afterhours-shift-manager/3cx-domain (and the client-id / client-secret names). Do not overwrite 3CX values.
  5. Retarget TFC_AWS_* to hcptf-afterhours-shift-manager / hcptf-afterhours-shift-manager-plan. Re-run the create script with no --allow-workspace.
  6. Second manual apply as the scoped role. Then seal auto-apply on.

GitHub Environment prod: reviewers, branch policy main and v*, Environment variable DEPLOY_ROLE_ARN = Terraform output github_deploy_role_arn. GitHub Environment dev: same DEPLOY_ROLE_ARN in the matching account, branch policy as needed for image CD.

Function zips: Actions → Deploy on push to main, or workflow_dispatch. Image CD: Actions → Deploy API (deploy-api.yaml). Keep schedules_enabled=false and ecs_schedules_enabled=false until the Fargate cutover below.

HCP outputs to copy: slack_request_url, api_origin, fargate_origin, holiday_scheduler_role_arn, github_deploy_role_arn, jobs_queue_arn.

4. Set the Slack Request URL

Reuse the existing Slack app. After the zip deploy, copy slack_request_url from HCP outputs:

  • Slash Commands → edit /oncall → set Request URL to that URL
  • Interactivity & Shortcuts → set Request URL to the same URL

Do this in the cutover window, not before DynamoDB is copied.

5. Seed the Schedule

python scripts/seed_schedule.py

This populates the DynamoDB table with the employee roster and default weekly schedule.

6. Invite the Bot & Register Users

  1. Create a channel (e.g. #after-hours-shifts) and invite the bot: /invite @After-Hours Shift Manager
  2. Each employee links their Slack account by running:
    /oncall register 114
    
    (using their own extension number)

7. Update the 3CX Scheduler (optional)

To have the 3CX scheduler read overrides from DynamoDB (so Slack-driven changes apply to future dates automatically), deploy the updated 3CX scheduler from the feature/dynamodb-shift-integration branch. See that branch's changes for details.

Without this step, the Slack bot still works — it invokes the 3CX scheduler Lambda directly for same-day changes. Future-date overrides would only take effect if the scheduler reads DynamoDB.

8. Prod cutover (PLAT-74)

Avoid Monday 06:00-08:00 ET and any holiday 08:00/17:00 ET window. Dry-run the scripts first (--execute is required for writes).

  1. Merge this repo's PR (SAM CD is gone). First HCP apply is the bootstrap window above with schedules_enabled=false.
  2. Copy DynamoDB afterhours-shifts mgmt → prod. Verify item counts: python scripts/cutover/copy_dynamodb.py --src-profile mgmt --dst-profile prod then --execute.
  3. Confirm secrets in prod. copy_secrets.py writes Slack bot token, Slack signing secret, and roster-api-token into empty Terraform shells and skips dest names that already have a value. It never writes 3CX secrets: python scripts/cutover/copy_secrets.py --src-profile mgmt --dst-profile prod then --execute. Strip trailing newlines is built in.
  4. GHA workflow_dispatch (or the merge deploy; re-run if it raced apply) to overwrite stubs.
  5. Recreate outstanding future holiday-activate-* / holiday-deactivate-* in prod against the new router ARN and scheduler role: python scripts/cutover/recreate_holiday_schedules.py --src-profile mgmt --dst-profile prod
  6. Instant cut: Slack Request URL → prod /slack/events; Paychex HCP variable afterhours_base_url on paychex-integrations-prod → prod api_origin; schedules_enabled=true via a terraform-only merge; disable mgmt EventBridge. Smoke: Slack /oncall, roster PUT/DELETE, weekly-post SendMessage (or simulate-principal-policy plus one smoke message), ring-scheduler invoke, holiday GetSchedule.
  7. Seal auto-apply on. Delete the mgmt SAM stack. Remove the mgmt SQS principal from paychex-checkcomponents. Update Confluence AWS Architecture Map and check PLAT-71 item 4.

Do not dual-run 3CX writers. Do not flip afterhours_base_url before DynamoDB is copied.

9. Fargate cutover (PLAT-216)

Dual-run ECS beside API Gateway. Do not dual-write 3CX. Do not flip Slack without the afterhours-shift-manager-dev workspace already serving /api/health.

  1. Image deploy via deploy-api.yaml. GET /api/health reports the real sha.
  2. Recreate outstanding holiday-activate-* / holiday-deactivate-* onto jobs SQS (same class of work as recreate_holiday_schedules.py): python scripts/cutover/retarget_holiday_schedules_to_sqs.py --profile prod --queue-arn <jobs_queue_arn> then --execute.
  3. Instant cut: Slack Request URL, Paychex AFTERHOURS_BASE_URL, portal VITE_SHIFTS_API_BASE → https://afterhours.seahaven.com (dev hostname in portal-dev). Smoke /oncall, roster PUT/DELETE, GET /api/shifts, one holiday GetSchedule, ring job.
  4. Set ecs_schedules_enabled=true and keep schedules_enabled=false. Confirm no 3CX writers remain on Lambda.
  5. After smoke, remove API Gateway, the eight Lambdas, zip packaging, and Lambda alarms in a follow-up apply. ALB 5xx/latency and unhealthy-host alarms stay.

Commands Reference

Command Description
/oncall Show this week's schedule
/oncall next Show next week's schedule
/oncall pick <date> Pick up a shift
/oncall drop <date> Drop your shift (marks it open)
/oncall swap <date> @person Hand your shift to someone else
/oncall register <ext> Link your Slack to your extension
/oncall roster Show all employees and their link status
/oncall help Show help

Dates can be: today, tomorrow, monday–sunday, 4/5, 2026-04-05