afterhours-shift-manager/README.md

286 lines
18 KiB
Markdown
Raw Normal View History

# After-Hours Shift Manager
![Python](https://img.shields.io/badge/Python-3776AB?logo=python&logoColor=white)
![Terraform](https://img.shields.io/badge/Terraform-844FBA?logo=terraform&logoColor=white)
![Slack](https://img.shields.io/badge/Slack-integration-4A154B?logo=slack&logoColor=white)
![CI](https://github.com/Sea-Haven-Industries/afterhours-shift-manager/actions/workflows/ci.yaml/badge.svg)
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
Slack bot for managing after-hours on-call shifts at Sea Haven Industries. Employees can pick up, drop, and swap shifts directly from Slack. Changes automatically update 3CX queue routing via the integrated ring scheduler.
## How It Works
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
A recurring weekly schedule assigns employees to after-hours phone duty. Weekend shifts are split into Day (8am-5pm) and Night (5pm-8am). Any unassigned shift shows as **Available** in Slack with a pickup button. When someone picks up or drops a shift for today, the 3CX queue is updated immediately. Future changes take effect when the ring scheduler runs at 8am daily and 5pm on weekends.
The weekly schedule post is updated live when shifts change, and the previous week's post is automatically deleted when the new one goes out.
Add changelog-driven releases and App Home tab (#112) * Add changelog-driven releases and App Home tab Version the bot continuously from CHANGELOG.md (the single source of truth for both the version and the staff-readable notes) and surface changes to users in two ways: - A new afterhours-release-notifier Lambda posts a "What's New" message to the shift channel on minor/major releases (patches stay silent). - The bot gains an App Home "About" tab showing what it does, the command list, and the current version's notes. release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN events don't start downstream workflows), checks out the deployed commit, and tags + publishes a GitHub Release + invokes the notifier. It assumes a dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the notifier; the account's cfn role gates role creation on that boundary. The manual Version Bump workflow is retired. A CI guard enforces that a CHANGELOG edit is a clean SemVer bump and that the in-package copy matches. * Harden release workflow and regex against CodeQL findings Address three code-scanning alerts on the PR: - Critical (actions/untrusted-checkout): split release.yaml into a read-only `prepare` job that checks out and runs repo code, and a privileged `publish` job (contents:write + OIDC) that never checks out repo code — it tags, releases, and invokes purely through the GitHub and AWS APIs. Also assert head_branch == main. - High x2 (py/polynomial-redos): rewrite the italic and link regexes in markdown_to_mrkdwn with possessive quantifiers and exclusive character classes so they run in linear time on adversarial input. Adds a regression test. * Move release/announce into Deploy workflow to clear CodeQL The workflow_run-triggered release.yaml kept tripping CodeQL's privileged-context rules (untrusted-checkout, then cache-poisoning) — CodeQL distrusts any workflow_run that checks out a ref, regardless of the main-only guarantee, and there is no autofix. Fold the release job into deploy.yaml gated on `needs: deploy`. A push-to-main run is a trusted context, so checking out and running repo code with write/OIDC is safe there. This still gates on deploy success and serializes via the deploy concurrency group, and removes the separate workflow entirely.
2026-06-11 19:41:31 -04:00
The bot also has an **About** page: open the bot in Slack and click its **Home** tab to see what it does, the full command list, and the latest "What's New" (see [Releases & Versioning](#releases--versioning)).
## Slack Commands
| Command | Description |
|---|---|
| `/oncall` | Show this week's schedule |
| `/oncall next` | Show next week's schedule |
| `/oncall pick <date>` | Pick up an available shift — instant if the shift hasn't started; if it's already underway (but not ended) it needs admin approval (see [Late-pickup approval](#late-pickup-approval)) |
| `/oncall drop <date>` | Drop your shift (marks it available) — blocked within 24h of shift start; swap or ask an admin instead |
| `/oncall swap <date> @person` | Request a swap — the other person gets an Accept/Decline DM and the shift only moves once they accept |
| `/oncall register <ext>` | Link your Slack account to your phone extension |
| `/oncall roster` | Show all employees and their link status |
| `/oncall pay` | Show last week's bonus pay summary |
| `/oncall rate` | Show current shift pay rates |
| `/oncall rate default <amount>` | Set the default per-shift rate |
| `/oncall rate <ext> <amount>` | Set a per-person shift rate |
| `/oncall help` | Show help |
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
### Admin Commands
Available to users listed in `admin_users` in the CONFIG record:
| Command | Description |
|---|---|
| `/oncall admin override <date> <ext>` | Assign a shift to an extension |
| `/oncall admin open <date>` | Mark a shift as open |
| `/oncall admin clear <date>` | Remove override (revert to weekly) |
| `/oncall admin roster add <ext> <name>` | Add an employee to the roster |
| `/oncall admin roster remove <ext>` | Remove an employee |
| `/oncall admin roster rename <ext> <name>` | Rename an employee |
| `/oncall admin holiday add <date> <slots> [x<mult>] <label>` | Schedule a holiday day shift (8am-5pm ET) with N slots, optional pay multiplier override (e.g. `x2`), and a label |
| `/oncall admin holiday remove <date>` | Remove a scheduled holiday and its activate/deactivate schedules |
| `/oncall admin holiday list` | List today-and-future scheduled holidays |
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
Dates accept: `today`, `tomorrow`, `monday`-`sunday`, `4/5`, `2026-04-05`
Example: `/oncall admin holiday add 2026-07-04 2 x2 Independence Day` schedules a
2-slot holiday paying 2x. Omitting the `x<mult>` token uses the default
`holiday_multiplier` from CONFIG (1.5). See [Holidays](#holidays) for the full flow.
## Architecture
- **Runtime**: Python 3.12 on AWS Lambda (arm64), seahaven-prod `011934824531`
- **Data**: DynamoDB single-table (`afterhours-shifts`)
- **IaC**: HCP Terraform workspace `afterhours-shift-manager-prod` (containers) plus GitHub Actions `deploy.yaml` (zips). `src/shared` is bundled into each function zip. Terraform does not package `src/`.
- **Slack**: Slack Bolt framework with `/oncall` slash command
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
- **3CX Integration**: Queue routing updated directly via 3CX Queue XAPI
- **Secrets**: AWS Secrets Manager (`afterhours-shift-manager/*`)
### Lambda Functions
| Function | Trigger | Purpose |
|---|---|---|
| `afterhours-shift-manager` | API Gateway (POST /slack/events) | Slack bot — handles `/oncall` commands and interactive buttons |
| `afterhours-weekly-post` | EventBridge (Monday 7am ET) | Posts weekly schedule to Slack, sends pay report email |
| `afterhours-roster-sync` | EventBridge (daily 6am ET) | Syncs employee roster from 3CX |
| `afterhours-roster-api` | API Gateway (PUT /roster, DELETE /roster/{extension}) | Bearer-authenticated roster upsert/delete for the identity processor |
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
| `afterhours-ring-scheduler` | EventBridge (daily 8am ET + weekend 5pm ET) | Updates 3CX queue routing based on who's on shift |
| `afterhours-holiday-router` | EventBridge Scheduler (per-holiday one-off: 8am activate / 5pm deactivate ET) | Repoints the IVR to the holiday queue and sets queue agents for a holiday day shift; reverts at 5pm (see [Holidays](#holidays)) |
| `afterhours-release-notifier` | Skeleton only until tagging exists | Posts a "What's New" announcement to the shift channel |
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
### Project Layout
```
src/
Add changelog-driven releases and App Home tab (#112) * Add changelog-driven releases and App Home tab Version the bot continuously from CHANGELOG.md (the single source of truth for both the version and the staff-readable notes) and surface changes to users in two ways: - A new afterhours-release-notifier Lambda posts a "What's New" message to the shift channel on minor/major releases (patches stay silent). - The bot gains an App Home "About" tab showing what it does, the command list, and the current version's notes. release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN events don't start downstream workflows), checks out the deployed commit, and tags + publishes a GitHub Release + invokes the notifier. It assumes a dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the notifier; the account's cfn role gates role creation on that boundary. The manual Version Bump workflow is retired. A CI guard enforces that a CHANGELOG edit is a clean SemVer bump and that the in-package copy matches. * Harden release workflow and regex against CodeQL findings Address three code-scanning alerts on the PR: - Critical (actions/untrusted-checkout): split release.yaml into a read-only `prepare` job that checks out and runs repo code, and a privileged `publish` job (contents:write + OIDC) that never checks out repo code — it tags, releases, and invokes purely through the GitHub and AWS APIs. Also assert head_branch == main. - High x2 (py/polynomial-redos): rewrite the italic and link regexes in markdown_to_mrkdwn with possessive quantifiers and exclusive character classes so they run in linear time on adversarial input. Adds a regression test. * Move release/announce into Deploy workflow to clear CodeQL The workflow_run-triggered release.yaml kept tripping CodeQL's privileged-context rules (untrusted-checkout, then cache-poisoning) — CodeQL distrusts any workflow_run that checks out a ref, regardless of the main-only guarantee, and there is no autofix. Fold the release job into deploy.yaml gated on `needs: deploy`. A push-to-main run is a trusted context, so checking out and running repo code with write/OIDC is safe there. This still gates on deploy success and serializes via the deploy concurrency group, and removes the separate workflow entirely.
2026-06-11 19:41:31 -04:00
slack-bot/ Slack Bolt Lambda (handler + app); ships CHANGELOG.md for App Home
weekly-post/ Monday schedule + pay post
roster-sync/ Daily 3CX roster sync
roster-api/ HTTP PUT/DELETE /roster for identity hire/offboard
Add changelog-driven releases and App Home tab (#112) * Add changelog-driven releases and App Home tab Version the bot continuously from CHANGELOG.md (the single source of truth for both the version and the staff-readable notes) and surface changes to users in two ways: - A new afterhours-release-notifier Lambda posts a "What's New" message to the shift channel on minor/major releases (patches stay silent). - The bot gains an App Home "About" tab showing what it does, the command list, and the current version's notes. release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN events don't start downstream workflows), checks out the deployed commit, and tags + publishes a GitHub Release + invokes the notifier. It assumes a dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the notifier; the account's cfn role gates role creation on that boundary. The manual Version Bump workflow is retired. A CI guard enforces that a CHANGELOG edit is a clean SemVer bump and that the in-package copy matches. * Harden release workflow and regex against CodeQL findings Address three code-scanning alerts on the PR: - Critical (actions/untrusted-checkout): split release.yaml into a read-only `prepare` job that checks out and runs repo code, and a privileged `publish` job (contents:write + OIDC) that never checks out repo code — it tags, releases, and invokes purely through the GitHub and AWS APIs. Also assert head_branch == main. - High x2 (py/polynomial-redos): rewrite the italic and link regexes in markdown_to_mrkdwn with possessive quantifiers and exclusive character classes so they run in linear time on adversarial input. Adds a regression test. * Move release/announce into Deploy workflow to clear CodeQL The workflow_run-triggered release.yaml kept tripping CodeQL's privileged-context rules (untrusted-checkout, then cache-poisoning) — CodeQL distrusts any workflow_run that checks out a ref, regardless of the main-only guarantee, and there is no autofix. Fold the release job into deploy.yaml gated on `needs: deploy`. A push-to-main run is a trusted context, so checking out and running repo code with write/OIDC is safe there. This still gates on deploy success and serializes via the deploy concurrency group, and removes the separate workflow entirely.
2026-06-11 19:41:31 -04:00
ring-scheduler/ 3CX queue routing updates
holiday-router/ 3CX IVR/queue repoint for holiday day shifts (activate/deactivate)
Add changelog-driven releases and App Home tab (#112) * Add changelog-driven releases and App Home tab Version the bot continuously from CHANGELOG.md (the single source of truth for both the version and the staff-readable notes) and surface changes to users in two ways: - A new afterhours-release-notifier Lambda posts a "What's New" message to the shift channel on minor/major releases (patches stay silent). - The bot gains an App Home "About" tab showing what it does, the command list, and the current version's notes. release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN events don't start downstream workflows), checks out the deployed commit, and tags + publishes a GitHub Release + invokes the notifier. It assumes a dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the notifier; the account's cfn role gates role creation on that boundary. The manual Version Bump workflow is retired. A CI guard enforces that a CHANGELOG edit is a clean SemVer bump and that the in-package copy matches. * Harden release workflow and regex against CodeQL findings Address three code-scanning alerts on the PR: - Critical (actions/untrusted-checkout): split release.yaml into a read-only `prepare` job that checks out and runs repo code, and a privileged `publish` job (contents:write + OIDC) that never checks out repo code — it tags, releases, and invokes purely through the GitHub and AWS APIs. Also assert head_branch == main. - High x2 (py/polynomial-redos): rewrite the italic and link regexes in markdown_to_mrkdwn with possessive quantifiers and exclusive character classes so they run in linear time on adversarial input. Adds a regression test. * Move release/announce into Deploy workflow to clear CodeQL The workflow_run-triggered release.yaml kept tripping CodeQL's privileged-context rules (untrusted-checkout, then cache-poisoning) — CodeQL distrusts any workflow_run that checks out a ref, regardless of the main-only guarantee, and there is no autofix. Fold the release job into deploy.yaml gated on `needs: deploy`. A push-to-main run is a trusted context, so checking out and running repo code with write/OIDC is safe there. This still gates on deploy success and serializes via the deploy concurrency group, and removes the separate workflow entirely.
2026-06-11 19:41:31 -04:00
release-notifier/ Posts release announcements to Slack
shared/ Bundled into each function zip (schedule, blocks, changelog, 3CX client, secrets)
terraform/ HCP Terraform (function skeletons, API, DDB, IAM, schedules)
Add changelog-driven releases and App Home tab (#112) * Add changelog-driven releases and App Home tab Version the bot continuously from CHANGELOG.md (the single source of truth for both the version and the staff-readable notes) and surface changes to users in two ways: - A new afterhours-release-notifier Lambda posts a "What's New" message to the shift channel on minor/major releases (patches stay silent). - The bot gains an App Home "About" tab showing what it does, the command list, and the current version's notes. release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN events don't start downstream workflows), checks out the deployed commit, and tags + publishes a GitHub Release + invokes the notifier. It assumes a dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the notifier; the account's cfn role gates role creation on that boundary. The manual Version Bump workflow is retired. A CI guard enforces that a CHANGELOG edit is a clean SemVer bump and that the in-package copy matches. * Harden release workflow and regex against CodeQL findings Address three code-scanning alerts on the PR: - Critical (actions/untrusted-checkout): split release.yaml into a read-only `prepare` job that checks out and runs repo code, and a privileged `publish` job (contents:write + OIDC) that never checks out repo code — it tags, releases, and invokes purely through the GitHub and AWS APIs. Also assert head_branch == main. - High x2 (py/polynomial-redos): rewrite the italic and link regexes in markdown_to_mrkdwn with possessive quantifiers and exclusive character classes so they run in linear time on adversarial input. Adds a regression test. * Move release/announce into Deploy workflow to clear CodeQL The workflow_run-triggered release.yaml kept tripping CodeQL's privileged-context rules (untrusted-checkout, then cache-poisoning) — CodeQL distrusts any workflow_run that checks out a ref, regardless of the main-only guarantee, and there is no autofix. Fold the release job into deploy.yaml gated on `needs: deploy`. A push-to-main run is a trusted context, so checking out and running repo code with write/OIDC is safe there. This still gates on deploy success and serializes via the deploy concurrency group, and removes the separate workflow entirely.
2026-06-11 19:41:31 -04:00
scripts/ changelog CLI + CI guard + in-package copy sync
tests/ pytest suite (mirrors src/, one dir per Lambda + shared)
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
```
### DynamoDB Schema
Single table with `PK` / `SK` keys:
| PK | SK | Description |
|---|---|---|
| `ROSTER` | `<extension>` | Employee: name, extension, slack_user_id |
| `WEEKLY` | `<DayName>` | Default weekly schedule: extension, name |
| `OVERRIDE` | `<YYYY-MM-DD>` | Date override from pickup/drop (or `OPEN`) |
| `SWAP` | `<YYYY-MM-DD>` | Pending/verified swap request: requester, target, status, `expires_at` (TTL) |
| `HOLIDAY` | `<YYYY-MM-DD>` | Holiday day shift (one per date): `slots` (int), `assignees` (MAP keyed by extension — `{"114": {name, claimed_at}}`), `multiplier` (Decimal, defaults to `CONFIG.holiday_multiplier` = 1.5, overridable per holiday), `label`, `created_at`, `created_by`, `activated` (bool), `schedule_names` (list) |
| `PICKUP_REQUEST` | `<YYYY-MM-DD>[-DAY]#<ext>` | Pending late-pickup awaiting admin approval: `requester_ext`, `requester_name`, `requester_slack`, `shift_type`, `is_holiday`, `status`, `created_at`, `expires_at` (TTL = shift end) |
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
| `SCHEDULE_POST` | `<channel_id>` | Current schedule message timestamp |
| `PAY` | `<YYYY-MM-DD>` | Weekly pay record (Monday date key) |
| `CONFIG` | `CONFIG` | Settings: shift_rate, fallback_extension, admin_users, `ring_group`, `holiday_multiplier` (default holiday pay multiplier, 1.5), `holiday_queue` (3CX queue repointed during holidays, default 802), `ivr_number` (3CX IVR repointed during holidays, default 800), `captured_ivr_routes` (original IVR routes saved at holiday activation, restored at deactivation) |
Weekend day-shift rows use a `-DAY` suffix on the SK (e.g. `OVERRIDE` / `2026-04-05-DAY`). The table has TTL enabled on `expires_at` so abandoned pending swaps and pickup requests self-clean.
Shift priority for any date is **HOLIDAY > OVERRIDE > WEEKLY** — a holiday record wins over a regular override, which wins over the standing weekly schedule.
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
Holiday slot claims are atomic Map updates so concurrent pickers can't oversubscribe:
- **Claim** — `SET assignees.#ext` guarded by `attribute_not_exists(assignees.#ext) AND size(assignees) < :slots`.
- **Release** — `REMOVE assignees.#ext` guarded by `attribute_exists`.
- **Swap** — `REMOVE #from SET #to` guarded by `attribute_exists(#from) AND attribute_not_exists(#to)`.
**Swap flow:** `/oncall swap` writes a `pending` `SWAP` record and DMs the target Accept/Decline buttons; it does **not** reassign the shift. On Accept, the override is written, 3CX is repointed if it's the active shift, and the record is marked `verified`. On Decline (or once the shift has started) the request is dropped and the shift stays with the original owner.
### Holidays
A **holiday** is a single day-only shift (08:00-17:00 ET), one `HOLIDAY` record per
date, that can hold multiple people (`slots`). Admins manage holidays with
`/oncall admin holiday add|remove|list`. A holiday takes priority over a regular
override and the weekly schedule for that date, and pays at its `multiplier`
(per-holiday override, else `CONFIG.holiday_multiplier`, default 1.5). Open slots
show in the schedule with a pickup button; claims, releases, and swaps are atomic
Map updates on the record (see above) so the slot count can't be oversubscribed.
**Holiday-router + Scheduler flow.** When an admin adds a holiday, the slack-bot
creates two **one-off EventBridge Scheduler** schedules for that date —
`holiday-activate-<YYYYMMDD>` at 08:00 ET and `holiday-deactivate-<YYYYMMDD>` at
17:00 ET — whose names are stored on the record's `schedule_names`. Scheduler
assumes `afterhours-shift-manager-holiday-scheduler` to invoke `afterhours-holiday-router`:
- **Activate (08:00):** capture both IVR `ivr_number` (800) routes — key-0 **and**
no-input/timeout — into `CONFIG.captured_ivr_routes` (skipped if they already
point at the holiday queue, so re-runs don't clobber the originals), set
`holiday_queue` (802) agents to the holiday's assignees (or `[fallback_extension]`
= `[100]` when no slots are filled — set **once** here), repoint **both** IVR 800
routes to queue 802, and mark `activated = True`. Idempotent.
- **Deactivate (17:00):** restore both IVR routes from `captured_ivr_routes` (only
routes still pointing at the queue, defensive against manual changes), clear them,
empty queue 802's agents, and mark `activated = False`. Idempotent.
Queue 801 (the daily ring-scheduler queue) is left untouched. If an admin adds a
holiday whose 08:00-17:00 window is already open, the slack-bot **inline-activates**
it immediately (invoking the router) rather than waiting for the 08:00 schedule.
Removing a holiday deletes the record and any outstanding schedules.
### Late-pickup approval
Picking up a shift behaves differently depending on timing, for **both** regular
and holiday shifts:
- **Before the shift starts** — immediate pickup (the prior behaviour, unchanged).
- **After the shift has started but before it ends** (08:00 for a day/holiday shift,
17:00 for a night shift) — the shift is **not** claimed yet. The bot writes a
`pending` `PICKUP_REQUEST` and DMs **every admin** Approve/Deny buttons (mirroring
the verified-swap flow). The **first admin to approve wins** (the claim is
conditional, so a second approval is a safe no-op). On approve, the shift is
claimed (regular: `set_override`; holiday: atomic `claim_holiday_slot`), the
requester and channel are notified, and 3CX is fired if the window is live —
regular shifts call `_update_3cx_routing(picker)` when it's today's active shift;
holidays refresh queue 802's agents to the current assignees. On deny, the request
is cleared and the requester is told.
- **After the shift has ended** — rejected outright; it's too late to pick up.
A slot claimed after the shift has started always needs an admin to approve it.
Merge ring-scheduler-3cx and resolve all open issues (#62) * Add arm64, log retention, and compliance fixes - Set arm64 architecture globally for all Lambda functions - Add explicit CloudWatch log groups with 60-day retention - Add missing WeeklyPostFunctionArn to stack outputs - Add Dependabot assignees for both ecosystems - Add samconfig.toml.example for onboarding * Restructure src/ to per-function layout with shared Layer Move from flat src/ to per-function directories: - src/slack-bot/ — Slack Bolt Lambda handler - src/weekly-post/ — Monday schedule + pay post - src/roster-sync/ — Daily 3CX roster sync - src/shared/ — Lambda Layer with schedule, blocks, three_cx_client Each function has its own requirements.txt and CodeUri. Shared modules are deployed as a SAM Layer (afterhours-shared) importable as `from shared.X import Y`. * Migrate secrets from SSM Parameter Store to Secrets Manager - Slack bot token and signing secret now read from Secrets Manager - 3CX credentials (domain, client-id, client-secret) moved to Secrets Manager under afterhours-shift-manager/3cx-* prefix - Channel ID is now a non-secret CloudFormation parameter (ShiftChannel) - Add shared secrets.py helper for Secrets Manager reads - Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue * Merge ring-scheduler-3cx as 4th Lambda function - Add afterhours-ring-scheduler Lambda with 4 EventBridge rules (daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group routing updates - Extract shared ring_scheduler.py module for direct ring group updates from both the scheduled Lambda and the Slack bot - Replace cross-Lambda invoke with direct update_ring_group() call in the Slack bot — eliminates lambda:InvokeFunction dependency - Use RingGroup API (correct) instead of Queue API (was wrong in the original ring-scheduler repo) - Eliminate YAML config fallback — DynamoDB is the sole schedule source - Add RingGroupNumber CloudFormation parameter * Add schedule post live-update and old post deletion (#40, #41) - Store schedule message timestamp in DynamoDB (SCHEDULE_POST record) - Delete previous week's schedule post before posting the new one - Live-update the schedule post via chat_update after any pick/drop/swap/button-pickup so it always reflects current state * Disallow past shifts and add day/night labels (#43, #42) - Reject /oncall pick and /oncall drop for past dates - Show ephemeral error when stale pickup buttons are clicked - Hide pickup buttons for dates in the past - Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to schedule lines, pickup buttons, and shift change notifications * Add admin slash commands for shift and roster management (#39) - /oncall admin override <date> <ext> — assign a shift - /oncall admin open <date> — mark shift as open - /oncall admin clear <date> — remove override, revert to weekly - /oncall admin roster add/remove/rename — manage roster entries - Admin access gated by admin_users list in DynamoDB CONFIG - Help message shows admin commands for admin users * Update README for merged architecture and new features * Switch from RingGroup API to Queue API at extension 801 The 3CX routing was changed from ring group 800 to queue 801 in a previous PR on ring-scheduler-3cx. Updates all callers and the SAM template parameter default accordingly. * Pass SAM parameter overrides in deploy workflow * Fix review findings: IAM, routing guards, past-date check, roster safety - Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query) - Button pickup: update 3CX for active shift type, not just night - Pick/drop/swap commands: only update 3CX when shift type is active - Swap command: add missing past-date guard - add_roster_entry: reject if extension already exists - Apply ruff formatting * Add error handling to ring scheduler 3CX call * Fix weekend day shift commands and admin 3CX routing - Add _find_employee_shift() to check both day/night on weekends - Drop/swap now correctly find and operate on weekend day shifts - Pick finds first available shift type on weekends - Admin override/open/clear update 3CX for same-day active shifts * Fix dependabot directories and admin weekend shift handling Dependabot now scans per-function requirement directories instead of the repo root. Admin override/open/clear commands accept an optional day/night parameter for weekend day shift management. * Fix weekend day shift active window to 8am-5pm Before midnight-8am on weekends incorrectly reported the day shift as active when the previous night shift is still running. * Show shift type label for both weekend shifts in notifications Night shift notifications on weekends were missing the type label, making them ambiguous. Also fix schedule post text fallback to use this_monday instead of now for the start date. * Extract determine_shift_type into shared layer Eliminates duplicated weekend day/night boundary logic between the ring scheduler and Slack bot Lambdas. * Fix weekly schedule fallback start date Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Include weekend shift type in command confirmations Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com> * Apply ruff formatting to app.py * Only show day/night shift labels on weekends in schedule display Weekday shifts are always night — the label was redundant clutter. * Deduplicate 3CX forwarding payload and add shift type to pick command Extract _update_forwarding helper in ThreeCXClient to share the payload between queue and ring group methods. Add optional day/night argument to /oncall pick so users can target a specific weekend shift. * Consolidate WEEKEND_DAYS and fix weekday pickup button labels Import WEEKEND_DAYS from shared.schedule instead of redefining in blocks.py and weekly-post/app.py. Gate pickup button day/night labels on weekends only, matching all other display surfaces. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com> Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
2026-05-12 19:55:39 -04:00
### Secrets Manager
| Secret | Description |
|---|---|
| `afterhours-shift-manager/slack-bot-token` | Slack bot OAuth token (`xoxb-...`) |
| `afterhours-shift-manager/slack-signing-secret` | Slack app signing secret |
| `afterhours-shift-manager/3cx-domain` | 3CX FQDN (e.g. `company.3cx.us`) |
| `afterhours-shift-manager/3cx-client-id` | 3CX OAuth2 client ID |
| `afterhours-shift-manager/3cx-client-secret` | 3CX OAuth2 client secret |
| `afterhours-shift-manager/roster-api-token` | Bearer token for PUT/DELETE `/roster`. Duplicate the same value into the seahaven-prod secret `paychex-integrations/afterhours-roster-token`. |
### Roster HTTP API
Identity hire/offboard in `paychex-integrations` calls this API. It is a separate Lambda on the same HTTP API as Slack (`POST /slack/events` is unchanged).
| Method | Path | Body | Success |
|---|---|---|---|
| PUT | `/roster` | `{"name","extension","slack_user_id"}` (all required strings) | 200 `{"ok":true}` |
| DELETE | `/roster/{extension}` | none | 204 empty body, including when the row is already gone |
Header: `Authorization: Bearer {token}`. Missing or wrong token is 401. Invalid JSON or fields is 400. A secret-read failure is 503.
Set processor `AFTERHOURS_BASE_URL` to the Terraform output `api_origin` (HCP variable `afterhours_base_url` on `paychex-integrations-prod`). That value is the API origin only. Do not append `/roster`. Flip it at cutover after DynamoDB is copied, not before.
Daily `afterhours-roster-sync` still owns the 3CX `DEFAULT` group at 6am ET: rows absent from that group are deleted. Hire is safe because 3CX create (into `DEFAULT`) happens before the roster PUT. An HTTP-only row that is not in that group will be removed on the next sync. Sync preserves `slack_user_id` on existing rows and does not overwrite a just-created API row's Slack id.
Token rotation is a maintenance-window action. The Lambda caches the token per execution environment. Update `afterhours-shift-manager/roster-api-token` and `paychex-integrations/afterhours-roster-token` together, then recycle `afterhours-roster-api`. Updating only one copy, or recycling environments out of order, causes 401s until both sides match.
## Documentation
The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This project's `afterhours-shift-manager` stack is represented there as a Mermaid subgraph.
- **[AWS Architecture Map](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098)** (Confluence, IT space, page 1540098)
## Deployment
Infrastructure is applied by HCP Terraform workspace `afterhours-shift-manager-prod` (VCS on `main`, working directory `terraform/`, file trigger `terraform/**` only). Function code is shipped by `.github/workflows/deploy.yaml` on push to `main` (`environment: prod`). A terraform-only merge does not run the zip deploy. A mixed app+terraform merge may race the apply; re-run the deploy job if the functions are still stubs.
Manual zip redeploy: Actions → Deploy → Run workflow (`workflow_dispatch`, always prod). Do not `terraform apply` locally to prod.
Add CloudWatch alarm coverage for all functions, table, and HTTP API (#126) * Add CloudWatch alarm coverage for all functions, table, and HTTP API Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests is not emitted at the TableName-only dimension (only TableName+Operation), so the table-level alarm would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Correct DynamoDB alarm docs and drop sign-off wording README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
2026-06-17 14:45:47 -04:00
## Monitoring & Alarms
All CloudWatch alarms are defined in `terraform/alarms.tf` and notify the shared
Add CloudWatch alarm coverage for all functions, table, and HTTP API (#126) * Add CloudWatch alarm coverage for all functions, table, and HTTP API Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests is not emitted at the TableName-only dimension (only TableName+Operation), so the table-level alarm would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Correct DynamoDB alarm docs and drop sign-off wording README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
2026-06-17 14:45:47 -04:00
`site-alerts` SNS topic (→ AWS Chatbot → Slack). None set `OKActions` — recovery
is not paged. Alarm names follow `Lambda-<Metric>-<fn>`
Add CloudWatch alarm coverage for all functions, table, and HTTP API (#126) * Add CloudWatch alarm coverage for all functions, table, and HTTP API Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests is not emitted at the TableName-only dimension (only TableName+Operation), so the table-level alarm would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Correct DynamoDB alarm docs and drop sign-off wording README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
2026-06-17 14:45:47 -04:00
(e.g. `Lambda-Errors-afterhours-ring-scheduler`).
**Lambda alarms** (all seven functions: `afterhours-shift-manager`,
`afterhours-weekly-post`, `afterhours-roster-sync`, `afterhours-roster-api`,
`afterhours-ring-scheduler`, `afterhours-holiday-router`,
`afterhours-release-notifier`):
Add CloudWatch alarm coverage for all functions, table, and HTTP API (#126) * Add CloudWatch alarm coverage for all functions, table, and HTTP API Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests is not emitted at the TableName-only dimension (only TableName+Operation), so the table-level alarm would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Correct DynamoDB alarm docs and drop sign-off wording README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
2026-06-17 14:45:47 -04:00
| Alarm | Metric | Condition | Notes |
|---|---|---|---|
| `Lambda-Errors-<fn>` | `Errors` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
| `Lambda-Duration-<fn>` | `Duration` (Maximum, ms) | `>= ~80% of timeout`, 2 of 3 5-min periods | Thresholds: 24000 ms (30s-timeout fns) / 48000 ms (60s-timeout fns) |
| `Lambda-Throttles-<fn>` | `Throttles` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
**DynamoDB alarm** (`afterhours-shifts` table):
| Alarm | Metric | Condition | Notes |
|---|---|---|---|
| `DDB-ReadThrottle-afterhours-shifts` | `ReadThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
| `DDB-WriteThrottle-afterhours-shifts` | `WriteThrottleEvents` (Sum) | `>= 1` over one 5-min period | `TreatMissingData: notBreaching` |
`ReadThrottleEvents` / `WriteThrottleEvents` are the table-level throttle
signals: AWS/DynamoDB emits them at the `TableName` dimension, so these alarms
transition normally. `ThrottledRequests` and `SystemErrors` are intentionally
**not** alarmed: AWS emits them only at `TableName`+`Operation` granularity, so a
`TableName`-only alarm would sit permanently in `INSUFFICIENT_DATA`.
**API Gateway alarms** (HTTP API, `AWS/ApiGateway` v2 metrics, `ApiId` dimension):
Add CloudWatch alarm coverage for all functions, table, and HTTP API (#126) * Add CloudWatch alarm coverage for all functions, table, and HTTP API Extend the in-template Lambda-<Metric>-<fn> alarm convention to full coverage: - Errors (Sum, >=1/5min) for roster-sync and release-notifier, plus orphan adoption of slack-bot and weekly-post (live alarms of those exact names already exist outside the stack and must be deleted before deploy). - Duration (Maximum, ~80% of timeout, 2-of-3) for all six functions. - Throttles (Sum, >=1/5min) for all six functions. - DynamoDB ThrottledRequests (Sum, >=1/5min) on afterhours-shifts. SystemErrors omitted: AWS emits it only per-Operation, so a TableName-only alarm would sit permanently in INSUFFICIENT_DATA. - API Gateway v2 4xx (>=5), 5xx (>=1), and p99 Latency (~3000ms, 2-of-3) on the implicit ServerlessHttpApi. All alarms page the shared site-alerts SNS topic, no OKActions, TreatMissingData notBreaching. README updated with a Monitoring & Alarms section. Duration and API latency thresholds pending sign-off. * Fix DynamoDB throttle alarm metric: use Read/WriteThrottleEvents ThrottledRequests is not emitted at the TableName-only dimension (only TableName+Operation), so the table-level alarm would sit permanently in INSUFFICIENT_DATA and never fire. Replace with ReadThrottleEvents and WriteThrottleEvents, which AWS/DynamoDB emits at the TableName dimension. * Correct DynamoDB alarm docs and drop sign-off wording README DynamoDB section now lists the alarms actually shipped (DDB-ReadThrottle / DDB-WriteThrottle on Read/WriteThrottleEvents) instead of the stale ThrottledRequests alarm. Thresholds are owner-approved, so remove PENDING ADAM SIGN-OFF wording from template.yaml comments.
2026-06-17 14:45:47 -04:00
| Alarm | Metric | Condition |
|---|---|---|
| `ApiGateway-4xx-<apiId>` | `4xx` (Sum) | `>= 5` over one 5-min period |
| `ApiGateway-5xx-<apiId>` | `5xx` (Sum) | `>= 1` over one 5-min period |
| `ApiGateway-Latency-<apiId>` | `Latency` (p99, ms) | `>= 3000` ms, 2 of 3 5-min periods |
> Duration and API latency thresholds are starting points and may be tuned after
> observing real traffic.
Add changelog-driven releases and App Home tab (#112) * Add changelog-driven releases and App Home tab Version the bot continuously from CHANGELOG.md (the single source of truth for both the version and the staff-readable notes) and surface changes to users in two ways: - A new afterhours-release-notifier Lambda posts a "What's New" message to the shift channel on minor/major releases (patches stay silent). - The bot gains an App Home "About" tab showing what it does, the command list, and the current version's notes. release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN events don't start downstream workflows), checks out the deployed commit, and tags + publishes a GitHub Release + invokes the notifier. It assumes a dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the notifier; the account's cfn role gates role creation on that boundary. The manual Version Bump workflow is retired. A CI guard enforces that a CHANGELOG edit is a clean SemVer bump and that the in-package copy matches. * Harden release workflow and regex against CodeQL findings Address three code-scanning alerts on the PR: - Critical (actions/untrusted-checkout): split release.yaml into a read-only `prepare` job that checks out and runs repo code, and a privileged `publish` job (contents:write + OIDC) that never checks out repo code — it tags, releases, and invokes purely through the GitHub and AWS APIs. Also assert head_branch == main. - High x2 (py/polynomial-redos): rewrite the italic and link regexes in markdown_to_mrkdwn with possessive quantifiers and exclusive character classes so they run in linear time on adversarial input. Adds a regression test. * Move release/announce into Deploy workflow to clear CodeQL The workflow_run-triggered release.yaml kept tripping CodeQL's privileged-context rules (untrusted-checkout, then cache-poisoning) — CodeQL distrusts any workflow_run that checks out a ref, regardless of the main-only guarantee, and there is no autofix. Fold the release job into deploy.yaml gated on `needs: deploy`. A push-to-main run is a trusted context, so checking out and running repo code with write/OIDC is safe there. This still gates on deploy success and serializes via the deploy concurrency group, and removes the separate workflow entirely.
2026-06-11 19:41:31 -04:00
## Releases & Versioning
The bot is versioned with SemVer, driven entirely by **`CHANGELOG.md`**. The
**App Home** tab reads the copy that ships in the slack-bot zip. Run
`python scripts/sync_changelog.py` after editing the root file. Changelog Guard
enforces that the in-package copy matches.
GitHub Releases and git tags are not cut by `deploy.yaml`. `afterhours-release-notifier`
exists as a function skeleton; CD does not invoke it until tagging exists.
Add changelog-driven releases and App Home tab (#112) * Add changelog-driven releases and App Home tab Version the bot continuously from CHANGELOG.md (the single source of truth for both the version and the staff-readable notes) and surface changes to users in two ways: - A new afterhours-release-notifier Lambda posts a "What's New" message to the shift channel on minor/major releases (patches stay silent). - The bot gains an App Home "About" tab showing what it does, the command list, and the current version's notes. release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN events don't start downstream workflows), checks out the deployed commit, and tags + publishes a GitHub Release + invokes the notifier. It assumes a dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the notifier; the account's cfn role gates role creation on that boundary. The manual Version Bump workflow is retired. A CI guard enforces that a CHANGELOG edit is a clean SemVer bump and that the in-package copy matches. * Harden release workflow and regex against CodeQL findings Address three code-scanning alerts on the PR: - Critical (actions/untrusted-checkout): split release.yaml into a read-only `prepare` job that checks out and runs repo code, and a privileged `publish` job (contents:write + OIDC) that never checks out repo code — it tags, releases, and invokes purely through the GitHub and AWS APIs. Also assert head_branch == main. - High x2 (py/polynomial-redos): rewrite the italic and link regexes in markdown_to_mrkdwn with possessive quantifiers and exclusive character classes so they run in linear time on adversarial input. Adds a regression test. * Move release/announce into Deploy workflow to clear CodeQL The workflow_run-triggered release.yaml kept tripping CodeQL's privileged-context rules (untrusted-checkout, then cache-poisoning) — CodeQL distrusts any workflow_run that checks out a ref, regardless of the main-only guarantee, and there is no autofix. Fold the release job into deploy.yaml gated on `needs: deploy`. A push-to-main run is a trusted context, so checking out and running repo code with write/OIDC is safe there. This still gates on deploy success and serializes via the deploy concurrency group, and removes the separate workflow entirely.
2026-06-11 19:41:31 -04:00
Add pytest suite and wire it into CI (#85) (#86) * Add pytest suite and wire it into CI Stands up the first automated tests for the repo (151 tests) and turns on the CI test step. - Lift slack-bot handlers out of create_app() closures to module level so they're unit-testable; create_app is now a thin Bolt-wiring layer. No behavior change (handler entrypoints and create_app signature unchanged). - tests/ mirrors src/: shared layer (schedule, blocks, 3CX client, ring_scheduler, secrets) + all four Lambdas (pay math, drop/swap/pick/ admin/register/rate, pickup button, roster sync, queue scheduler). - All boundaries mocked: DynamoDB/SES/Secrets via moto, 3CX HTTP via responses, Slack via fakes, time via freezegun. No real network/AWS. - pyproject.toml pytest config (pythonpath=src/shared, importlib mode); per-package conftest loads each app.py under a unique name to avoid the four-app.py collision. tests/requirements.txt for test-only deps. - ci.yaml: run-tests: true (reusable workflow auto-installs deps) and lint the tests dir too. - README Testing section. Closes #85 * Add least-privilege permissions block to CI workflow Resolves the CodeQL actions/missing-workflow-permissions alert: the CI workflow now restricts GITHUB_TOKEN to contents: read (it only checks out, lints, and runs tests). * Stop logging extension numbers in 3CX queue updates Resolves 3 high CodeQL py/clear-text-logging-sensitive-data alerts: the queue/ring-group forwarding logs no longer include the routed extension values (closed/holiday/extension). Non-sensitive context (resource id, queue number) is retained.
2026-06-01 19:07:08 -04:00
## Testing
Unit tests use `pytest` with all external boundaries mocked — DynamoDB / SES /
Secrets Manager via `moto`, 3CX HTTP via `responses`, Slack via fakes, and time
via `freezegun`. No test touches the network or real AWS.
```bash
python -m venv .venv && source .venv/bin/activate
pip install -r tests/requirements.txt # test-only deps
pip install -r src/slack-bot/requirements.txt \
-r src/weekly-post/requirements.txt \
-r src/shared/requirements.txt # runtime deps the imports need
pytest
```
Each Lambda has its own `app.py`, so the per-package `conftest.py` loads each one
under a unique module name (importlib mode) to avoid collisions. CI runs the same
suite on every PR via pytest plus `terraform fmt` / `init -backend=false` / `validate`.
Add pytest suite and wire it into CI (#85) (#86) * Add pytest suite and wire it into CI Stands up the first automated tests for the repo (151 tests) and turns on the CI test step. - Lift slack-bot handlers out of create_app() closures to module level so they're unit-testable; create_app is now a thin Bolt-wiring layer. No behavior change (handler entrypoints and create_app signature unchanged). - tests/ mirrors src/: shared layer (schedule, blocks, 3CX client, ring_scheduler, secrets) + all four Lambdas (pay math, drop/swap/pick/ admin/register/rate, pickup button, roster sync, queue scheduler). - All boundaries mocked: DynamoDB/SES/Secrets via moto, 3CX HTTP via responses, Slack via fakes, time via freezegun. No real network/AWS. - pyproject.toml pytest config (pythonpath=src/shared, importlib mode); per-package conftest loads each app.py under a unique name to avoid the four-app.py collision. tests/requirements.txt for test-only deps. - ci.yaml: run-tests: true (reusable workflow auto-installs deps) and lint the tests dir too. - README Testing section. Closes #85 * Add least-privilege permissions block to CI workflow Resolves the CodeQL actions/missing-workflow-permissions alert: the CI workflow now restricts GITHUB_TOKEN to contents: read (it only checks out, lints, and runs tests). * Stop logging extension numbers in 3CX queue updates Resolves 3 high CodeQL py/clear-text-logging-sensitive-data alerts: the queue/ring-group forwarding logs no longer include the routed extension values (closed/holiday/extension). Non-sensitive context (resource id, queue number) is retained.
2026-06-01 19:07:08 -04:00
See [SETUP.md](SETUP.md) for full deployment and Slack app creation instructions.