* ci(workflows): call org reusable CI and Fargate CD
Local CI and the image deploy duplicated the org workflows and still required ci / ci. Pin the callers to those workflows and trust the reusable deploy ref.
* test(ci): probe ruff with an undefined name
* test(ci): remove the undefined-name ruff probe
* ci: retrigger checks after removing the ruff probe
* test(ci): probe ruff with an unused import
* style: apply formatter
* test(ci): remove the autofix probe
---------
Co-authored-by: sea-haven-auto-fix[bot] <332630863+sea-haven-auto-fix[bot]@users.noreply.github.com>
* feat(infra): export vpc_id and public subnet outputs (DEV-289)
Portal Fargate and meals already attach to this VPC. These outputs are
the HCP existing_vpc_id / existing_public_subnet_ids values.
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
* feat(api): add OpenAPI 3.1 and Redocly lint in CI (DEV-289)
Same extends: recommended ruleset and @redocly/cli 2.52.1 as
internal-portal. Covers health, roster, and portal /api/shifts.
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
* fix(api): document 4xx and treat 302 as success in Redocly (DEV-289)
Health and CORS preflight document 400. Recommended only counted 2XX,
so login-style 302s use a shared 2XX-or-3XX rule.
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
* fix(api): fail Redocly on missing 4xx and 2xx/3xx (DEV-289)
Promote operation-4xx-response and the 2xx-or-3xx success rule to error.
Drop unused 400s on health and CORS OPTIONS. Health documents 403 like the
portal. CORS stays in Flask and is not part of the employee contract.
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
Prod is already a human GitHub Release. Remove the SAM tagging path so CHANGELOG.md stays App Home copy and leftover notifier IAM is destroyed on the next apply.
Origins already point at the Fargate hostnames. Drop the HTTP API, eight
functions, zip CD, and Lambda/API Gateway alarms while keeping leftover
Lambda IAM so Paychex can still name weekly-post.
* feat(api): collapse Slack, portal, and jobs onto Fargate (PLAT-216)
Move HTTP and scheduled work onto one always-on Flask task so after-hours
loses Lambda cold start without changing the Cognito or roster contracts.
* fix(portal-api): keep CORS headers on unexpected 500s
Portal SPA error handling needs Access-Control-Allow-Origin even when
DynamoDB or other internals fail, otherwise the browser hides the 500.
* fix(api): retarget holidays per account and ship App Home changelog (PLAT-216)
* fix(iam): list ECS tasks and fail closed on non-prod Paychex (PLAT-216)
* fix(portal-api): serve portal JSON with an explicit JSON content type
* feat(portal-api): add Cognito shift API for the employee portal (DEV-287)
Employees and admins can pick, drop, swap, and manage coverage through
GET/POST/DELETE /api/shifts. Roster PUT accepts optional email for portal
identity. Slack slash commands and App Home admin modals stay in place.
Co-authored-by: adam <adam@seahavenind.com>
* fix(portal-api): preserve shift and deployment invariants (DEV-287)
Co-authored-by: adam <adam@seahavenind.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
* fix(cutover): write Slack secrets into empty Terraform shells
DescribeSecret succeeds on HCP-created shells with no version, so skip-if-exists left roster and Slack tokens unset.
* feat(infra): migrate afterhours to HCP Terraform (PLAT-74)
Replace the mgmt SAM stack with a prod-only HCP workspace, in-repo hcptf IAM, stub Lambdas, and zip CD on push to main.
* fix(cutover): retry DDB unprocessed items and skip past at() holidays
Unprocessed BatchWriteItem rows and leftover past at() schedules would drop roster data or abort holiday recreation during prod cutover.
* feat(roster): add Bearer PUT/DELETE roster API
Identity hire needs to write Slack IDs onto roster rows without a stale
daily 3CX sync clearing them, using the existing HTTP client contract.
* fix(roster): strip Secrets Manager token whitespace
A file:// secret commonly includes a trailing newline, so compare_digest
must strip the cached value the same way it strips the Bearer header.
* ci(dependency-review): set explicit read-only GITHUB_TOKEN permissions
Resolves code-scanning alert 11 (actions/missing-workflow-permissions).
The callable workflow only needs contents: read.
* fix(logging): remove taint-flagged values from 3CX and roster-sync logs
Resolves code-scanning alerts 12-15 (py/clear-text-logging-sensitive-data).
CodeQL taints the 3CX response dicts via the Secrets Manager-sourced
domain in the request URL, so entity IDs subscripted from those
responses (ivr_id, resource_id, queue_id) and the roster result dict
trip the query. None of the flagged values are secrets, but the log
lines are rewritten so the pattern cannot trip: entity IDs are dropped
in favor of the untainted destination DNs, and the roster summary logs
counts instead of the member-derived dict (which also keeps employee
names out of the logs).
* fix: update ci workflow SHA to latest version
* fix(logging): drop employee-derived DNs from forwarding log
Resolves new code-scanning alerts 16/17. The closed/holiday DNs added
in the previous commit derive from roster employee lookups in the
Slack bot, so CodeQL classifies them as private data. Log only the
resource type; ring_scheduler already logs the queue number.
* Add changelog-driven releases and App Home tab
Version the bot continuously from CHANGELOG.md (the single source of
truth for both the version and the staff-readable notes) and surface
changes to users in two ways:
- A new afterhours-release-notifier Lambda posts a "What's New" message
to the shift channel on minor/major releases (patches stay silent).
- The bot gains an App Home "About" tab showing what it does, the
command list, and the current version's notes.
release.yaml runs on Deploy success (not release:published — GITHUB_TOKEN
events don't start downstream workflows), checks out the deployed commit,
and tags + publishes a GitHub Release + invokes the notifier. It assumes a
dedicated, boundary-carrying OIDC role scoped to InvokeFunction on the
notifier; the account's cfn role gates role creation on that boundary.
The manual Version Bump workflow is retired. A CI guard enforces that a
CHANGELOG edit is a clean SemVer bump and that the in-package copy matches.
* Harden release workflow and regex against CodeQL findings
Address three code-scanning alerts on the PR:
- Critical (actions/untrusted-checkout): split release.yaml into a
read-only `prepare` job that checks out and runs repo code, and a
privileged `publish` job (contents:write + OIDC) that never checks out
repo code — it tags, releases, and invokes purely through the GitHub
and AWS APIs. Also assert head_branch == main.
- High x2 (py/polynomial-redos): rewrite the italic and link regexes in
markdown_to_mrkdwn with possessive quantifiers and exclusive character
classes so they run in linear time on adversarial input. Adds a
regression test.
* Move release/announce into Deploy workflow to clear CodeQL
The workflow_run-triggered release.yaml kept tripping CodeQL's
privileged-context rules (untrusted-checkout, then cache-poisoning) —
CodeQL distrusts any workflow_run that checks out a ref, regardless of
the main-only guarantee, and there is no autofix.
Fold the release job into deploy.yaml gated on `needs: deploy`. A
push-to-main run is a trusted context, so checking out and running repo
code with write/OIDC is safe there. This still gates on deploy success
and serializes via the deploy concurrency group, and removes the
separate workflow entirely.
* Add dependency-review caller workflow
Add a pull_request-triggered caller that invokes the org-level
callable-dependency-review workflow to scan dependency changes and
fail on high-severity advisories.
* chore: retrigger checks
* chore: retrigger dep review (post-fix)
We now apply semantic version tags deliberately (see v1.7.19–v1.9.2), so the
daily auto-bump is no longer wanted.
- Drop the `schedule:` cron (and the now-unneeded DST guard) — the workflow
runs only on `workflow_dispatch`, with patch/minor/major options.
- Remove the Slack notification entirely: the "Update changelog canvas" and
"Post to Slack" steps (and the PR/bullet collection that fed them) are gone,
along with their SLACK_* secret usage.
- Keep the core behavior: compute the next version from the latest tag + chosen
bump and push an annotated tag.
Renames the workflow "Daily Version Bump" -> "Version Bump".
* Add pytest suite and wire it into CI
Stands up the first automated tests for the repo (151 tests) and turns on
the CI test step.
- Lift slack-bot handlers out of create_app() closures to module level so
they're unit-testable; create_app is now a thin Bolt-wiring layer. No
behavior change (handler entrypoints and create_app signature unchanged).
- tests/ mirrors src/: shared layer (schedule, blocks, 3CX client,
ring_scheduler, secrets) + all four Lambdas (pay math, drop/swap/pick/
admin/register/rate, pickup button, roster sync, queue scheduler).
- All boundaries mocked: DynamoDB/SES/Secrets via moto, 3CX HTTP via
responses, Slack via fakes, time via freezegun. No real network/AWS.
- pyproject.toml pytest config (pythonpath=src/shared, importlib mode);
per-package conftest loads each app.py under a unique name to avoid the
four-app.py collision. tests/requirements.txt for test-only deps.
- ci.yaml: run-tests: true (reusable workflow auto-installs deps) and lint
the tests dir too.
- README Testing section.
Closes#85
* Add least-privilege permissions block to CI workflow
Resolves the CodeQL actions/missing-workflow-permissions alert: the CI
workflow now restricts GITHUB_TOKEN to contents: read (it only checks out,
lints, and runs tests).
* Stop logging extension numbers in 3CX queue updates
Resolves 3 high CodeQL py/clear-text-logging-sensitive-data alerts: the
queue/ring-group forwarding logs no longer include the routed extension
values (closed/holiday/extension). Non-sensitive context (resource id,
queue number) is retained.
* Fix shared layer packaging — remove python/ wrapper that caused double nesting
SAM BuildMethod: python3.12 wraps layer content in python/ during build.
The source had an extra python/ directory, resulting in the shared package
landing at python/python/shared/ instead of python/shared/. All 4 Lambdas
are failing with ImportModuleError since the PR #62 merge.
* Update CI source-dirs to match new shared layer path
* Add arm64, log retention, and compliance fixes
- Set arm64 architecture globally for all Lambda functions
- Add explicit CloudWatch log groups with 60-day retention
- Add missing WeeklyPostFunctionArn to stack outputs
- Add Dependabot assignees for both ecosystems
- Add samconfig.toml.example for onboarding
* Restructure src/ to per-function layout with shared Layer
Move from flat src/ to per-function directories:
- src/slack-bot/ — Slack Bolt Lambda handler
- src/weekly-post/ — Monday schedule + pay post
- src/roster-sync/ — Daily 3CX roster sync
- src/shared/ — Lambda Layer with schedule, blocks, three_cx_client
Each function has its own requirements.txt and CodeUri. Shared
modules are deployed as a SAM Layer (afterhours-shared) importable
as `from shared.X import Y`.
* Migrate secrets from SSM Parameter Store to Secrets Manager
- Slack bot token and signing secret now read from Secrets Manager
- 3CX credentials (domain, client-id, client-secret) moved to
Secrets Manager under afterhours-shift-manager/3cx-* prefix
- Channel ID is now a non-secret CloudFormation parameter (ShiftChannel)
- Add shared secrets.py helper for Secrets Manager reads
- Remove SSM and KMS IAM policies, add secretsmanager:GetSecretValue
* Merge ring-scheduler-3cx as 4th Lambda function
- Add afterhours-ring-scheduler Lambda with 4 EventBridge rules
(daily 8am EST/EDT + weekend 5pm EST/EDT) for 3CX ring group
routing updates
- Extract shared ring_scheduler.py module for direct ring group
updates from both the scheduled Lambda and the Slack bot
- Replace cross-Lambda invoke with direct update_ring_group() call
in the Slack bot — eliminates lambda:InvokeFunction dependency
- Use RingGroup API (correct) instead of Queue API (was wrong in
the original ring-scheduler repo)
- Eliminate YAML config fallback — DynamoDB is the sole schedule
source
- Add RingGroupNumber CloudFormation parameter
* Add schedule post live-update and old post deletion (#40, #41)
- Store schedule message timestamp in DynamoDB (SCHEDULE_POST record)
- Delete previous week's schedule post before posting the new one
- Live-update the schedule post via chat_update after any
pick/drop/swap/button-pickup so it always reflects current state
* Disallow past shifts and add day/night labels (#43, #42)
- Reject /oncall pick and /oncall drop for past dates
- Show ephemeral error when stale pickup buttons are clicked
- Hide pickup buttons for dates in the past
- Add explicit "Day (8am-5pm)" and "Night (5pm-8am)" labels to
schedule lines, pickup buttons, and shift change notifications
* Add admin slash commands for shift and roster management (#39)
- /oncall admin override <date> <ext> — assign a shift
- /oncall admin open <date> — mark shift as open
- /oncall admin clear <date> — remove override, revert to weekly
- /oncall admin roster add/remove/rename — manage roster entries
- Admin access gated by admin_users list in DynamoDB CONFIG
- Help message shows admin commands for admin users
* Update README for merged architecture and new features
* Switch from RingGroup API to Queue API at extension 801
The 3CX routing was changed from ring group 800 to queue 801 in a
previous PR on ring-scheduler-3cx. Updates all callers and the SAM
template parameter default accordingly.
* Pass SAM parameter overrides in deploy workflow
* Fix review findings: IAM, routing guards, past-date check, roster safety
- Ring scheduler: use DynamoDBCrudPolicy (resolve_shift needs Query)
- Button pickup: update 3CX for active shift type, not just night
- Pick/drop/swap commands: only update 3CX when shift type is active
- Swap command: add missing past-date guard
- add_roster_entry: reject if extension already exists
- Apply ruff formatting
* Add error handling to ring scheduler 3CX call
* Fix weekend day shift commands and admin 3CX routing
- Add _find_employee_shift() to check both day/night on weekends
- Drop/swap now correctly find and operate on weekend day shifts
- Pick finds first available shift type on weekends
- Admin override/open/clear update 3CX for same-day active shifts
* Fix dependabot directories and admin weekend shift handling
Dependabot now scans per-function requirement directories instead
of the repo root. Admin override/open/clear commands accept an
optional day/night parameter for weekend day shift management.
* Fix weekend day shift active window to 8am-5pm
Before midnight-8am on weekends incorrectly reported the day shift
as active when the previous night shift is still running.
* Show shift type label for both weekend shifts in notifications
Night shift notifications on weekends were missing the type label,
making them ambiguous. Also fix schedule post text fallback to use
this_monday instead of now for the start date.
* Extract determine_shift_type into shared layer
Eliminates duplicated weekend day/night boundary logic between
the ring scheduler and Slack bot Lambdas.
* Fix weekly schedule fallback start date
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
* Include weekend shift type in command confirmations
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
* Apply ruff formatting to app.py
* Only show day/night shift labels on weekends in schedule display
Weekday shifts are always night — the label was redundant clutter.
* Deduplicate 3CX forwarding payload and add shift type to pick command
Extract _update_forwarding helper in ThreeCXClient to share the
payload between queue and ring group methods. Add optional day/night
argument to /oncall pick so users can target a specific weekend shift.
* Consolidate WEEKEND_DAYS and fix weekday pickup button labels
Import WEEKEND_DAYS from shared.schedule instead of redefining in
blocks.py and weekly-post/app.py. Gate pickup button day/night
labels on weekends only, matching all other display surfaces.
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Adam Moussa <amoussa1229@users.noreply.github.com>
The date comparison used git's author date in its stored timezone
(e.g. -04:00) against GitHub API mergedAt values in UTC (Z suffix).
Lexicographic string comparison across different timezone formats
caused every PR from the tagged commit's day to be re-included in
subsequent versions. Normalize to UTC with format-local so both
sides match.
Also replace the hardcoded canvas URL with the CANVAS_ID env var
already available in the step.