procurement-ingest/infra/deploy-role/README.md

25 lines
1.1 KiB
Markdown
Raw Normal View History

# Deploy role: githubdeploy-procurement-ingest (seahaven-prod) — PARKED
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125) * feat(migration): prepare stacks and tooling for the seahaven-prod account move Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531) migration. No behavior change in-account; everything here is additive or account-portability hygiene: - infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod (trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped to exactly the two email-processor fn ARNs). Codifies the previously out-of-band smoke-invoke grant. - Table resource policies: make_slack_bot_read_policy in cdk/common.py, applied to purchase-orders, verified-sites, WorkOrders, WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the mgmt-resident seahaven-slack-bot roles read-only cross-account access post-move (bot-side identity grants land in the slack-bot repo). - scripts/migrate_tables.py: dry-run-default backfill tool implementing the plan's per-table semantics (superset overwrite, ingested_at cutoff for WorkOrderComments, backup-gated truncate-and-load for the two site tables) plus a verify subcommand (count parity, spot checks, sticky-Cancelled drift check). - tests/test_resource_policy_helper.py: statement-shape unit tests + static pins that exactly the four bot-read tables carry the policy. - Account-literal fixes: account-agnostic fixture bucket in test_reprocess_contract; runbook/README/po-template-parser account references updated to prod with historical mgmt notes; README gains the account-prerequisites list (imported-by-name dependencies). deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept). Merge is held until migration Phase 0 completes; flipping the AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod deploy. * fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT) * refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS, consumer Lambdas gone); its successor sh-mcp is undeployed and uses same-account DynamoDB access. So no live consumer reads these tables cross-account. Per Adam's call, drop the cross-account grants entirely and re-add correctly-scoped ones if/when sh-mcp deploys to a different account. - Remove the four table resource policies + make_slack_bot_read_policy helper + its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's unit test. Both stacks synth with zero table ResourcePolicy. - scripts/migrate_tables.py hardening (fixes from the sh-security-review fan-out on the destructive backfill tool): * validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit canonical second-precision form) so a malformed cutoff can't silently copy dual-window rows or drop history; * reject `copy --all` up front (must run tables individually, in order, with the stream-drain wait) instead of writing three tables then erroring; * truncate backup gate now also checks recency (<1h) and TableId, not just status+name; * verify requires --cutoff whenever a cutoff table is in scope (else it false-flags dual-window rows as MISSING); * sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item never overwrites a dest-Cancelled PO), and the verify comment no longer overstates what its source-side scan covers; * spot-check all modes (truncate_load keys are verbatim, so key-existence is sound there too). - Deploy role: scope cloudformation:DescribeStacks to this repo's stacks + CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod account); add a drift check warning on unexpected role policies and drop the dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role accepted risk in the deploy-role README. * docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale) * ci: update workflow to use new workflow tag (ruff versioning fix) * fix(migration): address Open SWE review findings on migrate_tables.py - Validate the truncate backup on dry-run as well as --execute so a missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal run (finding f_24a48b8900). - Assert configured keys match the live key schema of both tables before any key projection, turning config/schema drift into a descriptive abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59). - Clarify why key-existence spot-checks are sound for WorkOrderComments: the copy Puts source items verbatim and the sample uses the same cutoff filter, so per-account comment_id divergence never enters the check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
> **PLAT-86 (2026-08-07):** CDK CD is retired. HCP Terraform workspace
> `procurement-ingest-prod` is the sole mutate path. This OIDC role is an
> **unused orphan** retained for a separate IAM-reviewed cleanup (same pattern
> as afi-backup-monitor / front-integrations). Do not use it for deploys. Do
> not recreate a mgmt twin (deleted in PLAT-67).
OIDC deploy role historically used by this repo's GitHub Actions CDK pipeline in
AWS account `011934824531` (seahaven-prod), us-east-1.
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125) * feat(migration): prepare stacks and tooling for the seahaven-prod account move Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531) migration. No behavior change in-account; everything here is additive or account-portability hygiene: - infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod (trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped to exactly the two email-processor fn ARNs). Codifies the previously out-of-band smoke-invoke grant. - Table resource policies: make_slack_bot_read_policy in cdk/common.py, applied to purchase-orders, verified-sites, WorkOrders, WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the mgmt-resident seahaven-slack-bot roles read-only cross-account access post-move (bot-side identity grants land in the slack-bot repo). - scripts/migrate_tables.py: dry-run-default backfill tool implementing the plan's per-table semantics (superset overwrite, ingested_at cutoff for WorkOrderComments, backup-gated truncate-and-load for the two site tables) plus a verify subcommand (count parity, spot checks, sticky-Cancelled drift check). - tests/test_resource_policy_helper.py: statement-shape unit tests + static pins that exactly the four bot-read tables carry the policy. - Account-literal fixes: account-agnostic fixture bucket in test_reprocess_contract; runbook/README/po-template-parser account references updated to prod with historical mgmt notes; README gains the account-prerequisites list (imported-by-name dependencies). deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept). Merge is held until migration Phase 0 completes; flipping the AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod deploy. * fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT) * refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS, consumer Lambdas gone); its successor sh-mcp is undeployed and uses same-account DynamoDB access. So no live consumer reads these tables cross-account. Per Adam's call, drop the cross-account grants entirely and re-add correctly-scoped ones if/when sh-mcp deploys to a different account. - Remove the four table resource policies + make_slack_bot_read_policy helper + its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's unit test. Both stacks synth with zero table ResourcePolicy. - scripts/migrate_tables.py hardening (fixes from the sh-security-review fan-out on the destructive backfill tool): * validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit canonical second-precision form) so a malformed cutoff can't silently copy dual-window rows or drop history; * reject `copy --all` up front (must run tables individually, in order, with the stream-drain wait) instead of writing three tables then erroring; * truncate backup gate now also checks recency (<1h) and TableId, not just status+name; * verify requires --cutoff whenever a cutoff table is in scope (else it false-flags dual-window rows as MISSING); * sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item never overwrites a dest-Cancelled PO), and the verify comment no longer overstates what its source-side scan covers; * spot-check all modes (truncate_load keys are verbatim, so key-existence is sound there too). - Deploy role: scope cloudformation:DescribeStacks to this repo's stacks + CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod account); add a drift check warning on unexpected role policies and drop the dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role accepted risk in the deploy-role README. * docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale) * ci: update workflow to use new workflow tag (ruff versioning fix) * fix(migration): address Open SWE review findings on migrate_tables.py - Validate the truncate backup on dry-run as well as --execute so a missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal run (finding f_24a48b8900). - Assert configured keys match the live key schema of both tables before any key projection, turning config/schema drift into a descriptive abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59). - Clarify why key-existence spot-checks are sound for WorkOrderComments: the copy Puts source items verbatim and the sample uses the same cutoff filter, so per-account comment_id divergence never enters the check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
## Files
| File | Purpose |
|---|---|
| `trust-policy.json` | OIDC trust: `repo:Sea-Haven-Industries/procurement-ingest:ref:refs/heads/main` only |
| `permissions-policy.json` | Historical CDK bootstrap AssumeRole + smoke InvokeFunction grants |
| `create-deploy-role.sh` | Idempotent create-or-update (do not run unless deliberately restoring) |
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125) * feat(migration): prepare stacks and tooling for the seahaven-prod account move Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531) migration. No behavior change in-account; everything here is additive or account-portability hygiene: - infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod (trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped to exactly the two email-processor fn ARNs). Codifies the previously out-of-band smoke-invoke grant. - Table resource policies: make_slack_bot_read_policy in cdk/common.py, applied to purchase-orders, verified-sites, WorkOrders, WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the mgmt-resident seahaven-slack-bot roles read-only cross-account access post-move (bot-side identity grants land in the slack-bot repo). - scripts/migrate_tables.py: dry-run-default backfill tool implementing the plan's per-table semantics (superset overwrite, ingested_at cutoff for WorkOrderComments, backup-gated truncate-and-load for the two site tables) plus a verify subcommand (count parity, spot checks, sticky-Cancelled drift check). - tests/test_resource_policy_helper.py: statement-shape unit tests + static pins that exactly the four bot-read tables carry the policy. - Account-literal fixes: account-agnostic fixture bucket in test_reprocess_contract; runbook/README/po-template-parser account references updated to prod with historical mgmt notes; README gains the account-prerequisites list (imported-by-name dependencies). deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept). Merge is held until migration Phase 0 completes; flipping the AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod deploy. * fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT) * refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS, consumer Lambdas gone); its successor sh-mcp is undeployed and uses same-account DynamoDB access. So no live consumer reads these tables cross-account. Per Adam's call, drop the cross-account grants entirely and re-add correctly-scoped ones if/when sh-mcp deploys to a different account. - Remove the four table resource policies + make_slack_bot_read_policy helper + its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's unit test. Both stacks synth with zero table ResourcePolicy. - scripts/migrate_tables.py hardening (fixes from the sh-security-review fan-out on the destructive backfill tool): * validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit canonical second-precision form) so a malformed cutoff can't silently copy dual-window rows or drop history; * reject `copy --all` up front (must run tables individually, in order, with the stream-drain wait) instead of writing three tables then erroring; * truncate backup gate now also checks recency (<1h) and TableId, not just status+name; * verify requires --cutoff whenever a cutoff table is in scope (else it false-flags dual-window rows as MISSING); * sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item never overwrites a dest-Cancelled PO), and the verify comment no longer overstates what its source-side scan covers; * spot-check all modes (truncate_load keys are verbatim, so key-existence is sound there too). - Deploy role: scope cloudformation:DescribeStacks to this repo's stacks + CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod account); add a drift check warning on unexpected role policies and drop the dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role accepted risk in the deploy-role README. * docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale) * ci: update workflow to use new workflow tag (ruff versioning fix) * fix(migration): address Open SWE review findings on migrate_tables.py - Validate the truncate backup on dry-run as well as --execute so a missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal run (finding f_24a48b8900). - Assert configured keys match the live key schema of both tables before any key projection, turning config/schema drift into a descriptive abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59). - Clarify why key-existence spot-checks are sound for WorkOrderComments: the copy Puts source items verbatim and the sample uses the same cutoff filter, so per-account comment_id divergence never enters the check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
## Retirement follow-up
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125) * feat(migration): prepare stacks and tooling for the seahaven-prod account move Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531) migration. No behavior change in-account; everything here is additive or account-portability hygiene: - infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod (trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped to exactly the two email-processor fn ARNs). Codifies the previously out-of-band smoke-invoke grant. - Table resource policies: make_slack_bot_read_policy in cdk/common.py, applied to purchase-orders, verified-sites, WorkOrders, WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the mgmt-resident seahaven-slack-bot roles read-only cross-account access post-move (bot-side identity grants land in the slack-bot repo). - scripts/migrate_tables.py: dry-run-default backfill tool implementing the plan's per-table semantics (superset overwrite, ingested_at cutoff for WorkOrderComments, backup-gated truncate-and-load for the two site tables) plus a verify subcommand (count parity, spot checks, sticky-Cancelled drift check). - tests/test_resource_policy_helper.py: statement-shape unit tests + static pins that exactly the four bot-read tables carry the policy. - Account-literal fixes: account-agnostic fixture bucket in test_reprocess_contract; runbook/README/po-template-parser account references updated to prod with historical mgmt notes; README gains the account-prerequisites list (imported-by-name dependencies). deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept). Merge is held until migration Phase 0 completes; flipping the AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod deploy. * fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT) * refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS, consumer Lambdas gone); its successor sh-mcp is undeployed and uses same-account DynamoDB access. So no live consumer reads these tables cross-account. Per Adam's call, drop the cross-account grants entirely and re-add correctly-scoped ones if/when sh-mcp deploys to a different account. - Remove the four table resource policies + make_slack_bot_read_policy helper + its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's unit test. Both stacks synth with zero table ResourcePolicy. - scripts/migrate_tables.py hardening (fixes from the sh-security-review fan-out on the destructive backfill tool): * validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit canonical second-precision form) so a malformed cutoff can't silently copy dual-window rows or drop history; * reject `copy --all` up front (must run tables individually, in order, with the stream-drain wait) instead of writing three tables then erroring; * truncate backup gate now also checks recency (<1h) and TableId, not just status+name; * verify requires --cutoff whenever a cutoff table is in scope (else it false-flags dual-window rows as MISSING); * sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item never overwrites a dest-Cancelled PO), and the verify comment no longer overstates what its source-side scan covers; * spot-check all modes (truncate_load keys are verbatim, so key-existence is sound there too). - Deploy role: scope cloudformation:DescribeStacks to this repo's stacks + CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod account); add a drift check warning on unexpected role policies and drop the dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role accepted risk in the deploy-role README. * docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale) * ci: update workflow to use new workflow tag (ruff versioning fix) * fix(migration): address Open SWE review findings on migrate_tables.py - Validate the truncate backup on dry-run as well as --execute so a missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal run (finding f_24a48b8900). - Assert configured keys match the live key schema of both tables before any key projection, turning config/schema drift into a descriptive abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59). - Clarify why key-existence spot-checks are sound for WorkOrderComments: the copy Puts source items verbatim and the sample uses the same cutoff filter, so per-account comment_id divergence never enters the check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
1. Confirm no workflow references `AWS_DEPLOY_ROLE_ARN` / `cd-cdk.yaml`.
2. IAM + security review, then delete the role and drop related substrate entries.
3. Remove this directory in the same cleanup PR.