Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
# Deploy role: githubdeploy-procurement-ingest (seahaven-prod)
OIDC deploy role for this repo's GitHub Actions pipeline in AWS account
`011934824531` (seahaven-prod), us-east-1. Created as part of the migration
from the management account (328440206208); the mgmt role of the same name
stays untouched until decommission as the emergency mgmt deploy path.
## Files
| File | Purpose |
|---|---|
| `trust-policy.json` | OIDC trust: `repo:Sea-Haven-Industries/procurement-ingest:ref:refs/heads/main` only |
feat(api): procurement-api read stack + OpenAPI docs (SHOC reconciliation path) (#127)
* feat(api): add procurement-api stack - read API + OpenAPI docs page
Third CDK stack: API Gateway REST API (IAM SigV4) over both pipelines'
tables, replacing SHOC's retired SyncController cross-account DynamoDB
scan as the reconciliation/backfill path.
- lambdas/api/: handler (healthcheck + docs-token gate + router dispatch),
router (single route table), pagination (opaque cursor, hostile -> 400),
Decimal-safe serialization, wo_repo/po_repo reads. No VendorReplies.
- OpenAPI 3.1 spec as source of truth incl. top-level webhooks section
documenting the outbound SHOC feed; phase-2 write endpoints x-planned
(router answers 501). Self-contained /docs page, no CDN.
- Auth: AWS_IAM on data routes + resource policy scoped to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev on GET/*; /docs and
/openapi.json carve-out is token-gated in the Lambda via shared
web_ui_auth (fail-closed, INFRA-74 posture).
- KMS: explicit Decrypt/DescribeKey on the DynamoDB CMK from SSM
(name-imported table drops the key association - INFRA-104 class).
- Alarms: errors/throttles/duration(p99>=22.5s) + gateway 5xx, ALARM-only
to site-alerts. No access logging in v1 (docs ?token= shim stays out of
logs); cloud_watch_role=False.
- Tests: handler auth-seam + routing + Decimal round-trip; moto cursor
pagination incl. hostile cursors; spec<->router drift gate; bundle
AST pins for the api command; pytest.ini --cov + loader siblings.
- Deploy role: third stack DescribeStacks ARN + procurement-api smoke
invoke ARN (re-run create-deploy-role.sh before merge).
* harden(api): apply sh-security-review findings to procurement-api
Fan-out (6 detectors) + review findings resolved:
Correctness / DoS:
- pagination: require EXACT key-set match (was subset) so a partial/foreign
composite cursor can't reach DynamoDB as an inconsistent ExclusiveStartKey
-> ValidationException -> 500; comments Query now pins the cursor's
work_order_id to the path entity.
- handler: map botocore ValidationException to 400 (defense in depth) so a
crafted cursor can't drive the zero-threshold 5xx alarm.
- web_ui_auth: compare tokens as bytes; a non-ASCII presented token now fails
closed (401) instead of crashing hmac.compare_digest into a 500. Resolves the
pre-existing xfail(strict) follow-up test; hardens the web UIs too.
Docs page:
- typeStr() now escapes the one spec-derived string that reached innerHTML.
- spec inlined into the docs <script> block escapes "<" -> < (</script>
breakout guard); /openapi.json still served byte-faithful.
- Cache-Control: no-store + Referrer-Policy: no-referrer on docs responses so
the ?token= URL stays out of caches/Referer.
- spec-drift test asserts the committed spec carries no "</" / "<!--".
IAM / IaC:
- resource policy enumerates the 7 data GET resources instead of GET/* so a
future GET route can't silently inherit SHOC cross-account reach.
- kms:Decrypt grant gains a kms:ViaService=dynamodb condition.
- stage throttling (50 rps / 100 burst) bounds the unauthenticated /docs blast
radius below the 10k account default.
- corrected the PATCH/POST comment (same-account callers aren't blocked by the
resource policy; 501 handler + absent write grant are the gate).
- documented the RETAIN log-group first-deploy rollback trap and the
resource-policy-needs-redeploy gotcha in-stack.
Mandatory GPT-4.1 cross-family review of the full policy surface: no BLOCK/FIX.
675 tests pass, ruff clean, cdk synth green.
2026-07-23 19:32:20 -04:00
| `permissions-policy.json` | `sts:AssumeRole` on the four `cdk-hnb659fds-*` bootstrap roles, `cloudformation:DescribeStacks` scoped to this repo's stacks + `CDKToolkit` (cd-cdk health check), and `lambda:InvokeFunction` on exactly the three smoke-gated function ARNs (post-deploy smoke gate) |
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
| `create-deploy-role.sh` | Idempotent create-or-update from the two JSON files, profile `seahaven-prod` |
feat(api): procurement-api read stack + OpenAPI docs (SHOC reconciliation path) (#127)
* feat(api): add procurement-api stack - read API + OpenAPI docs page
Third CDK stack: API Gateway REST API (IAM SigV4) over both pipelines'
tables, replacing SHOC's retired SyncController cross-account DynamoDB
scan as the reconciliation/backfill path.
- lambdas/api/: handler (healthcheck + docs-token gate + router dispatch),
router (single route table), pagination (opaque cursor, hostile -> 400),
Decimal-safe serialization, wo_repo/po_repo reads. No VendorReplies.
- OpenAPI 3.1 spec as source of truth incl. top-level webhooks section
documenting the outbound SHOC feed; phase-2 write endpoints x-planned
(router answers 501). Self-contained /docs page, no CDN.
- Auth: AWS_IAM on data routes + resource policy scoped to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev on GET/*; /docs and
/openapi.json carve-out is token-gated in the Lambda via shared
web_ui_auth (fail-closed, INFRA-74 posture).
- KMS: explicit Decrypt/DescribeKey on the DynamoDB CMK from SSM
(name-imported table drops the key association - INFRA-104 class).
- Alarms: errors/throttles/duration(p99>=22.5s) + gateway 5xx, ALARM-only
to site-alerts. No access logging in v1 (docs ?token= shim stays out of
logs); cloud_watch_role=False.
- Tests: handler auth-seam + routing + Decimal round-trip; moto cursor
pagination incl. hostile cursors; spec<->router drift gate; bundle
AST pins for the api command; pytest.ini --cov + loader siblings.
- Deploy role: third stack DescribeStacks ARN + procurement-api smoke
invoke ARN (re-run create-deploy-role.sh before merge).
* harden(api): apply sh-security-review findings to procurement-api
Fan-out (6 detectors) + review findings resolved:
Correctness / DoS:
- pagination: require EXACT key-set match (was subset) so a partial/foreign
composite cursor can't reach DynamoDB as an inconsistent ExclusiveStartKey
-> ValidationException -> 500; comments Query now pins the cursor's
work_order_id to the path entity.
- handler: map botocore ValidationException to 400 (defense in depth) so a
crafted cursor can't drive the zero-threshold 5xx alarm.
- web_ui_auth: compare tokens as bytes; a non-ASCII presented token now fails
closed (401) instead of crashing hmac.compare_digest into a 500. Resolves the
pre-existing xfail(strict) follow-up test; hardens the web UIs too.
Docs page:
- typeStr() now escapes the one spec-derived string that reached innerHTML.
- spec inlined into the docs <script> block escapes "<" -> < (</script>
breakout guard); /openapi.json still served byte-faithful.
- Cache-Control: no-store + Referrer-Policy: no-referrer on docs responses so
the ?token= URL stays out of caches/Referer.
- spec-drift test asserts the committed spec carries no "</" / "<!--".
IAM / IaC:
- resource policy enumerates the 7 data GET resources instead of GET/* so a
future GET route can't silently inherit SHOC cross-account reach.
- kms:Decrypt grant gains a kms:ViaService=dynamodb condition.
- stage throttling (50 rps / 100 burst) bounds the unauthenticated /docs blast
radius below the 10k account default.
- corrected the PATCH/POST comment (same-account callers aren't blocked by the
resource policy; 501 handler + absent write grant are the gate).
- documented the RETAIN log-group first-deploy rollback trap and the
resource-policy-needs-redeploy gotcha in-stack.
Mandatory GPT-4.1 cross-family review of the full policy surface: no BLOCK/FIX.
675 tests pass, ruff clean, cdk synth green.
2026-07-23 19:32:20 -04:00
> **Maintenance note:** `DescribeStacks` is scoped to `stack/po-ingest/*`, `stack/WorkorderIngestStack/*`, `stack/procurement-api/*` (added with the procurement-api stack), and `stack/CDKToolkit/*`. If another stack is ever added to this CDK app, add its ARN pattern here and re-run the review-then-apply flow — otherwise the cd-cdk health check on the new stack will `AccessDenied`.
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
## Why the SmokeInvokeLambda statement exists
`deploy.yaml` runs `scripts/post-deploy-smoke.sh` under the deploy role's own
session, not the assumed `cdk-*` roles. Without `lambda:InvokeFunction` on the
feat(api): procurement-api read stack + OpenAPI docs (SHOC reconciliation path) (#127)
* feat(api): add procurement-api stack - read API + OpenAPI docs page
Third CDK stack: API Gateway REST API (IAM SigV4) over both pipelines'
tables, replacing SHOC's retired SyncController cross-account DynamoDB
scan as the reconciliation/backfill path.
- lambdas/api/: handler (healthcheck + docs-token gate + router dispatch),
router (single route table), pagination (opaque cursor, hostile -> 400),
Decimal-safe serialization, wo_repo/po_repo reads. No VendorReplies.
- OpenAPI 3.1 spec as source of truth incl. top-level webhooks section
documenting the outbound SHOC feed; phase-2 write endpoints x-planned
(router answers 501). Self-contained /docs page, no CDN.
- Auth: AWS_IAM on data routes + resource policy scoped to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev on GET/*; /docs and
/openapi.json carve-out is token-gated in the Lambda via shared
web_ui_auth (fail-closed, INFRA-74 posture).
- KMS: explicit Decrypt/DescribeKey on the DynamoDB CMK from SSM
(name-imported table drops the key association - INFRA-104 class).
- Alarms: errors/throttles/duration(p99>=22.5s) + gateway 5xx, ALARM-only
to site-alerts. No access logging in v1 (docs ?token= shim stays out of
logs); cloud_watch_role=False.
- Tests: handler auth-seam + routing + Decimal round-trip; moto cursor
pagination incl. hostile cursors; spec<->router drift gate; bundle
AST pins for the api command; pytest.ini --cov + loader siblings.
- Deploy role: third stack DescribeStacks ARN + procurement-api smoke
invoke ARN (re-run create-deploy-role.sh before merge).
* harden(api): apply sh-security-review findings to procurement-api
Fan-out (6 detectors) + review findings resolved:
Correctness / DoS:
- pagination: require EXACT key-set match (was subset) so a partial/foreign
composite cursor can't reach DynamoDB as an inconsistent ExclusiveStartKey
-> ValidationException -> 500; comments Query now pins the cursor's
work_order_id to the path entity.
- handler: map botocore ValidationException to 400 (defense in depth) so a
crafted cursor can't drive the zero-threshold 5xx alarm.
- web_ui_auth: compare tokens as bytes; a non-ASCII presented token now fails
closed (401) instead of crashing hmac.compare_digest into a 500. Resolves the
pre-existing xfail(strict) follow-up test; hardens the web UIs too.
Docs page:
- typeStr() now escapes the one spec-derived string that reached innerHTML.
- spec inlined into the docs <script> block escapes "<" -> < (</script>
breakout guard); /openapi.json still served byte-faithful.
- Cache-Control: no-store + Referrer-Policy: no-referrer on docs responses so
the ?token= URL stays out of caches/Referer.
- spec-drift test asserts the committed spec carries no "</" / "<!--".
IAM / IaC:
- resource policy enumerates the 7 data GET resources instead of GET/* so a
future GET route can't silently inherit SHOC cross-account reach.
- kms:Decrypt grant gains a kms:ViaService=dynamodb condition.
- stage throttling (50 rps / 100 burst) bounds the unauthenticated /docs blast
radius below the 10k account default.
- corrected the PATCH/POST comment (same-account callers aren't blocked by the
resource policy; 501 handler + absent write grant are the gate).
- documented the RETAIN log-group first-deploy rollback trap and the
resource-policy-needs-redeploy gotcha in-stack.
Mandatory GPT-4.1 cross-family review of the full policy surface: no BLOCK/FIX.
675 tests pass, ruff clean, cdk synth green.
2026-07-23 19:32:20 -04:00
smoke-gated function ARNs (the two email processors + `procurement-api` ) the
smoke gate hits AccessDenied and every deploy fails closed. The mgmt-era grant
was applied out-of-band and undocumented; keeping it in these reviewed
artifacts closes that gap. Scope it to exactly the named ARNs, never
`Resource: "*"` .
Migrate to seahaven-prod: deploy role, backfill tooling, account-portability fixes (#125)
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).
2026-07-23 17:08:47 -04:00
## Change process
1. Edit the JSON artifacts on a branch; both gates must pass on the exact
files before anything is applied: GPT-4.1 cross-family review
(cross_review.py) and /sh-security-review.
2. Run `./create-deploy-role.sh` (idempotent) with the seahaven-prod profile.
3. Verify: `aws iam simulate-principal-policy` for the bootstrap-role
AssumeRole and both InvokeFunction ARNs, then a real pipeline run.