Writers already use kebab tables. Allow HCP to destroy PascalCase
WorkOrders / WorkOrderComments after the ≥24h kebab soak. Docs scrub
to kebab as the live physical names; GitHub #24 noted superseded.
Cutover apply updated Lambda env/IAM but ESM recreate failed until the
lambda execution boundary included kebab table ARNs. Import the
out-of-band Enabled mappings so HCP state matches live.
* feat(infra): retarget WO writers and ESMs to kebab Dynamo tables
Point Lambda env, IAM, emitter streams, and DDB alarms at work-orders /
work-order-comments for the PLAT-11 freeze cutover. Legacy PascalCase
tables remain in state until the decommission soak.
* style(infra): terraform fmt wo_ddb alarm map alignment
Tables were created in seahaven-prod after hcptf-procurement-ingest lacked
CreateTable on kebab ARNs; import blocks bring work-orders and
work-order-comments under Terraform ownership.
* feat(infra): add kebab WO tables and PLAT-11 cutover tooling
Create empty work-orders/work-order-comments under Terraform while writers
stay on PascalCase; make emitter table classification env-driven and add
same-account migrate/verify plus cutover runbook.
* style(test): ruff-format shoc emitter envelope tests
* chore(infra): remove cdk tree after hcp cutover
Delete retired CDK sources, retarget bundle/principal contract tests to
Terraform packaging, disable CDK synth in CI, and scrub deploy-adjacent docs.
* fix(test): restore exact SHOC principal pin in terraform
Pin shoc_consumer_role_arn's Terraform default and example to the trusted
ARN, and require grant sites to consume local.shoc_consumer_role_arn only.
* fix(shoc-emitter): skip empty-text comment webhooks
Stop delivering #nocomment# WorkOrderComments rows as comment_added events
so they no longer 400-park in the rejected queue.
* style(shoc-emitter): ruff-format blank-comment handler test
* feat(webhook): ACTIVATE SHOC WO webhook emitter (enabled=True)
The deliberate one-line activation flip (plan Phase 3). Turns on both
DynamoDB stream event-source mappings for workorder-shoc-emitter, which
shipped dark (enabled=False) in PR-2. LATEST start position => the live
feed begins at deploy, no historical flood; SHOC loads history via the
procurement read API first.
DO NOT MERGE until: (1) PR-2 (feat/shoc-wo-webhook) is merged to main;
(2) SHOC's receiver passes the shared HMAC test vectors
(docs/shoc-webhook-test-vectors.json) at the target endpoint; (3) the
cross-account secret read from shoc-backend-dev is confirmed working.
Draft, gated on Luby.
* docs(webhook): stamp activation date and correct backfill wording
* docs(webhook): adjust activation comment date to 2026-07-30.
Signed-off-by: Adam Moussa <adam@seahavenind.com>
* docs(readme): mark SHOC webhook emitter active as of 2026-07-30
---------
Signed-off-by: Adam Moussa <adam@seahavenind.com>
* iac(access): temporary shoc-assessment-dynamo-reader role for Luby initial assessment
30-day, read-only (GetItem/Query/Scan/DescribeTable) cross-account role in
seahaven-prod trusting seahaven-external-dev, trust-policy hard expiry
2026-08-23. Steady state remains procurement-api + webhook; teardown script
included.
* harden(access): resolve sh-security-review findings on assessment reader
C1 (confirmed medium): DateLessThan expiry condition duplicated into both
permissions statements so in-flight sessions die at the deadline, not +1h.
C2 (confirmed medium): teardown now strips ALL inline/attached policies and
instance profiles before DeleteRole (kill-switch semantics restored,
idempotent), emergency-revocation section added to README.
Cheap hardenings: CDPATH-immune SCRIPT_DIR, MaxSessionDuration re-asserted
on update path, data-handling expectations documented.
* harden(access): address Open SWE review on #141
- Trust now requires ArnLike aws:PrincipalArn on the Identity Center role
path (human SSO sessions only; string condition survives permission-set
reprovisioning, unlike a role-ARN Principal pin)
- README extend instructions cover BOTH expiry sites (trust + permissions)
- Create script logs caller ARN + timestamp and validates the policy's KMS
ARN against SSM /seahaven/dynamodb/cmk-arn before applying
* feat(api): custom domain procurement-api.seahaven.com for the read API
Stacked on feat/shoc-wo-webhook. Gives the SHOC-facing read API a stable,
brandable endpoint instead of the opaque execute-api URL.
- procurement_api_stack.py: REGIONAL API Gateway DomainName (TLS 1.2) +
empty base-path mapping to the prod stage, so callers hit
https://procurement-api.seahaven.com/work-orders (no /prod segment). The
ACM cert ARN is read from SSM (/procurement-api/custom-domain/certificate-arn)
via value_for_string_parameter, because the seahaven.com zone is in the
mgmt account (cross-account DNS) and the cert is issued out of band. Outputs
expose the regional alias target + hosted-zone id for the mgmt A-record.
- scripts/setup_procurement_api_domain.sh: idempotent two-step runbook
(cert: request + mgmt-zone validation + wait + SSM; alias: post-deploy
A-record from stack outputs). Verifies both account identities.
- handler._base_url: omit the /{stage} segment for a custom-domain request
(the base-path mapping serves the stage at the root) so the docs never
advertise a broken server URL; execute-api hosts keep /{stage}.
- openapi.json: custom domain added as servers[0] (recommended), execute-api
kept as the direct fallback + the per-request injection target.
No IAM/auth/policy change (same API id + resource policy), so the SigV4
surface and the mandatory cross-family gates are unaffected. 751 pytest,
ruff, cdk synth, redocly lint all green.
* fix(api): use .endswith('.amazonaws.com') instead of substring check for execute-api detection
The prior '.execute-api.' in domain substring check is fragile and
triggers CodeQL incomplete-sanitization warnings. All API Gateway default
domains end with .amazonaws.com, so a suffix check is more precise and
also silences the false-positive alert.
Refs: https://github.com/Sea-Haven-Industries/procurement-ingest/security/code-scanning/6
* docs(webhook): revise SHOC webhook contract and plan for post-migration reality
Branch re-cut on main 2026-07-23 (old base carried stale PR #99 commits).
Contract Rev 2026-07-23:
- Producer account corrected: seahaven-prod (011934824531); mgmt frozen
- Reconciliation backstop is the new procurement read API, not SyncController
- wo_status "unknown" is real; SHOC must map it (checklist item added)
- write_origin forward-compat note for phase-2 write-back echo suppression
- SyncVendorReplies retirement flagged (dead table, no vendor_reply event)
Plan updates:
- Account gate: seahaven-prod only; never enable streams on mgmt tables
- Emitter ships DARK (ESMs enabled=False); activation is a deliberate flip
after the SHOC receiver passes shared HMAC vectors
- Post-refactor conventions: common.py helpers, bundle-consistency AST pins,
pytest.ini --cov additions, consolidated test roots
- Dedicated-CMK rationale, secret-ARN handooff step, consumer audit refreshed
(slack-bot decommissioned), enum golden test, write_origin skip-branch test
* feat(webhook): SHOC WO webhook emitter — dark-ship streams, HMAC secret + rotation
Implements docs/shoc-webhook-plan.md Phases 1-5 (PR-2 of the SHOC
call-and-be-called effort). Everything ships DARK: both DynamoDB event
source mappings deploy enabled=False; activation is a deliberate
one-line follow-up PR gated on the SHOC receiver passing the shared
HMAC test vectors.
- Streams: NEW_AND_OLD_IMAGES on WorkOrders + WorkOrderComments
(in-place update, RETAIN + logical IDs untouched; no existing
consumers — verified live, neither table had a stream).
- workorder-shoc-emitter (Py3.12/ARM64): stream -> envelope ->
HMAC-signed POST per docs/shoc-webhook-contract.md; strict per-shard
ordering (parallelization 1, bisect off, retry until 24h age,
ReportBatchItemFailures); 429/5xx/timeout block the shard in order,
other 4xx park to workorder-shoc-emitter-rejected; ESM failures ->
workorder-shoc-emitter-failures (metadata; replay rebuilds from
DynamoDB). Echo guard skips write_origin=shoc-write-api.
- Secret workorder-ingest/shoc-webhook-hmac on a dedicated CMK
(alias workorder-ingest-shoc-webhook-kms); cross-account
GetSecretValue/DescribeSecret + kms:Decrypt granted to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev. RemovalPolicy
DESTROY deliberately (machine-generated material; avoids the
fixed-name RETAIN-orphan deadlock).
- workorder-shoc-hmac-rotator: 30-day rotation, dual-key overlap,
64-hex keys, kid = UTC %Y-%m-%dT%H.
- Alarms (ALARM-only -> site-alerts): emitter errors/throttles/
duration + iterator-age (>=10 min) + failures/rejected queue
depth; rotator standard trio.
- scripts/replay_shoc_webhooks.py: dry-run-default operator replay
(rebuilds from tables, replay:true envelopes).
- Tests: 742 passing, 85.56% aggregate; golden HMAC vectors shared
with SHOC in docs/shoc-webhook-test-vectors.json (emitter + replay
signing pinned to identical vectors); bundle-consistency AST pins
for both new bundles.
- README: WO stack + webhook feed section, alarm table, runbooks;
removed stale seahaven-slack-bot consumer references.
* fix(webhook): kms:ViaService pins, https-only delivery, cross-account principal CI pin
GPT-4.1 cross-family review of the policy surface (no BLOCK): FIX applied
to the cross-account shoc-backend-dev Decrypt statement and both Lambda
role KMS grants (the key is only ever used via Secrets Manager); its
invariant-enforcement QUESTION answered durably with
tests/test_cross_account_principal_pin.py (any new foreign IAM principal
in cdk/ fails CI). Scanner mediums fixed: delivery.py and the replay
script now refuse non-https URLs (urllib follows file:// and http://).
SQS metadata-action and dynamodb:ListStreams NITs skipped: standard CDK
grant shapes; ListStreams has no resource-level scoping. The 4 gitleaks
HIGHs on docs/shoc-webhook-test-vectors.json are deliberate non-secrets
(shared receiver-verification vectors) suppressed machine-level with
justification.
* harden(webhook): resolve /sh-security-review findings (1 confirmed medium + cheap fixes)
High-recall detector fan-out (injection/authz/secrets-crypto/iac-iam/logic)
+ proof-or-kill verifier. Gate PASSES: 1 confirmed medium, 0 confirmed
critical/high. Confirmed finding fixed; several unverified-but-cheap
hardenings applied since the emitter ships dark and activation is weeks out.
- CONFIRMED medium (confused deputy): the rotation Lambda's generated
invoke permission for secretsmanager.amazonaws.com carried no
SourceAccount/SourceArn, so any account's Secrets Manager could invoke
the rotator. Patched the generated CfnPermission in place (a second
permission would be additive, not restrictive) to pin account + this
secret ARN.
- delivery + replay: refuse to follow receiver 3xx redirects (no-redirect
opener) so live X-SH-* auth headers can't be forwarded to a
receiver-chosen Location and an http:// Location can't slip past the
https guard. Fixed the "unfollowed 3xx" comment that was factually wrong.
- delivery: classify 401/403 as retryable (invalidate key cache + retry in
order) instead of parking -- transient auth failures (rotation outran the
TTL cache, clock skew) are availability events, not contract bugs.
- envelope: build_event now genuinely total (guarded eventID /
ApproximateCreationDateTime subscripts) per its own never-raise contract.
- handler: catch-all so an unexpected per-record error (e.g. SQS park
failure) reports only that record instead of failing the whole batch
(which would re-deliver every earlier success for 24h); per-invocation
emit/skip batch summary so a systemic silent drop is queryable/alarmable.
- rotator: narrow the AWSCURRENT-read except to ResourceNotFound/JSONDecode
(transient SM/KMS errors re-raise so the overlap key isn't silently
dropped); kid uniqueness checked against ALL retained kids with a random
suffix on collision (never reissue a kid for a different secret).
- contract: skeleton-upsert required on ANY unknown work_order_id (not just
comment-before-create) + monotonicity guard (ignore older updated_at), so
a parked created or an out-of-order replay can't corrupt receiver state.
Unverified/refuted findings left as-is with rationale: the two "high" logic
claims (whole-batch crash triggers, ordering violation) were refuted on
reachability (real stream records carry required fields; persistence writes
strings only; full-state idempotent upsert absorbs the ordering gap). Signed
kid/version binding (AUTHZ-002) declined: coordinated contract change, not
cheap, no exploit with one algorithm/key.
* fix(webhook): drop kid from rotator test_ok log (CodeQL clear-text-logging FP)
GHAS CodeQL flagged py/clear-text-logging-sensitive-data (high) at
_test_secret's success log because head["kid"] is subscripted from the
same parsed-secret dict that holds head["secret"] — the taint tracker
can't tell the non-secret key id from the secret. The secret value is
never logged. Rather than dismiss the alert (fragile; re-alerts on line
moves), remove the flow: kid is already logged at stage time in
_create_secret and version_id correlates the steps, so the test_ok log
keeps only event + version_id. Also hardens against a future edit that
swaps the logged field.
* feat(api): Add @redocly/cli as a dev dependency
Signed-off-by: Adam Moussa <adam@seahavenind.com>
* feat(api): Add Redocly configuration file with custom rules
Signed-off-by: Adam Moussa <adam@seahavenind.com>
* chore(api): Redocly lint config + bring openapi.json into compliance
redocly.yaml from the Redocly guidelines builder, with three generated
rules corrected: response-contains-property had the status codes as the
required body fields (intent was the Error schema's top-level 'error';
403 exempt since API Gateway emits AWS's {message} shape, 501 not 503);
operation-4xx-problem-details-rfc7807 off (adopting RFC 7807 would be a
runtime + SHOC-contract change, decided against); the two inert casing
rules (parameter names, schema properties) removed because both name
sets are contract-pinned (gateway resource paths, DynamoDB items).
Spec changes, no runtime impact: operationIds renamed to method-prefixed
kebab-case (get-work-orders, post-work-order-comment, ...); tags added to
all 15 operations + root tags object (groups the Redoc sidebar); examples
on all six parameters; license field; server description punctuation; two
descriptions reworded to start capitalized. Real linter catches fixed:
the two x-planned ops were missing their {workOrderId} path parameter
and any 4xx response (403 added - true today, gateway rejects unsigned).
.redocly.lint-ignore.yaml pins the six deliberate exceptions: webhook
keys are the shipped SHOC contract event names (not renameable), and the
x-planned ops answer only 501 (no 2xx to document).
package.json: npm run lint:api. Verified: lint 0 errors, 675 pytest,
headless-Chrome render of the tagged docs page.
* feat(api): SHOC design-system theme for /docs (vendored fonts)
Themes the Redoc page with the canonical SHOC token set: Montserrat 600
headings / DM Sans body / JetBrains Mono code, primary #1c75bc, navy
#262262 sidebar text + right panel, #f9fafb background, 244px sidebar.
sortRequiredPropsFirst on; 200 responses pre-expanded.
Fonts ship as lambdas/api/fonts.css (latin woff2 subsets from
@fontsource 5.3.0, embedded as data URIs, ~90KB) and inline via a new
__FONTS_CSS__ placeholder with the same </style breakout guard --
the offline single-response invariant holds, nothing fetches Google
Fonts (test-pinned). Bundling cp + bundle-consistency pin + spec-drift
asset checks extended.
Verified: headless-Chrome render (theme + fonts applied), ruff, 675
pytest, cdk synth + staged-asset check.
* feat(api): SHOC gradient topbar on /docs
64px fixed header with the SHOC shell gradient token (#1b1f52 ->
#1c4f8f -> #1c75bc), Sea Haven wordmark in Montserrat 600, page name
right-aligned in DM Sans. Redoc's scrollYOffset: 64 keeps the sticky
sidebar and anchor scrolling clear of the fixed bar. Verified via
headless-Chrome render.
* style(api): normalize /docs header and right-panel blues
The right panel's #262262 is a purple-leaning navy that clashed with
the cyan-leaning gradient, and the bar's brightest point sat directly
over the dark panel. Right panel now uses #1b1f52 (the gradient's own
dark endpoint) and the gradient runs bright-to-dark so its dark end
lands flush on the panel -- no seam, one blue family. Verified via
headless-Chrome render.
* style(api): right-panel gradient on /docs via bundle-pinned override
Redoc's theme only takes solid colors (it derives shades from
rightPanel.backgroundColor), so the gradient (#1b3d79 -> #1b3068 ->
#1b1f52, continuing the topbar blend) rides as a CSS override on the
styled-components classes of the per-section right-panel divs
(.sc-iGgWBj.sc-gsFSXq + the .sc-dExYaf stub). Those names are
deterministic for the vendored 2.5.3 bundle (verified across loads) but
change on any Redoc bump: re-derive via headless probe (find elements
whose computed background equals the rightPanel color). If they stop
matching, the panel falls back to the solid #1b1f52 theme color --
cosmetic only. Verified via headless-Chrome render.
* feat(api): collapsible samples column on /docs
Redoc CE has no built-in panel toggle, so the topbar gains a Hide/Show
samples button that flips .samples-collapsed on <html>: the right-panel
divs hide (same bundle-pinned styled-components classes as the gradient
override) and each section's content half takes the full width. Choice
persists in localStorage; aria-pressed tracks state. If the pinned
classes stop matching after a Redoc bump the toggle goes inert --
cosmetic only. Both states verified via headless-Chrome render.
* ci(api): spec-lint CI gate + npm Dependabot coverage
New spec-lint job mirrors the local npm run lint:api so openapi.json
cannot drift from redocly.yaml with green CI. Dependabot gains the npm
ecosystem (package.json is new; nothing watched @redocly/cli).
* feat(api): docs finishing touches - x-tagGroups, favicon, docs:preview
x-tagGroups sections the Redoc sidebar (Read API / Meta / SHOC Feed);
inline data-URI SVG favicon (SHOC blue) stops the browser's follow-up
/favicon.ico request 403ing at the gateway; npm run docs:preview wraps
the real-handler local render (scripts/preview_docs.py); README gains a
docs-page architecture section covering the inline pattern, theme,
pinned-selector caveat, and tooling. Lint 0 errors, 675 pytest,
headless render verified.
---------
Signed-off-by: Adam Moussa <adam@seahavenind.com>
Redoc 2.5.3 standalone bundle (MIT) replaces the three swagger-ui-dist
assets: one ~1.05MB JS file instead of ~1.8MB of JS+CSS+preset, and the
layout traps (StandaloneLayout/BaseLayout) go away. Redoc is read-only by
design, which matches the existing posture: try-it-out was already
disabled since data routes need SigV4 (Postman for live calls).
Unchanged: single token-gated response, offline vendoring (no CDN),
per-request server-URL injection, script-breakout guards, cached shell
with per-request spec splice. Bundle self-containment verified: the
search worker is an inlined Blob, and the only new Worker(filename) path
is Prism's async mode, which Redoc never invokes.
Verified via headless-Chrome render of the real handler output: all
routes, the OpenAPI 3.1 webhooks section, and planned-route markers
render; no placeholder leakage.
* feat(api): use stock Swagger UI for the /docs page
Replaces the custom renderer with vendored stock Swagger UI
(swagger-ui-dist 5.17.14, Apache-2.0), kept offline (no CDN) and inlined
server-side into the single token-gated /docs response alongside the spec.
BaseLayout (topbar hidden); try-it-out disabled since data routes need SigV4
(use Postman for live calls). Breakout guards on the inlined css/js/spec.
Bundle-consistency + spec-drift + handler tests updated for the two vendored
assets. cdk diff = Lambda code asset only (no IAM/API/policy change).
* fix(api): render Swagger UI with the canonical StandaloneLayout recipe
The BaseLayout-only init (apis preset, no standalone preset) rendered
incorrectly. Switch to the canonical swagger-ui-dist recipe: vendor
swagger-ui-standalone-preset.js and init with
presets:[apis, SwaggerUIStandalonePreset] + layout:"StandaloneLayout"
(topbar hidden, try-it-out disabled). Verified via headless Chrome against
the live deployed page: all 9 endpoints + 4 webhooks + models render.
Tests/bundling updated for the third vendored asset.
* fix(api): declutter the Swagger UI docs page
The page rendered (stock Swagger UI, StandaloneLayout) but looked cramped:
a dense info.description wall of inline-code chips collided across Swagger
UI's tight default line-height, and the Servers box showed a SEE-STACK-OUTPUT
placeholder.
- Trim info.description to a few concise lines (detail lives in README + the
webhook contract doc).
- Inject the live invoke URL into servers[0].url per request (from the API
Gateway request context; committed literal is the fallback), so the Servers
box shows the real endpoint and drops the server-variables section.
- CSS: loosen description line-height + inline-code padding so chips never
overlap; tidy the scheme container spacing.
- Handler: cache the heavy CSS/JS/preset shell once; splice the (server-
injected) spec per request.
Verified via headless Chrome against the live deployed page.
* feat(api): add procurement-api stack - read API + OpenAPI docs page
Third CDK stack: API Gateway REST API (IAM SigV4) over both pipelines'
tables, replacing SHOC's retired SyncController cross-account DynamoDB
scan as the reconciliation/backfill path.
- lambdas/api/: handler (healthcheck + docs-token gate + router dispatch),
router (single route table), pagination (opaque cursor, hostile -> 400),
Decimal-safe serialization, wo_repo/po_repo reads. No VendorReplies.
- OpenAPI 3.1 spec as source of truth incl. top-level webhooks section
documenting the outbound SHOC feed; phase-2 write endpoints x-planned
(router answers 501). Self-contained /docs page, no CDN.
- Auth: AWS_IAM on data routes + resource policy scoped to exactly
arn:aws:iam::396287094661:role/shoc-backend-dev on GET/*; /docs and
/openapi.json carve-out is token-gated in the Lambda via shared
web_ui_auth (fail-closed, INFRA-74 posture).
- KMS: explicit Decrypt/DescribeKey on the DynamoDB CMK from SSM
(name-imported table drops the key association - INFRA-104 class).
- Alarms: errors/throttles/duration(p99>=22.5s) + gateway 5xx, ALARM-only
to site-alerts. No access logging in v1 (docs ?token= shim stays out of
logs); cloud_watch_role=False.
- Tests: handler auth-seam + routing + Decimal round-trip; moto cursor
pagination incl. hostile cursors; spec<->router drift gate; bundle
AST pins for the api command; pytest.ini --cov + loader siblings.
- Deploy role: third stack DescribeStacks ARN + procurement-api smoke
invoke ARN (re-run create-deploy-role.sh before merge).
* harden(api): apply sh-security-review findings to procurement-api
Fan-out (6 detectors) + review findings resolved:
Correctness / DoS:
- pagination: require EXACT key-set match (was subset) so a partial/foreign
composite cursor can't reach DynamoDB as an inconsistent ExclusiveStartKey
-> ValidationException -> 500; comments Query now pins the cursor's
work_order_id to the path entity.
- handler: map botocore ValidationException to 400 (defense in depth) so a
crafted cursor can't drive the zero-threshold 5xx alarm.
- web_ui_auth: compare tokens as bytes; a non-ASCII presented token now fails
closed (401) instead of crashing hmac.compare_digest into a 500. Resolves the
pre-existing xfail(strict) follow-up test; hardens the web UIs too.
Docs page:
- typeStr() now escapes the one spec-derived string that reached innerHTML.
- spec inlined into the docs <script> block escapes "<" -> < (</script>
breakout guard); /openapi.json still served byte-faithful.
- Cache-Control: no-store + Referrer-Policy: no-referrer on docs responses so
the ?token= URL stays out of caches/Referer.
- spec-drift test asserts the committed spec carries no "</" / "<!--".
IAM / IaC:
- resource policy enumerates the 7 data GET resources instead of GET/* so a
future GET route can't silently inherit SHOC cross-account reach.
- kms:Decrypt grant gains a kms:ViaService=dynamodb condition.
- stage throttling (50 rps / 100 burst) bounds the unauthenticated /docs blast
radius below the 10k account default.
- corrected the PATCH/POST comment (same-account callers aren't blocked by the
resource policy; 501 handler + absent write grant are the gate).
- documented the RETAIN log-group first-deploy rollback trap and the
resource-policy-needs-redeploy gotcha in-stack.
Mandatory GPT-4.1 cross-family review of the full policy surface: no BLOCK/FIX.
675 tests pass, ruff clean, cdk synth green.
First seahaven-prod deploy failed CREATE on the sender-auth MetricFilter:
it imported /aws/lambda/<fn> by name, which pre-existed in mgmt but not in
a fresh account. All 5 functions now get an explicit logs.LogGroup
(TWO_MONTHS, RETAIN) via common.make_function_log_group, and the metric
filter takes the construct so CFN orders it after the group exists.
Also removes the deprecated LogRetention custom resource and its
wildcard logs:PutRetentionPolicy role (CKV_AWS_111). Function roles keep
AWSLambdaBasicExecutionRole (verified in the synthesized template), so
log-write permissions are unchanged; cross-review's grant_write FIX was
a false positive on that basis. Supersedes PR #84, which hardcoded the
mgmt logs-CMK ARN and predates the common.py refactor.
mgmt collision note: these CREATEs would collide with the pre-existing
groups in mgmt; acceptable because the deploy secret now targets prod
and mgmt is frozen pending decommission.
* feat(migration): prepare stacks and tooling for the seahaven-prod account move
Phase 1 of the mgmt (328440206208) -> seahaven-prod (011934824531)
migration. No behavior change in-account; everything here is additive or
account-portability hygiene:
- infra/deploy-role/: reviewed OIDC deploy-role artifacts for prod
(trust main-only, cdk-hnb659fds-* AssumeRole, smoke-invoke-lambda scoped
to exactly the two email-processor fn ARNs). Codifies the previously
out-of-band smoke-invoke grant.
- Table resource policies: make_slack_bot_read_policy in cdk/common.py,
applied to purchase-orders, verified-sites, WorkOrders,
WorkOrderComments (NOT pending-site-review; no bot consumer). Grants the
mgmt-resident seahaven-slack-bot roles read-only cross-account access
post-move (bot-side identity grants land in the slack-bot repo).
- scripts/migrate_tables.py: dry-run-default backfill tool implementing
the plan's per-table semantics (superset overwrite, ingested_at cutoff
for WorkOrderComments, backup-gated truncate-and-load for the two
site tables) plus a verify subcommand (count parity, spot checks,
sticky-Cancelled drift check).
- tests/test_resource_policy_helper.py: statement-shape unit tests +
static pins that exactly the four bot-read tables carry the policy.
- Account-literal fixes: account-agnostic fixture bucket in
test_reprocess_contract; runbook/README/po-template-parser account
references updated to prod with historical mgmt notes; README gains the
account-prerequisites list (imported-by-name dependencies).
deploy.yaml is deliberately unchanged (push-to-main auto-deploy kept).
Merge is held until migration Phase 0 completes; flipping the
AWS_DEPLOY_ROLE_ARN repo secret and merging this PR IS the first prod
deploy.
* fix(migration): verify backup AVAILABLE pre-truncate; document wildcard risk acceptance (cross-review FIX/NIT)
* refactor(migration): drop cross-account read grants (slack-bot decommissioned); harden backfill + deploy role
seahaven-slack-bot was decommissioned 2026-07-23 (stack DELETE_IN_PROGRESS,
consumer Lambdas gone); its successor sh-mcp is undeployed and uses
same-account DynamoDB access. So no live consumer reads these tables
cross-account. Per Adam's call, drop the cross-account grants entirely and
re-add correctly-scoped ones if/when sh-mcp deploys to a different account.
- Remove the four table resource policies + make_slack_bot_read_policy helper
+ its constants (cdk/common.py, po_stack.py, wo_stack.py) and the helper's
unit test. Both stacks synth with zero table ResourcePolicy.
- scripts/migrate_tables.py hardening (fixes from the sh-security-review
fan-out on the destructive backfill tool):
* validate --cutoff strictly (parse ISO-8601, require aware UTC, re-emit
canonical second-precision form) so a malformed cutoff can't silently
copy dual-window rows or drop history;
* reject `copy --all` up front (must run tables individually, in order,
with the stream-drain wait) instead of writing three tables then erroring;
* truncate backup gate now also checks recency (<1h) and TableId, not just
status+name;
* verify requires --cutoff whenever a cutoff table is in scope (else it
false-flags dual-window rows as MISSING);
* sticky-cancel is now PREVENTED copy-side (a non-Cancelled source item
never overwrites a dest-Cancelled PO), and the verify comment no longer
overstates what its source-side scan covers;
* spot-check all modes (truncate_load keys are verbatim, so key-existence
is sound there too).
- Deploy role: scope cloudformation:DescribeStacks to this repo's stacks +
CDKToolkit (was Resource:*, disclosed all tenant stacks in the shared prod
account); add a drift check warning on unexpected role policies and drop the
dead SMOKE_POLICY_NAME var; document the shared-account bootstrap-role
accepted risk in the deploy-role README.
* docs(deploy-role): fold in cross-review NITs (DescribeStacks maintenance note, warn-only drift rationale)
* ci: update workflow to use new workflow tag (ruff versioning fix)
* fix(migration): address Open SWE review findings on migrate_tables.py
- Validate the truncate backup on dry-run as well as --execute so a
missing/stale/wrong-incarnation --backup-arn surfaces on the rehearsal
run (finding f_24a48b8900).
- Assert configured keys match the live key schema of both tables before
any key projection, turning config/schema drift into a descriptive
abort instead of a mid-backfill KeyError (finding f_cb6b5a6c59).
- Clarify why key-existence spot-checks are sound for WorkOrderComments:
the copy Puts source items verbatim and the sample uses the same
cutoff filter, so per-account comment_id divergence never enters the
check (finding f_390b7d6c3b is a false positive; comment hardened).