Commit graph

77 commits

Author SHA1 Message Date
Adam Moussa
2f5e5e6e66
fix(iam): drop unscoped door-unlock domain create (#127)
CreateDomainName cannot be hostname-pinned, and mgmt still holds doorunlock.seahaven.com. Attach the domain at cutover instead of granting collection POST.
2026-08-27 23:03:35 +00:00
Adam Moussa
23d954369d
fix(iam): allow door-unlock apply to create api domain (#126)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
CreateDomainName authorizes against the /domainnames collection, so the hostname-pinned ARN cannot complete first apply.
2026-08-27 22:43:48 +00:00
Adam Moussa
28a064966b
fix(iam): allow door-unlock plan to read 3cx secret metadata (#125)
The AWS secrets data source calls GetResourcePolicy; the first HCP plan failed without it on the three exact 3CX ARNs.
2026-08-27 22:16:11 +00:00
Adam Moussa
e5e7980508
feat(iam): add paychex-integrations hcptf roles and boundary (PLAT-120) (#124)
* feat(iam): add paychex-integrations hcptf roles and boundary

* fix(iam): split paychex plan lambda list onto Resource *
2026-08-27 21:50:11 +00:00
Adam Moussa
689ec147a3
feat(iam): add door-unlock-api hcptf roles and boundary (PLAT-76) (#123)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): add door-unlock-api hcptf roles and boundary

Give HCP Terraform a prod plan/apply pair, a per-workload Lambda boundary with exact SSM and 3CX ARNs, and API access-log delivery so PLAT-76 can leave the mgmt CDK stack.

* fix(iam): pin door-unlock apigw domain and ssm reads

Stop the apply role from managing every HTTP API custom domain, and keep SecureString door-unlock parameters off HCP plan and apply GetParameter.
2026-08-27 21:26:27 +00:00
9bffddfd7b
fix(iam): allow site plan role to describe CF function (PLAT-106) 2026-08-20 15:22:09 -04:00
5fa4267798
chore(iam): drop PascalCase WO Dynamo and alarm ARNs 2026-08-14 11:58:41 -04:00
0f84d7808b
feat(iam): add per-workload lambda execution boundaries
Shared seahaven-lambda-execution-boundary stays unchanged for live roles.
New named policies plus an enumerated StringEquals allow-list unblock the
next PLAT-71 widen without growing the 6144-character shared document.
2026-08-13 16:47:53 -04:00
81cc8eb4f9
fix(iam): consolidate meal-order boundary Sid under PolicySize cap 2026-08-10 15:57:08 -04:00
7c43c867e3
fix(iam): allow execute-api Invoke for meal-order weekly-menu boundary 2026-08-10 15:42:12 -04:00
b76d0578e0
fix(iam): drop PutResourcePolicy from meal-order apply role
Pre-grant delivery.logs write via MealOrderApiAccessLogResourcePolicy on
substrate so the HCP apply role cannot mutate account-wide log resource
policies.
2026-08-10 14:57:17 -04:00
e10462d25e
fix(iam): allow API GW Log Delivery on meal-order apply role 2026-08-10 14:47:30 -04:00
6bb2d54814
fix(iam): allow PassRole to apigateway for meal-order authorizer 2026-08-10 13:46:12 -04:00
1bb5e79ea0
fix(iam): allow ssm:ListTagsForResource on meal-order plan role 2026-08-10 13:23:56 -04:00
Adam Moussa
f44b88732e
fix(iam): allow ssm:DescribeParameters for meal-order hcptf roles (#99)
Some checks failed
Deploy / deploy-management (push) Has been cancelled
Deploy / deploy-external-dev (push) Has been cancelled
Deploy / deploy-security (push) Has been cancelled
Deploy / deploy-dev (push) Has been cancelled
Deploy / deploy-prod (push) Has been cancelled
2026-08-08 00:11:48 +00:00
Adam Moussa
3605215a28
refactor(iam): consolidate boundary statements under PolicySize cap (#98)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
Merge workload secret/DDB/S3 SIDs and trim meal-order extras so the
shared lambda execution boundary fits under the 6144-character limit.
2026-08-07 23:54:28 +00:00
Adam Moussa
44b672ae9a
feat(iam): add hcptf roles and boundary widen for meal-order-manager (PLAT-70) (#97)
* feat(iam): add hcptf roles and boundary widen for meal-order-manager

Append plan/apply OIDC roles for meal-order-manager-prod and widen the lambda execution boundary with exact prod secret ARN and data-plane statements.

* fix(iam): make meal-order plan role Lambda refresh read-only

Replace plan-role lambda:* with Get*/List* so plan-phase credentials cannot mutate functions or layers.
2026-08-07 19:36:20 -04:00
Adam Moussa
a6f22880db
feat(waf): add seahaven-prod shared CloudFront WebACL (PLAT-92) (#96)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(waf): add seahaven-prod shared CloudFront WebACL stack

Stand up AppWebAcl in a thin prod stack and widen seahaven-site HCP
roles to read the SSM ARN so CloudFront can associate the ACL in-account.

* fix(deploy): add app-web-acl-prod to deploy.yaml
2026-08-07 17:07:04 -04:00
Adam Moussa
35461d9267
feat(iam): add hcptf-seahaven-site plan/apply roles (PLAT-91) (#89)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): add hcptf-seahaven-site plan/apply roles

Static-site HCP substrate for seahaven-site-prod plus boundary widen for
the TF-managed content-deploy role (S3 origin + CloudFront invalidate).

* fix(iam): allow seahaven-site HCP roles to read GitHub OIDC provider

Plan refresh needs iam:GetOpenIDConnectProvider for the content-deploy
role trust data source (PLAT-91 first-plan AccessDenied).
2026-08-07 15:09:57 -04:00
Adam Moussa
e12944a42d
fix(iam): allow kebab WO tables on lambda execution boundary (#90)
Widen ProcurementIngestDynamoDB so workorder Lambdas can read/write
work-orders and work-order-comments after the PLAT-11 rename.
2026-08-07 14:20:40 -04:00
Adam Moussa
ff77ffd421
fix(iam): allow HCP procurement-ingest kebab WO Dynamo tables (#88)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
Widen hcptf-procurement-ingest apply and plan-refresh DynamoDB/CloudWatch
ARN pins for work-orders and work-order-comments (PLAT-11 rename).
2026-08-07 13:53:42 -04:00
202103041c
fix(iam): allow API GW and ESM tagging for procurement-ingest
Provider default_tags need apigateway /tags/* and unconditioned ESM
TagResource after import-in-place.
2026-08-07 11:09:06 -04:00
a5997f878b
fix(iam): widen procurement-ingest plan refresh for import
Add GetEventSourceMapping, SSM GetParameter pins, and Resource "*" for
kms:ListAliases so the first HCP import plan can refresh.
2026-08-07 11:00:49 -04:00
Adam Moussa
a5fa0b16a3
feat(iam): add hcptf roles and boundary for procurement-ingest (PLAT-86) (#85)
* feat(iam): add hcptf roles and boundary for procurement-ingest

* fix(iam): tighten procurement-ingest apply and plan scopes

Replace kms:* and secret-value writes on shell statements; split IAM
collection APIs onto Resource "*".
2026-08-07 10:41:12 -04:00
Adam Moussa
69f31842cb
feat(iam): add hcptf roles for sh-openswe-traces-prod (PLAT-73) (#80)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): add hcptf roles for sh-openswe-traces-prod

Storage/IAM-user apply and plan roles for the HCP workspace. No Lambda
boundary widen; explicit IAM user CRUD because hcptf-iam-management is
role-path-only.

* fix(iam): pin CreateSecret to exact export secret name

Remove CreateSecret and UpdateSecret from the ARN-prefix shell grant so
apply cannot create longer-named secrets or overwrite SecretString.
2026-08-05 22:46:32 +00:00
Adam Moussa
fc64b03e3d
feat(iam): hcptf front-integrations roles and boundary (PLAT-72) (#81)
* feat(iam): add hcptf front-integrations roles and boundary widen

Add plan/apply OIDC roles for front-integrations-prod and widen the
Lambda execution boundary with exact prod secret ARNs plus DynamoDB
CRUD on front-sla-alerts.

* fix(iam): restrict front-integrations plan role to lambda Get/List

Keep mutate APIs on the apply role so a compromised plan-phase
OIDC session cannot update or delete front-* functions.
2026-08-05 18:33:31 -04:00
Adam Moussa
9ee4d4a3d7
docs(iam): codify hcp terraform migration checklist from PLAT-56 (#79)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
Expand the README playbook to steps 0–10 and document the required
plan-refresh sidecar plus prefix-scoped apply-role wildcards so the
next workload copies afi patterns instead of relearning first-apply misses.
2026-08-05 16:14:16 -04:00
46829fc2c1
fix(iam): use lambda:* and events:* on afi hcptf apply scope
Provider refresh needs GetFunctionCodeSigningConfig and similar reads;
keep blast radius on afi-* function/layer/rule ARNs only.
2026-08-05 12:59:19 -04:00
f4292832fe
fix(iam): add plan-role refresh reads for afi terraform state
ViewOnlyAccess omits iam:GetRole and events:DescribeRule; without a
scoped refresh policy, HCP plans fail after the first partial apply.
2026-08-05 12:57:23 -04:00
3512214d14
fix(iam): allow s3:* on afi artifact bucket for provider reads
First HCP apply failed on s3:GetBucketAcl after CreateBucket; scope
remains the single artifact bucket ARN.
2026-08-05 12:56:24 -04:00
Adam Moussa
4e1cf4c0bd
feat(iam): add hcptf roles/boundary widen - afi-backup-monitor (PLAT-56) (#76)
* feat(iam): add hcptf roles and boundary widen for afi-backup-monitor

Provision plan/apply OIDC roles for workspace afi-backup-monitor-prod
and widen the prod Lambda boundary with the two exact secret ARNs.

* fix(iam): split DescribeLogGroups and allow afi artifact bucket

logs:DescribeLogGroups cannot be resource-scoped; grant it on *. Add
S3 permissions for the HCP Lambda artifact bucket used by PLAT-56.
2026-08-05 12:44:52 -04:00
Adam Moussa
c09cf1110d
fix(iam): allow API Gateway authorizer role passing
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
2026-08-03 14:38:39 -04:00
16a82c2a36
docs(iam): qualify mgmt-only verification claims per cross-review round 2 2026-07-31 13:46:29 -04:00
08a1d41b05
docs(iam): resolve confirmed review findings from both INFRA-186 gates
Cross-family round 1 plus the /sh-security-review verifier confirmed 11
findings on the floor reduction, all documentation defects; no policy
statement changes. The one HIGH: the Terraform migration checklist never
widened the boundary, so a Lambda-bearing Terraform migration would deploy
green and lose every data-plane call at first invoke. Checklist step 2 now
carries the widening requirement, step 3 verifies deployed boundary content,
and the terraform-substrate header no longer reads as 'Terraform path
unaffected'. Also corrected: Description is a REPLACEMENT property (a
Description edit wedges the custom-named policy and CFN's remedy is the
forbidden rename), the sanctioned-source contradiction, the false
AWSLambdaVPCAccessExecutionRole parity claim, the KMS log-group category
error, stale size numbers (691/5,453), the same-PR widening contradiction,
per-workload residue text, a LoggingConfig silent-log-loss note, the
us-east-1 region pin rationale, and ENI DoS deferral now tracked as
INFRA-200.
2026-07-31 13:44:21 -04:00
e791005af0
docs(iam): cite INFRA-187 as the per-workload boundary end state
Replaces the placeholder 'tracked as its own ticket' references with the real
key, and records the load-bearing constraint inline so the next reader does not
rediscover it: both guardrail policies pin ONE literal boundary ARN inside
StringEquals iam:PermissionsBoundary, and loosening that to a wildcard weakens
the gate rather than merely relaxing it.
2026-07-30 19:09:27 -04:00
32f06e74eb
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope
Adam's call after review: the security win of INFRA-186 comes from DELETING the
account-wide wildcards, not from enumerating replacements. Per-workload prefixes
add no security -- they only keep a workload functional -- and widening a
boundary is the safe direction (adding a resource never breaks a running Lambda;
only tightening does). So the per-workload scope moves to each migration PR,
which has the stack's real template open in front of it.

Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets
Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide
floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni
-- the statements every Lambda needs regardless of workload, and also the
silent-failure classes, which is why they belong in the floor.

KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups
(verified). A workload bringing a CMK adds the statement plus the matching
kms:ViaService principal in its own PR.

Why not keep the enumeration: it required predicting five stacks' needs from
this file's own permission-source comment block, and /sh-security-review found
SIX errors in the result -- three silent. The block is a secondary record, not
an authority. Deriving scope per-migration from the owning template removes the
whole error class.

Effect on the security objective: unchanged. secret:*, table/*, function:*,
sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the
amplifier is closed identically.

Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap
stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all
updated to match; widening path now leads with 'read the stack's own template',
names the silent-failure classes to check, and moves the version-budget check to
a precondition instead of a trailing step.

Verified unchanged: logical id and ManagedPolicyName, so all eight pinning
conditions across both guardrail policies still resolve. Both accounts synth
identically at 703 chars.
2026-07-30 19:07:09 -04:00
59852eff34
fix(iam): restore four permissions the boundary would have denied at migration
/sh-security-review (6 detectors + verifier) found four HIGH findings, all the
same defect class: the template's permission-source comment block was used as
the sanctioned scope source, but it is an incomplete and in places invented
secondary record. Each was verified against the real stack template before
fixing. None is live today (prod/dev boundary usage is 0); all four would have
been AccessDenied at first migration, three of them SILENTLY.

- SES configuration-set/seahaven-email-events restored. afterhours-shift-manager
  template.yaml:178-181 grants it with an in-repo comment stating the send is
  denied without it. An earlier revision dropped it after checking whether any
  config set exists in prod/dev today (none do) -- the wrong test. The right
  question is whether an enumerated stack's own IAM policy names it.
- SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults
  to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it.
  The prior 'unverified identity fails loudly anyway' argument holds only until
  the migration verifies the domain, which the migration procedure requires.
- scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com
  only) added. afterhours template.yaml:110-120 needs both; the block omitted
  them entirely. Failure is silent -- app.py wraps create_schedule in a bare
  except, so the Slack command reports success and no schedule exists.
- secret:afi-slack-webhook-* added. The block named
  'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret
  ARNs are deploy parameters, so the real names live only in that repo's
  README:48-49 (afi-api-key, afi-slack-webhook).

Also corrected, all comment-only:
- The permission-source block itself, at each of the four points it was wrong,
  with the correction and its evidence recorded inline.
- The false claim that SAM auto-names async DLQs (it does not -- all four
  payments queues are hand-written with explicit QueueNames). Replaced with the
  real invariant: any queue a boundary-carrying function sends to must be
  payments-* or the boundary widens in the same PR; a denied destination write
  is silent.
- SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these
  fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the
  header claimed. Flagged per-workload boundaries as the realistic next move.

Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName
seahaven-lambda-execution-boundary, so all eight pinning conditions across both
guardrail policies still resolve.
2026-07-30 18:17:16 -04:00
b264f74f01
fix(iam): correct two boundary-scoping defects found in review
Post-implementation verification of the INFRA-186 prod/dev scoping found two
functional defects that would have denied permissions the migrating stacks
actually need. Neither is live today (prod/dev boundary usage is 0), but both
would have surfaced as AccessDenied at first migration.

- KMS: the ViaService list omitted ssm., while SSMParameterRead in the same
  policy grants ssm:GetParameter*. A SecureString read decrypts via the SSM
  service principal, so the boundary denied reads it also granted.
- S3: payments-dashboard was classified read-only from the template's own
  permission-source comment, but that enumeration is incomplete -- the real
  stack grants s3:PutObject on BoaRawBucket (template.yaml:272, 1098-1099).
  Write is now allowed on seahaven-payments-boa-raw-* only; payroll-emails and
  payments-csv stay read-only, preserving the evidence-deletion protection.
  The seahaven-payments-* wildcard is replaced by the three literal bucket
  names, verified against payments-dashboard/template.yaml.

Not changed: SES configuration-set/*. Review claimed dropping it rested on a
false premise; verified live -- prod and dev both have ZERO configuration sets
and member-baseline-stack.ts:44 excludes SES monitoring. The drop is correct.

README: the 'substrate changes must edit both files' rule is now false for the
boundary specifically, and said so uniformly. Corrected to distinguish the
deliberately divergent boundary from the still-at-parity substrate resources.
2026-07-30 18:03:40 -04:00
274f933495
refactor(iam): scope Lambda execution boundary to per-workload prefixes
Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of
deploy-substrate.template.yaml from account-wide wildcards to per-workload
resource prefixes drawn from the template's own permission-source block.
This is the PROD/DEV HALF of INFRA-186.

What was scoped (wildcard -> per-workload prefix):
  - dynamodb  table/* + table/*/index/*  -> afterhours-*, front-*,
    meal-order-manager-*, PaymentsDashboard*, payments-dashboard-*
    (a trailing * after each prefix also covers the /index/* GSI ARNs, so the
    separate table/*/index/* entry is deleted rather than replaced)
  - s3        *-${AccountId}             -> meal-order-manager-*-${AccountId}
    (read/write) and seahaven-payments-* / seahaven-payroll-emails-*
    (read-only). The removed pattern was not an ownership check at all: S3 ARNs
    carry no account field, so it was a bare name-suffix filter that matched 8
    of 9 buckets in prod -- including the org's own Config and VPC-flow-log
    buckets -- with PutObject and DeleteObject.
  - secretsmanager  secret:*             -> five <stack>/ prefixes + the legacy
    bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing
    key, i.e. a webhook-forgery primitive.
  - ssm       parameter/*                -> afterhours-shift-manager and
    meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath
    authorises against the path, not the leaf)
  - sqs       :*                         -> payments-*
  - lambda    function:*                 -> afterhours-*, meal-order-manager-*,
    payments-*. Highest-leverage fix here: an invoked function runs under its
    OWN role, and every non-SAM function in prod is CDK-deployed with no
    boundary, so function:* was a boundary-escape primitive, not just lateral
    movement.
  - ses       identity/* + configuration-set/* -> the two verified prod
    identities; configuration-set dropped (zero exist)
  - logs      split into a scoped write half (/aws/lambda*) and a wildcard
    describe half (DescribeLogGroups is a collection action AWS authorises
    against "*" regardless of the ARN supplied)

Deliberately NOT tightened, each with written justification on the statement:
CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use
actions that support no resource-level permissions. KMS keeps key/* -- key ARNs
carry UUID key ids, not workload names -- and is constrained by a kms:ViaService
condition instead, which inherits the per-workload scoping of the services
above for free.

No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file
deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified
2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the
handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that
is slightly too tight is recoverable -- the migrating stack widens it in its own
PR before its first deploy -- whereas leaving it loose perpetuates the exposure.
The widening path and its ordering hazard are documented in the template.

mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The
management account (328440206208) uses a separate copy in
Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE
boundary-carrying roles, where tightening is a production change with a silent,
deploy-time-invisible failure mode; it needs its own validated rollout and is
explicitly out of scope. The header's parity rule is therefore now SCOPED, not
global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical
and must still change together, while LambdaExecutionBoundary must NOT be
reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a
future mechanical drift check does not "fix" it away, following the same pattern
terraform-substrate.template.yaml uses for its divergences.

Content-only change: ManagedPolicyName, the policy ARN and the logical id
LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary
conditions across this file and terraform-substrate.template.yaml pin the
boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming
a non-existent policy simply never matches.

Verification:
  - npx tsc --noEmit: clean
  - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds
  - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY
    changed resource; GitHubOIDCProvider, SamCfnExecutionRole and
    SamCfnIamManagementPolicy are byte-identical
  - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements,
    identical in both accounts
  - iam simulate-custom-policy against live prod, every deny re-checked against
    an Allow */* positive control: 11/11 cross-tenant denies are real (Config
    and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials,
    workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs,
    WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com)
    and 23/23 enumerated workload resources still allow

Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because
three statements legitimately retain Resource:"*", so the suppression is still
required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new
ids added, superseded ids retained, and the boundary justification's stale "OPEN
follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since
this commit is what closes it. Scanners: RESULT PASS.

Refs: INFRA-186
2026-07-30 17:46:37 -04:00
981960433f
fix(iam): scope Terraform guardrail role writes to a Terraform-owned path
Security review (6 detectors + proof-or-kill verifier) confirmed 1 critical and
1 high in the first revision, both inherited by mirroring the SAM copy's
Resource "*" role grants:

- C1 (critical): iam:UpdateAssumeRolePolicy on "*" with DenySelfMutation
  covering only three name patterns lets the principal repoint the
  AdministratorAccess CDK bootstrap role's trust policy to an external account.
- C2 (high): the SAM justification for role/* (SAM auto-roles land at path /
  with no settable RolePath) does not transfer -- Terraform's aws_iam_role
  supports path.

Fixes, closing the class at the root rather than by denylist:
- All role writes, boundary sets and PassRole confined to role/tf-managed/*;
  reads split into a separate statement that keeps Resource "*".
- DenySelfMutation extended to cdk-hnb659fds-*, OrganizationAccountAccessRole
  and seahaven-* as defense in depth.
- OIDC provider made conditional (CreateOIDCProvider), mirroring the sibling
  substrate, so a first-create rollback is recoverable rather than wedging the
  stack in ROLLBACK_COMPLETE against a Retained orphan.
- README corrected: the guardrail policy is NOT Retain (only the provider is),
  so the Deny backstops do not survive a stack delete.

checkov CKV_AWS_109 no longer fires on this template, so no suppression is
needed. The template header records every divergence from the SAM copy.
2026-07-30 16:55:45 -04:00
ea27635ef2
feat(iac): add per-account HCP Terraform deploy substrate for prod and dev
New stack seahaven-terraform-substrate (instances terraform-substrate-prod +
terraform-substrate-dev): app.terraform.io OIDC provider and the shared
boundary-gated guardrail policy seahaven-hcptf-iam-management that
per-workspace Terraform apply roles attach at migration time. No roles are
pre-provisioned (accumulator pattern, parallel to githubdeploy-*).

Guardrail statements mirror seahaven-cfn-exec-iam-management byte-identically
except DenySelfMutation, whose scope extends to hcptf-* alongside the
GitHub-substrate principals. Explicit stack dependency on the same-account
deploy-substrate stack (boundary ARN appears only in Condition strings, so
CFN infers no edge).
2026-07-30 16:31:34 -04:00
61a94da4fc
fix(iam): reconcile the remaining substrate divergences from the mgmt copy
Review of the DenySelfMutation port found the header's 'reconciled' claim
was not yet true: mgmt Phase A also added the CloudWatch Logs
metric-filter actions (afterhours-shift-manager creates an
AWS::Logs::MetricFilter through this role), and without them a migrating
SAM stack fails mid-deploy with AccessDenied. Ports those three actions
and corrects two stale header notes. Every IAM statement in the three
shared resources is now byte-identical across both files, verified
programmatically; the only delta left is the DependsOn ordering line.
2026-07-27 18:55:10 -04:00
62f6a76e6c
fix(iam): port DenySelfMutation self-protection into the prod/dev deploy substrate
The seahaven-cfn-exec-iam-management policy in prod and dev carried only
DenyBoundaryTampering + DenyBoundaryPolicyEdit: the mgmt Phase A review
later showed a Deny-in-a-managed-policy control is self-detachable
(iam:DetachRolePolicy on * is unconditioned), so without DenySelfMutation
the exec role can detach the very policy carrying the Denies and
reinstate the boundary-removal escalation. Latent today (no PassRole
grants, zero SAM stacks in prod/dev) but must be closed before the first
SAM workload migrates.

Ports verbatim from .github/oidc-deploy-roles.yaml (mgmt, PRs #95/#98):
- DenySelfMutation over role/github-cfn-execution-role + githubdeploy-*
- DenyBoundaryPolicyEdit widened to policy/seahaven-*

Statement set verified byte-identical to the mgmt copy (9 sids);
provenance header updated - the two copies are reconciled.
2026-07-27 18:37:39 -04:00
2cfc122269
fix(deploy-substrate): move boundary-gated IAM policy off the role's inline budget
The first deploy of seahaven-deploy-substrate failed in both prod and dev
with ServiceLimitExceeded: 'Maximum policy size of 10240 bytes exceeded
for role github-cfn-execution-role'. The role's inline policies already
sat ~94 bytes under IAM's hard 10,240-byte per-role limit, so the two
Deny statements added to close the boundary-removal escalation did not
fit (10,656 total).

Moves the whole boundary-gated IAM block (6 Allow + 2 Deny statements)
into an attached managed policy, which carries its own separate
6,144-byte budget. Inline drops to 8,285 with ~1.9 KB of headroom;
the managed policy sits at 2,371.

Effective permissions are unchanged: the union of role statements
(inline + attached) is byte-identical as a sorted set before and after
the move (27 statements both sides), identity policies are unioned, and
an explicit Deny still wins. Boundary and trust policy untouched.

Both failed stacks rolled back cleanly with zero orphaned resources and
were deleted before this retry.
2026-07-27 16:43:15 -04:00
d6bea33436
feat(deploy-substrate): per-account GitHub Actions deploy substrate for prod/dev
SAM repos migrating off the frozen management account need the shared
deploy plumbing (permissions boundary + github-cfn-execution-role) in
their target account; none of it existed outside mgmt, so there was no
OIDC SAM deploy path into seahaven-prod or seahaven-dev at all.

Adds a templated, per-account substrate stack so onboarding a future
account is one bin/app.ts instance plus one CD job, not a hand-rolled
copy. Per-repo githubdeploy-* roles stay out by design: they are
provisioned per repo at migration time so an account never accumulates
trust for repos that do not deploy to it.

The template is a verbatim extraction of the reviewed mgmt substrate,
with deliberate, documented divergences — notably the removal of
iam:DeleteRolePermissionsBoundary plus explicit Deny backstops, which
closes a confirmed privilege-escalation path (see PR body).
2026-07-27 16:24:09 -04:00
Adam Moussa
7cce026f4a
chore: drop deleted tables from Phase2 backup selection (#59)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
seahaven-conversations + seahaven-unanswered-questions (seahaven-slack-bot
teardown 2026-07-23), exec-aide (decommissioned 2026-07-10), and
internal-portal-data (internal-portal decommission) no longer exist;
their backup selections would fail nightly.
2026-07-23 16:21:04 -04:00
Adam Moussa
5d7ca83097
fix(scp): exempt chatbot:* from workloads-region-lock (global service, us-east-2 control plane) (#58)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
AWS Chatbot's control-plane API is homed in us-east-2, so every chatbot call
carries aws:RequestedRegion=us-east-2 and is denied by the workloads-region-lock
region deny (approved set = us-east-1/us-west-2). This blocked Slack
workspace/channel setup in seahaven-prod (chatbot:GetSlackOauthParameters
denied), which the prod site-alerts topic needs for Slack delivery. Adds
chatbot:* to the SCP's global-service NotAction exemption list alongside
iam/organizations/cloudfront/route53 — a region-agnostic full-prefix exemption,
the same shape as the other global services. targetIds unchanged (workloads OU);
regional services (s3/kms/logs) and the Bedrock carve-out untouched.

GPT-4.1 cross-review: SAFE TO MERGE. /sh-security-review: block=false (0 confirmed
critical/high). Security-OU region-lock deliberately NOT changed (runs no such
workloads, same asymmetry as its missing Bedrock carve-out).
2026-07-23 15:47:11 -04:00
Adam Moussa
4d3c846b88
feat(prod): seahaven-prod DynamoDB CMK + site-alerts alarm topic (procurement-ingest migration Phase 0a) (#57)
* feat(prod): add seahaven-prod DynamoDB CMK and site-alerts alarm-topic stacks

Provisions the two shared dependencies procurement-ingest imports by name,
ahead of its migration from mgmt to seahaven-prod:

- dynamodb-cmk-prod: second DynamoDbCmkStack instance (same stack name,
  prod account) creating alias/seahaven-dynamodb + the
  /seahaven/dynamodb/cmk-arn SSM param. Adds a cross-account key-policy
  statement so the mgmt seahaven-slack-bot roles can keep reading the
  CMK-encrypted purchase-orders table after it moves (ViaService +
  PrincipalArn-wildcard scoped; identity-policy half lands in the
  slack-bot repo's cutover PR).
- alarm-topic-prod: codified site-alerts SNS topic + seahaven-alarm-topics
  CMK with the cloudwatch.amazonaws.com publish grant (mirrors the working
  mgmt pattern; mgmt's topic remains CLI-managed debt).
- deploy.yaml: both appended to the deploy-prod job's explicit stack list
  (SH-ORG-005 rule: unlisted stacks silently never deploy).

* fix(scripts): account-id assertion in cfn-stack-decommission; complete the aws-cdk-lib 2.262.0 bump (patched brace-expansion); document CMK cutover trap

- cfn-stack-decommission.sh: --account-id is now REQUIRED and asserted
  against sts get-caller-identity before anything runs. Stack names are no
  longer org-unique (seahaven-dynamodb-cmk now exists in mgmt AND prod), so
  a name-only lookup under the wrong ambient profile could report or delete
  the wrong account's stack (security-review LOGIC-001).
- package.json/lock: PR #56's bump-for-patched-brace-expansion landed the
  commit title but not the pin; package.json still said 2.261.0 and the
  lockfile still resolved brace-expansion 5.0.6 (GHSA-3jxr-9vmj-r5cp HIGH,
  blocking the pre-commit scanner). Pin 2.262.0 and regenerate; npm audit
  now clean.
- bin/app.ts comments: slack-bot cutover MUST grant the PROD key ARN, never
  the account-local mgmt SSM param (LOGIC-005); failed-first-create orphan
  CMK recovery note (LOGIC-004).

* refactor(prod): drop cross-account CMK grant (slack-bot decommissioned 2026-07-23)

The AllowMgmtSlackBotReadViaDynamoDb key-policy statement targeted the
seahaven-slack-bot roles, which were decommissioned 2026-07-23. Its successor
sh-mcp is undeployed and uses same-account DynamoDB access, so no cross-account
reader of the CMK-encrypted purchase-orders table exists. The prod CMK + SSM
param + alarm-topic stacks remain (procurement-ingest still imports them). Add
a scoped cross-account grant if/when a real cross-account consumer deploys.
2026-07-23 15:29:55 -04:00
Adam Moussa
cc54b1e28b
chore(security): add explicit workflow permissions and bump aws-cdk-lib to 2.262.0 (#56)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* docs: update aws profile specified in script (local renaming)

* ci: add least-privilege permissions blocks to workflow callers

Resolves code scanning alerts #3 and #4 (actions/missing-workflow-permissions). Both callable workflows only need contents: read; the dependency-review callable already declares it internally, this caps the caller token to match."

* chore(deps): bump aws-cdk-lib to 2.262.0 for patched brace-expansion

Resolves Dependabot alert #4 (CVE-2026-13149, exponential-time DoS in brace-expansion expand()). The vulnerable 5.0.6 is a bundled dependency inside the aws-cdk-lib tarball, so it cannot be updated independently; 2.262.0 bundles the patched 5.0.7.

Also migrates Stack#addDependency to addStackDependency (deprecated in this release) in bin/app.ts.
2026-07-23 17:38:17 +00:00
seahaven-openswe[bot]
88fa5777d1
fix(scp): carve out Bedrock InvokeModel/Converse to us-east-2 for cross-region inference (#54)
Some checks failed
Deploy / deploy-management (push) Has been cancelled
Deploy / deploy-external-dev (push) Has been cancelled
Deploy / deploy-security (push) Has been cancelled
Deploy / deploy-dev (push) Has been cancelled
Deploy / deploy-prod (push) Has been cancelled
* fix(scp): carve out Bedrock InvokeModel/Converse to us-east-2 for cross-region inference

Add bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream,
bedrock:Converse, and bedrock:ConverseStream to the existing
DenyRegionsOutsideApproved NotAction list so the us-east-1/us-west-2
region condition no longer denies them. Add a companion
DenyBedrockInvokeOutsideInference statement that re-denies those same
four actions outside {us-east-1, us-west-2, us-east-2}, bounding the
carve-out to us-east-2 only.

Without this, Anthropic cross-region inference profiles (us.anthropic.*)
that route InvokeModel to us-east-2 are denied, blocking all Claude
generation in workload accounts.

Refs: #53

* fix: add ACCEPTED RISK disposition, hoist Bedrock actions to shared const, mark security-asymmetry

- ACCEPTED RISK: Bedrock carve-out is resource-unscoped (NotAction
  can't be resource-scoped); us-east-2 window admits four actions
  against any Bedrock resource. Per-account IAM and model-access
  enablement gate actual access.
- Hoist the four Bedrock invoke actions into BEDROCK_INVOKE_ACTIONS
  shared const referenced by both NotAction and DenyBedrockInvoke
  statements to prevent future one-sided edit divergence.
- Mark asymmetry in security-guardrails DenyRegionsOutsideApproved:
  no Bedrock carve-out by design — security account runs no Bedrock
  workloads.

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-15 18:52:05 -04:00