seahaven-org-baseline/lib/deploy-substrate
Adam Moussa 274f933495
refactor(iam): scope Lambda execution boundary to per-workload prefixes
Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of
deploy-substrate.template.yaml from account-wide wildcards to per-workload
resource prefixes drawn from the template's own permission-source block.
This is the PROD/DEV HALF of INFRA-186.

What was scoped (wildcard -> per-workload prefix):
  - dynamodb  table/* + table/*/index/*  -> afterhours-*, front-*,
    meal-order-manager-*, PaymentsDashboard*, payments-dashboard-*
    (a trailing * after each prefix also covers the /index/* GSI ARNs, so the
    separate table/*/index/* entry is deleted rather than replaced)
  - s3        *-${AccountId}             -> meal-order-manager-*-${AccountId}
    (read/write) and seahaven-payments-* / seahaven-payroll-emails-*
    (read-only). The removed pattern was not an ownership check at all: S3 ARNs
    carry no account field, so it was a bare name-suffix filter that matched 8
    of 9 buckets in prod -- including the org's own Config and VPC-flow-log
    buckets -- with PutObject and DeleteObject.
  - secretsmanager  secret:*             -> five <stack>/ prefixes + the legacy
    bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing
    key, i.e. a webhook-forgery primitive.
  - ssm       parameter/*                -> afterhours-shift-manager and
    meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath
    authorises against the path, not the leaf)
  - sqs       :*                         -> payments-*
  - lambda    function:*                 -> afterhours-*, meal-order-manager-*,
    payments-*. Highest-leverage fix here: an invoked function runs under its
    OWN role, and every non-SAM function in prod is CDK-deployed with no
    boundary, so function:* was a boundary-escape primitive, not just lateral
    movement.
  - ses       identity/* + configuration-set/* -> the two verified prod
    identities; configuration-set dropped (zero exist)
  - logs      split into a scoped write half (/aws/lambda*) and a wildcard
    describe half (DescribeLogGroups is a collection action AWS authorises
    against "*" regardless of the ARN supplied)

Deliberately NOT tightened, each with written justification on the statement:
CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use
actions that support no resource-level permissions. KMS keeps key/* -- key ARNs
carry UUID key ids, not workload names -- and is constrained by a kms:ViaService
condition instead, which inherits the per-workload scoping of the services
above for free.

No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file
deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified
2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the
handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that
is slightly too tight is recoverable -- the migrating stack widens it in its own
PR before its first deploy -- whereas leaving it loose perpetuates the exposure.
The widening path and its ordering hazard are documented in the template.

mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The
management account (328440206208) uses a separate copy in
Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE
boundary-carrying roles, where tightening is a production change with a silent,
deploy-time-invisible failure mode; it needs its own validated rollout and is
explicitly out of scope. The header's parity rule is therefore now SCOPED, not
global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical
and must still change together, while LambdaExecutionBoundary must NOT be
reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a
future mechanical drift check does not "fix" it away, following the same pattern
terraform-substrate.template.yaml uses for its divergences.

Content-only change: ManagedPolicyName, the policy ARN and the logical id
LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary
conditions across this file and terraform-substrate.template.yaml pin the
boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming
a non-existent policy simply never matches.

Verification:
  - npx tsc --noEmit: clean
  - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds
  - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY
    changed resource; GitHubOIDCProvider, SamCfnExecutionRole and
    SamCfnIamManagementPolicy are byte-identical
  - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements,
    identical in both accounts
  - iam simulate-custom-policy against live prod, every deny re-checked against
    an Allow */* positive control: 11/11 cross-tenant denies are real (Config
    and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials,
    workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs,
    WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com)
    and 23/23 enumerated workload resources still allow

Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because
three statements legitimately retain Resource:"*", so the suppression is still
required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new
ids added, superseded ids retained, and the boundary justification's stale "OPEN
follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since
this commit is what closes it. Scanners: RESULT PASS.

Refs: INFRA-186
2026-07-30 17:46:37 -04:00
..
deploy-substrate.template.yaml refactor(iam): scope Lambda execution boundary to per-workload prefixes 2026-07-30 17:46:37 -04:00