seahaven-org-baseline/lib/deploy-substrate/deploy-substrate.template.yaml

1196 lines
61 KiB
YAML
Raw Normal View History

AWSTemplateFormatVersion: "2010-09-09"
Description: >-
Per-account GitHub Actions deploy substrate for Sea Haven Industries:
the shared account-level resources every SAM deploy pipeline needs
(GitHub OIDC provider, Lambda execution permissions boundary, and the
shared CloudFormation execution role). Per-repo githubdeploy-* roles
are NOT here — they are provisioned per repo at migration/onboarding
time in the target account.
# PROVENANCE / DRIFT WARNING
# The Resources below are a VERBATIM extraction of the substrate section
# (OIDC provider + LambdaExecutionBoundary + SamCfnExecutionRole) of
# Sea-Haven-Industries/.github/oidc-deploy-roles.yaml at commit 786dcfe8,
# which remains the deployed source of truth for the management account
# (328440206208) until that account's stacks finish migrating out. If a
# substrate resource must change while both copies are live, change BOTH
# files in the same piece of work. Documented deltas from the source:
# - unused GitHubOrg parameter dropped (only serves the per-repo roles
# left behind),
# - DependsOn: LambdaExecutionBoundary added to SamCfnExecutionRole (the
# role only names the boundary ARN inside Condition strings, so CFN
# infers no edge; first-create needs the boundary to exist first — moot
# for mgmt where both resources already exist, so mgmt's copy is
# deliberately unchanged),
# - DeletionPolicy/UpdateReplacePolicy Retain on the OIDC provider,
# - the boundary-gated IAM block moved from an INLINE role policy into an
# attached managed policy (SamCfnIamManagementPolicy). Forced by IAM's
# 10,240-byte per-role inline limit: mgmt's inline set is ~10.1 KB, i.e.
# ~94 bytes from the cap, so the added Deny statements did not fit and the
# first deploy failed with ServiceLimitExceeded (2026-07-27). Effective
# permissions are unchanged — verified by comparing the full 27-statement
# set before and after the move (identical), since identity policies are
# unioned and an explicit Deny still wins. mgmt received this same
# restructure in Phase B (.github PR #98), so this is no longer a
# divergence,
# - SECURITY FIX (now in BOTH copies): iam:DeleteRolePermissionsBoundary
# removed from Sid IAMPutPermissionsBoundary and explicit Deny statements
# (DenyBoundaryTampering / DenyBoundaryPolicyEdit / DenySelfMutation)
# added. The mgmt copy was remediated 2026-07-27 (.github PRs #95 Phase A
# + #98 Phase B); DenySelfMutation and the widened policy/seahaven-*
# DenyBoundaryPolicyEdit scope were then ported back here, so the two
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# copies' GUARDRAIL statement sets were reconciled as of that date.
# SamCfnIamManagementPolicy and SamCfnExecutionRole remain byte-identical
# across the two files and MUST still be changed together; the only delta
# between them is the DependsOn line above, which is ordering, not
# permission. LambdaExecutionBoundary is NO LONGER byte-identical — see
# DELIBERATE DIVERGENCE below.
#
# DELIBERATE DIVERGENCE — LambdaExecutionBoundary (INFRA-186, 2026-07-30)
# The parity rule above is SCOPED, not global. LambdaExecutionBoundary in THIS
# file is DELIBERATELY STRICTER than the mgmt copy in
# Sea-Haven-Industries/.github/oidc-deploy-roles.yaml. Do not "reconcile" the two
# by copying mgmt's statements back over these, or vice versa; the divergence is
# load-bearing. A future mechanical drift check WILL read it as drift — it is not.
#
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# 1. WHAT DIVERGED. Every per-workload data-plane statement was REMOVED from
# LambdaExecutionBoundary in this file, leaving only the fleet-wide floor:
# CloudWatchLogsWrite (scoped to /aws/lambda*), CloudWatchLogsDescribe,
# XRay and Ec2Eni. The account-wide wildcards mgmt still carries — table/*,
# table/*/index/*, secret:*, parameter/*, sqs :*, function:*, ses
# identity/* + configuration-set/*, kms key/*, and an s3:::*-<accountid>
# pattern that was a bare name-suffix filter rather than an ownership
# check — are simply GONE here rather than re-scoped.
# The security win is the deletion: it is what closes the amplifier whereby
# a principal able to write an inline policy onto a boundary-carrying role
# could read every secret in the account. Per-workload prefixes add no
# security — they only keep a workload functional — so they are added by
# each migration PR, from that stack's own template, when the stack
# actually lands. See the note on the boundary resource for the full
# rationale and the six errors that the pre-loaded approach produced.
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
#
# 2. WHY MGMT'S RATIONALE IS LEGITIMATE THERE. The superset framing this file
# used to carry ("being slightly broad is the correct trade-off; a boundary
# that is too tight will break Lambda functions at runtime AFTER deploy") is
# a real constraint in the management account: 328440206208 has 26 LIVE
# roles carrying seahaven-lambda-execution-boundary, across all five SAM
# stacks. Tightening there is a production change to running workloads with
# a silent, deploy-time-invisible failure mode.
#
# 3. WHY IT DOES NOT TRANSFER HERE. This file deploys ONLY to seahaven-prod
# (011934824531) and seahaven-dev (710827005802), where
# PermissionsBoundaryUsageCount is 0 and 0 respectively (aws iam get-policy,
# verified 2026-07-30; corroborated by list-roles returning no role carrying
# any permissions boundary in either account). No live Lambda can break, so
# the risk that justifies mgmt's breadth is absent — while the exposure is
# strictly WORSE here than in mgmt, because prod is multi-tenant: the old
# wildcards reached proposal-system's, procurement-ingest's and
# workorder-ingest's CDK-owned tables, buckets, secrets and queues, the org's
# own Config and VPC-flow-log buckets, and — via function:* — CDK Lambdas
# that carry no boundary at all.
#
# 4. RECONCILIATION OBLIGATION, RESTATED. For LambdaExecutionBoundary the two
# copies are now INTENTIONALLY DIFFERENT and must NOT be synchronised:
# - A change to the per-workload Resource patterns in THIS file does NOT
# propagate to mgmt.
# - A change to mgmt's boundary does NOT propagate here.
# - Any change to the ACTION lists, or any new statement, is a substrate
# semantic change and DOES still require the same review in both copies.
# For SamCfnIamManagementPolicy and SamCfnExecutionRole the original rule is
# unchanged: change BOTH files in the same piece of work.
#
# KNOWN OPEN ITEM (deferred, not closed by INFRA-186): mgmt 328440206208 still
# carries the account-wide patterns. Tightening it needs its own validated
# rollout — enumerate what the 26 live roles actually call, stage it, and be
# ready to roll back — and is explicitly OUT OF SCOPE of INFRA-186. Until that
# lands, the two copies stay divergent and that is the intended state.
#
# COUPLING: the boundary's ManagedPolicyName and ARN are UNCHANGED and must stay
# unchanged. Four Conditions in SamCfnIamManagementPolicy below, and four more in
# HcptfIamManagementPolicy in lib/terraform-substrate/terraform-substrate.template.yaml,
# pin arn:aws:iam::<account>:policy/seahaven-lambda-execution-boundary by literal
# string inside StringEquals iam:PermissionsBoundary. A rename fails SILENTLY — an
# IAM condition naming a non-existent policy simply never matches, so the
# escalation control would evaporate rather than error — and would additionally
# force a CloudFormation REPLACEMENT that any role carrying the boundary would
# block. INFRA-186 is a CONTENT-ONLY change for exactly this reason;
# terraform-substrate.template.yaml already records that expectation and is
# correctly left untouched.
#
# SIZE BUDGET: an attached managed policy document is capped at 6,144 characters
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# (whitespace excluded). LambdaExecutionBoundary measures 703 characters across
# 4 statements as of 2026-07-30 — the fleet-wide floor only. Measure before
# widening — len(json.dumps(doc,separators=(',',':'))) on the synthesized
fix(iam): restore four permissions the boundary would have denied at migration /sh-security-review (6 detectors + verifier) found four HIGH findings, all the same defect class: the template's permission-source comment block was used as the sanctioned scope source, but it is an incomplete and in places invented secondary record. Each was verified against the real stack template before fixing. None is live today (prod/dev boundary usage is 0); all four would have been AccessDenied at first migration, three of them SILENTLY. - SES configuration-set/seahaven-email-events restored. afterhours-shift-manager template.yaml:178-181 grants it with an in-repo comment stating the send is denied without it. An earlier revision dropped it after checking whether any config set exists in prod/dev today (none do) -- the wrong test. The right question is whether an enumerated stack's own IAM policy names it. - SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it. The prior 'unverified identity fails loudly anyway' argument holds only until the migration verifies the domain, which the migration procedure requires. - scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com only) added. afterhours template.yaml:110-120 needs both; the block omitted them entirely. Failure is silent -- app.py wraps create_schedule in a bare except, so the Slack command reports success and no schedule exists. - secret:afi-slack-webhook-* added. The block named 'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret ARNs are deploy parameters, so the real names live only in that repo's README:48-49 (afi-api-key, afi-slack-webhook). Also corrected, all comment-only: - The permission-source block itself, at each of the four points it was wrong, with the correction and its evidence recorded inline. - The false claim that SAM auto-names async DLQs (it does not -- all four payments queues are hand-written with explicit QueueNames). Replaced with the real invariant: any queue a boundary-carrying function sends to must be payments-* or the boundary widens in the same PR; a denied destination write is silent. - SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the header claimed. Flagged per-workload boundaries as the realistic next move. Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName seahaven-lambda-execution-boundary, so all eight pinning conditions across both guardrail policies still resolve.
2026-07-30 18:17:16 -04:00
# PolicyDocument with ${AWS::AccountId} resolved, and UPDATE THESE TWO NUMBERS in
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# the same edit (they went stale twice inside this branch alone).
fix(iam): restore four permissions the boundary would have denied at migration /sh-security-review (6 detectors + verifier) found four HIGH findings, all the same defect class: the template's permission-source comment block was used as the sanctioned scope source, but it is an incomplete and in places invented secondary record. Each was verified against the real stack template before fixing. None is live today (prod/dev boundary usage is 0); all four would have been AccessDenied at first migration, three of them SILENTLY. - SES configuration-set/seahaven-email-events restored. afterhours-shift-manager template.yaml:178-181 grants it with an in-repo comment stating the send is denied without it. An earlier revision dropped it after checking whether any config set exists in prod/dev today (none do) -- the wrong test. The right question is whether an enumerated stack's own IAM policy names it. - SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it. The prior 'unverified identity fails loudly anyway' argument holds only until the migration verifies the domain, which the migration procedure requires. - scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com only) added. afterhours template.yaml:110-120 needs both; the block omitted them entirely. Failure is silent -- app.py wraps create_schedule in a bare except, so the Slack command reports success and no schedule exists. - secret:afi-slack-webhook-* added. The block named 'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret ARNs are deploy parameters, so the real names live only in that repo's README:48-49 (afi-api-key, afi-slack-webhook). Also corrected, all comment-only: - The permission-source block itself, at each of the four points it was wrong, with the correction and its evidence recorded inline. - The false claim that SAM auto-names async DLQs (it does not -- all four payments queues are hand-written with explicit QueueNames). Replaced with the real invariant: any queue a boundary-carrying function sends to must be payments-* or the boundary widens in the same PR; a denied destination write is silent. - SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the header claimed. Flagged per-workload boundaries as the realistic next move. Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName seahaven-lambda-execution-boundary, so all eight pinning conditions across both guardrail policies still resolve.
2026-07-30 18:17:16 -04:00
#
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# Headroom is 5,441 characters, roughly TWELVE workloads at ~450 each. That is a
# deliberate outcome, not luck: an earlier revision of this branch pre-loaded
# per-workload prefixes for all five mgmt SAM stacks and reached 5,457 characters
# with 687 left — about one workload of room — before any stack had actually
# migrated. Deferring per-workload scope to each migration PR (see the note on
# the boundary itself) removed that pressure entirely. If the budget tightens
# again as workloads land, the end-state fix is per-workload boundaries
# (seahaven-lambda-execution-boundary-<workload>), which also resolves the
# shared-ceiling residual — tracked as INFRA-187, do not improvise it.
fix(iam): restore four permissions the boundary would have denied at migration /sh-security-review (6 detectors + verifier) found four HIGH findings, all the same defect class: the template's permission-source comment block was used as the sanctioned scope source, but it is an incomplete and in places invented secondary record. Each was verified against the real stack template before fixing. None is live today (prod/dev boundary usage is 0); all four would have been AccessDenied at first migration, three of them SILENTLY. - SES configuration-set/seahaven-email-events restored. afterhours-shift-manager template.yaml:178-181 grants it with an in-repo comment stating the send is denied without it. An earlier revision dropped it after checking whether any config set exists in prod/dev today (none do) -- the wrong test. The right question is whether an enumerated stack's own IAM policy names it. - SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it. The prior 'unverified identity fails loudly anyway' argument holds only until the migration verifies the domain, which the migration procedure requires. - scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com only) added. afterhours template.yaml:110-120 needs both; the block omitted them entirely. Failure is silent -- app.py wraps create_schedule in a bare except, so the Slack command reports success and no schedule exists. - secret:afi-slack-webhook-* added. The block named 'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret ARNs are deploy parameters, so the real names live only in that repo's README:48-49 (afi-api-key, afi-slack-webhook). Also corrected, all comment-only: - The permission-source block itself, at each of the four points it was wrong, with the correction and its evidence recorded inline. - The false claim that SAM auto-names async DLQs (it does not -- all four payments queues are hand-written with explicit QueueNames). Replaced with the real invariant: any queue a boundary-carrying function sends to must be payments-* or the boundary widens in the same PR; a denied destination write is silent. - SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the header claimed. Flagged per-workload boundaries as the realistic next move. Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName seahaven-lambda-execution-boundary, so all eight pinning conditions across both guardrail policies still resolve.
2026-07-30 18:17:16 -04:00
# CRITICAL: unlike the 2026-07-27 inline-limit incident,
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# there is NO restructure available when this cap is reached — a role has exactly
# ONE permissions boundary, so statements cannot be spilled into a second attached
# managed policy. At the cap the only levers are prefix consolidation and dropping
# unused actions.
#
# This template is deployed via lib/deploy-substrate-stack.ts
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# (cloudformation-include) as stack seahaven-deploy-substrate, once per member
# account that hosts SAM workloads (currently seahaven-prod 011934824531 and
# seahaven-dev 710827005802 via bin/app.ts instances deploy-substrate-prod /
# deploy-substrate-dev; NEVER mgmt — 328440206208 is served by the .github copy
# named above until its stacks migrate out).
Parameters:
CreateOIDCProvider:
Type: String
Default: "false"
AllowedValues: ["true", "false"]
Description: Set to true only if the GitHub OIDC provider does not already exist in this account
Conditions:
ShouldCreateOIDCProvider: !Equals [!Ref CreateOIDCProvider, "true"]
Resources:
# ---------------------------------------------------------------------------
# OIDC Provider (conditional — most accounts already have it; seahaven-prod
# and seahaven-dev both do, from their githubdeploy-* role provisioning)
# ---------------------------------------------------------------------------
GitHubOIDCProvider:
Type: AWS::IAM::OIDCProvider
Condition: ShouldCreateOIDCProvider
Properties:
Url: https://token.actions.githubusercontent.com
ClientIdList:
- sts.amazonaws.com
ThumbprintList:
- 6938fd4d98bab03faadb97b34396831e3780aea1
# An account has exactly ONE provider per URL and every githubdeploy-* role
# trusts it. Retain so that flipping createOidcProvider back to false (or
# deleting this stack) can never delete the account's federation anchor and
# break every deploy into it.
DeletionPolicy: Retain
UpdateReplacePolicy: Retain
# ---------------------------------------------------------------------------
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# Lambda execution permissions boundary (INFRA-103, re-scoped by INFRA-186)
#
# This managed policy is the CEILING for every Lambda execution role that the
# five SAM stacks auto-generate via AWS::Serverless::Function. Applying it as
# PermissionsBoundary on those roles means the effective permissions are the
# intersection of the role's own policies and this boundary, so a misconfigured
# SAM role can never exceed what is listed here.
#
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# SCOPING RULE (INFRA-186, 2026-07-30). This boundary is the FLEET-WIDE FLOOR
# ONLY: what every Lambda execution role needs regardless of workload. It
# carries NO per-workload data-plane statements — each migrating stack adds its
# own, from its own template, in its own PR (see the note on the boundary
# resource and the WIDENING PATH below). It was previously a deliberate SUPERSET
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# with account-wide wildcards (table/*, secret:*, sqs :*, function:*,
# parameter/*, and an s3:::*-<accountid> pattern that was a name-suffix filter,
# not an ownership check). That trade-off was made when the only account
# carrying this policy hosted nothing but the five SAM stacks. It does not
# survive multi-tenancy: seahaven-prod now hosts CDK-owned tenants
# (proposal-system, procurement-ingest, workorder-ingest) whose tables, buckets,
# secrets and queues those wildcards reached, and whose Lambdas carry NO
# permissions boundary at all — making function:* a boundary-escape primitive.
#
# Tightening here carries ZERO runtime risk and was sequenced deliberately:
# PermissionsBoundaryUsageCount is 0 in BOTH accounts this template deploys to
# (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified
# 2026-07-30), so no live Lambda can break. A boundary that is slightly TOO
# TIGHT is recoverable here — the migrating stack widens it in its own PR before
# its first deploy — whereas leaving it loose perpetuates the exposure. Prefer
# tighter; the widening path is below.
#
# A resource pattern that genuinely CANNOT be scoped keeps its wildcard WITH a
# written justification on the statement: CloudWatchLogsDescribe, XRay and
# Ec2Eni name runtime-created resources or use actions AWS authorises against
# "*" regardless of the ARN supplied. Do not "tighten" those.
#
# WIDENING PATH — read this before migrating a stack into prod or dev.
# The boundary is never widened by the person who hits the AccessDenied. It is
# widened by the migrating stack's owner, in THIS repo, BEFORE the workload's
# first deploy into the target account:
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# 0. PRECONDITION — check the managed-policy VERSION budget BEFORE merging:
# max 5 versions, both
# accounts are on v1 today. Every widening — and every Description-only
# edit — burns one. Delete the oldest non-default version if at 5.
# 1. Derive the workload's needs from ITS OWN TEMPLATE — open the stack's
# template.yaml and read the actual IAM policy statements. The
# permission-source block below is a STARTING POINT, NOT THE AUTHORITY:
# the /sh-security-review pass on 2026-07-30 found SIX places where it was
# incomplete or simply invented a resource name, three of which would have
# failed silently. Update that block in the same edit with what you find.
# 2. Add the workload's statements. Group by service so a second workload can
# extend a Resource list rather than duplicate an action list (~250
# characters for zero new actions; an extra ARN costs ~60). Check for the
# SILENT classes specifically: a denied SQS destination/DLQ write discards
# the async event with no error and no alarm; a denied scheduler call may
# sit behind a bare except; a denied KMS decrypt for env-var encryption
# fails at cold-start INIT; and any CMK-encrypted resource needs the
# matching kms:ViaService principal, not just the kms action.
# 3. Measure. See SIZE BUDGET in the header. There is NO escape hatch.
# 4. Both review gates run and neither discharges the other: the GPT-4.1
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# cross-family review against the real diff, and /sh-security-review
# (IaC/IAM is on the mandatory surface). CLI down = review outstanding.
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# 5. Merge and let CI deploy deploy-substrate-prod / deploy-substrate-dev to
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# UPDATE_COMPLETE, THEN deploy the workload stack.
# ORDERING IS NOT ENFORCED BY CLOUDFORMATION AND THIS IS THE MOST IMPORTANT
# SENTENCE HERE: the workload's deploy SUCCEEDS even against a stale boundary,
# because seahaven-cfn-exec-iam-management's gate checks that the boundary ARN
# is attached, never its contents. The failure surfaces later, at first invoke,
# as AccessDenied. A stale boundary is a silent deploy-time pass and a loud
# production failure.
#
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# CONSIDERED AND REJECTED: a Deny statement reserving the seahaven-* namespace.
# With the Allow set now enumerated per workload it is fully redundant (verified
# 2026-07-30: seahaven-prod-config-* and seahaven-prod-vpc-flow-logs-* are
# already denied by the Allow set alone), and a Deny inside a BOUNDARY is the
# hardest failure mode in the estate to debug — it beats every Allow in every
# policy with no synth-time signal. Revisit only if a widening ever has to
# re-broaden a per-service Resource list back toward a wildcard.
#
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# NOTE ON WHAT THESE PREFIXES ARE. All five stacks below currently live in the
# MANAGEMENT account and none of their resources exists in seahaven-prod or
# seahaven-dev yet. These are MIGRATION-CANDIDATE prefixes for the accounts this
# template deploys to, not an inventory of what is deployed there. They are the
# sanctioned scope source because they are the enumerated permission sources of
# the stacks this boundary exists to cap.
#
# Permission sources per stack (verified live 2026-07-30; this block is the
# sanctioned source for every Resource pattern above — keep it accurate):
#
# afterhours-shift-manager (functions: afterhours-*, 6 live)
# - DynamoDB CRUD (afterhours-shifts table)
# - secretsmanager:GetSecretValue (afterhours-shift-manager/*)
fix(iam): restore four permissions the boundary would have denied at migration /sh-security-review (6 detectors + verifier) found four HIGH findings, all the same defect class: the template's permission-source comment block was used as the sanctioned scope source, but it is an incomplete and in places invented secondary record. Each was verified against the real stack template before fixing. None is live today (prod/dev boundary usage is 0); all four would have been AccessDenied at first migration, three of them SILENTLY. - SES configuration-set/seahaven-email-events restored. afterhours-shift-manager template.yaml:178-181 grants it with an in-repo comment stating the send is denied without it. An earlier revision dropped it after checking whether any config set exists in prod/dev today (none do) -- the wrong test. The right question is whether an enumerated stack's own IAM policy names it. - SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it. The prior 'unverified identity fails loudly anyway' argument holds only until the migration verifies the domain, which the migration procedure requires. - scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com only) added. afterhours template.yaml:110-120 needs both; the block omitted them entirely. Failure is silent -- app.py wraps create_schedule in a bare except, so the Slack command reports success and no schedule exists. - secret:afi-slack-webhook-* added. The block named 'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret ARNs are deploy parameters, so the real names live only in that repo's README:48-49 (afi-api-key, afi-slack-webhook). Also corrected, all comment-only: - The permission-source block itself, at each of the four points it was wrong, with the correction and its evidence recorded inline. - The false claim that SAM auto-names async DLQs (it does not -- all four payments queues are hand-written with explicit QueueNames). Replaced with the real invariant: any queue a boundary-carrying function sends to must be payments-* or the boundary widens in the same PR; a denied destination write is silent. - SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the header claimed. Flagged per-workload boundaries as the realistic next move. Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName seahaven-lambda-execution-boundary, so all eight pinning conditions across both guardrail policies still resolve.
2026-07-30 18:17:16 -04:00
# - ses:SendEmail on the identity AND on
# configuration-set/seahaven-email-events (template.yaml:178-181 — the
# send is DENIED without the config-set ARN when the identity has a
# default configuration set)
# - scheduler:CreateSchedule/DeleteSchedule/GetSchedule on
# schedule/default/holiday-* + iam:PassRole to scheduler.amazonaws.com
# (template.yaml:110-120). CORRECTED 2026-07-30: this block previously
# omitted both, and the omission is SILENT at runtime (bare except).
# - NO ssm. CORRECTED 2026-07-30: this block previously credited
# ssm:GetParameter to this stack; `grep -c 'ssm:' template.yaml` = 0.
# Its slack tokens come from Secrets Manager and the channel id from a
# CloudFormation parameter.
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - lambda:InvokeFunction (ReleaseNotifyInvokeRole,
# HolidaySchedulerExecutionRole — these two carry the boundary and need
# ONLY this action)
# - CloudWatch Logs (all functions)
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - UNRESOLVED: /3cx-scheduler/* ownership (this stack vs. the retired
# standalone 3CX scheduler). Deliberately NOT granted. Resolve at migration.
#
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# payments-dashboard (functions: payments-*)
# - DynamoDB CRUD / Read (PaymentsDashboard table — legacy PascalCase)
fix(iam): restore four permissions the boundary would have denied at migration /sh-security-review (6 detectors + verifier) found four HIGH findings, all the same defect class: the template's permission-source comment block was used as the sanctioned scope source, but it is an incomplete and in places invented secondary record. Each was verified against the real stack template before fixing. None is live today (prod/dev boundary usage is 0); all four would have been AccessDenied at first migration, three of them SILENTLY. - SES configuration-set/seahaven-email-events restored. afterhours-shift-manager template.yaml:178-181 grants it with an in-repo comment stating the send is denied without it. An earlier revision dropped it after checking whether any config set exists in prod/dev today (none do) -- the wrong test. The right question is whether an enumerated stack's own IAM policy names it. - SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it. The prior 'unverified identity fails loudly anyway' argument holds only until the migration verifies the domain, which the migration procedure requires. - scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com only) added. afterhours template.yaml:110-120 needs both; the block omitted them entirely. Failure is silent -- app.py wraps create_schedule in a bare except, so the Slack command reports success and no schedule exists. - secret:afi-slack-webhook-* added. The block named 'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret ARNs are deploy parameters, so the real names live only in that repo's README:48-49 (afi-api-key, afi-slack-webhook). Also corrected, all comment-only: - The permission-source block itself, at each of the four points it was wrong, with the correction and its evidence recorded inline. - The false claim that SAM auto-names async DLQs (it does not -- all four payments queues are hand-written with explicit QueueNames). Replaced with the real invariant: any queue a boundary-carrying function sends to must be payments-* or the boundary widens in the same PR; a denied destination write is silent. - SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the header claimed. Flagged per-workload boundaries as the realistic next move. Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName seahaven-lambda-execution-boundary, so all eight pinning conditions across both guardrail policies still resolve.
2026-07-30 18:17:16 -04:00
# - S3 GetObject on seahaven-payments-csv-* and seahaven-payroll-emails-*;
# GetObject + PutObject on seahaven-payments-boa-raw-*
# (template.yaml:272 and 1097-1099, fetchBoaTransactions raw archive).
# CORRECTED 2026-07-30: this block previously said "GetObject ONLY —
# no write intent enumerated", which was false and would have denied
# the raw-archive write at migration.
# - secretsmanager:GetSecretValue (payments-dashboard/*)
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - sqs Send/Receive/Delete etc. (payments-payroll-batch + DLQs)
# - lambda:InvokeFunction (ExpenseReceiver -> ExpenseProcessor)
# - ec2 ENI lifecycle (VPC-attached functions)
# - KMS via dynamodb (table CMK) — no SSM, no SES
# - CloudWatch Logs
#
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# meal-order-manager (functions: meal-order-manager-*, 7 live)
# - DynamoDB CRUD / Read (meal-order-manager-orders table)
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - S3 CRUD (meal-order-manager-reports-*, meal-order-manager-form-*)
# - secretsmanager:GetSecretValue (meal-order-manager/*)
# - ssm:GetParameter (/meal-order-manager/*)
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - lambda:InvokeFunction (submit-order -> slack-notifier,
# close-form -> aggregate-orders, plus AdminAuthorizerInvokeRole)
# - ses:SendRawEmail
# - CloudWatch Logs
#
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# front-integrations (functions: front-*)
# - DynamoDB CRUD (front-sla-alerts table)
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - secretsmanager:GetSecretValue (front-integrations/*)
# - CloudWatch Logs
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - no S3 / SQS / SSM / SES / KMS / VPC
#
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# afi-backup-monitor (functions: afi-*)
fix(iam): restore four permissions the boundary would have denied at migration /sh-security-review (6 detectors + verifier) found four HIGH findings, all the same defect class: the template's permission-source comment block was used as the sanctioned scope source, but it is an incomplete and in places invented secondary record. Each was verified against the real stack template before fixing. None is live today (prod/dev boundary usage is 0); all four would have been AccessDenied at first migration, three of them SILENTLY. - SES configuration-set/seahaven-email-events restored. afterhours-shift-manager template.yaml:178-181 grants it with an in-repo comment stating the send is denied without it. An earlier revision dropped it after checking whether any config set exists in prod/dev today (none do) -- the wrong test. The right question is whether an enumerated stack's own IAM policy names it. - SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it. The prior 'unverified identity fails loudly anyway' argument holds only until the migration verifies the domain, which the migration procedure requires. - scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com only) added. afterhours template.yaml:110-120 needs both; the block omitted them entirely. Failure is silent -- app.py wraps create_schedule in a bare except, so the Slack command reports success and no schedule exists. - secret:afi-slack-webhook-* added. The block named 'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret ARNs are deploy parameters, so the real names live only in that repo's README:48-49 (afi-api-key, afi-slack-webhook). Also corrected, all comment-only: - The permission-source block itself, at each of the four points it was wrong, with the correction and its evidence recorded inline. - The false claim that SAM auto-names async DLQs (it does not -- all four payments queues are hand-written with explicit QueueNames). Replaced with the real invariant: any queue a boundary-carrying function sends to must be payments-* or the boundary widens in the same PR; a denied destination write is silent. - SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the header claimed. Flagged per-workload boundaries as the realistic next move. Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName seahaven-lambda-execution-boundary, so all eight pinning conditions across both guardrail policies still resolve.
2026-07-30 18:17:16 -04:00
# - secretsmanager:GetSecretValue on TWO bare, unprefixed secrets:
# afi-api-key and afi-slack-webhook. CORRECTED 2026-07-30: this block
# previously named "afi-backup-monitor/slack-webhook-url", which does
# not exist. Both ARNs are deploy PARAMETERS in that stack
# (AfiApiKeySecretArn / SlackWebhookSecretArn), so no name is
# discoverable from its template — the names are in its README:48-49.
# - CloudWatch Logs
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# - nothing else
#
# ---------------------------------------------------------------------------
LambdaExecutionBoundary:
Type: AWS::IAM::ManagedPolicy
Properties:
ManagedPolicyName: seahaven-lambda-execution-boundary
Description: >-
Permissions boundary ceiling for all SAM-managed Lambda execution roles.
Applied via PermissionsBoundary on every Globals.Function in the five
SAM stacks (INFRA-103). Effective permissions are the intersection of
this policy and the role's own inline policies.
PolicyDocument:
Version: "2012-10-17"
Statement:
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# ── CloudWatch Logs — write (every Lambda) ──────────────────────────
# Scoped to the Lambda log-group namespace. Every SAM function's group
# is /aws/lambda/<function>, and the trailing * also covers the
# :log-stream:<name> suffix PutLogEvents authorises against, so one ARN
# serves CreateLogGroup, CreateLogStream, PutLogEvents and
# DescribeLogStreams. The * is deliberately NOT after a trailing slash:
# /aws/lambda* also matches the /aws/lambda-insights groups the Lambda
# Insights extension writes to, which /aws/lambda/* would have denied.
# VERIFICATION PROVENANCE, stated precisely (2026-07-30). What
# iam simulate-custom-policy DOES confirm: this pattern allows
# logs:CreateLogGroup / logs:PutLogEvents on the bare group ARN
# log-group:/aws/lambda/<fn> (and /aws/lambda/<a>/<b>), and DENIES
# log-group:seahaven-prod-vpc-flow-logs — the latter re-checked against
# an Allow */* positive control, which allows it, so the deny is real
# policy behaviour and not a simulator artifact.
# What the simulator CANNOT evaluate, so do NOT claim it was verified:
# log-stream-qualified ARNs (log-group:<g>:log-stream:<s>) and the bare
# /aws/lambda-insights group both return implicitDeny EVEN UNDER an
# Allow */* policy. That is a simulator resource-parsing limitation, not
# a denial. Coverage of those two rests on documented IAM wildcard
# semantics — "*" matches any sequence of characters including ":" and
# "/" — which is why the * is deliberately NOT placed after a trailing
# slash. If this ever needs true end-to-end proof, it must come from a
# real invoke in dev, not from the simulator.
#
# NOT scoped per workload, deliberately. A per-stack prefix
# (/aws/lambda/payments-* etc.) was considered and rejected: a Lambda
# denied PutLogEvents does not fail — it keeps running and silently
# produces no logs. Log denial is the one failure class in this policy
# that is NOT loud, so it must not depend on function-name discipline.
# ACCEPTED RESIDUAL RISK: a SAM Lambda can write into another tenant's
# /aws/lambda/* group (log poisoning). No read action is granted here, so
# this is not an exfiltration path. Tighten only once every
# boundary-carrying function is confirmed to set an explicit FunctionName.
- Sid: CloudWatchLogsWrite
Effect: Allow
Action:
- logs:CreateLogGroup
- logs:CreateLogStream
- logs:PutLogEvents
- logs:DescribeLogStreams
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
Resource:
- !Sub "arn:aws:logs:us-east-1:${AWS::AccountId}:log-group:/aws/lambda*"
# ── CloudWatch Logs — describe (UNSCOPABLE, kept "*" deliberately) ──
# logs:DescribeLogGroups is a COLLECTION action: AWS authorises it
# against "*" regardless of any resource ARN supplied. Scoping it would
# produce a policy that reads tighter and denies at runtime, so it is
# split into its own statement and keeps the wildcard. Read-only
# metadata; it cannot mutate anything or return log content.
- Sid: CloudWatchLogsDescribe
Effect: Allow
Action:
- logs:DescribeLogGroups
Resource: "*"
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# ── X-Ray tracing (UNSCOPABLE, kept "*" deliberately) ───────────────
# xray:PutTraceSegments / PutTelemetryRecords support no resource-level
# permissions — X-Ray exposes no ARN for them, which is why the AWS
# managed AWSXRayDaemonWriteAccess also uses "*". Any ARN written here
# would be inert and would falsely imply a control exists. Write-only
# into this account's own trace store; no cross-tenant read is
# expressible with this action set.
- Sid: XRay
Effect: Allow
Action:
- xray:PutTraceSegments
- xray:PutTelemetryRecords
Resource: "*"
refactor(iam): scope Lambda execution boundary to per-workload prefixes Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of deploy-substrate.template.yaml from account-wide wildcards to per-workload resource prefixes drawn from the template's own permission-source block. This is the PROD/DEV HALF of INFRA-186. What was scoped (wildcard -> per-workload prefix): - dynamodb table/* + table/*/index/* -> afterhours-*, front-*, meal-order-manager-*, PaymentsDashboard*, payments-dashboard-* (a trailing * after each prefix also covers the /index/* GSI ARNs, so the separate table/*/index/* entry is deleted rather than replaced) - s3 *-${AccountId} -> meal-order-manager-*-${AccountId} (read/write) and seahaven-payments-* / seahaven-payroll-emails-* (read-only). The removed pattern was not an ownership check at all: S3 ARNs carry no account field, so it was a bare name-suffix filter that matched 8 of 9 buckets in prod -- including the org's own Config and VPC-flow-log buckets -- with PutObject and DeleteObject. - secretsmanager secret:* -> five <stack>/ prefixes + the legacy bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing key, i.e. a webhook-forgery primitive. - ssm parameter/* -> afterhours-shift-manager and meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath authorises against the path, not the leaf) - sqs :* -> payments-* - lambda function:* -> afterhours-*, meal-order-manager-*, payments-*. Highest-leverage fix here: an invoked function runs under its OWN role, and every non-SAM function in prod is CDK-deployed with no boundary, so function:* was a boundary-escape primitive, not just lateral movement. - ses identity/* + configuration-set/* -> the two verified prod identities; configuration-set dropped (zero exist) - logs split into a scoped write half (/aws/lambda*) and a wildcard describe half (DescribeLogGroups is a collection action AWS authorises against "*" regardless of the ARN supplied) Deliberately NOT tightened, each with written justification on the statement: CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use actions that support no resource-level permissions. KMS keeps key/* -- key ARNs carry UUID key ids, not workload names -- and is constrained by a kms:ViaService condition instead, which inherits the per-workload scoping of the services above for free. No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified 2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that is slightly too tight is recoverable -- the migrating stack widens it in its own PR before its first deploy -- whereas leaving it loose perpetuates the exposure. The widening path and its ordering hazard are documented in the template. mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The management account (328440206208) uses a separate copy in Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE boundary-carrying roles, where tightening is a production change with a silent, deploy-time-invisible failure mode; it needs its own validated rollout and is explicitly out of scope. The header's parity rule is therefore now SCOPED, not global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical and must still change together, while LambdaExecutionBoundary must NOT be reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a future mechanical drift check does not "fix" it away, following the same pattern terraform-substrate.template.yaml uses for its divergences. Content-only change: ManagedPolicyName, the policy ARN and the logical id LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary conditions across this file and terraform-substrate.template.yaml pin the boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming a non-existent policy simply never matches. Verification: - npx tsc --noEmit: clean - npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds - synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY changed resource; GitHubOIDCProvider, SamCfnExecutionRole and SamCfnIamManagementPolicy are byte-identical - policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements, identical in both accounts - iam simulate-custom-policy against live prod, every deny re-checked against an Allow */* positive control: 11/11 cross-tenant denies are real (Config and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials, workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs, WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com) and 23/23 enumerated workload resources still allow Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because three statements legitimately retain Resource:"*", so the suppression is still required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new ids added, superseded ids retained, and the boundary justification's stale "OPEN follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since this commit is what closes it. Scanners: RESULT PASS. Refs: INFRA-186
2026-07-30 17:46:37 -04:00
# ── VPC / ENI management (UNSCOPABLE, kept "*" deliberately) ────────
# Matches AWSLambdaVPCAccessExecutionRole exactly, and for the same
# reasons. The three ec2:Describe* actions do not support resource-level
# permissions AT ALL — an ARN in Resource is ignored and the call is
# authorised against "*" — so narrowing them is cosmetic. The ENI in
# CreateNetworkInterface / DeleteNetworkInterface is created by the
# Lambda service at attach time with an id that cannot exist when this
# policy is written. Nothing here is scopable by resource name.
# AssignPrivateIpAddresses / UnassignPrivateIpAddresses are for EFA and
# secondary IPs — not part of the Lambda ENI lifecycle — omitted.
#
# KNOWN OPEN ITEM (pre-existing, NOT introduced by INFRA-186):
# ec2:DeleteNetworkInterface on "*" lets a bounded Lambda delete any ENI
# in the account, including NAT / VPC-endpoint / RDS ENIs — a
# denial-of-service primitive inherited from the AWS managed policy. The
# durable fix is a Condition on ec2:Subnet / ec2:Vpc naming the VPC the
# Lambda fleet attaches to. That VPC does not exist in seahaven-prod or
# seahaven-dev today (payments-dashboard's 10.20.0.0/16 VPC is in mgmt),
# so writing the condition now would encode an mgmt resource into a
# prod/dev template. Whoever brings the VPC across in payments-dashboard's
# migration PR adds the condition in the same PR.
- Sid: Ec2Eni
Effect: Allow
Action:
- ec2:CreateNetworkInterface
- ec2:DescribeNetworkInterfaces
- ec2:DeleteNetworkInterface
- ec2:DescribeSubnets
- ec2:DescribeSecurityGroups
- ec2:DescribeVpcs
Resource: "*"
refactor(iam): reduce boundary to the fleet-wide floor, defer per-workload scope Adam's call after review: the security win of INFRA-186 comes from DELETING the account-wide wildcards, not from enumerating replacements. Per-workload prefixes add no security -- they only keep a workload functional -- and widening a boundary is the safe direction (adding a resource never breaks a running Lambda; only tightening does). So the per-workload scope moves to each migration PR, which has the stack's real template open in front of it. Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni -- the statements every Lambda needs regardless of workload, and also the silent-failure classes, which is why they belong in the floor. KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups (verified). A workload bringing a CMK adds the statement plus the matching kms:ViaService principal in its own PR. Why not keep the enumeration: it required predicting five stacks' needs from this file's own permission-source comment block, and /sh-security-review found SIX errors in the result -- three silent. The block is a secondary record, not an authority. Deriving scope per-migration from the owning template removes the whole error class. Effect on the security objective: unchanged. secret:*, table/*, function:*, sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the amplifier is closed identically. Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all updated to match; widening path now leads with 'read the stack's own template', names the silent-failure classes to check, and moves the version-budget check to a precondition instead of a trailing step. Verified unchanged: logical id and ManagedPolicyName, so all eight pinning conditions across both guardrail policies still resolve. Both accounts synth identically at 703 chars.
2026-07-30 19:07:09 -04:00
# ── NO PER-WORKLOAD DATA-PLANE STATEMENTS — BY DESIGN ───────────────
# The statements above are the FLEET-WIDE FLOOR: what every Lambda
# execution role needs regardless of which workload it belongs to.
# There are deliberately NO DynamoDB, S3, Secrets Manager, SSM, SQS,
# SES, KMS, lambda:InvokeFunction or scheduler statements here.
#
# WHY (decided 2026-07-30, Adam):
# The security win of INFRA-186 comes from DELETION, not enumeration.
# Removing the account-wide secret:*, table/*, function:* and sqs:*
# wildcards is what closes the amplifier — the ability of a principal
# who can write an inline policy onto a boundary-carrying role to read
# every secret in the account. Per-workload prefixes add no security;
# they exist only to keep a workload FUNCTIONAL once it arrives.
#
# An earlier revision of this branch PRE-LOADED prefixes for all five
# mgmt SAM stacks before any of them had migrated. That required
# predicting five stacks' permission needs from the permission-source
# comment block above, and the /sh-security-review pass found SIX
# errors in the result — three of which would have failed SILENTLY at
# first migration (afterhours' SES config-set, its holiday scheduler
# behind a bare except, and afi's webhook secret under an invented
# name). The block is a secondary record and is not a substitute for
# reading the owning repo's template.
#
# WIDENING IS THE SAFE DIRECTION. Adding a resource to a boundary can
# never break a running Lambda; only tightening can. So there is no
# cost to deferring per-workload scope to the migration PR that has
# the real template open in front of it — and a large cost to
# guessing it years ahead of the migration.
#
# CONSEQUENCE FOR EVERY MIGRATION PR (mandatory, see WIDENING PATH in
# the header): a stack landing in prod or dev MUST add its own
# data-plane statements here, derived from ITS OWN template, in the
# same PR that deploys it. Without them its Lambdas get AccessDenied
# at first invoke. The permission-source block above is the starting
# point, NOT the authority — verify every entry against the stack.
#
# Prod/dev boundary usage is 0 (verified 2026-07-30), so this floor
# currently constrains nothing that exists. Fleet-wide statements that
# genuinely cannot be scoped (Logs, X-Ray, ENI) stay above with their
# justifications; they are also the SILENT-failure classes, which is
# why they belong in the floor rather than in per-workload widenings.
#
# KMS is absent deliberately: both accounts have ZERO CMK-encrypted
# log groups today (verified 2026-07-30). A workload bringing a
# CMK-encrypted resource adds a KMS statement with the matching
# kms:ViaService principal in its own migration PR — a missing
# ViaService entry denies, and for env-var encryption it fails at
# cold-start INIT.
#
# The end-state fix for the shared-ceiling residual (one boundary =
# every SAM workload reaches every other's data plane once they land)
# is per-workload boundaries — tracked as INFRA-187. Do not improvise
# it: the load-bearing problem there is that both guardrail policies
# pin ONE literal boundary ARN inside StringEquals conditions, and
# loosening that to a wildcard weakens the gate.
# ---------------------------------------------------------------------------
# Shared CloudFormation execution role (SAM stacks) — INFRA-97 scoped
#
# Replaces the previous blanket managed-policy set (IAMFullAccess +
# *FullAccess) with per-service inline statements that cover exactly
# what the five SAM stacks need during a CloudFormation deploy/update.
#
# PRIMARY ESCALATION CONTROL
# iam:CreateRole and iam:AttachRolePolicy / iam:PutRolePolicy are
# conditioned on iam:PermissionsBoundary StringEquals the boundary ARN
# (seahaven-lambda-execution-boundary, created in INFRA-103). That
# condition is what prevents the CFN execution role from minting an
# unconstrained admin role.
#
# SAM RolePath deviation note
# The original cross-review suggestion mentioned scoping IAM role
# creation to a specific path (/cfn-managed/). AWS::Serverless::Function
# does NOT support a custom RolePath on auto-generated execution roles —
# the PermissionsBoundary property is supported, but the role always lands
# at path /. Relying on a path condition (iam:ResourceTag or path-prefix)
# would therefore exclude the SAM auto-roles and break every deploy.
# The iam:PermissionsBoundary condition achieves the same security goal
# without requiring a path. For any explicit AWS::IAM::Role resources
# in SAM templates (e.g. AdminAuthorizerInvokeRole in meal-order-manager)
# where we can control the path, path scoping can be added in a follow-up.
#
# DEPLOY ORDER DEPENDENCY
# This role references the boundary ARN only as literal !Sub strings inside
# Condition values, so CloudFormation infers NO creation edge from the
# references alone. The explicit DependsOn below is what guarantees the
# boundary exists before the role on first create (IAM would otherwise
# accept the role, leaving a window where the role exists unbounded-gated
# against a not-yet-existing boundary policy).
# ---------------------------------------------------------------------------
SamCfnExecutionRole:
Type: AWS::IAM::Role
DependsOn: LambdaExecutionBoundary
Properties:
RoleName: github-cfn-execution-role
ManagedPolicyArns:
- !Ref SamCfnIamManagementPolicy
AssumeRolePolicyDocument:
Version: "2012-10-17"
Statement:
- Effect: Allow
Principal:
Service: cloudformation.amazonaws.com
Action: sts:AssumeRole
Policies:
# ── CloudFormation transforms (SAM macro) ─────────────────────────
- PolicyName: cloudformation-transforms
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: AllowSAMTransform
Effect: Allow
Action:
- cloudformation:CreateChangeSet
Resource:
- arn:aws:cloudformation:us-east-1:aws:transform/*
# ── Lambda management ─────────────────────────────────────────────
# Covers function create/update/delete, aliases, event source
# mappings, and Lambda layers — all needed for SAM deploys.
- PolicyName: lambda-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: LambdaFunctions
Effect: Allow
Action:
- lambda:AddPermission
- lambda:CreateFunction
- lambda:DeleteFunction
- lambda:GetFunction
- lambda:GetFunctionConfiguration
- lambda:ListFunctions
- lambda:RemovePermission
- lambda:UpdateFunctionCode
- lambda:UpdateFunctionConfiguration
- lambda:UpdateFunctionEventInvokeConfig
- lambda:PutFunctionEventInvokeConfig
- lambda:DeleteFunctionEventInvokeConfig
- lambda:GetFunctionEventInvokeConfig
- lambda:ListTags
- lambda:TagResource
- lambda:UntagResource
- lambda:GetPolicy
- lambda:ListVersionsByFunction
- lambda:PublishVersion
- lambda:CreateAlias
- lambda:DeleteAlias
- lambda:UpdateAlias
- lambda:GetAlias
Resource:
- !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:function:*"
- Sid: LambdaLayers
Effect: Allow
Action:
- lambda:PublishLayerVersion
- lambda:DeleteLayerVersion
- lambda:GetLayerVersion
- lambda:ListLayerVersions
- lambda:ListLayers
- lambda:AddLayerVersionPermission
- lambda:RemoveLayerVersionPermission
Resource:
- !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:layer:*"
- Sid: LambdaEventSourceMappings
Effect: Allow
Action:
- lambda:CreateEventSourceMapping
- lambda:DeleteEventSourceMapping
- lambda:GetEventSourceMapping
- lambda:ListEventSourceMappings
- lambda:UpdateEventSourceMapping
Resource: "*"
# ── API Gateway (HTTP APIs + REST APIs) ───────────────────────────
- PolicyName: apigateway-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: ApiGateway
Effect: Allow
Action:
- apigateway:GET
- apigateway:POST
- apigateway:PUT
- apigateway:PATCH
- apigateway:DELETE
Resource:
- "arn:aws:apigateway:us-east-1::*"
# ── DynamoDB ──────────────────────────────────────────────────────
- PolicyName: dynamodb-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: DynamoDBTables
Effect: Allow
Action:
- dynamodb:CreateTable
- dynamodb:DeleteTable
- dynamodb:DescribeTable
- dynamodb:UpdateTable
- dynamodb:ListTables
- dynamodb:TagResource
- dynamodb:UntagResource
- dynamodb:DescribeTimeToLive
- dynamodb:UpdateTimeToLive
- dynamodb:DescribeContinuousBackups
- dynamodb:UpdateContinuousBackups
Resource:
- !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/*"
# ── S3 ────────────────────────────────────────────────────────────
# Covers bucket create/configure + object operations for SAM
# artifact buckets and application buckets.
- PolicyName: s3-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: S3BucketOps
Effect: Allow
Action:
- s3:CreateBucket
- s3:DeleteBucket
- s3:GetBucketLocation
- s3:GetBucketPolicy
- s3:PutBucketPolicy
- s3:DeleteBucketPolicy
- s3:GetBucketTagging
- s3:PutBucketTagging
- s3:GetBucketVersioning
- s3:PutBucketVersioning
- s3:GetLifecycleConfiguration
- s3:PutLifecycleConfiguration
- s3:GetBucketPublicAccessBlock
- s3:PutBucketPublicAccessBlock
# Explicit BucketEncryption blocks (first: payments-dashboard
# BoaRawBucket, 2026-07-22) need the encryption config pair.
- s3:GetEncryptionConfiguration
- s3:PutEncryptionConfiguration
- s3:GetBucketNotification
- s3:PutBucketNotification
- s3:GetBucketWebsite
- s3:PutBucketWebsite
- s3:DeleteBucketWebsite
- s3:GetBucketAcl
- s3:PutBucketAcl
Resource:
- "arn:aws:s3:::*"
- Sid: S3ObjectOps
Effect: Allow
Action:
- s3:GetObject
- s3:PutObject
- s3:DeleteObject
- s3:ListBucket
- s3:ListBucketVersions
- s3:GetObjectVersion
Resource:
- "arn:aws:s3:::*"
- "arn:aws:s3:::*/*"
# ── CloudWatch Logs ───────────────────────────────────────────────
- PolicyName: cloudwatch-logs-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: CWLogs
Effect: Allow
Action:
- logs:CreateLogGroup
- logs:DeleteLogGroup
- logs:DescribeLogGroups
- logs:PutRetentionPolicy
- logs:DeleteRetentionPolicy
- logs:ListTagsLogGroup
- logs:TagLogGroup
- logs:UntagLogGroup
- logs:ListTagsForResource
- logs:TagResource
- logs:UntagResource
- logs:CreateLogDelivery
- logs:GetLogDelivery
- logs:UpdateLogDelivery
- logs:DeleteLogDelivery
- logs:ListLogDeliveries
- logs:PutResourcePolicy
- logs:DescribeResourcePolicies
- logs:PutDestination
- logs:DeleteDestination
- logs:DescribeDestinations
- logs:AssociateKmsKey
- logs:DisassociateKmsKey
# Ported from the mgmt copy (Phase A): afterhours-shift-manager
# creates an AWS::Logs::MetricFilter through this role, so a
# SAM stack migrating here fails mid-deploy without these.
- logs:PutMetricFilter
- logs:DeleteMetricFilter
- logs:DescribeMetricFilters
Resource: "*"
# ── EventBridge / CloudWatch Events (scheduled Lambdas) ───────────
- PolicyName: eventbridge-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: EventBridge
Effect: Allow
Action:
- events:DeleteRule
- events:DescribeRule
- events:EnableRule
- events:DisableRule
- events:ListRules
- events:ListTargetsByRule
- events:PutRule
- events:PutTargets
- events:RemoveTargets
- events:TagResource
- events:UntagResource
- events:ListTagsForResource
- events:PutPermission
- events:RemovePermission
Resource: "*"
# ── SES (afterhours weekly-post, meal-order email-report) ─────────
- PolicyName: ses-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: SESRules
Effect: Allow
Action:
- ses:CreateReceiptRule
- ses:DeleteReceiptRule
- ses:DescribeReceiptRule
- ses:UpdateReceiptRule
- ses:CreateReceiptRuleSet
- ses:DescribeActiveReceiptRuleSet
- ses:DescribeReceiptRuleSet
- ses:SetActiveReceiptRuleSet
- ses:ReorderReceiptRuleSet
- ses:GetIdentityVerificationAttributes
- ses:ListIdentities
Resource: "*"
# ── SQS (payments-dashboard queues + DLQs) ────────────────────────
- PolicyName: sqs-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: SQSQueues
Effect: Allow
Action:
- sqs:CreateQueue
- sqs:DeleteQueue
- sqs:GetQueueAttributes
- sqs:SetQueueAttributes
- sqs:GetQueueUrl
- sqs:ListQueues
- sqs:TagQueue
- sqs:UntagQueue
- sqs:ListQueueTags
- sqs:AddPermission
- sqs:RemovePermission
Resource:
- !Sub "arn:aws:sqs:us-east-1:${AWS::AccountId}:*"
# ── SNS (validation / alarm notifications) ────────────────────────
- PolicyName: sns-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: SNS
Effect: Allow
Action:
- sns:CreateTopic
- sns:DeleteTopic
- sns:GetTopicAttributes
- sns:SetTopicAttributes
- sns:Subscribe
- sns:Unsubscribe
- sns:ListSubscriptionsByTopic
- sns:ListTopics
- sns:TagResource
- sns:UntagResource
Resource:
- !Sub "arn:aws:sns:us-east-1:${AWS::AccountId}:*"
# ── CloudWatch Alarms ─────────────────────────────────────────────
- PolicyName: cloudwatch-alarms-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: CWAlarms
Effect: Allow
Action:
- cloudwatch:PutMetricAlarm
- cloudwatch:DeleteAlarms
- cloudwatch:DescribeAlarms
- cloudwatch:EnableAlarmActions
- cloudwatch:DisableAlarmActions
- cloudwatch:ListTagsForResource
- cloudwatch:TagResource
- cloudwatch:UntagResource
Resource: "*"
# ── EC2 / VPC / NAT / EIP / Security Groups ───────────────────────
# payments-dashboard deploys a VPC, NAT gateway, EIP, route tables,
# subnets, security groups, and gateway VPC endpoints.
- PolicyName: ec2-vpc-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: EC2VPC
Effect: Allow
Action:
- ec2:AllocateAddress
- ec2:AssociateRouteTable
- ec2:AttachInternetGateway
- ec2:AuthorizeSecurityGroupEgress
- ec2:AuthorizeSecurityGroupIngress
- ec2:CreateInternetGateway
- ec2:CreateNatGateway
- ec2:CreateRoute
- ec2:CreateRouteTable
- ec2:CreateSecurityGroup
- ec2:CreateSubnet
- ec2:CreateVpc
- ec2:CreateVpcEndpoint
- ec2:CreateTags
- ec2:DeleteInternetGateway
- ec2:DeleteNatGateway
- ec2:DeleteRoute
- ec2:DeleteRouteTable
- ec2:DeleteSecurityGroup
- ec2:DeleteSubnet
- ec2:DeleteVpc
- ec2:DeleteVpcEndpoints
- ec2:DescribeAddresses
- ec2:DescribeAvailabilityZones
- ec2:DescribeInternetGateways
- ec2:DescribeNatGateways
- ec2:DescribeRouteTables
- ec2:DescribeSecurityGroups
- ec2:DescribeSubnets
- ec2:DescribeVpcEndpoints
- ec2:DescribeVpcs
- ec2:DescribePrefixLists
- ec2:DetachInternetGateway
- ec2:DisassociateAddress
- ec2:DisassociateRouteTable
- ec2:ModifySubnetAttribute
- ec2:ModifyVpcAttribute
- ec2:ModifyVpcEndpoint
- ec2:ReleaseAddress
- ec2:RevokeSecurityGroupEgress
- ec2:RevokeSecurityGroupIngress
- ec2:UpdateSecurityGroupRuleDescriptionsEgress
- ec2:UpdateSecurityGroupRuleDescriptionsIngress
Resource: "*"
# ── CloudFront + OAC (meal-order-manager form distribution) ───────
- PolicyName: cloudfront-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: CloudFront
Effect: Allow
Action:
- cloudfront:CreateDistribution
- cloudfront:DeleteDistribution
- cloudfront:GetDistribution
- cloudfront:GetDistributionConfig
- cloudfront:UpdateDistribution
- cloudfront:TagResource
- cloudfront:UntagResource
- cloudfront:ListTagsForResource
- cloudfront:CreateOriginAccessControl
- cloudfront:DeleteOriginAccessControl
- cloudfront:GetOriginAccessControl
- cloudfront:GetOriginAccessControlConfig
- cloudfront:UpdateOriginAccessControl
- cloudfront:ListOriginAccessControls
- cloudfront:CreateInvalidation
- cloudfront:GetInvalidation
Resource: "*"
# ── SSM Parameter Store (meal-order-manager, afterhours) ──────────
# Write is needed because meal-order-manager creates
# /meal-order-manager/slack-channel-id via AWS::SSM::Parameter.
- PolicyName: ssm-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: SSMParameters
Effect: Allow
Action:
- ssm:GetParameter
- ssm:GetParameters
- ssm:GetParametersByPath
- ssm:PutParameter
- ssm:DeleteParameter
- ssm:DeleteParameters
- ssm:DescribeParameters
- ssm:AddTagsToResource
- ssm:RemoveTagsFromResource
- ssm:ListTagsForResource
Resource:
- !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/*"
# WAF association needs SSM parameter read at deploy time
# (/seahaven/waf/app-web-acl-arn value lookup)
- Sid: SSMParameterDescribe
Effect: Allow
Action:
- ssm:DescribeParameters
Resource: "*"
# ── WAF (meal-order-manager CloudFront WebACL association) ────────
- PolicyName: waf-management
PolicyDocument:
Version: "2012-10-17"
Statement:
- Sid: WAF
Effect: Allow
Action:
- wafv2:GetWebACL
- wafv2:GetWebACLForResource
- wafv2:ListWebACLs
- wafv2:AssociateWebACL
- wafv2:DisassociateWebACL
- wafv2:ListResourcesForWebACL
Resource: "*"
# ---------------------------------------------------------------------------
# IAM role lifecycle - BOUNDARY-GATED (attached managed policy)
#
# Lives in a MANAGED policy, not inline on the role, because the role's
# inline policies total ~10.1 KB against IAM's hard 10,240-byte per-role
# inline limit - adding the Deny statements below inline exceeds it and
# fails the deploy (ServiceLimitExceeded, hit live 2026-07-27). Attached
# managed policies have their own separate 6,144-byte budget, so moving this
# block out both fits the Denies and leaves ~1.9 KB of inline headroom for
# future statements. Identity policies are unioned and an explicit Deny still
# wins, so effective permissions are unchanged by the relocation.
#
# This is the PRIMARY escalation control for INFRA-97.
#
# iam:CreateRole / iam:AttachRolePolicy / iam:PutRolePolicy are
# conditioned on iam:PermissionsBoundary StringEquals the
# seahaven-lambda-execution-boundary ARN. That condition means
# any role this execution role creates must have the boundary
# applied, so it can never exceed what the boundary allows
# (which is scoped to the services the five stacks actually use).
#
# iam:PassRole is also included here so CloudFormation can pass
# the auto-generated Lambda execution role to the Lambda service.
#
# Why not path-scoped (e.g. iam:ResourceTag / path /cfn-managed/)?
# SAM's AWS::Serverless::Function auto-generates execution roles at
# path / — there is no supported way to set a custom RolePath on
# SAM auto-roles. A path condition would therefore exclude the
# SAM auto-roles and break every deploy. The PermissionsBoundary
# condition achieves the same security goal without a path requirement.
# ---------------------------------------------------------------------------
SamCfnIamManagementPolicy:
Type: AWS::IAM::ManagedPolicy
Properties:
# Fixed name: changing it makes CloudFormation create a replacement policy
# and detach this one, which briefly drops the role's IAM permissions
# mid-update. Treat a rename as a coordinated migration, not an edit. This
# is the role's FIRST attached managed policy (per-role quota is 10).
ManagedPolicyName: seahaven-cfn-exec-iam-management
Description: >-
Boundary-gated IAM role lifecycle for github-cfn-execution-role, plus the
explicit Deny backstops that keep the permissions boundary from being
detached or rewritten. Separated from the role's inline policies to stay
under IAM's 10,240-byte inline limit.
PolicyDocument:
Version: "2012-10-17"
Statement:
# Create role — MUST attach boundary
- Sid: IAMCreateRoleWithBoundary
Effect: Allow
Action:
- iam:CreateRole
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:role/*"
Condition:
StringEquals:
"iam:PermissionsBoundary": !Sub "arn:aws:iam::${AWS::AccountId}:policy/seahaven-lambda-execution-boundary"
# Attach managed policies — MUST have boundary already on role
- Sid: IAMAttachPolicyWithBoundary
Effect: Allow
Action:
- iam:AttachRolePolicy
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:role/*"
Condition:
StringEquals:
"iam:PermissionsBoundary": !Sub "arn:aws:iam::${AWS::AccountId}:policy/seahaven-lambda-execution-boundary"
# Put inline policy — MUST have boundary already on role
- Sid: IAMPutRolePolicyWithBoundary
Effect: Allow
Action:
- iam:PutRolePolicy
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:role/*"
Condition:
StringEquals:
"iam:PermissionsBoundary": !Sub "arn:aws:iam::${AWS::AccountId}:policy/seahaven-lambda-execution-boundary"
# Boundary management — SET the boundary only. DELETE is NOT
# granted: for a delete, the iam:PermissionsBoundary condition key
# reflects the boundary CURRENTLY attached to the target role, so
# a StringEquals condition on the boundary ARN MATCHES exactly the
# roles the gate protects. Granting delete under that condition
# lets this role create a boundary-gated role with an inline *:*
# policy, strip the boundary, and pass the now-unbounded role to
# Lambda — defeating the primary escalation control. Verified live
# against the mgmt copy 2026-07-27 (simulate-principal-policy:
# iam:DeleteRolePermissionsBoundary = allowed). SAM never needs
# the delete: it only SETS the boundary on roles it creates, and
# stack teardown calls DeleteRole, not DeleteRolePermissionsBoundary.
- Sid: IAMPutPermissionsBoundary
Effect: Allow
Action:
- iam:PutRolePermissionsBoundary
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:role/*"
Condition:
StringEquals:
"iam:PermissionsBoundary": !Sub "arn:aws:iam::${AWS::AccountId}:policy/seahaven-lambda-execution-boundary"
# Explicit Deny backstop (AWS's documented NoBoundaryPolicyEdit /
# NoBoundaryDelete delegation pattern). A Deny is required, not
# merely omitting the Allow: without it, any future Allow added to
# this role — or a broader managed policy attached to it — silently
# reopens the escalation. Covers both removing a boundary from a
# role and rewriting the boundary POLICY DOCUMENT itself (the
# latter is only implicitly denied today).
- Sid: DenyBoundaryTampering
Effect: Deny
Action:
- iam:DeleteRolePermissionsBoundary
- iam:DeleteUserPermissionsBoundary
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:role/*"
- !Sub "arn:aws:iam::${AWS::AccountId}:user/*"
# Scoped to the whole seahaven-* policy family, not just the boundary:
# this policy carries the Deny statements, so it is now a
# higher-value target than the boundary it protects. Safe to scope
# broadly — the role holds no iam:CreatePolicy anywhere and no SAM
# stack manages a managed policy through it (both verified
# 2026-07-27), so nothing legitimate writes policy versions here.
- Sid: DenyBoundaryPolicyEdit
Effect: Deny
Action:
- iam:CreatePolicyVersion
- iam:SetDefaultPolicyVersion
- iam:DeletePolicyVersion
- iam:DeletePolicy
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:policy/seahaven-*"
# Self-protection. Without this the whole control is one API call
# from being undone: IAMRoleReadAndDelete below grants
# iam:DetachRolePolicy on Resource "*" with no condition, so this
# role could detach the very policy carrying these Denies from
# itself and reinstate the escalation. Verified live 2026-07-27:
# simulate-principal-policy returned "allowed" for DetachRolePolicy,
# DeleteRolePolicy and DeleteRole against this role's own ARN and
# against githubdeploy-* roles.
#
# Also closes a denial-of-service and a self-elevation precondition:
# iam:PutRolePermissionsBoundary is condition-pinned to the Lambda
# boundary ARN but NOT scoped by target, so this role could apply
# that runtime boundary to itself or to a githubdeploy-* role —
# bricking the pipelines, unrecoverable without an admin because
# removing a boundary is denied above, and making the otherwise-inert
# AttachRolePolicy/PutRolePolicy self-elevation conditions start
# matching.
#
# Costs nothing operationally: this role is only ever passed to
# CloudFormation for SAM application stacks. The substrate's own
# roles are managed by THIS stack (deployed through the CDK
# bootstrap execution role), and per-repo githubdeploy-* roles are
# provisioned at onboarding time outside any stack this role
# executes — so CloudFormation never exercises these actions
# against them as this role. SAM-generated roles are named
# <stack>-<Function>Role-<hash> and are unaffected.
- Sid: DenySelfMutation
Effect: Deny
Action:
- iam:AttachRolePolicy
- iam:DeleteRole
- iam:DeleteRolePolicy
- iam:DeleteRolePermissionsBoundary
- iam:DetachRolePolicy
- iam:PutRolePolicy
- iam:PutRolePermissionsBoundary
- iam:UpdateAssumeRolePolicy
- iam:UpdateRole
- iam:UpdateRoleDescription
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:role/github-cfn-execution-role"
- !Sub "arn:aws:iam::${AWS::AccountId}:role/githubdeploy-*"
# Read / tag / delete role and policy — no boundary condition needed
- Sid: IAMRoleReadAndDelete
Effect: Allow
Action:
- iam:DeleteRole
- iam:DeleteRolePolicy
- iam:DetachRolePolicy
- iam:GetRole
- iam:GetRolePolicy
- iam:ListAttachedRolePolicies
- iam:ListRolePolicies
- iam:ListRoles
- iam:TagRole
- iam:UntagRole
- iam:UpdateRole
- iam:UpdateRoleDescription
- iam:UpdateAssumeRolePolicy
- iam:GetPolicy
- iam:GetPolicyVersion
- iam:ListPolicies
- iam:ListPolicyVersions
Resource: "*"
# PassRole — CloudFormation passes the Lambda execution role
# to the Lambda service. Scoped to SAM-generated role pattern.
- Sid: IAMPassRole
Effect: Allow
Action:
- iam:PassRole
Resource:
- !Sub "arn:aws:iam::${AWS::AccountId}:role/*"
Condition:
StringEquals:
"iam:PassedToService": "lambda.amazonaws.com"