diff --git a/lib/deploy-substrate/deploy-substrate.template.yaml b/lib/deploy-substrate/deploy-substrate.template.yaml index 0a569f9..9a91de2 100644 --- a/lib/deploy-substrate/deploy-substrate.template.yaml +++ b/lib/deploy-substrate/deploy-substrate.template.yaml @@ -53,17 +53,21 @@ Description: >- # by copying mgmt's statements back over these, or vice versa; the divergence is # load-bearing. A future mechanical drift check WILL read it as drift — it is not. # -# 1. WHAT DIVERGED. Every scopable Resource pattern in LambdaExecutionBoundary -# was re-scoped from account-wide wildcards (table/*, table/*/index/*, -# secret:*, parameter/*, sqs :*, function:*, ses identity/* + -# configuration-set/*, and an s3:::*- pattern that was a bare -# name-suffix filter rather than an ownership check) to per-workload -# prefixes drawn from the permission-source block above. The three genuinely -# unscopable statements (CloudWatchLogsDescribe, XRay, Ec2Eni) keep -# Resource "*" with written justification on each. CloudWatch Logs was split -# into a scoped write half and a wildcard describe half. KMS keeps key/* and -# is scoped by a kms:ViaService condition instead, because key ARNs carry -# UUID key ids that cannot be prefix-scoped. +# 1. WHAT DIVERGED. Every per-workload data-plane statement was REMOVED from +# LambdaExecutionBoundary in this file, leaving only the fleet-wide floor: +# CloudWatchLogsWrite (scoped to /aws/lambda*), CloudWatchLogsDescribe, +# XRay and Ec2Eni. The account-wide wildcards mgmt still carries — table/*, +# table/*/index/*, secret:*, parameter/*, sqs :*, function:*, ses +# identity/* + configuration-set/*, kms key/*, and an s3:::*- +# pattern that was a bare name-suffix filter rather than an ownership +# check — are simply GONE here rather than re-scoped. +# The security win is the deletion: it is what closes the amplifier whereby +# a principal able to write an inline policy onto a boundary-carrying role +# could read every secret in the account. Per-workload prefixes add no +# security — they only keep a workload functional — so they are added by +# each migration PR, from that stack's own template, when the stack +# actually lands. See the note on the boundary resource for the full +# rationale and the six errors that the pre-loaded approach produced. # # 2. WHY MGMT'S RATIONALE IS LEGITIMATE THERE. The superset framing this file # used to carry ("being slightly broad is the correct trade-off; a boundary @@ -114,19 +118,21 @@ Description: >- # correctly left untouched. # # SIZE BUDGET: an attached managed policy document is capped at 6,144 characters -# (whitespace excluded). LambdaExecutionBoundary measures 5,457 characters across -# 16 statements as of 2026-07-30 (was 2,507 across 11 before this scoping pass — -# TIGHTENING COSTS CHARACTERS, and the review corrections cost more). Measure -# before widening — len(json.dumps(doc,separators=(',',':'))) on the synthesized +# (whitespace excluded). LambdaExecutionBoundary measures 703 characters across +# 4 statements as of 2026-07-30 — the fleet-wide floor only. Measure before +# widening — len(json.dumps(doc,separators=(',',':'))) on the synthesized # PolicyDocument with ${AWS::AccountId} resolved, and UPDATE THESE TWO NUMBERS in -# the same edit (they went stale once already inside a single branch). +# the same edit (they went stale twice inside this branch alone). # -# ⚠ HEADROOM IS 687 CHARACTERS — roughly ONE more workload at ~450 each, NOT the -# "five" an earlier revision of this header claimed. The next stack to migrate is -# likely to exhaust it. Read the note below before assuming there is room: the -# realistic next move is per-workload boundaries -# (seahaven-lambda-execution-boundary-), which is also the durable fix -# for the shared-ceiling residual documented in the SCOPING RULE. +# Headroom is 5,441 characters, roughly TWELVE workloads at ~450 each. That is a +# deliberate outcome, not luck: an earlier revision of this branch pre-loaded +# per-workload prefixes for all five mgmt SAM stacks and reached 5,457 characters +# with 687 left — about one workload of room — before any stack had actually +# migrated. Deferring per-workload scope to each migration PR (see the note on +# the boundary itself) removed that pressure entirely. If the budget tightens +# again as workloads land, the end-state fix is per-workload boundaries +# (seahaven-lambda-execution-boundary-), which also resolves the +# shared-ceiling residual — tracked as its own ticket, do not improvise it. # CRITICAL: unlike the 2026-07-27 inline-limit incident, # there is NO restructure available when this cap is reached — a role has exactly # ONE permissions boundary, so statements cannot be spilled into a second attached @@ -181,8 +187,11 @@ Resources: # intersection of the role's own policies and this boundary, so a misconfigured # SAM role can never exceed what is listed here. # - # SCOPING RULE (INFRA-186, 2026-07-30). This boundary is a CLOSED ENUMERATION - # of per-workload resource prefixes. It was previously a deliberate SUPERSET + # SCOPING RULE (INFRA-186, 2026-07-30). This boundary is the FLEET-WIDE FLOOR + # ONLY: what every Lambda execution role needs regardless of workload. It + # carries NO per-workload data-plane statements — each migrating stack adds its + # own, from its own template, in its own PR (see the note on the boundary + # resource and the WIDENING PATH below). It was previously a deliberate SUPERSET # with account-wide wildcards (table/*, secret:*, sqs :*, function:*, # parameter/*, and an s3:::*- pattern that was a name-suffix filter, # not an ownership check). That trade-off was made when the only account @@ -209,21 +218,30 @@ Resources: # The boundary is never widened by the person who hits the AccessDenied. It is # widened by the migrating stack's owner, in THIS repo, BEFORE the workload's # first deploy into the target account: - # 1. Add the workload's prefixes to the relevant per-service Resource lists - # above — never add a new per-workload statement (a duplicated action list - # costs ~250 characters for zero new actions; an extra ARN costs ~60). - # Add its entry to the permission-source block below in the same edit: that - # block is the sanctioned scope source, and a prefix added without one will - # be "reconciled" away later. - # 2. Measure. See SIZE BUDGET in the header. There is NO escape hatch. - # 3. Both review gates run and neither discharges the other: the GPT-4.1 - # cross-family review against the real diff, and /sh-security-review - # (IaC/IAM is on the mandatory surface). CLI down = review outstanding. - # 4. Merge and let CI deploy deploy-substrate-prod / deploy-substrate-dev to - # UPDATE_COMPLETE, THEN deploy the workload stack. - # 5. Check the managed-policy VERSION budget first: max 5 versions, both + # 0. PRECONDITION — check the managed-policy VERSION budget BEFORE merging: + # max 5 versions, both # accounts are on v1 today. Every widening — and every Description-only # edit — burns one. Delete the oldest non-default version if at 5. + # 1. Derive the workload's needs from ITS OWN TEMPLATE — open the stack's + # template.yaml and read the actual IAM policy statements. The + # permission-source block below is a STARTING POINT, NOT THE AUTHORITY: + # the /sh-security-review pass on 2026-07-30 found SIX places where it was + # incomplete or simply invented a resource name, three of which would have + # failed silently. Update that block in the same edit with what you find. + # 2. Add the workload's statements. Group by service so a second workload can + # extend a Resource list rather than duplicate an action list (~250 + # characters for zero new actions; an extra ARN costs ~60). Check for the + # SILENT classes specifically: a denied SQS destination/DLQ write discards + # the async event with no error and no alarm; a denied scheduler call may + # sit behind a bare except; a denied KMS decrypt for env-var encryption + # fails at cold-start INIT; and any CMK-encrypted resource needs the + # matching kms:ViaService principal, not just the kms action. + # 3. Measure. See SIZE BUDGET in the header. There is NO escape hatch. + # 4. Both review gates run and neither discharges the other: the GPT-4.1 + # cross-family review against the real diff, and /sh-security-review + # (IaC/IAM is on the mandatory surface). CLI down = review outstanding. + # 5. Merge and let CI deploy deploy-substrate-prod / deploy-substrate-dev to + # UPDATE_COMPLETE, THEN deploy the workload stack. # ORDERING IS NOT ENFORCED BY CLOUDFORMATION AND THIS IS THE MOST IMPORTANT # SENTENCE HERE: the workload's deploy SUCCEEDS even against a stale boundary, # because seahaven-cfn-exec-iam-management's gate checks that the boundary ARN @@ -428,366 +446,59 @@ Resources: - ec2:DescribeVpcs Resource: "*" - # ── DynamoDB — one prefix per enumerated workload ─────────────────── - # Replaces table/* + table/*/index/*, which reached every table in the - # account. A trailing * on each workload prefix covers the base table - # AND its /index/* GSI ARNs in a single entry (IAM wildcards match "/"), - # verified 2026-07-30 with iam simulate-custom-policy against - # table/afterhours-shifts/index/gsi1 — so the separate table/*/index/* - # line is deleted, not replaced. Prod's CDK-owned tables (WorkOrders, - # WorkOrderComments, purchase-orders, verified-sites, pending-site-review) - # match none of these prefixes, which is the point. - # PaymentsDashboard is a legacy PascalCase table exempt from renaming - # under handbook naming-conventions.md; the kebab-case prefix is listed - # alongside it so a future rename cannot silently lock the stack out. - # Table names are NOT stack names (afterhours-shifts, front-sla-alerts) — - # scoping naively off stack name would deny at runtime. - - Sid: DynamoDB - Effect: Allow - Action: - - dynamodb:GetItem - - dynamodb:PutItem - - dynamodb:UpdateItem - - dynamodb:DeleteItem - - dynamodb:Query - - dynamodb:Scan - - dynamodb:BatchGetItem - - dynamodb:BatchWriteItem - - dynamodb:DescribeTable - - dynamodb:ConditionCheckItem - Resource: - - !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/afterhours-*" - - !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/front-*" - - !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/meal-order-manager-*" - - !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/PaymentsDashboard*" - - !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/payments-dashboard-*" - - # ── S3 — write access (meal-order-manager only) ───────────────────── - # THE LARGEST SECURITY WIN IN INFRA-186. The removed - # arn:aws:s3:::*-${AWS::AccountId} was not a per-workload scope at all: - # S3 ARNs carry no account field, so it was a bare NAME-SUFFIX FILTER - # matching every bucket in the account whose name ends in the account id. - # In seahaven-prod today that was 8 of 9 buckets — all four - # proposal-system-*, both ingest-email buckets, AND the org's own - # seahaven-prod-config-* and seahaven-prod-vpc-flow-logs-* — with - # PutObject and DeleteObject. That is anti-forensics capability over the - # org's own security telemetry, handed to the SAM fleet's ceiling. - # Verified 2026-07-30 with iam simulate-custom-policy: the patterns below - # deny seahaven-prod-config-*, seahaven-prod-vpc-flow-logs-* and - # proposal-system-uploads-*, and allow meal-order-manager-reports-*. - # The four meal-order-manager-*-${AWS::AccountId} entries previously - # listed here were strictly redundant — every one ends in - and - # was already matched by the wildcard above them; their comment claiming a - # "non-AccountId suffix pattern" was contradicted by the ARNs beneath it. - # Split by DIRECTION of access: meal-order-manager is the only stack the - # permission-source block gives write intent to (ReportsBucket / - # FormBucket CRUD). Both the bucket and object ARN forms are listed in - # each statement because s3:ListBucket authorises against the bucket ARN - # and s3:GetObject against the object ARN. - - Sid: S3WorkloadReadWrite - Effect: Allow - Action: - - s3:GetObject - - s3:PutObject - - s3:DeleteObject - - s3:ListBucket - - s3:GetBucketLocation - - s3:GetObjectVersion - - s3:GetObjectTagging - - s3:PutObjectTagging - Resource: - - !Sub "arn:aws:s3:::meal-order-manager-*-${AWS::AccountId}" - - !Sub "arn:aws:s3:::meal-order-manager-*-${AWS::AccountId}/*" - - # ── S3 — payments-dashboard ───────────────────────────────────────── - # Split by direction, but note the source of truth: the template - # comment above enumerates "S3 GetObject" for payments-dashboard, and - # that enumeration is INCOMPLETE. Verified against the real stack - # 2026-07-30 — payments-dashboard/template.yaml grants s3:PutObject on - # BoaRawBucket (lines 272 and 1098-1099), so the fetchBoaTransactions - # path writes raw BoA payloads. A read-only grant here would deny that - # write at migration time. Write is therefore allowed on the boa-raw - # bucket ONLY; payroll-emails and payments-csv stay read-only, so a - # compromised payments function still cannot delete payroll evidence. - # Buckets carry the org-wide seahaven- prefix rather than a payments- - # one, which is why bucket names cannot be derived from stack names. - - Sid: S3PaymentsBoaRawWrite - Effect: Allow - Action: - - s3:GetObject - - s3:PutObject - - s3:GetObjectVersion - - s3:ListBucket - - s3:GetBucketLocation - Resource: - - !Sub "arn:aws:s3:::seahaven-payments-boa-raw-${AWS::AccountId}" - - !Sub "arn:aws:s3:::seahaven-payments-boa-raw-${AWS::AccountId}/*" - - - Sid: S3WorkloadReadOnly - Effect: Allow - Action: - - s3:GetObject - - s3:GetObjectVersion - - s3:ListBucket - - s3:GetBucketLocation - Resource: - - !Sub "arn:aws:s3:::seahaven-payments-csv-${AWS::AccountId}" - - !Sub "arn:aws:s3:::seahaven-payments-csv-${AWS::AccountId}/*" - - !Sub "arn:aws:s3:::seahaven-payroll-emails-${AWS::AccountId}" - - !Sub "arn:aws:s3:::seahaven-payroll-emails-${AWS::AccountId}/*" - - # ── Secrets Manager — one / prefix per workload ────────────── - # Replaces secret:*, which in seahaven-prod today reads - # proposal-system/db-credentials, proposal-system/bedrock-user, - # procurement-ingest/web-ui-auth-token and workorder-ingest/shoc-webhook-hmac - # — the last of which is an HMAC SIGNING key, so the wildcard was a - # webhook-forgery primitive against the SHOC integration. - # The trailing * after each / is MANDATORY, not decorative: Secrets - # Manager appends a random 6-character suffix to every ARN, so an - # exact-name ARN never matches. Verified 2026-07-30 with - # iam simulate-custom-policy against a suffixed ARN. - # afi-api-key is a bare, unprefixed, account-root secret that predates the - # / convention (afi-backup-monitor's OTHER secret, - # afi-backup-monitor/slack-webhook-url, is correctly prefixed). It is - # listed explicitly because omitting it would encode a KNOWN-WRONG scope - # that fails silently months from now at migration. MIGRATION OBLIGATION: - # afi-backup-monitor's migration PR renames it to - # afi-backup-monitor/api-key and deletes this line. - # secretsmanager:ListSecrets and BatchGetSecretValue are deliberately - # ABSENT and must stay absent — AWS authorises both against "*" regardless - # of any resource list, so adding either would silently reinstate - # account-wide read. (Precedent: BatchGetSecretValue synthed clean, passed - # review, then AccessDenied in production and crash-looped open-swe.) - - Sid: SecretsManager - Effect: Allow - Action: - - secretsmanager:GetSecretValue - - secretsmanager:DescribeSecret - Resource: - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:afterhours-shift-manager/*" - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:afi-backup-monitor/*" - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:afi-api-key-*" - # CORRECTION (2026-07-30 review): the permission-source block - # named this secret "afi-backup-monitor/slack-webhook-url", which - # does not exist. afi-backup-monitor/template.yaml takes both - # secret ARNs as deploy PARAMETERS, so no name is discoverable - # from the template; the real names are in that repo's README - # (afi-api-key and afi-slack-webhook, both bare/unprefixed). Same - # rename-at-migration obligation as afi-api-key applies. - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:afi-slack-webhook-*" - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:front-integrations/*" - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:meal-order-manager/*" - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:payments-dashboard/*" - - # ── SSM Parameter Store (meal-order-manager, afterhours) ──────────── - # Replaces parameter/*, which in seahaven-prod reads - # /procurement-api/custom-domain/certificate-arn and - # /seahaven/dynamodb/cmk-arn. These are exactly the two hierarchies the - # statement's own heading already claimed to serve. - # The bare-path entries alongside the /* entries are REQUIRED, not - # duplicates: ssm:GetParametersByPath authorises against the PATH ARN, - # not the leaf, so a /*-only grant can deny the recursive read. - # Note the ARN form drops the parameter's leading slash. - # /3cx-scheduler/* is deliberately EXCLUDED: ownership between - # afterhours-shift-manager and the retired standalone 3CX ring-group - # scheduler is unresolved, and inventing a scope is not permitted. - # Resolve at afterhours' migration and widen then if it is genuinely ours. - - Sid: SSMParameterRead - Effect: Allow - Action: - - ssm:GetParameter - - ssm:GetParameters - - ssm:GetParametersByPath - Resource: - - !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/afterhours-shift-manager" - - !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/afterhours-shift-manager/*" - - !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/meal-order-manager" - - !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/meal-order-manager/*" - - # ── SQS (payments-dashboard batch queues + async DLQs) ────────────── - # payments-dashboard is the only enumerated stack with a queue - # dependency, and all of its live queues are payments-prefixed. Replacing - # :* costs four characters and removes sqs:ReceiveMessage / - # sqs:DeleteMessage on proposal-system-jobs, workorder-shoc-emitter-failures - # and workorder-shoc-emitter-rejected, where the wildcard was a silent - # message-drain (data-loss) primitive against another tenant's pipeline. - # Verified denied 2026-07-30 with iam simulate-custom-policy. - # A prefix rather than four literals is deliberate, BUT the original - # justification for it was wrong and is corrected here: SAM does NOT - # auto-create or auto-name async DLQs — all four payments queues are - # hand-written AWS::SQS::Queue resources with explicit QueueNames, and - # OnFailure destinations take an explicit ARN. The real invariant is - # therefore a naming rule, not a framework behaviour: ANY queue a - # boundary-carrying function sends to — including async OnFailure - # destinations and DeadLetterQueue targets — must be named payments-*, - # or the boundary must be widened in the SAME PR. A denied destination - # write is SILENT: the async event is discarded with no caller to - # error, no Errors datapoint and no DLQ contents. - # sqs:ListQueues is absent and must stay absent (authorised against "*"). - - Sid: SQS - Effect: Allow - Action: - - sqs:SendMessage - - sqs:ReceiveMessage - - sqs:DeleteMessage - - sqs:GetQueueAttributes - - sqs:GetQueueUrl - - sqs:ChangeMessageVisibility - Resource: - - !Sub "arn:aws:sqs:us-east-1:${AWS::AccountId}:payments-*" - - # ── Lambda invocation — HIGHEST-LEVERAGE FIX IN THIS CHANGE ───────── - # function:* was a BOUNDARY-ESCAPE primitive, not merely lateral - # movement: an invoked function executes under ITS OWN execution role, - # and every non-SAM function in seahaven-prod (proposal-system-*, - # procurement-api, workorder-*, po-*) is CDK-deployed and carries NO - # permissions boundary at all. A bounded SAM Lambda could therefore reach, - # by proxy, capability this ceiling exists to deny. Verified 2026-07-30 - # with iam simulate-custom-policy: proposal-system-api is now denied, - # payments-expenseProcessor still allowed. - # Scoped to the three workloads with enumerated inter-function calls: - # payments-dashboard (ExpenseReceiver -> ExpenseProcessor), - # meal-order-manager (submit-order -> slack-notifier, close-form -> - # aggregate-orders, plus AdminAuthorizerInvokeRole), and - # afterhours-shift-manager, whose boundary-carrying ReleaseNotifyInvokeRole - # and HolidaySchedulerExecutionRole exist solely to invoke. - # front-integrations and afi-backup-monitor have no enumerated invoke need - # and are deliberately absent. Prefixes (not literal function ARNs) are - # used because they also match the : / : qualified form. - - Sid: LambdaInvoke - Effect: Allow - Action: - - lambda:InvokeFunction - Resource: - - !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:function:afterhours-*" - - !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:function:meal-order-manager-*" - - !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:function:payments-*" - - # ── SES — pinned to account inventory, NOT per-workload scoping ───── - # Labelled honestly: an SES identity is a shared DOMAIN, so no - # per-workload prefix exists to scope to. These are seahaven-prod's two - # verified identities (re-verified live 2026-07-30); seahaven-dev has - # none, so both entries are simply inert there. The win is bounded but - # real: identity/* would let the SAM fleet send as ANY identity ever added - # to the account, including a customer or partner domain, from a - # legitimately-authenticated sender. This cannot break anything that would - # otherwise work — sending from an unverified identity fails with - # MessageRejected regardless of IAM — so the failure mode is LOUD. - # configuration-set/* is DROPPED: zero configuration sets exist in prod or - # dev and no enumerated stack uses one. - # MIGRATION-BLOCKING QUESTION: mgmt additionally has seahavenind.com, - # apfacilities.org, adam@seahaven.com and payroll@seahaven.com verified; - # prod does NOT. Each migrating stack must confirm its actual Source - # address, verify that domain in the target account, and add the identity - # ARN here in the same PR. - # A ses:FromAddress condition would be stronger but is not available: the - # permission-source block names payroll@seahavenind.com, which is not a - # verified identity in prod — writing that condition would invent a scope. - # CORRECTION (2026-07-30 review): an earlier revision DROPPED - # configuration-set/* on the reasoning that zero configuration sets - # exist in prod or dev today. That test was the wrong one. SES - # authorizes SendEmail against the CONFIGURATION-SET resource in - # addition to the identity whenever the identity has a default - # configuration set, and afterhours-shift-manager/template.yaml:178-181 - # grants exactly that ARN with an in-repo comment recording that - # omitting it DENIES the send. The correct question is not "does the - # resource exist in the target account yet" but "does an enumerated - # stack's own IAM policy name it". seahavenind.com is listed because - # meal-order-manager's SenderEmail parameter defaults to - # adam@seahavenind.com (meal-order-manager/template.yaml:20-22) and its - # email_report handler sends with that Source; the earlier "sending - # from an unverified identity fails loudly anyway" argument only holds - # until the migration verifies the domain, which the migration - # procedure itself requires. - - Sid: SES - Effect: Allow - Action: - - ses:SendEmail - - ses:SendRawEmail - Resource: - - !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:identity/int.seahaven.com" - - !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:identity/seahaven.com" - - !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:identity/seahavenind.com" - - !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:configuration-set/seahaven-email-events" - - # ── EventBridge Scheduler (afterhours-shift-manager holiday routing) ─ - # Omitted from the permission-source block entirely, which the block's - # "verified live" header did not catch. afterhours-shift-manager - # template.yaml:110-120 grants these three scheduler actions plus - # iam:PassRole on its holiday-scheduler execution role. Without them - # the failure is SILENT: src/slack-bot/app.py wraps create_schedule in - # a bare `except Exception`, so the Slack command returns success, the - # holiday record is written with no schedule, and no alarm fires. - - Sid: EventBridgeScheduler - Effect: Allow - Action: - - scheduler:CreateSchedule - - scheduler:DeleteSchedule - - scheduler:GetSchedule - Resource: - - !Sub "arn:aws:scheduler:us-east-1:${AWS::AccountId}:schedule/default/holiday-*" - - # PassRole is confined to the scheduler service principal, so this - # cannot be used to hand a role to Lambda or any other service. - - Sid: SchedulerPassRole - Effect: Allow - Action: - - iam:PassRole - Resource: - - !Sub "arn:aws:iam::${AWS::AccountId}:role/afterhours-shift-manager-*" - Condition: - StringEquals: - "iam:PassedToService": "scheduler.amazonaws.com" - - # ── KMS — key/* RETAINED, scoped by condition instead ─────────────── - # KMS is the one high-value data plane that CANNOT be scoped by resource - # name: key ARNs carry UUID key ids, not workload names, and alias ARNs - # are not valid in a Resource for these actions (alias scoping needs a - # kms:RequestAlias condition). The two key ids named in this statement's - # previous comment (key/0b660af3, key/b748750c) are MGMT keys that do not - # exist in seahaven-prod or seahaven-dev; hardcoding them — or prod's - # three live CMKs — would encode one account's inventory into a template - # shared by two, and would break on any key replacement. - # So Resource stays key/* and the scope is derived from kms:ViaService: - # the fleet may use a CMK ONLY as part of a request one of these services - # makes on its behalf. Decrypting a DynamoDB item or an S3 object still - # works; a direct kms:Decrypt on arbitrary ciphertext lifted from anywhere - # in the account is denied — the actual escalation path key/* opened. - # Because those services are themselves prefix-scoped above, the effective - # KMS scope INHERITS the per-workload scoping for free, with no key ids - # and no per-account parameterisation. - # logs.us-east-1.amazonaws.com is included as belt-and-braces: CloudWatch - # Logs is believed to decrypt log-group CMKs under its own service grant - # rather than the execution role's credentials, which would make this entry - # a no-op — but a KMS denial on the logging path would be SILENT, and the - # entry cannot grant anything meaningful on its own, so the insurance is - # bought deliberately. - # If a specific key must ever be named, use a kms:RequestAlias condition — - # never a literal key id. - - Sid: KMS - Effect: Allow - Action: - - kms:Decrypt - - kms:GenerateDataKey - - kms:DescribeKey - Resource: - - !Sub "arn:aws:kms:us-east-1:${AWS::AccountId}:key/*" - Condition: - StringEquals: - kms:ViaService: - - dynamodb.us-east-1.amazonaws.com - - logs.us-east-1.amazonaws.com - - s3.us-east-1.amazonaws.com - - secretsmanager.us-east-1.amazonaws.com - # REQUIRED for parity with SSMParameterRead above: a - # SecureString parameter decrypts via the SSM service - # principal, so omitting this denies reads that this same - # policy grants — a self-inconsistency caught by the - # 2026-07-30 review. Any ssm:GetParameter* grant in this - # boundary must keep this entry. - - ssm.us-east-1.amazonaws.com - - sqs.us-east-1.amazonaws.com - + # ── NO PER-WORKLOAD DATA-PLANE STATEMENTS — BY DESIGN ─────────────── + # The statements above are the FLEET-WIDE FLOOR: what every Lambda + # execution role needs regardless of which workload it belongs to. + # There are deliberately NO DynamoDB, S3, Secrets Manager, SSM, SQS, + # SES, KMS, lambda:InvokeFunction or scheduler statements here. + # + # WHY (decided 2026-07-30, Adam): + # The security win of INFRA-186 comes from DELETION, not enumeration. + # Removing the account-wide secret:*, table/*, function:* and sqs:* + # wildcards is what closes the amplifier — the ability of a principal + # who can write an inline policy onto a boundary-carrying role to read + # every secret in the account. Per-workload prefixes add no security; + # they exist only to keep a workload FUNCTIONAL once it arrives. + # + # An earlier revision of this branch PRE-LOADED prefixes for all five + # mgmt SAM stacks before any of them had migrated. That required + # predicting five stacks' permission needs from the permission-source + # comment block above, and the /sh-security-review pass found SIX + # errors in the result — three of which would have failed SILENTLY at + # first migration (afterhours' SES config-set, its holiday scheduler + # behind a bare except, and afi's webhook secret under an invented + # name). The block is a secondary record and is not a substitute for + # reading the owning repo's template. + # + # WIDENING IS THE SAFE DIRECTION. Adding a resource to a boundary can + # never break a running Lambda; only tightening can. So there is no + # cost to deferring per-workload scope to the migration PR that has + # the real template open in front of it — and a large cost to + # guessing it years ahead of the migration. + # + # CONSEQUENCE FOR EVERY MIGRATION PR (mandatory, see WIDENING PATH in + # the header): a stack landing in prod or dev MUST add its own + # data-plane statements here, derived from ITS OWN template, in the + # same PR that deploys it. Without them its Lambdas get AccessDenied + # at first invoke. The permission-source block above is the starting + # point, NOT the authority — verify every entry against the stack. + # + # Prod/dev boundary usage is 0 (verified 2026-07-30), so this floor + # currently constrains nothing that exists. Fleet-wide statements that + # genuinely cannot be scoped (Logs, X-Ray, ENI) stay above with their + # justifications; they are also the SILENT-failure classes, which is + # why they belong in the floor rather than in per-workload widenings. + # + # KMS is absent deliberately: both accounts have ZERO CMK-encrypted + # log groups today (verified 2026-07-30). A workload bringing a + # CMK-encrypted resource adds a KMS statement with the matching + # kms:ViaService principal in its own migration PR — a missing + # ViaService entry denies, and for env-var encryption it fails at + # cold-start INIT. + # + # The end-state fix for the shared-ceiling residual (one boundary = + # every SAM workload reaches every other's data plane once they land) + # is per-workload boundaries — tracked separately, see the header. # --------------------------------------------------------------------------- # Shared CloudFormation execution role (SAM stacks) — INFRA-97 scoped #