diff --git a/README.md b/README.md index d8f168c..9aaf5f9 100644 --- a/README.md +++ b/README.md @@ -133,11 +133,25 @@ account-level deploy plumbing: only for an account that does not already have one — one provider per URL per account). -The template is a verbatim extraction of the substrate section of -`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml` (see the provenance -header in the template — mgmt's copy remains source of truth for 328440206208 -until its stacks migrate out; substrate changes while both are live must edit -both files). Per-repo `githubdeploy-*` deploy roles are deliberately NOT part +The template began as a verbatim extraction of the substrate section of +`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml`, which remains the source +of truth for mgmt (328440206208) until its stacks migrate out. + +**The two copies are no longer at parity, and the old "edit both files" rule no +longer applies uniformly.** Under INFRA-186, `seahaven-lambda-execution-boundary` +in *this* copy was reduced to a fleet-wide floor for prod and dev (where +boundary usage was 0, so no live Lambda could break): CloudWatch Logs write on +`/aws/lambda*`, log-group describe, X-Ray, and ENI lifecycle — nothing else. +Each migrating stack adds its own data-plane statements, derived from its own +template, in its own PR (per-workload boundaries are the INFRA-187 end state). +mgmt's copy keeps the account-wide wildcards pending its own separately +validated rollout across 26 live boundary-carrying roles. So: **the boundary +resource is deliberately divergent**; every *other* substrate resource +(`github-cfn-execution-role`, `seahaven-cfn-exec-iam-management`) is still +expected to change in both files together. The template's provenance header +records which is which — read it before assuming either parity or divergence. + +Per-repo `githubdeploy-*` deploy roles are deliberately NOT part of the substrate — they are provisioned per repo at migration/onboarding time so an account never carries trust relationships for repos that do not deploy to it. @@ -258,14 +272,24 @@ split this substrate exists to enforce. `organization:seahaven:project:seahaven-:workspace::run_phase:plan` (or `:apply`). Exact `StringEquals` only — never `StringLike`, never a wildcarded `run_phase` (a speculative PR plan must never hold write - credentials). IAM roles = mandatory GPT-4.1 cross-review + + credentials). **If the stack creates Lambda execution roles, this same PR + must also widen `seahaven-lambda-execution-boundary`** per the WIDENING + PATH in `lib/deploy-substrate/deploy-substrate.template.yaml`: the + guardrail forces every Terraform-created role to carry that boundary, and + it is a fleet-wide floor with zero data-plane permissions until widened — + an unwidened migration deploys green, then every data-plane call is denied + at first invoke and async/DLQ writes are discarded silently. IAM roles and + boundary widenings = mandatory cross-family review + `/sh-security-review` on the diff. 3. After deploy, verify: both roles exist; `hcptf-` lists `seahaven-hcptf-iam-management` in `list-attached-role-policies`; trust subs match the live org/project/workspace names byte-for-byte; simulate the apply role against a `hcptf-*` ARN (expect `explicitDeny` from `DenySelfMutation`) and against a normal stack role name (expect - `allowed`). + `allowed`); and if step 2 widened the boundary, confirm the deployed + default version carries the stack's data-plane statements + (`aws iam get-policy-version`) — role verification alone never checks + boundary content. 4. Set **workspace-level** variables `TFC_AWS_PLAN_ROLE_ARN` + `TFC_AWS_APPLY_ROLE_ARN` (category env) to the verified role ARNs, plus `TFC_AWS_PROVIDER_AUTH=true`. Never project-scoped variable sets — the diff --git a/lib/deploy-substrate/deploy-substrate.template.yaml b/lib/deploy-substrate/deploy-substrate.template.yaml index ce16245..301af15 100644 --- a/lib/deploy-substrate/deploy-substrate.template.yaml +++ b/lib/deploy-substrate/deploy-substrate.template.yaml @@ -39,16 +39,115 @@ Description: >- # added. The mgmt copy was remediated 2026-07-27 (.github PRs #95 Phase A # + #98 Phase B); DenySelfMutation and the widened policy/seahaven-* # DenyBoundaryPolicyEdit scope were then ported back here, so the two -# copies' statement sets are reconciled as of that date — every IAM -# statement in LambdaExecutionBoundary, SamCfnIamManagementPolicy and -# SamCfnExecutionRole is byte-identical across the two files; the only -# remaining delta is the DependsOn line above, which is ordering, not -# permission. If a substrate statement changes again, change BOTH files -# in the same piece of work. +# copies' GUARDRAIL statement sets were reconciled as of that date. +# SamCfnIamManagementPolicy and SamCfnExecutionRole remain at parity on +# their IAM STATEMENT SETS across the two files and MUST still be changed +# together. Parity covers statements, not surrounding comments — a comment +# may diverge where it describes boundary content, which now differs +# between the files. The only functional delta +# between them is the DependsOn line above, which is ordering, not +# permission. LambdaExecutionBoundary is NO LONGER byte-identical — see +# DELIBERATE DIVERGENCE below. +# +# DELIBERATE DIVERGENCE — LambdaExecutionBoundary (INFRA-186, 2026-07-30) +# The parity rule above is SCOPED, not global. LambdaExecutionBoundary in THIS +# file is DELIBERATELY STRICTER than the mgmt copy in +# Sea-Haven-Industries/.github/oidc-deploy-roles.yaml. Do not "reconcile" the two +# by copying mgmt's statements back over these, or vice versa; the divergence is +# load-bearing. A future mechanical drift check WILL read it as drift — it is not. +# +# 1. WHAT DIVERGED. Every per-workload data-plane statement was REMOVED from +# LambdaExecutionBoundary in this file, leaving only the fleet-wide floor: +# CloudWatchLogsWrite (scoped to /aws/lambda*), CloudWatchLogsDescribe, +# XRay and Ec2Eni. The account-wide wildcards mgmt still carries — table/*, +# table/*/index/*, secret:*, parameter/*, sqs :*, function:*, ses +# identity/* + configuration-set/*, kms key/*, and an s3:::*- +# pattern that was a bare name-suffix filter rather than an ownership +# check — are simply GONE here rather than re-scoped. +# The security win is the deletion: it is what closes the amplifier whereby +# a principal able to write an inline policy onto a boundary-carrying role +# could read every secret in the account. Per-workload prefixes add no +# security — they only keep a workload functional — so they are added by +# each migration PR, from that stack's own template, when the stack +# actually lands. See the note on the boundary resource for the full +# rationale and the six errors that the pre-loaded approach produced. +# +# 2. WHY MGMT'S RATIONALE IS LEGITIMATE THERE. The superset framing this file +# used to carry ("being slightly broad is the correct trade-off; a boundary +# that is too tight will break Lambda functions at runtime AFTER deploy") is +# a real constraint in the management account: 328440206208 has 26 LIVE +# roles carrying seahaven-lambda-execution-boundary, across all five SAM +# stacks. Tightening there is a production change to running workloads with +# a silent, deploy-time-invisible failure mode. +# +# 3. WHY IT DOES NOT TRANSFER HERE. This file deploys ONLY to seahaven-prod +# (011934824531) and seahaven-dev (710827005802), where +# PermissionsBoundaryUsageCount is 0 and 0 respectively (aws iam get-policy, +# verified 2026-07-30; corroborated by list-roles returning no role carrying +# any permissions boundary in either account). No live Lambda can break, so +# the risk that justifies mgmt's breadth is absent — while the exposure is +# strictly WORSE here than in mgmt, because prod is multi-tenant: the old +# wildcards reached proposal-system's, procurement-ingest's and +# workorder-ingest's CDK-owned tables, buckets, secrets and queues, the org's +# own Config and VPC-flow-log buckets, and — via function:* — CDK Lambdas +# that carry no boundary at all. +# +# 4. RECONCILIATION OBLIGATION, RESTATED. For LambdaExecutionBoundary the two +# copies are now INTENTIONALLY DIFFERENT and must NOT be synchronised: +# - A change to the per-workload Resource patterns in THIS file does NOT +# propagate to mgmt. +# - A change to mgmt's boundary does NOT propagate here. +# - Any change to the ACTION lists, or any new statement, is a substrate +# semantic change and DOES still require the same review in both copies. +# For SamCfnIamManagementPolicy and SamCfnExecutionRole the original rule is +# unchanged: change BOTH files in the same piece of work. +# +# KNOWN OPEN ITEM (deferred, not closed by INFRA-186): mgmt 328440206208 still +# carries the account-wide patterns. Tightening it needs its own validated +# rollout — enumerate what the 26 live roles actually call, stage it, and be +# ready to roll back — and is explicitly OUT OF SCOPE of INFRA-186. Until that +# lands, the two copies stay divergent and that is the intended state. +# +# COUPLING: the boundary's ManagedPolicyName and ARN are UNCHANGED and must stay +# unchanged. Four Conditions in SamCfnIamManagementPolicy below, and four more in +# HcptfIamManagementPolicy in lib/terraform-substrate/terraform-substrate.template.yaml, +# pin arn:aws:iam:::policy/seahaven-lambda-execution-boundary by literal +# string inside StringEquals iam:PermissionsBoundary. A rename fails SILENTLY — an +# IAM condition naming a non-existent policy simply never matches, so the +# escalation control would evaporate rather than error — and would additionally +# force a CloudFormation REPLACEMENT that any role carrying the boundary would +# block. INFRA-186 is a CONTENT-ONLY change for exactly this reason; +# terraform-substrate.template.yaml already records that expectation and is +# correctly left untouched. +# +# SIZE BUDGET: an attached managed policy document is capped at 6,144 characters +# (whitespace excluded). LambdaExecutionBoundary measures 691 characters across +# 4 statements as of 2026-07-31 — the fleet-wide floor only. Measure before +# widening — len(json.dumps(doc,separators=(',',':'))) on the synthesized +# PolicyDocument with ${AWS::AccountId} resolved, and UPDATE THESE TWO NUMBERS in +# the same edit (they went stale twice inside this branch alone). +# +# Headroom is 5,453 characters, roughly TWELVE workloads at ~450 each. That is a +# deliberate outcome, not luck: an earlier revision of this branch pre-loaded +# per-workload prefixes for all five mgmt SAM stacks and reached 5,457 characters +# with 687 left — about one workload of room — before any stack had actually +# migrated. Deferring per-workload scope to each migration PR (see the note on +# the boundary itself) removed that pressure entirely. If the budget tightens +# again as workloads land, the end-state fix is per-workload boundaries +# (seahaven-lambda-execution-boundary-), which also resolves the +# shared-ceiling residual — tracked as INFRA-187, do not improvise it. +# CRITICAL: unlike the 2026-07-27 inline-limit incident, +# there is NO restructure available when this cap is reached — a role has exactly +# ONE permissions boundary, so statements cannot be spilled into a second attached +# managed policy. At the cap the only levers are prefix consolidation and dropping +# unused actions. # # This template is deployed via lib/deploy-substrate-stack.ts -# (cloudformation-include) as stack seahaven-deploy-substrate, once per -# member account that hosts SAM workloads. +# (cloudformation-include) as stack seahaven-deploy-substrate, once per member +# account that hosts SAM workloads (currently seahaven-prod 011934824531 and +# seahaven-dev 710827005802 via bin/app.ts instances deploy-substrate-prod / +# deploy-substrate-dev; NEVER mgmt — 328440206208 is served by the .github copy +# named above until its stacks migrate out). Parameters: CreateOIDCProvider: @@ -83,7 +182,7 @@ Resources: UpdateReplacePolicy: Retain # --------------------------------------------------------------------------- - # Lambda execution permissions boundary (INFRA-103) + # Lambda execution permissions boundary (INFRA-103, re-scoped by INFRA-186) # # This managed policy is the CEILING for every Lambda execution role that the # five SAM stacks auto-generate via AWS::Serverless::Function. Applying it as @@ -91,49 +190,173 @@ Resources: # intersection of the role's own policies and this boundary, so a misconfigured # SAM role can never exceed what is listed here. # - # The boundary is intentionally a SUPERSET of the union of all runtime - # permissions currently granted across the five stacks. Being slightly broad - # is the correct trade-off at this stage — a boundary that is too tight will - # break Lambda functions at runtime after deploy, which is worse than a slightly - # loose boundary that is tightened in a follow-up. + # SCOPING RULE (INFRA-186, 2026-07-30). This boundary is the FLEET-WIDE FLOOR + # ONLY: what every Lambda execution role needs regardless of workload. It + # carries NO per-workload data-plane statements — each migrating stack adds its + # own, from its own template, in its own PR (see the note on the boundary + # resource and the WIDENING PATH below). It was previously a deliberate SUPERSET + # with account-wide wildcards (table/*, secret:*, sqs :*, function:*, + # parameter/*, and an s3:::*- pattern that was a name-suffix filter, + # not an ownership check). That trade-off was made when the only account + # carrying this policy hosted nothing but the five SAM stacks. It does not + # survive multi-tenancy: seahaven-prod now hosts CDK-owned tenants + # (proposal-system, procurement-ingest, workorder-ingest) whose tables, buckets, + # secrets and queues those wildcards reached, and whose Lambdas carry NO + # permissions boundary at all — making function:* a boundary-escape primitive. # - # Permission sources per stack: + # Tightening here carries ZERO runtime risk and was sequenced deliberately: + # PermissionsBoundaryUsageCount is 0 in BOTH accounts this template deploys to + # (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified + # 2026-07-30), so no live Lambda can break. A boundary that is slightly TOO + # TIGHT is recoverable here — the migrating stack widens it in its own PR before + # its first deploy — whereas leaving it loose perpetuates the exposure. Prefer + # tighter; the widening path is below. # - # afterhours-shift-manager + # A resource pattern that genuinely CANNOT be scoped keeps its wildcard WITH a + # written justification on the statement: CloudWatchLogsDescribe, XRay and + # Ec2Eni name runtime-created resources or use actions AWS authorises against + # "*" regardless of the ARN supplied. Do not "tighten" those. + # + # Every "verified " annotation in this file is a POINT-IN-TIME + # observation, not live state. Re-validate (usage counts, log-group CMK state, + # per-stack permission sources) before citing one as justification for a + # future change. + # + # WIDENING PATH — read this before migrating a stack into prod or dev. + # The boundary is never widened by the person who hits the AccessDenied. It is + # widened by the migrating stack's owner, in THIS repo, BEFORE the workload's + # first deploy into the target account: + # 0. PRECONDITION — check the managed-policy VERSION budget BEFORE merging: + # max 5 versions, both + # accounts are on v1 today. Every widening (PolicyDocument edit) burns + # one. Description, ManagedPolicyName and Path are REPLACEMENT + # properties per the CFN resource reference — CloudFormation cannot + # replace a custom-named policy, so a Description-only edit FAILS the + # stack update, and the error's suggested remedy (rename) is exactly + # the forbidden rename in the COUPLING note above. Never edit those + # three properties. Delete the oldest non-default version if at 5: + # aws iam list-policy-versions --policy-arn \ + # arn:aws:iam:::policy/seahaven-lambda-execution-boundary + # aws iam delete-policy-version --version-id v ... + # 1. Derive the workload's needs from ITS OWN TEMPLATE — open the stack's + # template.yaml and read the actual IAM policy statements. The + # permission-source block below is a STARTING POINT, NOT THE AUTHORITY: + # the /sh-security-review pass on 2026-07-30 found SIX places where it was + # incomplete or simply invented a resource name, three of which would have + # failed silently. Update that block in the same edit with what you find. + # 2. Add the workload's statements. Group by service so a second workload can + # extend a Resource list rather than duplicate an action list (~250 + # characters for zero new actions; an extra ARN costs ~60). Check for the + # SILENT classes specifically: a denied SQS destination/DLQ write discards + # the async event with no error and no alarm; a denied scheduler call may + # sit behind a bare except; a denied KMS decrypt for env-var encryption + # fails at cold-start INIT; any CMK-encrypted resource needs the + # matching kms:ViaService principal, not just the kms action; and a + # function using LoggingConfig with a custom log-group name outside + # /aws/lambda* silently loses ALL logs — add a scoped logs statement + # for the custom group or keep the default group name. + # 3. Measure. See SIZE BUDGET in the header. There is NO escape hatch. + # 4. Both review gates run and neither discharges the other: the GPT-4.1 + # cross-family review against the real diff, and /sh-security-review + # (IaC/IAM is on the mandatory surface). CLI down = review outstanding. + # 5. Merge and let CI deploy deploy-substrate-prod / deploy-substrate-dev to + # UPDATE_COMPLETE, THEN deploy the workload stack. + # ORDERING IS NOT ENFORCED BY CLOUDFORMATION AND THIS IS THE MOST IMPORTANT + # SENTENCE HERE: the workload's deploy SUCCEEDS even against a stale boundary, + # because seahaven-cfn-exec-iam-management's gate checks that the boundary ARN + # is attached, never its contents. The failure surfaces later, at first invoke, + # as AccessDenied. A stale boundary is a silent deploy-time pass and a loud + # production failure. + # + # CONSIDERED AND REJECTED: a Deny statement reserving the seahaven-* namespace. + # With the Allow set reduced to the fleet-wide floor it is fully redundant (verified + # 2026-07-30: seahaven-prod-config-* and seahaven-prod-vpc-flow-logs-* are + # already denied by the Allow set alone), and a Deny inside a BOUNDARY is the + # hardest failure mode in the estate to debug — it beats every Allow in every + # policy with no synth-time signal. Revisit only if a widening ever has to + # re-broaden a per-service Resource list back toward a wildcard. + # + # NOTE ON WHAT THESE PREFIXES ARE. All five stacks below currently live in the + # MANAGEMENT account and none of their resources exists in seahaven-prod or + # seahaven-dev yet. These are MIGRATION-CANDIDATE prefixes for the accounts this + # template deploys to, not an inventory of what is deployed there. They are a + # SECONDARY RECORD and a starting point for widening PRs — the authority is + # each stack's own template (WIDENING PATH step 1). No Resource pattern in + # the floor above derives from this block. + # + # Permission sources per stack (verified live 2026-07-30 IN MGMT — these + # stacks and resources do not exist in prod/dev yet, so nothing below is a + # prod/dev observation; starting point only — verify every entry against + # the owning repo before use. PARAMETERIZED resources (ARNs passed as + # deploy parameters) must be re-derived from the live stack configuration + # at migration time, as exact ARNs — never inferred from these names into + # broad patterns like secret:afi-*): + # + # afterhours-shift-manager (functions: afterhours-*, 6 live) # - DynamoDB CRUD (afterhours-shifts table) # - secretsmanager:GetSecretValue (afterhours-shift-manager/*) - # - ses:SendEmail (SES identity) + # - ses:SendEmail on the identity AND on + # configuration-set/seahaven-email-events (template.yaml:178-181 — the + # send is DENIED without the config-set ARN when the identity has a + # default configuration set) + # - scheduler:CreateSchedule/DeleteSchedule/GetSchedule on + # schedule/default/holiday-* + iam:PassRole to scheduler.amazonaws.com + # (template.yaml:110-120). CORRECTED 2026-07-30: this block previously + # omitted both, and the omission is SILENT at runtime (bare except). + # - NO ssm. CORRECTED 2026-07-30: this block previously credited + # ssm:GetParameter to this stack; `grep -c 'ssm:' template.yaml` = 0. + # Its slack tokens come from Secrets Manager and the channel id from a + # CloudFormation parameter. + # - lambda:InvokeFunction (ReleaseNotifyInvokeRole, + # HolidaySchedulerExecutionRole — these two carry the boundary and need + # ONLY this action) # - CloudWatch Logs (all functions) + # - UNRESOLVED: /3cx-scheduler/* ownership (this stack vs. the retired + # standalone 3CX scheduler). Deliberately NOT granted. Resolve at + # migration in that stack's own repo — do NOT add any 3cx-scheduler + # resource here until ownership is resolved. # - # payments-dashboard - # - DynamoDB CRUD / Read (PaymentsDashboard table) - # - S3 GetObject (payroll-emails, payments-csv buckets) + # payments-dashboard (functions: payments-*) + # - DynamoDB CRUD / Read (PaymentsDashboard table — legacy PascalCase) + # - S3 GetObject on seahaven-payments-csv-* and seahaven-payroll-emails-*; + # GetObject + PutObject on seahaven-payments-boa-raw-* + # (template.yaml:272 and 1097-1099, fetchBoaTransactions raw archive). + # CORRECTED 2026-07-30: this block previously said "GetObject ONLY — + # no write intent enumerated", which was false and would have denied + # the raw-archive write at migration. # - secretsmanager:GetSecretValue (payments-dashboard/*) - # - sqs:SendMessage + sqs:ReceiveMessage + sqs:DeleteMessage etc. - # (PayrollBatchQueue + DLQs) - # - lambda:InvokeFunction (ExpenseReceiver → ExpenseProcessor) - # - ec2:CreateNetworkInterface / DescribeNetworkInterfaces / - # DeleteNetworkInterface (VPC-attached functions) + # - sqs Send/Receive/Delete etc. (payments-payroll-batch + DLQs) + # - lambda:InvokeFunction (ExpenseReceiver -> ExpenseProcessor) + # - ec2 ENI lifecycle (VPC-attached functions) + # - KMS via dynamodb (table CMK) — no SSM, no SES # - CloudWatch Logs # - # meal-order-manager + # meal-order-manager (functions: meal-order-manager-*, 7 live) # - DynamoDB CRUD / Read (meal-order-manager-orders table) - # - S3 CRUD (ReportsBucket) + s3:GetObject (ReportsBucket presigned URLs) + # - S3 CRUD (meal-order-manager-reports-*, meal-order-manager-form-*) # - secretsmanager:GetSecretValue (meal-order-manager/*) # - ssm:GetParameter (/meal-order-manager/*) - # - lambda:InvokeFunction (submit-order → slack-notifier, - # close-form → aggregate-orders) + # - lambda:InvokeFunction (submit-order -> slack-notifier, + # close-form -> aggregate-orders, plus AdminAuthorizerInvokeRole) # - ses:SendRawEmail # - CloudWatch Logs # - # front-integrations + # front-integrations (functions: front-*) # - DynamoDB CRUD (front-sla-alerts table) - # - secretsmanager:GetSecretValue (by ARN, various) + # - secretsmanager:GetSecretValue (front-integrations/*) # - CloudWatch Logs + # - no IAM permissions for S3 / SQS / SSM / SES / KMS / VPC in this + # stack's template as of 2026-07-30 # - # afi-backup-monitor - # - secretsmanager:GetSecretValue (by ARN) + # afi-backup-monitor (functions: afi-*) + # - secretsmanager:GetSecretValue on TWO bare, unprefixed secrets: + # afi-api-key and afi-slack-webhook. CORRECTED 2026-07-30: this block + # previously named "afi-backup-monitor/slack-webhook-url", which does + # not exist. Both ARNs are deploy PARAMETERS in that stack + # (AfiApiKeySecretArn / SlackWebhookSecretArn), so no name is + # discoverable from its template — the names are in its README:48-49. # - CloudWatch Logs + # - nothing else # # --------------------------------------------------------------------------- LambdaExecutionBoundary: @@ -149,18 +372,76 @@ Resources: Version: "2012-10-17" Statement: - # ── CloudWatch Logs (every Lambda) ────────────────────────────────── - - Sid: CloudWatchLogs + # ── CloudWatch Logs — write (every Lambda) ────────────────────────── + # Scoped to the Lambda log-group namespace. Every SAM function's group + # is /aws/lambda/, and the trailing * also covers the + # :log-stream: suffix PutLogEvents authorises against, so one ARN + # serves CreateLogGroup, CreateLogStream, PutLogEvents and + # DescribeLogStreams. The * is deliberately NOT after a trailing slash: + # /aws/lambda* also matches the /aws/lambda-insights groups the Lambda + # Insights extension writes to, which /aws/lambda/* would have denied. + # VERIFICATION PROVENANCE, stated precisely (2026-07-30). What + # iam simulate-custom-policy DOES confirm: this pattern allows + # logs:CreateLogGroup / logs:PutLogEvents on the bare group ARN + # log-group:/aws/lambda/ (and /aws/lambda//), and DENIES + # log-group:seahaven-prod-vpc-flow-logs — the latter re-checked against + # an Allow */* positive control, which allows it, so the deny is real + # policy behaviour and not a simulator artifact. + # What the simulator CANNOT evaluate, so do NOT claim it was verified: + # log-stream-qualified ARNs (log-group::log-stream:) and the bare + # /aws/lambda-insights group both return implicitDeny EVEN UNDER an + # Allow */* policy. That is a simulator resource-parsing limitation, not + # a denial. Coverage of those two rests on documented IAM wildcard + # semantics — "*" matches any sequence of characters including ":" and + # "/" — which is why the * is deliberately NOT placed after a trailing + # slash. If this ever needs true end-to-end proof, it must come from a + # real invoke in dev, not from the simulator. + # + # NOT scoped per workload, deliberately. A per-stack prefix + # (/aws/lambda/payments-* etc.) was considered and rejected: a Lambda + # denied PutLogEvents does not fail — it keeps running and silently + # produces no logs. Log denial is the one failure class in this policy + # that is NOT loud, so it must not depend on function-name discipline. + # ACCEPTED RESIDUAL RISK: a SAM Lambda can write into another tenant's + # /aws/lambda/* group (log poisoning). No read action is granted here, so + # this is not an exfiltration path. Tighten only once every + # boundary-carrying function is confirmed to set an explicit FunctionName. + # REGION IS PINNED TO us-east-1 DELIBERATELY: every Sea Haven workload + # deploys to us-east-1, and this template itself only ever deploys there. + # ${AWS::Region} would resolve to the identical string, so it would + # document nothing. A future stack in another region carries this + # boundary but CANNOT write its logs (the silent class above) — so a + # cross-region migration MUST add region-scoped statements in its + # widening PR, same as any other data-plane need. + - Sid: CloudWatchLogsWrite Effect: Allow Action: - logs:CreateLogGroup - logs:CreateLogStream - logs:PutLogEvents - - logs:DescribeLogGroups - logs:DescribeLogStreams + Resource: + - !Sub "arn:aws:logs:us-east-1:${AWS::AccountId}:log-group:/aws/lambda*" + + # ── CloudWatch Logs — describe (UNSCOPABLE, kept "*" deliberately) ── + # logs:DescribeLogGroups is a COLLECTION action: AWS authorises it + # against "*" regardless of any resource ARN supplied. Scoping it would + # produce a policy that reads tighter and denies at runtime, so it is + # split into its own statement and keeps the wildcard. Read-only + # metadata; it cannot mutate anything or return log content. + - Sid: CloudWatchLogsDescribe + Effect: Allow + Action: + - logs:DescribeLogGroups Resource: "*" - # ── X-Ray tracing (standard Lambda execution) ──────────────────── + # ── X-Ray tracing (UNSCOPABLE, kept "*" deliberately) ─────────────── + # xray:PutTraceSegments / PutTelemetryRecords support no resource-level + # permissions — X-Ray exposes no ARN for them, which is why the AWS + # managed AWSXRayDaemonWriteAccess also uses "*". Any ARN written here + # would be inert and would falsely imply a control exists. Write-only + # into this account's own trace store; no cross-tenant read is + # expressible with this action set. - Sid: XRay Effect: Allow Action: @@ -168,10 +449,33 @@ Resources: - xray:PutTelemetryRecords Resource: "*" - # ── VPC / ENI management (payments-dashboard VPC functions) ──────── - # Matches AWSLambdaVPCAccessExecutionRole exactly. - # AssignPrivateIpAddresses / UnassignPrivateIpAddresses are for EFA - # and secondary IPs — not part of the Lambda ENI lifecycle — omitted. + # ── VPC / ENI management (UNSCOPABLE, kept "*" deliberately) ──────── + # Derived from AWSLambdaVPCAccessExecutionRole, NOT an exact match: + # DescribeSecurityGroups and DescribeVpcs exceed that managed policy + # (kept for CFN/SAM VpcConfig validation; read-only). The four + # ec2:Describe* actions do not support resource-level + # permissions AT ALL — an ARN in Resource is ignored and the call is + # authorised against "*" — so narrowing them is cosmetic. The ENI in + # CreateNetworkInterface / DeleteNetworkInterface is created by the + # Lambda service at attach time with an id that cannot exist when this + # policy is written. Nothing here is scopable by resource name. + # AssignPrivateIpAddresses / UnassignPrivateIpAddresses are for EFA and + # secondary IPs — not part of the Lambda ENI lifecycle — omitted. + # + # KNOWN OPEN ITEM (pre-existing, NOT introduced by INFRA-186; tracked + # as INFRA-200): + # ec2:DeleteNetworkInterface on "*" lets a bounded Lambda delete any ENI + # in the account, including NAT / VPC-endpoint / RDS ENIs — a + # denial-of-service primitive inherited from the AWS managed policy. The + # durable fix is a Condition on ec2:Subnet / ec2:Vpc naming THE SET OF + # VPCs that boundary-carrying Lambdas attach to — not a single VPC id; + # the list must be extended whenever a workload introduces a new VPC + # (tag-based conditions are the alternative if the set churns). No such + # VPC exists in seahaven-prod or seahaven-dev today (payments-dashboard's + # 10.20.0.0/16 VPC is in mgmt), so writing the condition now would encode + # an mgmt resource into a prod/dev template. Whoever brings the first VPC + # across in payments-dashboard's migration PR adds the condition in the + # same PR. - Sid: Ec2Eni Effect: Allow Action: @@ -183,115 +487,70 @@ Resources: - ec2:DescribeVpcs Resource: "*" - # ── DynamoDB (afterhours, payments, meal-order, front-integrations) ─ - # Table/* covers base-table operations; table/*/index/* is required for - # Query/Scan on Global Secondary Indexes. - - Sid: DynamoDB - Effect: Allow - Action: - - dynamodb:GetItem - - dynamodb:PutItem - - dynamodb:UpdateItem - - dynamodb:DeleteItem - - dynamodb:Query - - dynamodb:Scan - - dynamodb:BatchGetItem - - dynamodb:BatchWriteItem - - dynamodb:DescribeTable - - dynamodb:ConditionCheckItem - Resource: - - !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/*" - - !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/*/index/*" - - # ── S3 (payments-dashboard read, meal-order-manager CRUD) ────────── - - Sid: S3 - Effect: Allow - Action: - - s3:GetObject - - s3:PutObject - - s3:DeleteObject - - s3:ListBucket - - s3:GetBucketLocation - - s3:GetObjectVersion - - s3:GetObjectTagging - - s3:PutObjectTagging - Resource: - - !Sub "arn:aws:s3:::*-${AWS::AccountId}" - - !Sub "arn:aws:s3:::*-${AWS::AccountId}/*" - # meal-order-manager ReportsBucket (non-AccountId suffix pattern) - - !Sub "arn:aws:s3:::meal-order-manager-reports-${AWS::AccountId}" - - !Sub "arn:aws:s3:::meal-order-manager-reports-${AWS::AccountId}/*" - - !Sub "arn:aws:s3:::meal-order-manager-form-${AWS::AccountId}" - - !Sub "arn:aws:s3:::meal-order-manager-form-${AWS::AccountId}/*" - - # ── Secrets Manager (all stacks) ────────────────────────────────── - - Sid: SecretsManager - Effect: Allow - Action: - - secretsmanager:GetSecretValue - - secretsmanager:DescribeSecret - Resource: - - !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:*" - - # ── SSM Parameter Store (meal-order-manager, afterhours) ────────── - - Sid: SSMParameterRead - Effect: Allow - Action: - - ssm:GetParameter - - ssm:GetParameters - - ssm:GetParametersByPath - Resource: - - !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/*" - - # ── SQS (payments-dashboard batch queues) ───────────────────────── - - Sid: SQS - Effect: Allow - Action: - - sqs:SendMessage - - sqs:ReceiveMessage - - sqs:DeleteMessage - - sqs:GetQueueAttributes - - sqs:GetQueueUrl - - sqs:ChangeMessageVisibility - Resource: - - !Sub "arn:aws:sqs:us-east-1:${AWS::AccountId}:*" - - # ── Lambda invocation (payments, meal-order inter-function calls) ── - - Sid: LambdaInvoke - Effect: Allow - Action: - - lambda:InvokeFunction - Resource: - - !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:function:*" - - # ── SES (afterhours weekly-post, meal-order email-report) ────────── - - Sid: SES - Effect: Allow - Action: - - ses:SendEmail - - ses:SendRawEmail - Resource: - - !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:identity/*" - - !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:configuration-set/*" - - # ── KMS (CMK-encrypted resources) ───────────────────────────────── - # Required for Lambda functions that read/write CMK-encrypted AWS - # resources. Verified live state: - # - PaymentsDashboard DynamoDB table: CMK key/0b660af3 (KMS:ENABLED) - # - payments-dashboard CloudWatch log groups: CMK key/b748750c - # Secrets Manager + SQS queues in these stacks use AWS-managed keys - # (aws/secretsmanager, aws/sqs) which do not require explicit kms:* - # actions in the execution role policy. The CMK keys are scoped to - # this account to prevent cross-account KMS calls. - - Sid: KMS - Effect: Allow - Action: - - kms:Decrypt - - kms:GenerateDataKey - - kms:DescribeKey - Resource: - - !Sub "arn:aws:kms:us-east-1:${AWS::AccountId}:key/*" - + # ── NO PER-WORKLOAD DATA-PLANE STATEMENTS — BY DESIGN ─────────────── + # The statements above are the FLEET-WIDE FLOOR: what every Lambda + # execution role needs regardless of which workload it belongs to. + # There are deliberately NO DynamoDB, S3, Secrets Manager, SSM, SQS, + # SES, KMS, lambda:InvokeFunction or scheduler statements here. + # + # WHY (decided 2026-07-30, Adam): + # The security win of INFRA-186 comes from DELETION, not enumeration. + # Removing the account-wide secret:*, table/*, function:* and sqs:* + # wildcards is what closes the amplifier — the ability of a principal + # who can write an inline policy onto a boundary-carrying role to read + # every secret in the account. Per-workload prefixes add no security; + # they exist only to keep a workload FUNCTIONAL once it arrives. + # + # An earlier revision of this branch PRE-LOADED prefixes for all five + # mgmt SAM stacks before any of them had migrated. That required + # predicting five stacks' permission needs from the permission-source + # comment block above, and the /sh-security-review pass found SIX + # errors in the result — three of which would have failed SILENTLY at + # first migration (afterhours' SES config-set, its holiday scheduler + # behind a bare except, and afi's webhook secret under an invented + # name). The block is a secondary record and is not a substitute for + # reading the owning repo's template. + # + # WIDENING IS THE SAFE DIRECTION. Adding a resource to a boundary can + # never break a running Lambda; only tightening can. So there is no + # cost to deferring per-workload scope to the migration PR that has + # the real template open in front of it — and a large cost to + # guessing it years ahead of the migration. + # + # CONSEQUENCE FOR EVERY MIGRATION PR (mandatory, see WIDENING PATH in + # the header): a stack landing in prod or dev MUST add its own + # data-plane statements here, derived from ITS OWN template, in its + # own PR to THIS repo, deployed to UPDATE_COMPLETE before the + # workload's first deploy from its own repo (the workload PR and the + # widening PR cannot be the same PR — they live in different repos). + # Without them its Lambdas get AccessDenied + # at first invoke. The permission-source block above is the starting + # point, NOT the authority — verify every entry against the stack. + # + # Prod/dev boundary usage is 0 (verified 2026-07-30), so this floor + # currently constrains nothing that exists. Fleet-wide statements that + # genuinely cannot be scoped (Logs, X-Ray, ENI) stay above with their + # justifications; they are also the SILENT-failure classes, which is + # why they belong in the floor rather than in per-workload widenings. + # + # KMS is absent deliberately: zero boundary-carrying roles exist, so + # nothing bounded decrypts anything yet. (Log-group CMKs are the one + # KMS case that never hits an execution role — CloudWatch Logs + # decrypts via the key policy's grant to the Logs service principal — + # so log-group CMK state is not the gate here. seahaven-prod already + # hosts the seahaven-dynamodb-cmk table key; the first migrating + # stack touching that table must cover it.) A workload bringing a + # CMK-encrypted resource adds a KMS statement with the matching + # kms:ViaService principal in its own migration PR — a missing + # ViaService entry denies, and for env-var encryption it fails at + # cold-start INIT. + # + # The end-state fix for the shared-ceiling residual (one boundary = + # every SAM workload reaches every other's data plane once they land) + # is per-workload boundaries — tracked as INFRA-187. Do not improvise + # it: the load-bearing problem there is that both guardrail policies + # pin ONE literal boundary ARN inside StringEquals conditions, and + # loosening that to a wildcard weakens the gate. # --------------------------------------------------------------------------- # Shared CloudFormation execution role (SAM stacks) — INFRA-97 scoped # @@ -789,8 +1048,9 @@ Resources: # conditioned on iam:PermissionsBoundary StringEquals the # seahaven-lambda-execution-boundary ARN. That condition means # any role this execution role creates must have the boundary - # applied, so it can never exceed what the boundary allows - # (which is scoped to the services the five stacks actually use). + # applied, so it can never exceed what the boundary allows (in + # THIS file the fleet-wide floor plus per-migration widenings; + # in mgmt's copy the five SAM stacks' service wildcards). # # iam:PassRole is also included here so CloudFormation can pass # the auto-generated Lambda execution role to the Lambda service. diff --git a/lib/terraform-substrate/terraform-substrate.template.yaml b/lib/terraform-substrate/terraform-substrate.template.yaml index 538f6da..4875cd7 100644 --- a/lib/terraform-substrate/terraform-substrate.template.yaml +++ b/lib/terraform-substrate/terraform-substrate.template.yaml @@ -65,8 +65,14 @@ Description: >- # deploy-substrate stack instead. The coupling is by NAME: if the boundary # policy is ever renamed or replaced, every Condition below (and the # deploy-substrate copy) must change in the same piece of work. INFRA-186 -# (per-workload boundary scoping) changes the boundary's CONTENT, not its ARN, -# and does not touch this file. +# (boundary reduced to a fleet-wide floor; per-workload boundaries are +# INFRA-187) changed the boundary's CONTENT, not its ARN, so this file is +# textually untouched — but the Terraform path IS affected: the Conditions +# below FORCE every role a Terraform apply creates onto that boundary, and +# the floor carries zero data-plane permissions. A migrating stack that +# creates Lambda execution roles must widen the boundary per the WIDENING +# PATH in lib/deploy-substrate/deploy-substrate.template.yaml, deployed +# before its first apply (README migration checklist step 2). # # SIZE BUDGET: an attached managed policy document is capped at 6,144 # characters (whitespace excluded). The statement set below is ~2.5 KB.