Merge pull request #68 from Sea-Haven-Industries/feature/INFRA-186-boundary-prod-dev-scoping
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run

refactor(iam): reduce prod/dev Lambda execution boundary to a fleet-wide floor (INFRA-186)
This commit is contained in:
Adam Moussa 2026-07-31 13:49:55 -04:00 • committed by GitHub
commit 9faf0295d5
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
3 changed files with 451 additions and 161 deletions

View file

@ -133,11 +133,25 @@ account-level deploy plumbing:
only for an account that does not already have one — one provider per URL
per account).
The template is a verbatim extraction of the substrate section of
`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml` (see the provenance
header in the template — mgmt's copy remains source of truth for 328440206208
until its stacks migrate out; substrate changes while both are live must edit
both files). Per-repo `githubdeploy-*` deploy roles are deliberately NOT part
The template began as a verbatim extraction of the substrate section of
`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml`, which remains the source
of truth for mgmt (328440206208) until its stacks migrate out.
**The two copies are no longer at parity, and the old "edit both files" rule no
longer applies uniformly.** Under INFRA-186, `seahaven-lambda-execution-boundary`
in *this* copy was reduced to a fleet-wide floor for prod and dev (where
boundary usage was 0, so no live Lambda could break): CloudWatch Logs write on
`/aws/lambda*`, log-group describe, X-Ray, and ENI lifecycle — nothing else.
Each migrating stack adds its own data-plane statements, derived from its own
template, in its own PR (per-workload boundaries are the INFRA-187 end state).
mgmt's copy keeps the account-wide wildcards pending its own separately
validated rollout across 26 live boundary-carrying roles. So: **the boundary
resource is deliberately divergent**; every *other* substrate resource
(`github-cfn-execution-role`, `seahaven-cfn-exec-iam-management`) is still
expected to change in both files together. The template's provenance header
records which is which — read it before assuming either parity or divergence.
Per-repo `githubdeploy-*` deploy roles are deliberately NOT part
of the substrate — they are provisioned per repo at migration/onboarding time
so an account never carries trust relationships for repos that do not deploy
to it.
@ -258,14 +272,24 @@ split this substrate exists to enforce.
`organization:seahaven:project:seahaven-<env>:workspace:<workspace>:run_phase:plan`
(or `:apply`). Exact `StringEquals` only — never `StringLike`, never a
wildcarded `run_phase` (a speculative PR plan must never hold write
credentials). IAM roles = mandatory GPT-4.1 cross-review +
credentials). **If the stack creates Lambda execution roles, this same PR
must also widen `seahaven-lambda-execution-boundary`** per the WIDENING
PATH in `lib/deploy-substrate/deploy-substrate.template.yaml`: the
guardrail forces every Terraform-created role to carry that boundary, and
it is a fleet-wide floor with zero data-plane permissions until widened —
an unwidened migration deploys green, then every data-plane call is denied
at first invoke and async/DLQ writes are discarded silently. IAM roles and
boundary widenings = mandatory cross-family review +
`/sh-security-review` on the diff.
3. After deploy, verify: both roles exist; `hcptf-<stack>` lists
`seahaven-hcptf-iam-management` in `list-attached-role-policies`; trust
subs match the live org/project/workspace names byte-for-byte; simulate
the apply role against a `hcptf-*` ARN (expect `explicitDeny` from
`DenySelfMutation`) and against a normal stack role name (expect
`allowed`).
`allowed`); and if step 2 widened the boundary, confirm the deployed
default version carries the stack's data-plane statements
(`aws iam get-policy-version`) — role verification alone never checks
boundary content.
4. Set **workspace-level** variables `TFC_AWS_PLAN_ROLE_ARN` +
`TFC_AWS_APPLY_ROLE_ARN` (category env) to the verified role ARNs, plus
`TFC_AWS_PROVIDER_AUTH=true`. Never project-scoped variable sets — the

View file

@ -39,16 +39,115 @@ Description: >-
# added. The mgmt copy was remediated 2026-07-27 (.github PRs #95 Phase A
# + #98 Phase B); DenySelfMutation and the widened policy/seahaven-*
# DenyBoundaryPolicyEdit scope were then ported back here, so the two
# copies' statement sets are reconciled as of that date — every IAM
# statement in LambdaExecutionBoundary, SamCfnIamManagementPolicy and
# SamCfnExecutionRole is byte-identical across the two files; the only
# remaining delta is the DependsOn line above, which is ordering, not
# permission. If a substrate statement changes again, change BOTH files
# in the same piece of work.
# copies' GUARDRAIL statement sets were reconciled as of that date.
# SamCfnIamManagementPolicy and SamCfnExecutionRole remain at parity on
# their IAM STATEMENT SETS across the two files and MUST still be changed
# together. Parity covers statements, not surrounding comments — a comment
# may diverge where it describes boundary content, which now differs
# between the files. The only functional delta
# between them is the DependsOn line above, which is ordering, not
# permission. LambdaExecutionBoundary is NO LONGER byte-identical — see
# DELIBERATE DIVERGENCE below.
#
# DELIBERATE DIVERGENCE — LambdaExecutionBoundary (INFRA-186, 2026-07-30)
# The parity rule above is SCOPED, not global. LambdaExecutionBoundary in THIS
# file is DELIBERATELY STRICTER than the mgmt copy in
# Sea-Haven-Industries/.github/oidc-deploy-roles.yaml. Do not "reconcile" the two
# by copying mgmt's statements back over these, or vice versa; the divergence is
# load-bearing. A future mechanical drift check WILL read it as drift — it is not.
#
# 1. WHAT DIVERGED. Every per-workload data-plane statement was REMOVED from
# LambdaExecutionBoundary in this file, leaving only the fleet-wide floor:
# CloudWatchLogsWrite (scoped to /aws/lambda*), CloudWatchLogsDescribe,
# XRay and Ec2Eni. The account-wide wildcards mgmt still carries — table/*,
# table/*/index/*, secret:*, parameter/*, sqs :*, function:*, ses
# identity/* + configuration-set/*, kms key/*, and an s3:::*-<accountid>
# pattern that was a bare name-suffix filter rather than an ownership
# check — are simply GONE here rather than re-scoped.
# The security win is the deletion: it is what closes the amplifier whereby
# a principal able to write an inline policy onto a boundary-carrying role
# could read every secret in the account. Per-workload prefixes add no
# security — they only keep a workload functional — so they are added by
# each migration PR, from that stack's own template, when the stack
# actually lands. See the note on the boundary resource for the full
# rationale and the six errors that the pre-loaded approach produced.
#
# 2. WHY MGMT'S RATIONALE IS LEGITIMATE THERE. The superset framing this file
# used to carry ("being slightly broad is the correct trade-off; a boundary
# that is too tight will break Lambda functions at runtime AFTER deploy") is
# a real constraint in the management account: 328440206208 has 26 LIVE
# roles carrying seahaven-lambda-execution-boundary, across all five SAM
# stacks. Tightening there is a production change to running workloads with
# a silent, deploy-time-invisible failure mode.
#
# 3. WHY IT DOES NOT TRANSFER HERE. This file deploys ONLY to seahaven-prod
# (011934824531) and seahaven-dev (710827005802), where
# PermissionsBoundaryUsageCount is 0 and 0 respectively (aws iam get-policy,
# verified 2026-07-30; corroborated by list-roles returning no role carrying
# any permissions boundary in either account). No live Lambda can break, so
# the risk that justifies mgmt's breadth is absent — while the exposure is
# strictly WORSE here than in mgmt, because prod is multi-tenant: the old
# wildcards reached proposal-system's, procurement-ingest's and
# workorder-ingest's CDK-owned tables, buckets, secrets and queues, the org's
# own Config and VPC-flow-log buckets, and — via function:* — CDK Lambdas
# that carry no boundary at all.
#
# 4. RECONCILIATION OBLIGATION, RESTATED. For LambdaExecutionBoundary the two
# copies are now INTENTIONALLY DIFFERENT and must NOT be synchronised:
# - A change to the per-workload Resource patterns in THIS file does NOT
# propagate to mgmt.
# - A change to mgmt's boundary does NOT propagate here.
# - Any change to the ACTION lists, or any new statement, is a substrate
# semantic change and DOES still require the same review in both copies.
# For SamCfnIamManagementPolicy and SamCfnExecutionRole the original rule is
# unchanged: change BOTH files in the same piece of work.
#
# KNOWN OPEN ITEM (deferred, not closed by INFRA-186): mgmt 328440206208 still
# carries the account-wide patterns. Tightening it needs its own validated
# rollout — enumerate what the 26 live roles actually call, stage it, and be
# ready to roll back — and is explicitly OUT OF SCOPE of INFRA-186. Until that
# lands, the two copies stay divergent and that is the intended state.
#
# COUPLING: the boundary's ManagedPolicyName and ARN are UNCHANGED and must stay
# unchanged. Four Conditions in SamCfnIamManagementPolicy below, and four more in
# HcptfIamManagementPolicy in lib/terraform-substrate/terraform-substrate.template.yaml,
# pin arn:aws:iam::<account>:policy/seahaven-lambda-execution-boundary by literal
# string inside StringEquals iam:PermissionsBoundary. A rename fails SILENTLY — an
# IAM condition naming a non-existent policy simply never matches, so the
# escalation control would evaporate rather than error — and would additionally
# force a CloudFormation REPLACEMENT that any role carrying the boundary would
# block. INFRA-186 is a CONTENT-ONLY change for exactly this reason;
# terraform-substrate.template.yaml already records that expectation and is
# correctly left untouched.
#
# SIZE BUDGET: an attached managed policy document is capped at 6,144 characters
# (whitespace excluded). LambdaExecutionBoundary measures 691 characters across
# 4 statements as of 2026-07-31 — the fleet-wide floor only. Measure before
# widening — len(json.dumps(doc,separators=(',',':'))) on the synthesized
# PolicyDocument with ${AWS::AccountId} resolved, and UPDATE THESE TWO NUMBERS in
# the same edit (they went stale twice inside this branch alone).
#
# Headroom is 5,453 characters, roughly TWELVE workloads at ~450 each. That is a
# deliberate outcome, not luck: an earlier revision of this branch pre-loaded
# per-workload prefixes for all five mgmt SAM stacks and reached 5,457 characters
# with 687 left — about one workload of room — before any stack had actually
# migrated. Deferring per-workload scope to each migration PR (see the note on
# the boundary itself) removed that pressure entirely. If the budget tightens
# again as workloads land, the end-state fix is per-workload boundaries
# (seahaven-lambda-execution-boundary-<workload>), which also resolves the
# shared-ceiling residual — tracked as INFRA-187, do not improvise it.
# CRITICAL: unlike the 2026-07-27 inline-limit incident,
# there is NO restructure available when this cap is reached — a role has exactly
# ONE permissions boundary, so statements cannot be spilled into a second attached
# managed policy. At the cap the only levers are prefix consolidation and dropping
# unused actions.
#
# This template is deployed via lib/deploy-substrate-stack.ts
# (cloudformation-include) as stack seahaven-deploy-substrate, once per
# member account that hosts SAM workloads.
# (cloudformation-include) as stack seahaven-deploy-substrate, once per member
# account that hosts SAM workloads (currently seahaven-prod 011934824531 and
# seahaven-dev 710827005802 via bin/app.ts instances deploy-substrate-prod /
# deploy-substrate-dev; NEVER mgmt — 328440206208 is served by the .github copy
# named above until its stacks migrate out).
Parameters:
CreateOIDCProvider:
@ -83,7 +182,7 @@ Resources:
UpdateReplacePolicy: Retain
# ---------------------------------------------------------------------------
# Lambda execution permissions boundary (INFRA-103)
# Lambda execution permissions boundary (INFRA-103, re-scoped by INFRA-186)
#
# This managed policy is the CEILING for every Lambda execution role that the
# five SAM stacks auto-generate via AWS::Serverless::Function. Applying it as
@ -91,49 +190,173 @@ Resources:
# intersection of the role's own policies and this boundary, so a misconfigured
# SAM role can never exceed what is listed here.
#
# The boundary is intentionally a SUPERSET of the union of all runtime
# permissions currently granted across the five stacks. Being slightly broad
# is the correct trade-off at this stage — a boundary that is too tight will
# break Lambda functions at runtime after deploy, which is worse than a slightly
# loose boundary that is tightened in a follow-up.
# SCOPING RULE (INFRA-186, 2026-07-30). This boundary is the FLEET-WIDE FLOOR
# ONLY: what every Lambda execution role needs regardless of workload. It
# carries NO per-workload data-plane statements — each migrating stack adds its
# own, from its own template, in its own PR (see the note on the boundary
# resource and the WIDENING PATH below). It was previously a deliberate SUPERSET
# with account-wide wildcards (table/*, secret:*, sqs :*, function:*,
# parameter/*, and an s3:::*-<accountid> pattern that was a name-suffix filter,
# not an ownership check). That trade-off was made when the only account
# carrying this policy hosted nothing but the five SAM stacks. It does not
# survive multi-tenancy: seahaven-prod now hosts CDK-owned tenants
# (proposal-system, procurement-ingest, workorder-ingest) whose tables, buckets,
# secrets and queues those wildcards reached, and whose Lambdas carry NO
# permissions boundary at all — making function:* a boundary-escape primitive.
#
# Permission sources per stack:
# Tightening here carries ZERO runtime risk and was sequenced deliberately:
# PermissionsBoundaryUsageCount is 0 in BOTH accounts this template deploys to
# (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified
# 2026-07-30), so no live Lambda can break. A boundary that is slightly TOO
# TIGHT is recoverable here — the migrating stack widens it in its own PR before
# its first deploy — whereas leaving it loose perpetuates the exposure. Prefer
# tighter; the widening path is below.
#
# afterhours-shift-manager
# A resource pattern that genuinely CANNOT be scoped keeps its wildcard WITH a
# written justification on the statement: CloudWatchLogsDescribe, XRay and
# Ec2Eni name runtime-created resources or use actions AWS authorises against
# "*" regardless of the ARN supplied. Do not "tighten" those.
#
# Every "verified <date>" annotation in this file is a POINT-IN-TIME
# observation, not live state. Re-validate (usage counts, log-group CMK state,
# per-stack permission sources) before citing one as justification for a
# future change.
#
# WIDENING PATH — read this before migrating a stack into prod or dev.
# The boundary is never widened by the person who hits the AccessDenied. It is
# widened by the migrating stack's owner, in THIS repo, BEFORE the workload's
# first deploy into the target account:
# 0. PRECONDITION — check the managed-policy VERSION budget BEFORE merging:
# max 5 versions, both
# accounts are on v1 today. Every widening (PolicyDocument edit) burns
# one. Description, ManagedPolicyName and Path are REPLACEMENT
# properties per the CFN resource reference — CloudFormation cannot
# replace a custom-named policy, so a Description-only edit FAILS the
# stack update, and the error's suggested remedy (rename) is exactly
# the forbidden rename in the COUPLING note above. Never edit those
# three properties. Delete the oldest non-default version if at 5:
# aws iam list-policy-versions --policy-arn \
# arn:aws:iam::<acct>:policy/seahaven-lambda-execution-boundary
# aws iam delete-policy-version --version-id v<oldest-non-default> ...
# 1. Derive the workload's needs from ITS OWN TEMPLATE — open the stack's
# template.yaml and read the actual IAM policy statements. The
# permission-source block below is a STARTING POINT, NOT THE AUTHORITY:
# the /sh-security-review pass on 2026-07-30 found SIX places where it was
# incomplete or simply invented a resource name, three of which would have
# failed silently. Update that block in the same edit with what you find.
# 2. Add the workload's statements. Group by service so a second workload can
# extend a Resource list rather than duplicate an action list (~250
# characters for zero new actions; an extra ARN costs ~60). Check for the
# SILENT classes specifically: a denied SQS destination/DLQ write discards
# the async event with no error and no alarm; a denied scheduler call may
# sit behind a bare except; a denied KMS decrypt for env-var encryption
# fails at cold-start INIT; any CMK-encrypted resource needs the
# matching kms:ViaService principal, not just the kms action; and a
# function using LoggingConfig with a custom log-group name outside
# /aws/lambda* silently loses ALL logs — add a scoped logs statement
# for the custom group or keep the default group name.
# 3. Measure. See SIZE BUDGET in the header. There is NO escape hatch.
# 4. Both review gates run and neither discharges the other: the GPT-4.1
# cross-family review against the real diff, and /sh-security-review
# (IaC/IAM is on the mandatory surface). CLI down = review outstanding.
# 5. Merge and let CI deploy deploy-substrate-prod / deploy-substrate-dev to
# UPDATE_COMPLETE, THEN deploy the workload stack.
# ORDERING IS NOT ENFORCED BY CLOUDFORMATION AND THIS IS THE MOST IMPORTANT
# SENTENCE HERE: the workload's deploy SUCCEEDS even against a stale boundary,
# because seahaven-cfn-exec-iam-management's gate checks that the boundary ARN
# is attached, never its contents. The failure surfaces later, at first invoke,
# as AccessDenied. A stale boundary is a silent deploy-time pass and a loud
# production failure.
#
# CONSIDERED AND REJECTED: a Deny statement reserving the seahaven-* namespace.
# With the Allow set reduced to the fleet-wide floor it is fully redundant (verified
# 2026-07-30: seahaven-prod-config-* and seahaven-prod-vpc-flow-logs-* are
# already denied by the Allow set alone), and a Deny inside a BOUNDARY is the
# hardest failure mode in the estate to debug — it beats every Allow in every
# policy with no synth-time signal. Revisit only if a widening ever has to
# re-broaden a per-service Resource list back toward a wildcard.
#
# NOTE ON WHAT THESE PREFIXES ARE. All five stacks below currently live in the
# MANAGEMENT account and none of their resources exists in seahaven-prod or
# seahaven-dev yet. These are MIGRATION-CANDIDATE prefixes for the accounts this
# template deploys to, not an inventory of what is deployed there. They are a
# SECONDARY RECORD and a starting point for widening PRs — the authority is
# each stack's own template (WIDENING PATH step 1). No Resource pattern in
# the floor above derives from this block.
#
# Permission sources per stack (verified live 2026-07-30 IN MGMT — these
# stacks and resources do not exist in prod/dev yet, so nothing below is a
# prod/dev observation; starting point only — verify every entry against
# the owning repo before use. PARAMETERIZED resources (ARNs passed as
# deploy parameters) must be re-derived from the live stack configuration
# at migration time, as exact ARNs — never inferred from these names into
# broad patterns like secret:afi-*):
#
# afterhours-shift-manager (functions: afterhours-*, 6 live)
# - DynamoDB CRUD (afterhours-shifts table)
# - secretsmanager:GetSecretValue (afterhours-shift-manager/*)
# - ses:SendEmail (SES identity)
# - ses:SendEmail on the identity AND on
# configuration-set/seahaven-email-events (template.yaml:178-181 — the
# send is DENIED without the config-set ARN when the identity has a
# default configuration set)
# - scheduler:CreateSchedule/DeleteSchedule/GetSchedule on
# schedule/default/holiday-* + iam:PassRole to scheduler.amazonaws.com
# (template.yaml:110-120). CORRECTED 2026-07-30: this block previously
# omitted both, and the omission is SILENT at runtime (bare except).
# - NO ssm. CORRECTED 2026-07-30: this block previously credited
# ssm:GetParameter to this stack; `grep -c 'ssm:' template.yaml` = 0.
# Its slack tokens come from Secrets Manager and the channel id from a
# CloudFormation parameter.
# - lambda:InvokeFunction (ReleaseNotifyInvokeRole,
# HolidaySchedulerExecutionRole — these two carry the boundary and need
# ONLY this action)
# - CloudWatch Logs (all functions)
# - UNRESOLVED: /3cx-scheduler/* ownership (this stack vs. the retired
# standalone 3CX scheduler). Deliberately NOT granted. Resolve at
# migration in that stack's own repo — do NOT add any 3cx-scheduler
# resource here until ownership is resolved.
#
# payments-dashboard
# - DynamoDB CRUD / Read (PaymentsDashboard table)
# - S3 GetObject (payroll-emails, payments-csv buckets)
# payments-dashboard (functions: payments-*)
# - DynamoDB CRUD / Read (PaymentsDashboard table — legacy PascalCase)
# - S3 GetObject on seahaven-payments-csv-* and seahaven-payroll-emails-*;
# GetObject + PutObject on seahaven-payments-boa-raw-*
# (template.yaml:272 and 1097-1099, fetchBoaTransactions raw archive).
# CORRECTED 2026-07-30: this block previously said "GetObject ONLY —
# no write intent enumerated", which was false and would have denied
# the raw-archive write at migration.
# - secretsmanager:GetSecretValue (payments-dashboard/*)
# - sqs:SendMessage + sqs:ReceiveMessage + sqs:DeleteMessage etc.
# (PayrollBatchQueue + DLQs)
# - lambda:InvokeFunction (ExpenseReceiver → ExpenseProcessor)
# - ec2:CreateNetworkInterface / DescribeNetworkInterfaces /
# DeleteNetworkInterface (VPC-attached functions)
# - sqs Send/Receive/Delete etc. (payments-payroll-batch + DLQs)
# - lambda:InvokeFunction (ExpenseReceiver -> ExpenseProcessor)
# - ec2 ENI lifecycle (VPC-attached functions)
# - KMS via dynamodb (table CMK) — no SSM, no SES
# - CloudWatch Logs
#
# meal-order-manager
# meal-order-manager (functions: meal-order-manager-*, 7 live)
# - DynamoDB CRUD / Read (meal-order-manager-orders table)
# - S3 CRUD (ReportsBucket) + s3:GetObject (ReportsBucket presigned URLs)
# - S3 CRUD (meal-order-manager-reports-*, meal-order-manager-form-*)
# - secretsmanager:GetSecretValue (meal-order-manager/*)
# - ssm:GetParameter (/meal-order-manager/*)
# - lambda:InvokeFunction (submit-order → slack-notifier,
# close-form → aggregate-orders)
# - lambda:InvokeFunction (submit-order -> slack-notifier,
# close-form -> aggregate-orders, plus AdminAuthorizerInvokeRole)
# - ses:SendRawEmail
# - CloudWatch Logs
#
# front-integrations
# front-integrations (functions: front-*)
# - DynamoDB CRUD (front-sla-alerts table)
# - secretsmanager:GetSecretValue (by ARN, various)
# - secretsmanager:GetSecretValue (front-integrations/*)
# - CloudWatch Logs
# - no IAM permissions for S3 / SQS / SSM / SES / KMS / VPC in this
# stack's template as of 2026-07-30
#
# afi-backup-monitor
# - secretsmanager:GetSecretValue (by ARN)
# afi-backup-monitor (functions: afi-*)
# - secretsmanager:GetSecretValue on TWO bare, unprefixed secrets:
# afi-api-key and afi-slack-webhook. CORRECTED 2026-07-30: this block
# previously named "afi-backup-monitor/slack-webhook-url", which does
# not exist. Both ARNs are deploy PARAMETERS in that stack
# (AfiApiKeySecretArn / SlackWebhookSecretArn), so no name is
# discoverable from its template — the names are in its README:48-49.
# - CloudWatch Logs
# - nothing else
#
# ---------------------------------------------------------------------------
LambdaExecutionBoundary:
@ -149,18 +372,76 @@ Resources:
Version: "2012-10-17"
Statement:
# ── CloudWatch Logs (every Lambda) ──────────────────────────────────
- Sid: CloudWatchLogs
# ── CloudWatch Logs — write (every Lambda) ──────────────────────────
# Scoped to the Lambda log-group namespace. Every SAM function's group
# is /aws/lambda/<function>, and the trailing * also covers the
# :log-stream:<name> suffix PutLogEvents authorises against, so one ARN
# serves CreateLogGroup, CreateLogStream, PutLogEvents and
# DescribeLogStreams. The * is deliberately NOT after a trailing slash:
# /aws/lambda* also matches the /aws/lambda-insights groups the Lambda
# Insights extension writes to, which /aws/lambda/* would have denied.
# VERIFICATION PROVENANCE, stated precisely (2026-07-30). What
# iam simulate-custom-policy DOES confirm: this pattern allows
# logs:CreateLogGroup / logs:PutLogEvents on the bare group ARN
# log-group:/aws/lambda/<fn> (and /aws/lambda/<a>/<b>), and DENIES
# log-group:seahaven-prod-vpc-flow-logs — the latter re-checked against
# an Allow */* positive control, which allows it, so the deny is real
# policy behaviour and not a simulator artifact.
# What the simulator CANNOT evaluate, so do NOT claim it was verified:
# log-stream-qualified ARNs (log-group:<g>:log-stream:<s>) and the bare
# /aws/lambda-insights group both return implicitDeny EVEN UNDER an
# Allow */* policy. That is a simulator resource-parsing limitation, not
# a denial. Coverage of those two rests on documented IAM wildcard
# semantics — "*" matches any sequence of characters including ":" and
# "/" — which is why the * is deliberately NOT placed after a trailing
# slash. If this ever needs true end-to-end proof, it must come from a
# real invoke in dev, not from the simulator.
#
# NOT scoped per workload, deliberately. A per-stack prefix
# (/aws/lambda/payments-* etc.) was considered and rejected: a Lambda
# denied PutLogEvents does not fail — it keeps running and silently
# produces no logs. Log denial is the one failure class in this policy
# that is NOT loud, so it must not depend on function-name discipline.
# ACCEPTED RESIDUAL RISK: a SAM Lambda can write into another tenant's
# /aws/lambda/* group (log poisoning). No read action is granted here, so
# this is not an exfiltration path. Tighten only once every
# boundary-carrying function is confirmed to set an explicit FunctionName.
# REGION IS PINNED TO us-east-1 DELIBERATELY: every Sea Haven workload
# deploys to us-east-1, and this template itself only ever deploys there.
# ${AWS::Region} would resolve to the identical string, so it would
# document nothing. A future stack in another region carries this
# boundary but CANNOT write its logs (the silent class above) — so a
# cross-region migration MUST add region-scoped statements in its
# widening PR, same as any other data-plane need.
- Sid: CloudWatchLogsWrite
Effect: Allow
Action:
- logs:CreateLogGroup
- logs:CreateLogStream
- logs:PutLogEvents
- logs:DescribeLogGroups
- logs:DescribeLogStreams
Resource:
- !Sub "arn:aws:logs:us-east-1:${AWS::AccountId}:log-group:/aws/lambda*"
# ── CloudWatch Logs — describe (UNSCOPABLE, kept "*" deliberately) ──
# logs:DescribeLogGroups is a COLLECTION action: AWS authorises it
# against "*" regardless of any resource ARN supplied. Scoping it would
# produce a policy that reads tighter and denies at runtime, so it is
# split into its own statement and keeps the wildcard. Read-only
# metadata; it cannot mutate anything or return log content.
- Sid: CloudWatchLogsDescribe
Effect: Allow
Action:
- logs:DescribeLogGroups
Resource: "*"
# ── X-Ray tracing (standard Lambda execution) ────────────────────
# ── X-Ray tracing (UNSCOPABLE, kept "*" deliberately) ───────────────
# xray:PutTraceSegments / PutTelemetryRecords support no resource-level
# permissions — X-Ray exposes no ARN for them, which is why the AWS
# managed AWSXRayDaemonWriteAccess also uses "*". Any ARN written here
# would be inert and would falsely imply a control exists. Write-only
# into this account's own trace store; no cross-tenant read is
# expressible with this action set.
- Sid: XRay
Effect: Allow
Action:
@ -168,10 +449,33 @@ Resources:
- xray:PutTelemetryRecords
Resource: "*"
# ── VPC / ENI management (payments-dashboard VPC functions) ────────
# Matches AWSLambdaVPCAccessExecutionRole exactly.
# AssignPrivateIpAddresses / UnassignPrivateIpAddresses are for EFA
# and secondary IPs — not part of the Lambda ENI lifecycle — omitted.
# ── VPC / ENI management (UNSCOPABLE, kept "*" deliberately) ────────
# Derived from AWSLambdaVPCAccessExecutionRole, NOT an exact match:
# DescribeSecurityGroups and DescribeVpcs exceed that managed policy
# (kept for CFN/SAM VpcConfig validation; read-only). The four
# ec2:Describe* actions do not support resource-level
# permissions AT ALL — an ARN in Resource is ignored and the call is
# authorised against "*" — so narrowing them is cosmetic. The ENI in
# CreateNetworkInterface / DeleteNetworkInterface is created by the
# Lambda service at attach time with an id that cannot exist when this
# policy is written. Nothing here is scopable by resource name.
# AssignPrivateIpAddresses / UnassignPrivateIpAddresses are for EFA and
# secondary IPs — not part of the Lambda ENI lifecycle — omitted.
#
# KNOWN OPEN ITEM (pre-existing, NOT introduced by INFRA-186; tracked
# as INFRA-200):
# ec2:DeleteNetworkInterface on "*" lets a bounded Lambda delete any ENI
# in the account, including NAT / VPC-endpoint / RDS ENIs — a
# denial-of-service primitive inherited from the AWS managed policy. The
# durable fix is a Condition on ec2:Subnet / ec2:Vpc naming THE SET OF
# VPCs that boundary-carrying Lambdas attach to — not a single VPC id;
# the list must be extended whenever a workload introduces a new VPC
# (tag-based conditions are the alternative if the set churns). No such
# VPC exists in seahaven-prod or seahaven-dev today (payments-dashboard's
# 10.20.0.0/16 VPC is in mgmt), so writing the condition now would encode
# an mgmt resource into a prod/dev template. Whoever brings the first VPC
# across in payments-dashboard's migration PR adds the condition in the
# same PR.
- Sid: Ec2Eni
Effect: Allow
Action:
@ -183,115 +487,70 @@ Resources:
- ec2:DescribeVpcs
Resource: "*"
# ── DynamoDB (afterhours, payments, meal-order, front-integrations) ─
# Table/* covers base-table operations; table/*/index/* is required for
# Query/Scan on Global Secondary Indexes.
- Sid: DynamoDB
Effect: Allow
Action:
- dynamodb:GetItem
- dynamodb:PutItem
- dynamodb:UpdateItem
- dynamodb:DeleteItem
- dynamodb:Query
- dynamodb:Scan
- dynamodb:BatchGetItem
- dynamodb:BatchWriteItem
- dynamodb:DescribeTable
- dynamodb:ConditionCheckItem
Resource:
- !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/*"
- !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/*/index/*"
# ── S3 (payments-dashboard read, meal-order-manager CRUD) ──────────
- Sid: S3
Effect: Allow
Action:
- s3:GetObject
- s3:PutObject
- s3:DeleteObject
- s3:ListBucket
- s3:GetBucketLocation
- s3:GetObjectVersion
- s3:GetObjectTagging
- s3:PutObjectTagging
Resource:
- !Sub "arn:aws:s3:::*-${AWS::AccountId}"
- !Sub "arn:aws:s3:::*-${AWS::AccountId}/*"
# meal-order-manager ReportsBucket (non-AccountId suffix pattern)
- !Sub "arn:aws:s3:::meal-order-manager-reports-${AWS::AccountId}"
- !Sub "arn:aws:s3:::meal-order-manager-reports-${AWS::AccountId}/*"
- !Sub "arn:aws:s3:::meal-order-manager-form-${AWS::AccountId}"
- !Sub "arn:aws:s3:::meal-order-manager-form-${AWS::AccountId}/*"
# ── Secrets Manager (all stacks) ──────────────────────────────────
- Sid: SecretsManager
Effect: Allow
Action:
- secretsmanager:GetSecretValue
- secretsmanager:DescribeSecret
Resource:
- !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:*"
# ── SSM Parameter Store (meal-order-manager, afterhours) ──────────
- Sid: SSMParameterRead
Effect: Allow
Action:
- ssm:GetParameter
- ssm:GetParameters
- ssm:GetParametersByPath
Resource:
- !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/*"
# ── SQS (payments-dashboard batch queues) ─────────────────────────
- Sid: SQS
Effect: Allow
Action:
- sqs:SendMessage
- sqs:ReceiveMessage
- sqs:DeleteMessage
- sqs:GetQueueAttributes
- sqs:GetQueueUrl
- sqs:ChangeMessageVisibility
Resource:
- !Sub "arn:aws:sqs:us-east-1:${AWS::AccountId}:*"
# ── Lambda invocation (payments, meal-order inter-function calls) ──
- Sid: LambdaInvoke
Effect: Allow
Action:
- lambda:InvokeFunction
Resource:
- !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:function:*"
# ── SES (afterhours weekly-post, meal-order email-report) ──────────
- Sid: SES
Effect: Allow
Action:
- ses:SendEmail
- ses:SendRawEmail
Resource:
- !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:identity/*"
- !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:configuration-set/*"
# ── KMS (CMK-encrypted resources) ─────────────────────────────────
# Required for Lambda functions that read/write CMK-encrypted AWS
# resources. Verified live state:
# - PaymentsDashboard DynamoDB table: CMK key/0b660af3 (KMS:ENABLED)
# - payments-dashboard CloudWatch log groups: CMK key/b748750c
# Secrets Manager + SQS queues in these stacks use AWS-managed keys
# (aws/secretsmanager, aws/sqs) which do not require explicit kms:*
# actions in the execution role policy. The CMK keys are scoped to
# this account to prevent cross-account KMS calls.
- Sid: KMS
Effect: Allow
Action:
- kms:Decrypt
- kms:GenerateDataKey
- kms:DescribeKey
Resource:
- !Sub "arn:aws:kms:us-east-1:${AWS::AccountId}:key/*"
# ── NO PER-WORKLOAD DATA-PLANE STATEMENTS — BY DESIGN ───────────────
# The statements above are the FLEET-WIDE FLOOR: what every Lambda
# execution role needs regardless of which workload it belongs to.
# There are deliberately NO DynamoDB, S3, Secrets Manager, SSM, SQS,
# SES, KMS, lambda:InvokeFunction or scheduler statements here.
#
# WHY (decided 2026-07-30, Adam):
# The security win of INFRA-186 comes from DELETION, not enumeration.
# Removing the account-wide secret:*, table/*, function:* and sqs:*
# wildcards is what closes the amplifier — the ability of a principal
# who can write an inline policy onto a boundary-carrying role to read
# every secret in the account. Per-workload prefixes add no security;
# they exist only to keep a workload FUNCTIONAL once it arrives.
#
# An earlier revision of this branch PRE-LOADED prefixes for all five
# mgmt SAM stacks before any of them had migrated. That required
# predicting five stacks' permission needs from the permission-source
# comment block above, and the /sh-security-review pass found SIX
# errors in the result — three of which would have failed SILENTLY at
# first migration (afterhours' SES config-set, its holiday scheduler
# behind a bare except, and afi's webhook secret under an invented
# name). The block is a secondary record and is not a substitute for
# reading the owning repo's template.
#
# WIDENING IS THE SAFE DIRECTION. Adding a resource to a boundary can
# never break a running Lambda; only tightening can. So there is no
# cost to deferring per-workload scope to the migration PR that has
# the real template open in front of it — and a large cost to
# guessing it years ahead of the migration.
#
# CONSEQUENCE FOR EVERY MIGRATION PR (mandatory, see WIDENING PATH in
# the header): a stack landing in prod or dev MUST add its own
# data-plane statements here, derived from ITS OWN template, in its
# own PR to THIS repo, deployed to UPDATE_COMPLETE before the
# workload's first deploy from its own repo (the workload PR and the
# widening PR cannot be the same PR — they live in different repos).
# Without them its Lambdas get AccessDenied
# at first invoke. The permission-source block above is the starting
# point, NOT the authority — verify every entry against the stack.
#
# Prod/dev boundary usage is 0 (verified 2026-07-30), so this floor
# currently constrains nothing that exists. Fleet-wide statements that
# genuinely cannot be scoped (Logs, X-Ray, ENI) stay above with their
# justifications; they are also the SILENT-failure classes, which is
# why they belong in the floor rather than in per-workload widenings.
#
# KMS is absent deliberately: zero boundary-carrying roles exist, so
# nothing bounded decrypts anything yet. (Log-group CMKs are the one
# KMS case that never hits an execution role — CloudWatch Logs
# decrypts via the key policy's grant to the Logs service principal —
# so log-group CMK state is not the gate here. seahaven-prod already
# hosts the seahaven-dynamodb-cmk table key; the first migrating
# stack touching that table must cover it.) A workload bringing a
# CMK-encrypted resource adds a KMS statement with the matching
# kms:ViaService principal in its own migration PR — a missing
# ViaService entry denies, and for env-var encryption it fails at
# cold-start INIT.
#
# The end-state fix for the shared-ceiling residual (one boundary =
# every SAM workload reaches every other's data plane once they land)
# is per-workload boundaries — tracked as INFRA-187. Do not improvise
# it: the load-bearing problem there is that both guardrail policies
# pin ONE literal boundary ARN inside StringEquals conditions, and
# loosening that to a wildcard weakens the gate.
# ---------------------------------------------------------------------------
# Shared CloudFormation execution role (SAM stacks) — INFRA-97 scoped
#
@ -789,8 +1048,9 @@ Resources:
# conditioned on iam:PermissionsBoundary StringEquals the
# seahaven-lambda-execution-boundary ARN. That condition means
# any role this execution role creates must have the boundary
# applied, so it can never exceed what the boundary allows
# (which is scoped to the services the five stacks actually use).
# applied, so it can never exceed what the boundary allows (in
# THIS file the fleet-wide floor plus per-migration widenings;
# in mgmt's copy the five SAM stacks' service wildcards).
#
# iam:PassRole is also included here so CloudFormation can pass
# the auto-generated Lambda execution role to the Lambda service.

View file

@ -65,8 +65,14 @@ Description: >-
# deploy-substrate stack instead. The coupling is by NAME: if the boundary
# policy is ever renamed or replaced, every Condition below (and the
# deploy-substrate copy) must change in the same piece of work. INFRA-186
# (per-workload boundary scoping) changes the boundary's CONTENT, not its ARN,
# and does not touch this file.
# (boundary reduced to a fleet-wide floor; per-workload boundaries are
# INFRA-187) changed the boundary's CONTENT, not its ARN, so this file is
# textually untouched — but the Terraform path IS affected: the Conditions
# below FORCE every role a Terraform apply creates onto that boundary, and
# the floor carries zero data-plane permissions. A migrating stack that
# creates Lambda execution roles must widen the boundary per the WIDENING
# PATH in lib/deploy-substrate/deploy-substrate.template.yaml, deployed
# before its first apply (README migration checklist step 2).
#
# SIZE BUDGET: an attached managed policy document is capped at 6,144
# characters (whitespace excluded). The statement set below is ~2.5 KB.