mirror of
https://github.com/Sea-Haven-Industries/seahaven-account-baseline.git
synced 2026-08-04 16:56:14 +00:00
Merge pull request #68 from Sea-Haven-Industries/feature/INFRA-186-boundary-prod-dev-scoping
refactor(iam): reduce prod/dev Lambda execution boundary to a fleet-wide floor (INFRA-186)
This commit is contained in:
commit
9faf0295d5
3 changed files with 451 additions and 161 deletions
38
README.md
38
README.md
|
|
@ -133,11 +133,25 @@ account-level deploy plumbing:
|
|||
only for an account that does not already have one — one provider per URL
|
||||
per account).
|
||||
|
||||
The template is a verbatim extraction of the substrate section of
|
||||
`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml` (see the provenance
|
||||
header in the template — mgmt's copy remains source of truth for 328440206208
|
||||
until its stacks migrate out; substrate changes while both are live must edit
|
||||
both files). Per-repo `githubdeploy-*` deploy roles are deliberately NOT part
|
||||
The template began as a verbatim extraction of the substrate section of
|
||||
`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml`, which remains the source
|
||||
of truth for mgmt (328440206208) until its stacks migrate out.
|
||||
|
||||
**The two copies are no longer at parity, and the old "edit both files" rule no
|
||||
longer applies uniformly.** Under INFRA-186, `seahaven-lambda-execution-boundary`
|
||||
in *this* copy was reduced to a fleet-wide floor for prod and dev (where
|
||||
boundary usage was 0, so no live Lambda could break): CloudWatch Logs write on
|
||||
`/aws/lambda*`, log-group describe, X-Ray, and ENI lifecycle — nothing else.
|
||||
Each migrating stack adds its own data-plane statements, derived from its own
|
||||
template, in its own PR (per-workload boundaries are the INFRA-187 end state).
|
||||
mgmt's copy keeps the account-wide wildcards pending its own separately
|
||||
validated rollout across 26 live boundary-carrying roles. So: **the boundary
|
||||
resource is deliberately divergent**; every *other* substrate resource
|
||||
(`github-cfn-execution-role`, `seahaven-cfn-exec-iam-management`) is still
|
||||
expected to change in both files together. The template's provenance header
|
||||
records which is which — read it before assuming either parity or divergence.
|
||||
|
||||
Per-repo `githubdeploy-*` deploy roles are deliberately NOT part
|
||||
of the substrate — they are provisioned per repo at migration/onboarding time
|
||||
so an account never carries trust relationships for repos that do not deploy
|
||||
to it.
|
||||
|
|
@ -258,14 +272,24 @@ split this substrate exists to enforce.
|
|||
`organization:seahaven:project:seahaven-<env>:workspace:<workspace>:run_phase:plan`
|
||||
(or `:apply`). Exact `StringEquals` only — never `StringLike`, never a
|
||||
wildcarded `run_phase` (a speculative PR plan must never hold write
|
||||
credentials). IAM roles = mandatory GPT-4.1 cross-review +
|
||||
credentials). **If the stack creates Lambda execution roles, this same PR
|
||||
must also widen `seahaven-lambda-execution-boundary`** per the WIDENING
|
||||
PATH in `lib/deploy-substrate/deploy-substrate.template.yaml`: the
|
||||
guardrail forces every Terraform-created role to carry that boundary, and
|
||||
it is a fleet-wide floor with zero data-plane permissions until widened —
|
||||
an unwidened migration deploys green, then every data-plane call is denied
|
||||
at first invoke and async/DLQ writes are discarded silently. IAM roles and
|
||||
boundary widenings = mandatory cross-family review +
|
||||
`/sh-security-review` on the diff.
|
||||
3. After deploy, verify: both roles exist; `hcptf-<stack>` lists
|
||||
`seahaven-hcptf-iam-management` in `list-attached-role-policies`; trust
|
||||
subs match the live org/project/workspace names byte-for-byte; simulate
|
||||
the apply role against a `hcptf-*` ARN (expect `explicitDeny` from
|
||||
`DenySelfMutation`) and against a normal stack role name (expect
|
||||
`allowed`).
|
||||
`allowed`); and if step 2 widened the boundary, confirm the deployed
|
||||
default version carries the stack's data-plane statements
|
||||
(`aws iam get-policy-version`) — role verification alone never checks
|
||||
boundary content.
|
||||
4. Set **workspace-level** variables `TFC_AWS_PLAN_ROLE_ARN` +
|
||||
`TFC_AWS_APPLY_ROLE_ARN` (category env) to the verified role ARNs, plus
|
||||
`TFC_AWS_PROVIDER_AUTH=true`. Never project-scoped variable sets — the
|
||||
|
|
|
|||
|
|
@ -39,16 +39,115 @@ Description: >-
|
|||
# added. The mgmt copy was remediated 2026-07-27 (.github PRs #95 Phase A
|
||||
# + #98 Phase B); DenySelfMutation and the widened policy/seahaven-*
|
||||
# DenyBoundaryPolicyEdit scope were then ported back here, so the two
|
||||
# copies' statement sets are reconciled as of that date — every IAM
|
||||
# statement in LambdaExecutionBoundary, SamCfnIamManagementPolicy and
|
||||
# SamCfnExecutionRole is byte-identical across the two files; the only
|
||||
# remaining delta is the DependsOn line above, which is ordering, not
|
||||
# permission. If a substrate statement changes again, change BOTH files
|
||||
# in the same piece of work.
|
||||
# copies' GUARDRAIL statement sets were reconciled as of that date.
|
||||
# SamCfnIamManagementPolicy and SamCfnExecutionRole remain at parity on
|
||||
# their IAM STATEMENT SETS across the two files and MUST still be changed
|
||||
# together. Parity covers statements, not surrounding comments — a comment
|
||||
# may diverge where it describes boundary content, which now differs
|
||||
# between the files. The only functional delta
|
||||
# between them is the DependsOn line above, which is ordering, not
|
||||
# permission. LambdaExecutionBoundary is NO LONGER byte-identical — see
|
||||
# DELIBERATE DIVERGENCE below.
|
||||
#
|
||||
# DELIBERATE DIVERGENCE — LambdaExecutionBoundary (INFRA-186, 2026-07-30)
|
||||
# The parity rule above is SCOPED, not global. LambdaExecutionBoundary in THIS
|
||||
# file is DELIBERATELY STRICTER than the mgmt copy in
|
||||
# Sea-Haven-Industries/.github/oidc-deploy-roles.yaml. Do not "reconcile" the two
|
||||
# by copying mgmt's statements back over these, or vice versa; the divergence is
|
||||
# load-bearing. A future mechanical drift check WILL read it as drift — it is not.
|
||||
#
|
||||
# 1. WHAT DIVERGED. Every per-workload data-plane statement was REMOVED from
|
||||
# LambdaExecutionBoundary in this file, leaving only the fleet-wide floor:
|
||||
# CloudWatchLogsWrite (scoped to /aws/lambda*), CloudWatchLogsDescribe,
|
||||
# XRay and Ec2Eni. The account-wide wildcards mgmt still carries — table/*,
|
||||
# table/*/index/*, secret:*, parameter/*, sqs :*, function:*, ses
|
||||
# identity/* + configuration-set/*, kms key/*, and an s3:::*-<accountid>
|
||||
# pattern that was a bare name-suffix filter rather than an ownership
|
||||
# check — are simply GONE here rather than re-scoped.
|
||||
# The security win is the deletion: it is what closes the amplifier whereby
|
||||
# a principal able to write an inline policy onto a boundary-carrying role
|
||||
# could read every secret in the account. Per-workload prefixes add no
|
||||
# security — they only keep a workload functional — so they are added by
|
||||
# each migration PR, from that stack's own template, when the stack
|
||||
# actually lands. See the note on the boundary resource for the full
|
||||
# rationale and the six errors that the pre-loaded approach produced.
|
||||
#
|
||||
# 2. WHY MGMT'S RATIONALE IS LEGITIMATE THERE. The superset framing this file
|
||||
# used to carry ("being slightly broad is the correct trade-off; a boundary
|
||||
# that is too tight will break Lambda functions at runtime AFTER deploy") is
|
||||
# a real constraint in the management account: 328440206208 has 26 LIVE
|
||||
# roles carrying seahaven-lambda-execution-boundary, across all five SAM
|
||||
# stacks. Tightening there is a production change to running workloads with
|
||||
# a silent, deploy-time-invisible failure mode.
|
||||
#
|
||||
# 3. WHY IT DOES NOT TRANSFER HERE. This file deploys ONLY to seahaven-prod
|
||||
# (011934824531) and seahaven-dev (710827005802), where
|
||||
# PermissionsBoundaryUsageCount is 0 and 0 respectively (aws iam get-policy,
|
||||
# verified 2026-07-30; corroborated by list-roles returning no role carrying
|
||||
# any permissions boundary in either account). No live Lambda can break, so
|
||||
# the risk that justifies mgmt's breadth is absent — while the exposure is
|
||||
# strictly WORSE here than in mgmt, because prod is multi-tenant: the old
|
||||
# wildcards reached proposal-system's, procurement-ingest's and
|
||||
# workorder-ingest's CDK-owned tables, buckets, secrets and queues, the org's
|
||||
# own Config and VPC-flow-log buckets, and — via function:* — CDK Lambdas
|
||||
# that carry no boundary at all.
|
||||
#
|
||||
# 4. RECONCILIATION OBLIGATION, RESTATED. For LambdaExecutionBoundary the two
|
||||
# copies are now INTENTIONALLY DIFFERENT and must NOT be synchronised:
|
||||
# - A change to the per-workload Resource patterns in THIS file does NOT
|
||||
# propagate to mgmt.
|
||||
# - A change to mgmt's boundary does NOT propagate here.
|
||||
# - Any change to the ACTION lists, or any new statement, is a substrate
|
||||
# semantic change and DOES still require the same review in both copies.
|
||||
# For SamCfnIamManagementPolicy and SamCfnExecutionRole the original rule is
|
||||
# unchanged: change BOTH files in the same piece of work.
|
||||
#
|
||||
# KNOWN OPEN ITEM (deferred, not closed by INFRA-186): mgmt 328440206208 still
|
||||
# carries the account-wide patterns. Tightening it needs its own validated
|
||||
# rollout — enumerate what the 26 live roles actually call, stage it, and be
|
||||
# ready to roll back — and is explicitly OUT OF SCOPE of INFRA-186. Until that
|
||||
# lands, the two copies stay divergent and that is the intended state.
|
||||
#
|
||||
# COUPLING: the boundary's ManagedPolicyName and ARN are UNCHANGED and must stay
|
||||
# unchanged. Four Conditions in SamCfnIamManagementPolicy below, and four more in
|
||||
# HcptfIamManagementPolicy in lib/terraform-substrate/terraform-substrate.template.yaml,
|
||||
# pin arn:aws:iam::<account>:policy/seahaven-lambda-execution-boundary by literal
|
||||
# string inside StringEquals iam:PermissionsBoundary. A rename fails SILENTLY — an
|
||||
# IAM condition naming a non-existent policy simply never matches, so the
|
||||
# escalation control would evaporate rather than error — and would additionally
|
||||
# force a CloudFormation REPLACEMENT that any role carrying the boundary would
|
||||
# block. INFRA-186 is a CONTENT-ONLY change for exactly this reason;
|
||||
# terraform-substrate.template.yaml already records that expectation and is
|
||||
# correctly left untouched.
|
||||
#
|
||||
# SIZE BUDGET: an attached managed policy document is capped at 6,144 characters
|
||||
# (whitespace excluded). LambdaExecutionBoundary measures 691 characters across
|
||||
# 4 statements as of 2026-07-31 — the fleet-wide floor only. Measure before
|
||||
# widening — len(json.dumps(doc,separators=(',',':'))) on the synthesized
|
||||
# PolicyDocument with ${AWS::AccountId} resolved, and UPDATE THESE TWO NUMBERS in
|
||||
# the same edit (they went stale twice inside this branch alone).
|
||||
#
|
||||
# Headroom is 5,453 characters, roughly TWELVE workloads at ~450 each. That is a
|
||||
# deliberate outcome, not luck: an earlier revision of this branch pre-loaded
|
||||
# per-workload prefixes for all five mgmt SAM stacks and reached 5,457 characters
|
||||
# with 687 left — about one workload of room — before any stack had actually
|
||||
# migrated. Deferring per-workload scope to each migration PR (see the note on
|
||||
# the boundary itself) removed that pressure entirely. If the budget tightens
|
||||
# again as workloads land, the end-state fix is per-workload boundaries
|
||||
# (seahaven-lambda-execution-boundary-<workload>), which also resolves the
|
||||
# shared-ceiling residual — tracked as INFRA-187, do not improvise it.
|
||||
# CRITICAL: unlike the 2026-07-27 inline-limit incident,
|
||||
# there is NO restructure available when this cap is reached — a role has exactly
|
||||
# ONE permissions boundary, so statements cannot be spilled into a second attached
|
||||
# managed policy. At the cap the only levers are prefix consolidation and dropping
|
||||
# unused actions.
|
||||
#
|
||||
# This template is deployed via lib/deploy-substrate-stack.ts
|
||||
# (cloudformation-include) as stack seahaven-deploy-substrate, once per
|
||||
# member account that hosts SAM workloads.
|
||||
# (cloudformation-include) as stack seahaven-deploy-substrate, once per member
|
||||
# account that hosts SAM workloads (currently seahaven-prod 011934824531 and
|
||||
# seahaven-dev 710827005802 via bin/app.ts instances deploy-substrate-prod /
|
||||
# deploy-substrate-dev; NEVER mgmt — 328440206208 is served by the .github copy
|
||||
# named above until its stacks migrate out).
|
||||
|
||||
Parameters:
|
||||
CreateOIDCProvider:
|
||||
|
|
@ -83,7 +182,7 @@ Resources:
|
|||
UpdateReplacePolicy: Retain
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Lambda execution permissions boundary (INFRA-103)
|
||||
# Lambda execution permissions boundary (INFRA-103, re-scoped by INFRA-186)
|
||||
#
|
||||
# This managed policy is the CEILING for every Lambda execution role that the
|
||||
# five SAM stacks auto-generate via AWS::Serverless::Function. Applying it as
|
||||
|
|
@ -91,49 +190,173 @@ Resources:
|
|||
# intersection of the role's own policies and this boundary, so a misconfigured
|
||||
# SAM role can never exceed what is listed here.
|
||||
#
|
||||
# The boundary is intentionally a SUPERSET of the union of all runtime
|
||||
# permissions currently granted across the five stacks. Being slightly broad
|
||||
# is the correct trade-off at this stage — a boundary that is too tight will
|
||||
# break Lambda functions at runtime after deploy, which is worse than a slightly
|
||||
# loose boundary that is tightened in a follow-up.
|
||||
# SCOPING RULE (INFRA-186, 2026-07-30). This boundary is the FLEET-WIDE FLOOR
|
||||
# ONLY: what every Lambda execution role needs regardless of workload. It
|
||||
# carries NO per-workload data-plane statements — each migrating stack adds its
|
||||
# own, from its own template, in its own PR (see the note on the boundary
|
||||
# resource and the WIDENING PATH below). It was previously a deliberate SUPERSET
|
||||
# with account-wide wildcards (table/*, secret:*, sqs :*, function:*,
|
||||
# parameter/*, and an s3:::*-<accountid> pattern that was a name-suffix filter,
|
||||
# not an ownership check). That trade-off was made when the only account
|
||||
# carrying this policy hosted nothing but the five SAM stacks. It does not
|
||||
# survive multi-tenancy: seahaven-prod now hosts CDK-owned tenants
|
||||
# (proposal-system, procurement-ingest, workorder-ingest) whose tables, buckets,
|
||||
# secrets and queues those wildcards reached, and whose Lambdas carry NO
|
||||
# permissions boundary at all — making function:* a boundary-escape primitive.
|
||||
#
|
||||
# Permission sources per stack:
|
||||
# Tightening here carries ZERO runtime risk and was sequenced deliberately:
|
||||
# PermissionsBoundaryUsageCount is 0 in BOTH accounts this template deploys to
|
||||
# (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified
|
||||
# 2026-07-30), so no live Lambda can break. A boundary that is slightly TOO
|
||||
# TIGHT is recoverable here — the migrating stack widens it in its own PR before
|
||||
# its first deploy — whereas leaving it loose perpetuates the exposure. Prefer
|
||||
# tighter; the widening path is below.
|
||||
#
|
||||
# afterhours-shift-manager
|
||||
# A resource pattern that genuinely CANNOT be scoped keeps its wildcard WITH a
|
||||
# written justification on the statement: CloudWatchLogsDescribe, XRay and
|
||||
# Ec2Eni name runtime-created resources or use actions AWS authorises against
|
||||
# "*" regardless of the ARN supplied. Do not "tighten" those.
|
||||
#
|
||||
# Every "verified <date>" annotation in this file is a POINT-IN-TIME
|
||||
# observation, not live state. Re-validate (usage counts, log-group CMK state,
|
||||
# per-stack permission sources) before citing one as justification for a
|
||||
# future change.
|
||||
#
|
||||
# WIDENING PATH — read this before migrating a stack into prod or dev.
|
||||
# The boundary is never widened by the person who hits the AccessDenied. It is
|
||||
# widened by the migrating stack's owner, in THIS repo, BEFORE the workload's
|
||||
# first deploy into the target account:
|
||||
# 0. PRECONDITION — check the managed-policy VERSION budget BEFORE merging:
|
||||
# max 5 versions, both
|
||||
# accounts are on v1 today. Every widening (PolicyDocument edit) burns
|
||||
# one. Description, ManagedPolicyName and Path are REPLACEMENT
|
||||
# properties per the CFN resource reference — CloudFormation cannot
|
||||
# replace a custom-named policy, so a Description-only edit FAILS the
|
||||
# stack update, and the error's suggested remedy (rename) is exactly
|
||||
# the forbidden rename in the COUPLING note above. Never edit those
|
||||
# three properties. Delete the oldest non-default version if at 5:
|
||||
# aws iam list-policy-versions --policy-arn \
|
||||
# arn:aws:iam::<acct>:policy/seahaven-lambda-execution-boundary
|
||||
# aws iam delete-policy-version --version-id v<oldest-non-default> ...
|
||||
# 1. Derive the workload's needs from ITS OWN TEMPLATE — open the stack's
|
||||
# template.yaml and read the actual IAM policy statements. The
|
||||
# permission-source block below is a STARTING POINT, NOT THE AUTHORITY:
|
||||
# the /sh-security-review pass on 2026-07-30 found SIX places where it was
|
||||
# incomplete or simply invented a resource name, three of which would have
|
||||
# failed silently. Update that block in the same edit with what you find.
|
||||
# 2. Add the workload's statements. Group by service so a second workload can
|
||||
# extend a Resource list rather than duplicate an action list (~250
|
||||
# characters for zero new actions; an extra ARN costs ~60). Check for the
|
||||
# SILENT classes specifically: a denied SQS destination/DLQ write discards
|
||||
# the async event with no error and no alarm; a denied scheduler call may
|
||||
# sit behind a bare except; a denied KMS decrypt for env-var encryption
|
||||
# fails at cold-start INIT; any CMK-encrypted resource needs the
|
||||
# matching kms:ViaService principal, not just the kms action; and a
|
||||
# function using LoggingConfig with a custom log-group name outside
|
||||
# /aws/lambda* silently loses ALL logs — add a scoped logs statement
|
||||
# for the custom group or keep the default group name.
|
||||
# 3. Measure. See SIZE BUDGET in the header. There is NO escape hatch.
|
||||
# 4. Both review gates run and neither discharges the other: the GPT-4.1
|
||||
# cross-family review against the real diff, and /sh-security-review
|
||||
# (IaC/IAM is on the mandatory surface). CLI down = review outstanding.
|
||||
# 5. Merge and let CI deploy deploy-substrate-prod / deploy-substrate-dev to
|
||||
# UPDATE_COMPLETE, THEN deploy the workload stack.
|
||||
# ORDERING IS NOT ENFORCED BY CLOUDFORMATION AND THIS IS THE MOST IMPORTANT
|
||||
# SENTENCE HERE: the workload's deploy SUCCEEDS even against a stale boundary,
|
||||
# because seahaven-cfn-exec-iam-management's gate checks that the boundary ARN
|
||||
# is attached, never its contents. The failure surfaces later, at first invoke,
|
||||
# as AccessDenied. A stale boundary is a silent deploy-time pass and a loud
|
||||
# production failure.
|
||||
#
|
||||
# CONSIDERED AND REJECTED: a Deny statement reserving the seahaven-* namespace.
|
||||
# With the Allow set reduced to the fleet-wide floor it is fully redundant (verified
|
||||
# 2026-07-30: seahaven-prod-config-* and seahaven-prod-vpc-flow-logs-* are
|
||||
# already denied by the Allow set alone), and a Deny inside a BOUNDARY is the
|
||||
# hardest failure mode in the estate to debug — it beats every Allow in every
|
||||
# policy with no synth-time signal. Revisit only if a widening ever has to
|
||||
# re-broaden a per-service Resource list back toward a wildcard.
|
||||
#
|
||||
# NOTE ON WHAT THESE PREFIXES ARE. All five stacks below currently live in the
|
||||
# MANAGEMENT account and none of their resources exists in seahaven-prod or
|
||||
# seahaven-dev yet. These are MIGRATION-CANDIDATE prefixes for the accounts this
|
||||
# template deploys to, not an inventory of what is deployed there. They are a
|
||||
# SECONDARY RECORD and a starting point for widening PRs — the authority is
|
||||
# each stack's own template (WIDENING PATH step 1). No Resource pattern in
|
||||
# the floor above derives from this block.
|
||||
#
|
||||
# Permission sources per stack (verified live 2026-07-30 IN MGMT — these
|
||||
# stacks and resources do not exist in prod/dev yet, so nothing below is a
|
||||
# prod/dev observation; starting point only — verify every entry against
|
||||
# the owning repo before use. PARAMETERIZED resources (ARNs passed as
|
||||
# deploy parameters) must be re-derived from the live stack configuration
|
||||
# at migration time, as exact ARNs — never inferred from these names into
|
||||
# broad patterns like secret:afi-*):
|
||||
#
|
||||
# afterhours-shift-manager (functions: afterhours-*, 6 live)
|
||||
# - DynamoDB CRUD (afterhours-shifts table)
|
||||
# - secretsmanager:GetSecretValue (afterhours-shift-manager/*)
|
||||
# - ses:SendEmail (SES identity)
|
||||
# - ses:SendEmail on the identity AND on
|
||||
# configuration-set/seahaven-email-events (template.yaml:178-181 — the
|
||||
# send is DENIED without the config-set ARN when the identity has a
|
||||
# default configuration set)
|
||||
# - scheduler:CreateSchedule/DeleteSchedule/GetSchedule on
|
||||
# schedule/default/holiday-* + iam:PassRole to scheduler.amazonaws.com
|
||||
# (template.yaml:110-120). CORRECTED 2026-07-30: this block previously
|
||||
# omitted both, and the omission is SILENT at runtime (bare except).
|
||||
# - NO ssm. CORRECTED 2026-07-30: this block previously credited
|
||||
# ssm:GetParameter to this stack; `grep -c 'ssm:' template.yaml` = 0.
|
||||
# Its slack tokens come from Secrets Manager and the channel id from a
|
||||
# CloudFormation parameter.
|
||||
# - lambda:InvokeFunction (ReleaseNotifyInvokeRole,
|
||||
# HolidaySchedulerExecutionRole — these two carry the boundary and need
|
||||
# ONLY this action)
|
||||
# - CloudWatch Logs (all functions)
|
||||
# - UNRESOLVED: /3cx-scheduler/* ownership (this stack vs. the retired
|
||||
# standalone 3CX scheduler). Deliberately NOT granted. Resolve at
|
||||
# migration in that stack's own repo — do NOT add any 3cx-scheduler
|
||||
# resource here until ownership is resolved.
|
||||
#
|
||||
# payments-dashboard
|
||||
# - DynamoDB CRUD / Read (PaymentsDashboard table)
|
||||
# - S3 GetObject (payroll-emails, payments-csv buckets)
|
||||
# payments-dashboard (functions: payments-*)
|
||||
# - DynamoDB CRUD / Read (PaymentsDashboard table — legacy PascalCase)
|
||||
# - S3 GetObject on seahaven-payments-csv-* and seahaven-payroll-emails-*;
|
||||
# GetObject + PutObject on seahaven-payments-boa-raw-*
|
||||
# (template.yaml:272 and 1097-1099, fetchBoaTransactions raw archive).
|
||||
# CORRECTED 2026-07-30: this block previously said "GetObject ONLY —
|
||||
# no write intent enumerated", which was false and would have denied
|
||||
# the raw-archive write at migration.
|
||||
# - secretsmanager:GetSecretValue (payments-dashboard/*)
|
||||
# - sqs:SendMessage + sqs:ReceiveMessage + sqs:DeleteMessage etc.
|
||||
# (PayrollBatchQueue + DLQs)
|
||||
# - lambda:InvokeFunction (ExpenseReceiver → ExpenseProcessor)
|
||||
# - ec2:CreateNetworkInterface / DescribeNetworkInterfaces /
|
||||
# DeleteNetworkInterface (VPC-attached functions)
|
||||
# - sqs Send/Receive/Delete etc. (payments-payroll-batch + DLQs)
|
||||
# - lambda:InvokeFunction (ExpenseReceiver -> ExpenseProcessor)
|
||||
# - ec2 ENI lifecycle (VPC-attached functions)
|
||||
# - KMS via dynamodb (table CMK) — no SSM, no SES
|
||||
# - CloudWatch Logs
|
||||
#
|
||||
# meal-order-manager
|
||||
# meal-order-manager (functions: meal-order-manager-*, 7 live)
|
||||
# - DynamoDB CRUD / Read (meal-order-manager-orders table)
|
||||
# - S3 CRUD (ReportsBucket) + s3:GetObject (ReportsBucket presigned URLs)
|
||||
# - S3 CRUD (meal-order-manager-reports-*, meal-order-manager-form-*)
|
||||
# - secretsmanager:GetSecretValue (meal-order-manager/*)
|
||||
# - ssm:GetParameter (/meal-order-manager/*)
|
||||
# - lambda:InvokeFunction (submit-order → slack-notifier,
|
||||
# close-form → aggregate-orders)
|
||||
# - lambda:InvokeFunction (submit-order -> slack-notifier,
|
||||
# close-form -> aggregate-orders, plus AdminAuthorizerInvokeRole)
|
||||
# - ses:SendRawEmail
|
||||
# - CloudWatch Logs
|
||||
#
|
||||
# front-integrations
|
||||
# front-integrations (functions: front-*)
|
||||
# - DynamoDB CRUD (front-sla-alerts table)
|
||||
# - secretsmanager:GetSecretValue (by ARN, various)
|
||||
# - secretsmanager:GetSecretValue (front-integrations/*)
|
||||
# - CloudWatch Logs
|
||||
# - no IAM permissions for S3 / SQS / SSM / SES / KMS / VPC in this
|
||||
# stack's template as of 2026-07-30
|
||||
#
|
||||
# afi-backup-monitor
|
||||
# - secretsmanager:GetSecretValue (by ARN)
|
||||
# afi-backup-monitor (functions: afi-*)
|
||||
# - secretsmanager:GetSecretValue on TWO bare, unprefixed secrets:
|
||||
# afi-api-key and afi-slack-webhook. CORRECTED 2026-07-30: this block
|
||||
# previously named "afi-backup-monitor/slack-webhook-url", which does
|
||||
# not exist. Both ARNs are deploy PARAMETERS in that stack
|
||||
# (AfiApiKeySecretArn / SlackWebhookSecretArn), so no name is
|
||||
# discoverable from its template — the names are in its README:48-49.
|
||||
# - CloudWatch Logs
|
||||
# - nothing else
|
||||
#
|
||||
# ---------------------------------------------------------------------------
|
||||
LambdaExecutionBoundary:
|
||||
|
|
@ -149,18 +372,76 @@ Resources:
|
|||
Version: "2012-10-17"
|
||||
Statement:
|
||||
|
||||
# ── CloudWatch Logs (every Lambda) ──────────────────────────────────
|
||||
- Sid: CloudWatchLogs
|
||||
# ── CloudWatch Logs — write (every Lambda) ──────────────────────────
|
||||
# Scoped to the Lambda log-group namespace. Every SAM function's group
|
||||
# is /aws/lambda/<function>, and the trailing * also covers the
|
||||
# :log-stream:<name> suffix PutLogEvents authorises against, so one ARN
|
||||
# serves CreateLogGroup, CreateLogStream, PutLogEvents and
|
||||
# DescribeLogStreams. The * is deliberately NOT after a trailing slash:
|
||||
# /aws/lambda* also matches the /aws/lambda-insights groups the Lambda
|
||||
# Insights extension writes to, which /aws/lambda/* would have denied.
|
||||
# VERIFICATION PROVENANCE, stated precisely (2026-07-30). What
|
||||
# iam simulate-custom-policy DOES confirm: this pattern allows
|
||||
# logs:CreateLogGroup / logs:PutLogEvents on the bare group ARN
|
||||
# log-group:/aws/lambda/<fn> (and /aws/lambda/<a>/<b>), and DENIES
|
||||
# log-group:seahaven-prod-vpc-flow-logs — the latter re-checked against
|
||||
# an Allow */* positive control, which allows it, so the deny is real
|
||||
# policy behaviour and not a simulator artifact.
|
||||
# What the simulator CANNOT evaluate, so do NOT claim it was verified:
|
||||
# log-stream-qualified ARNs (log-group:<g>:log-stream:<s>) and the bare
|
||||
# /aws/lambda-insights group both return implicitDeny EVEN UNDER an
|
||||
# Allow */* policy. That is a simulator resource-parsing limitation, not
|
||||
# a denial. Coverage of those two rests on documented IAM wildcard
|
||||
# semantics — "*" matches any sequence of characters including ":" and
|
||||
# "/" — which is why the * is deliberately NOT placed after a trailing
|
||||
# slash. If this ever needs true end-to-end proof, it must come from a
|
||||
# real invoke in dev, not from the simulator.
|
||||
#
|
||||
# NOT scoped per workload, deliberately. A per-stack prefix
|
||||
# (/aws/lambda/payments-* etc.) was considered and rejected: a Lambda
|
||||
# denied PutLogEvents does not fail — it keeps running and silently
|
||||
# produces no logs. Log denial is the one failure class in this policy
|
||||
# that is NOT loud, so it must not depend on function-name discipline.
|
||||
# ACCEPTED RESIDUAL RISK: a SAM Lambda can write into another tenant's
|
||||
# /aws/lambda/* group (log poisoning). No read action is granted here, so
|
||||
# this is not an exfiltration path. Tighten only once every
|
||||
# boundary-carrying function is confirmed to set an explicit FunctionName.
|
||||
# REGION IS PINNED TO us-east-1 DELIBERATELY: every Sea Haven workload
|
||||
# deploys to us-east-1, and this template itself only ever deploys there.
|
||||
# ${AWS::Region} would resolve to the identical string, so it would
|
||||
# document nothing. A future stack in another region carries this
|
||||
# boundary but CANNOT write its logs (the silent class above) — so a
|
||||
# cross-region migration MUST add region-scoped statements in its
|
||||
# widening PR, same as any other data-plane need.
|
||||
- Sid: CloudWatchLogsWrite
|
||||
Effect: Allow
|
||||
Action:
|
||||
- logs:CreateLogGroup
|
||||
- logs:CreateLogStream
|
||||
- logs:PutLogEvents
|
||||
- logs:DescribeLogGroups
|
||||
- logs:DescribeLogStreams
|
||||
Resource:
|
||||
- !Sub "arn:aws:logs:us-east-1:${AWS::AccountId}:log-group:/aws/lambda*"
|
||||
|
||||
# ── CloudWatch Logs — describe (UNSCOPABLE, kept "*" deliberately) ──
|
||||
# logs:DescribeLogGroups is a COLLECTION action: AWS authorises it
|
||||
# against "*" regardless of any resource ARN supplied. Scoping it would
|
||||
# produce a policy that reads tighter and denies at runtime, so it is
|
||||
# split into its own statement and keeps the wildcard. Read-only
|
||||
# metadata; it cannot mutate anything or return log content.
|
||||
- Sid: CloudWatchLogsDescribe
|
||||
Effect: Allow
|
||||
Action:
|
||||
- logs:DescribeLogGroups
|
||||
Resource: "*"
|
||||
|
||||
# ── X-Ray tracing (standard Lambda execution) ────────────────────
|
||||
# ── X-Ray tracing (UNSCOPABLE, kept "*" deliberately) ───────────────
|
||||
# xray:PutTraceSegments / PutTelemetryRecords support no resource-level
|
||||
# permissions — X-Ray exposes no ARN for them, which is why the AWS
|
||||
# managed AWSXRayDaemonWriteAccess also uses "*". Any ARN written here
|
||||
# would be inert and would falsely imply a control exists. Write-only
|
||||
# into this account's own trace store; no cross-tenant read is
|
||||
# expressible with this action set.
|
||||
- Sid: XRay
|
||||
Effect: Allow
|
||||
Action:
|
||||
|
|
@ -168,10 +449,33 @@ Resources:
|
|||
- xray:PutTelemetryRecords
|
||||
Resource: "*"
|
||||
|
||||
# ── VPC / ENI management (payments-dashboard VPC functions) ────────
|
||||
# Matches AWSLambdaVPCAccessExecutionRole exactly.
|
||||
# AssignPrivateIpAddresses / UnassignPrivateIpAddresses are for EFA
|
||||
# and secondary IPs — not part of the Lambda ENI lifecycle — omitted.
|
||||
# ── VPC / ENI management (UNSCOPABLE, kept "*" deliberately) ────────
|
||||
# Derived from AWSLambdaVPCAccessExecutionRole, NOT an exact match:
|
||||
# DescribeSecurityGroups and DescribeVpcs exceed that managed policy
|
||||
# (kept for CFN/SAM VpcConfig validation; read-only). The four
|
||||
# ec2:Describe* actions do not support resource-level
|
||||
# permissions AT ALL — an ARN in Resource is ignored and the call is
|
||||
# authorised against "*" — so narrowing them is cosmetic. The ENI in
|
||||
# CreateNetworkInterface / DeleteNetworkInterface is created by the
|
||||
# Lambda service at attach time with an id that cannot exist when this
|
||||
# policy is written. Nothing here is scopable by resource name.
|
||||
# AssignPrivateIpAddresses / UnassignPrivateIpAddresses are for EFA and
|
||||
# secondary IPs — not part of the Lambda ENI lifecycle — omitted.
|
||||
#
|
||||
# KNOWN OPEN ITEM (pre-existing, NOT introduced by INFRA-186; tracked
|
||||
# as INFRA-200):
|
||||
# ec2:DeleteNetworkInterface on "*" lets a bounded Lambda delete any ENI
|
||||
# in the account, including NAT / VPC-endpoint / RDS ENIs — a
|
||||
# denial-of-service primitive inherited from the AWS managed policy. The
|
||||
# durable fix is a Condition on ec2:Subnet / ec2:Vpc naming THE SET OF
|
||||
# VPCs that boundary-carrying Lambdas attach to — not a single VPC id;
|
||||
# the list must be extended whenever a workload introduces a new VPC
|
||||
# (tag-based conditions are the alternative if the set churns). No such
|
||||
# VPC exists in seahaven-prod or seahaven-dev today (payments-dashboard's
|
||||
# 10.20.0.0/16 VPC is in mgmt), so writing the condition now would encode
|
||||
# an mgmt resource into a prod/dev template. Whoever brings the first VPC
|
||||
# across in payments-dashboard's migration PR adds the condition in the
|
||||
# same PR.
|
||||
- Sid: Ec2Eni
|
||||
Effect: Allow
|
||||
Action:
|
||||
|
|
@ -183,115 +487,70 @@ Resources:
|
|||
- ec2:DescribeVpcs
|
||||
Resource: "*"
|
||||
|
||||
# ── DynamoDB (afterhours, payments, meal-order, front-integrations) ─
|
||||
# Table/* covers base-table operations; table/*/index/* is required for
|
||||
# Query/Scan on Global Secondary Indexes.
|
||||
- Sid: DynamoDB
|
||||
Effect: Allow
|
||||
Action:
|
||||
- dynamodb:GetItem
|
||||
- dynamodb:PutItem
|
||||
- dynamodb:UpdateItem
|
||||
- dynamodb:DeleteItem
|
||||
- dynamodb:Query
|
||||
- dynamodb:Scan
|
||||
- dynamodb:BatchGetItem
|
||||
- dynamodb:BatchWriteItem
|
||||
- dynamodb:DescribeTable
|
||||
- dynamodb:ConditionCheckItem
|
||||
Resource:
|
||||
- !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/*"
|
||||
- !Sub "arn:aws:dynamodb:us-east-1:${AWS::AccountId}:table/*/index/*"
|
||||
|
||||
# ── S3 (payments-dashboard read, meal-order-manager CRUD) ──────────
|
||||
- Sid: S3
|
||||
Effect: Allow
|
||||
Action:
|
||||
- s3:GetObject
|
||||
- s3:PutObject
|
||||
- s3:DeleteObject
|
||||
- s3:ListBucket
|
||||
- s3:GetBucketLocation
|
||||
- s3:GetObjectVersion
|
||||
- s3:GetObjectTagging
|
||||
- s3:PutObjectTagging
|
||||
Resource:
|
||||
- !Sub "arn:aws:s3:::*-${AWS::AccountId}"
|
||||
- !Sub "arn:aws:s3:::*-${AWS::AccountId}/*"
|
||||
# meal-order-manager ReportsBucket (non-AccountId suffix pattern)
|
||||
- !Sub "arn:aws:s3:::meal-order-manager-reports-${AWS::AccountId}"
|
||||
- !Sub "arn:aws:s3:::meal-order-manager-reports-${AWS::AccountId}/*"
|
||||
- !Sub "arn:aws:s3:::meal-order-manager-form-${AWS::AccountId}"
|
||||
- !Sub "arn:aws:s3:::meal-order-manager-form-${AWS::AccountId}/*"
|
||||
|
||||
# ── Secrets Manager (all stacks) ──────────────────────────────────
|
||||
- Sid: SecretsManager
|
||||
Effect: Allow
|
||||
Action:
|
||||
- secretsmanager:GetSecretValue
|
||||
- secretsmanager:DescribeSecret
|
||||
Resource:
|
||||
- !Sub "arn:aws:secretsmanager:us-east-1:${AWS::AccountId}:secret:*"
|
||||
|
||||
# ── SSM Parameter Store (meal-order-manager, afterhours) ──────────
|
||||
- Sid: SSMParameterRead
|
||||
Effect: Allow
|
||||
Action:
|
||||
- ssm:GetParameter
|
||||
- ssm:GetParameters
|
||||
- ssm:GetParametersByPath
|
||||
Resource:
|
||||
- !Sub "arn:aws:ssm:us-east-1:${AWS::AccountId}:parameter/*"
|
||||
|
||||
# ── SQS (payments-dashboard batch queues) ─────────────────────────
|
||||
- Sid: SQS
|
||||
Effect: Allow
|
||||
Action:
|
||||
- sqs:SendMessage
|
||||
- sqs:ReceiveMessage
|
||||
- sqs:DeleteMessage
|
||||
- sqs:GetQueueAttributes
|
||||
- sqs:GetQueueUrl
|
||||
- sqs:ChangeMessageVisibility
|
||||
Resource:
|
||||
- !Sub "arn:aws:sqs:us-east-1:${AWS::AccountId}:*"
|
||||
|
||||
# ── Lambda invocation (payments, meal-order inter-function calls) ──
|
||||
- Sid: LambdaInvoke
|
||||
Effect: Allow
|
||||
Action:
|
||||
- lambda:InvokeFunction
|
||||
Resource:
|
||||
- !Sub "arn:aws:lambda:us-east-1:${AWS::AccountId}:function:*"
|
||||
|
||||
# ── SES (afterhours weekly-post, meal-order email-report) ──────────
|
||||
- Sid: SES
|
||||
Effect: Allow
|
||||
Action:
|
||||
- ses:SendEmail
|
||||
- ses:SendRawEmail
|
||||
Resource:
|
||||
- !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:identity/*"
|
||||
- !Sub "arn:aws:ses:us-east-1:${AWS::AccountId}:configuration-set/*"
|
||||
|
||||
# ── KMS (CMK-encrypted resources) ─────────────────────────────────
|
||||
# Required for Lambda functions that read/write CMK-encrypted AWS
|
||||
# resources. Verified live state:
|
||||
# - PaymentsDashboard DynamoDB table: CMK key/0b660af3 (KMS:ENABLED)
|
||||
# - payments-dashboard CloudWatch log groups: CMK key/b748750c
|
||||
# Secrets Manager + SQS queues in these stacks use AWS-managed keys
|
||||
# (aws/secretsmanager, aws/sqs) which do not require explicit kms:*
|
||||
# actions in the execution role policy. The CMK keys are scoped to
|
||||
# this account to prevent cross-account KMS calls.
|
||||
- Sid: KMS
|
||||
Effect: Allow
|
||||
Action:
|
||||
- kms:Decrypt
|
||||
- kms:GenerateDataKey
|
||||
- kms:DescribeKey
|
||||
Resource:
|
||||
- !Sub "arn:aws:kms:us-east-1:${AWS::AccountId}:key/*"
|
||||
|
||||
# ── NO PER-WORKLOAD DATA-PLANE STATEMENTS — BY DESIGN ───────────────
|
||||
# The statements above are the FLEET-WIDE FLOOR: what every Lambda
|
||||
# execution role needs regardless of which workload it belongs to.
|
||||
# There are deliberately NO DynamoDB, S3, Secrets Manager, SSM, SQS,
|
||||
# SES, KMS, lambda:InvokeFunction or scheduler statements here.
|
||||
#
|
||||
# WHY (decided 2026-07-30, Adam):
|
||||
# The security win of INFRA-186 comes from DELETION, not enumeration.
|
||||
# Removing the account-wide secret:*, table/*, function:* and sqs:*
|
||||
# wildcards is what closes the amplifier — the ability of a principal
|
||||
# who can write an inline policy onto a boundary-carrying role to read
|
||||
# every secret in the account. Per-workload prefixes add no security;
|
||||
# they exist only to keep a workload FUNCTIONAL once it arrives.
|
||||
#
|
||||
# An earlier revision of this branch PRE-LOADED prefixes for all five
|
||||
# mgmt SAM stacks before any of them had migrated. That required
|
||||
# predicting five stacks' permission needs from the permission-source
|
||||
# comment block above, and the /sh-security-review pass found SIX
|
||||
# errors in the result — three of which would have failed SILENTLY at
|
||||
# first migration (afterhours' SES config-set, its holiday scheduler
|
||||
# behind a bare except, and afi's webhook secret under an invented
|
||||
# name). The block is a secondary record and is not a substitute for
|
||||
# reading the owning repo's template.
|
||||
#
|
||||
# WIDENING IS THE SAFE DIRECTION. Adding a resource to a boundary can
|
||||
# never break a running Lambda; only tightening can. So there is no
|
||||
# cost to deferring per-workload scope to the migration PR that has
|
||||
# the real template open in front of it — and a large cost to
|
||||
# guessing it years ahead of the migration.
|
||||
#
|
||||
# CONSEQUENCE FOR EVERY MIGRATION PR (mandatory, see WIDENING PATH in
|
||||
# the header): a stack landing in prod or dev MUST add its own
|
||||
# data-plane statements here, derived from ITS OWN template, in its
|
||||
# own PR to THIS repo, deployed to UPDATE_COMPLETE before the
|
||||
# workload's first deploy from its own repo (the workload PR and the
|
||||
# widening PR cannot be the same PR — they live in different repos).
|
||||
# Without them its Lambdas get AccessDenied
|
||||
# at first invoke. The permission-source block above is the starting
|
||||
# point, NOT the authority — verify every entry against the stack.
|
||||
#
|
||||
# Prod/dev boundary usage is 0 (verified 2026-07-30), so this floor
|
||||
# currently constrains nothing that exists. Fleet-wide statements that
|
||||
# genuinely cannot be scoped (Logs, X-Ray, ENI) stay above with their
|
||||
# justifications; they are also the SILENT-failure classes, which is
|
||||
# why they belong in the floor rather than in per-workload widenings.
|
||||
#
|
||||
# KMS is absent deliberately: zero boundary-carrying roles exist, so
|
||||
# nothing bounded decrypts anything yet. (Log-group CMKs are the one
|
||||
# KMS case that never hits an execution role — CloudWatch Logs
|
||||
# decrypts via the key policy's grant to the Logs service principal —
|
||||
# so log-group CMK state is not the gate here. seahaven-prod already
|
||||
# hosts the seahaven-dynamodb-cmk table key; the first migrating
|
||||
# stack touching that table must cover it.) A workload bringing a
|
||||
# CMK-encrypted resource adds a KMS statement with the matching
|
||||
# kms:ViaService principal in its own migration PR — a missing
|
||||
# ViaService entry denies, and for env-var encryption it fails at
|
||||
# cold-start INIT.
|
||||
#
|
||||
# The end-state fix for the shared-ceiling residual (one boundary =
|
||||
# every SAM workload reaches every other's data plane once they land)
|
||||
# is per-workload boundaries — tracked as INFRA-187. Do not improvise
|
||||
# it: the load-bearing problem there is that both guardrail policies
|
||||
# pin ONE literal boundary ARN inside StringEquals conditions, and
|
||||
# loosening that to a wildcard weakens the gate.
|
||||
# ---------------------------------------------------------------------------
|
||||
# Shared CloudFormation execution role (SAM stacks) — INFRA-97 scoped
|
||||
#
|
||||
|
|
@ -789,8 +1048,9 @@ Resources:
|
|||
# conditioned on iam:PermissionsBoundary StringEquals the
|
||||
# seahaven-lambda-execution-boundary ARN. That condition means
|
||||
# any role this execution role creates must have the boundary
|
||||
# applied, so it can never exceed what the boundary allows
|
||||
# (which is scoped to the services the five stacks actually use).
|
||||
# applied, so it can never exceed what the boundary allows (in
|
||||
# THIS file the fleet-wide floor plus per-migration widenings;
|
||||
# in mgmt's copy the five SAM stacks' service wildcards).
|
||||
#
|
||||
# iam:PassRole is also included here so CloudFormation can pass
|
||||
# the auto-generated Lambda execution role to the Lambda service.
|
||||
|
|
|
|||
|
|
@ -65,8 +65,14 @@ Description: >-
|
|||
# deploy-substrate stack instead. The coupling is by NAME: if the boundary
|
||||
# policy is ever renamed or replaced, every Condition below (and the
|
||||
# deploy-substrate copy) must change in the same piece of work. INFRA-186
|
||||
# (per-workload boundary scoping) changes the boundary's CONTENT, not its ARN,
|
||||
# and does not touch this file.
|
||||
# (boundary reduced to a fleet-wide floor; per-workload boundaries are
|
||||
# INFRA-187) changed the boundary's CONTENT, not its ARN, so this file is
|
||||
# textually untouched — but the Terraform path IS affected: the Conditions
|
||||
# below FORCE every role a Terraform apply creates onto that boundary, and
|
||||
# the floor carries zero data-plane permissions. A migrating stack that
|
||||
# creates Lambda execution roles must widen the boundary per the WIDENING
|
||||
# PATH in lib/deploy-substrate/deploy-substrate.template.yaml, deployed
|
||||
# before its first apply (README migration checklist step 2).
|
||||
#
|
||||
# SIZE BUDGET: an attached managed policy document is capped at 6,144
|
||||
# characters (whitespace excluded). The statement set below is ~2.5 KB.
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue