Cross-family round 1 plus the /sh-security-review verifier confirmed 11
findings on the floor reduction, all documentation defects; no policy
statement changes. The one HIGH: the Terraform migration checklist never
widened the boundary, so a Lambda-bearing Terraform migration would deploy
green and lose every data-plane call at first invoke. Checklist step 2 now
carries the widening requirement, step 3 verifies deployed boundary content,
and the terraform-substrate header no longer reads as 'Terraform path
unaffected'. Also corrected: Description is a REPLACEMENT property (a
Description edit wedges the custom-named policy and CFN's remedy is the
forbidden rename), the sanctioned-source contradiction, the false
AWSLambdaVPCAccessExecutionRole parity claim, the KMS log-group category
error, stale size numbers (691/5,453), the same-PR widening contradiction,
per-workload residue text, a LoggingConfig silent-log-loss note, the
us-east-1 region pin rationale, and ENI DoS deferral now tracked as
INFRA-200.
Replaces the placeholder 'tracked as its own ticket' references with the real
key, and records the load-bearing constraint inline so the next reader does not
rediscover it: both guardrail policies pin ONE literal boundary ARN inside
StringEquals iam:PermissionsBoundary, and loosening that to a wildcard weakens
the gate rather than merely relaxing it.
Adam's call after review: the security win of INFRA-186 comes from DELETING the
account-wide wildcards, not from enumerating replacements. Per-workload prefixes
add no security -- they only keep a workload functional -- and widening a
boundary is the safe direction (adding a resource never breaks a running Lambda;
only tightening does). So the per-workload scope moves to each migration PR,
which has the stack's real template open in front of it.
Removed all nine per-workload data-plane statements (DynamoDB, S3 x3, Secrets
Manager, SSM, SQS, Lambda invoke, SES, scheduler x2, KMS). Kept the fleet-wide
floor: CloudWatchLogsWrite (/aws/lambda*), CloudWatchLogsDescribe, XRay, Ec2Eni
-- the statements every Lambda needs regardless of workload, and also the
silent-failure classes, which is why they belong in the floor.
KMS dropped entirely: both accounts have ZERO CMK-encrypted log groups
(verified). A workload bringing a CMK adds the statement plus the matching
kms:ViaService principal in its own PR.
Why not keep the enumeration: it required predicting five stacks' needs from
this file's own permission-source comment block, and /sh-security-review found
SIX errors in the result -- three silent. The block is a secondary record, not
an authority. Deriving scope per-migration from the owning template removes the
whole error class.
Effect on the security objective: unchanged. secret:*, table/*, function:*,
sqs:* and the s3:::*-<acct> name-suffix filter are gone either way, so the
amplifier is closed identically.
Size: 5,457 chars / 16 statements -> 703 / 4. Headroom 687 -> 5,441, so the cap
stops being a forcing function. Header, SCOPING RULE and WIDENING PATH all
updated to match; widening path now leads with 'read the stack's own template',
names the silent-failure classes to check, and moves the version-budget check to
a precondition instead of a trailing step.
Verified unchanged: logical id and ManagedPolicyName, so all eight pinning
conditions across both guardrail policies still resolve. Both accounts synth
identically at 703 chars.
/sh-security-review (6 detectors + verifier) found four HIGH findings, all the
same defect class: the template's permission-source comment block was used as
the sanctioned scope source, but it is an incomplete and in places invented
secondary record. Each was verified against the real stack template before
fixing. None is live today (prod/dev boundary usage is 0); all four would have
been AccessDenied at first migration, three of them SILENTLY.
- SES configuration-set/seahaven-email-events restored. afterhours-shift-manager
template.yaml:178-181 grants it with an in-repo comment stating the send is
denied without it. An earlier revision dropped it after checking whether any
config set exists in prod/dev today (none do) -- the wrong test. The right
question is whether an enumerated stack's own IAM policy names it.
- SES identity/seahavenind.com added. meal-order-manager's SenderEmail defaults
to adam@seahavenind.com (template.yaml:20-22) and email_report sends with it.
The prior 'unverified identity fails loudly anyway' argument holds only until
the migration verifies the domain, which the migration procedure requires.
- scheduler:Create/Delete/GetSchedule + iam:PassRole (scheduler.amazonaws.com
only) added. afterhours template.yaml:110-120 needs both; the block omitted
them entirely. Failure is silent -- app.py wraps create_schedule in a bare
except, so the Slack command reports success and no schedule exists.
- secret:afi-slack-webhook-* added. The block named
'afi-backup-monitor/slack-webhook-url', which does not exist; both afi secret
ARNs are deploy parameters, so the real names live only in that repo's
README:48-49 (afi-api-key, afi-slack-webhook).
Also corrected, all comment-only:
- The permission-source block itself, at each of the four points it was wrong,
with the correction and its evidence recorded inline.
- The false claim that SAM auto-names async DLQs (it does not -- all four
payments queues are hand-written with explicit QueueNames). Replaced with the
real invariant: any queue a boundary-carrying function sends to must be
payments-* or the boundary widens in the same PR; a denied destination write
is silent.
- SIZE BUDGET: was 13 statements / 4,060 chars, actually 16 / 5,457 after these
fixes. Headroom is 687 chars, roughly ONE more workload -- not the five the
header claimed. Flagged per-workload boundaries as the realistic next move.
Verified unchanged: logical id LambdaExecutionBoundary and ManagedPolicyName
seahaven-lambda-execution-boundary, so all eight pinning conditions across both
guardrail policies still resolve.
Post-implementation verification of the INFRA-186 prod/dev scoping found two
functional defects that would have denied permissions the migrating stacks
actually need. Neither is live today (prod/dev boundary usage is 0), but both
would have surfaced as AccessDenied at first migration.
- KMS: the ViaService list omitted ssm., while SSMParameterRead in the same
policy grants ssm:GetParameter*. A SecureString read decrypts via the SSM
service principal, so the boundary denied reads it also granted.
- S3: payments-dashboard was classified read-only from the template's own
permission-source comment, but that enumeration is incomplete -- the real
stack grants s3:PutObject on BoaRawBucket (template.yaml:272, 1098-1099).
Write is now allowed on seahaven-payments-boa-raw-* only; payroll-emails and
payments-csv stay read-only, preserving the evidence-deletion protection.
The seahaven-payments-* wildcard is replaced by the three literal bucket
names, verified against payments-dashboard/template.yaml.
Not changed: SES configuration-set/*. Review claimed dropping it rested on a
false premise; verified live -- prod and dev both have ZERO configuration sets
and member-baseline-stack.ts:44 excludes SES monitoring. The drop is correct.
README: the 'substrate changes must edit both files' rule is now false for the
boundary specifically, and said so uniformly. Corrected to distinguish the
deliberately divergent boundary from the still-at-parity substrate resources.
Re-scope seahaven-lambda-execution-boundary in the prod/dev copy of
deploy-substrate.template.yaml from account-wide wildcards to per-workload
resource prefixes drawn from the template's own permission-source block.
This is the PROD/DEV HALF of INFRA-186.
What was scoped (wildcard -> per-workload prefix):
- dynamodb table/* + table/*/index/* -> afterhours-*, front-*,
meal-order-manager-*, PaymentsDashboard*, payments-dashboard-*
(a trailing * after each prefix also covers the /index/* GSI ARNs, so the
separate table/*/index/* entry is deleted rather than replaced)
- s3 *-${AccountId} -> meal-order-manager-*-${AccountId}
(read/write) and seahaven-payments-* / seahaven-payroll-emails-*
(read-only). The removed pattern was not an ownership check at all: S3 ARNs
carry no account field, so it was a bare name-suffix filter that matched 8
of 9 buckets in prod -- including the org's own Config and VPC-flow-log
buckets -- with PutObject and DeleteObject.
- secretsmanager secret:* -> five <stack>/ prefixes + the legacy
bare afi-api-key-*. The wildcard reached workorder-ingest's HMAC signing
key, i.e. a webhook-forgery primitive.
- ssm parameter/* -> afterhours-shift-manager and
meal-order-manager, each as both the bare path ARN and /* (GetParametersByPath
authorises against the path, not the leaf)
- sqs :* -> payments-*
- lambda function:* -> afterhours-*, meal-order-manager-*,
payments-*. Highest-leverage fix here: an invoked function runs under its
OWN role, and every non-SAM function in prod is CDK-deployed with no
boundary, so function:* was a boundary-escape primitive, not just lateral
movement.
- ses identity/* + configuration-set/* -> the two verified prod
identities; configuration-set dropped (zero exist)
- logs split into a scoped write half (/aws/lambda*) and a wildcard
describe half (DescribeLogGroups is a collection action AWS authorises
against "*" regardless of the ARN supplied)
Deliberately NOT tightened, each with written justification on the statement:
CloudWatchLogsDescribe, XRay and Ec2Eni name runtime-created resources or use
actions that support no resource-level permissions. KMS keeps key/* -- key ARNs
carry UUID key ids, not workload names -- and is constrained by a kms:ViaService
condition instead, which inherits the per-workload scoping of the services
above for free.
No runtime risk. PermissionsBoundaryUsageCount is 0 in BOTH accounts this file
deploys to (seahaven-prod 011934824531 and seahaven-dev 710827005802, verified
2026-07-30 via aws iam get-policy), so no live Lambda can break. Adam scoped the
handoff to prod/dev for exactly this reason. Since usage is 0, a boundary that
is slightly too tight is recoverable -- the migrating stack widens it in its own
PR before its first deploy -- whereas leaving it loose perpetuates the exposure.
The widening path and its ordering hazard are documented in the template.
mgmt is DELIBERATELY UNTOUCHED and the two copies are now DIVERGENT. The
management account (328440206208) uses a separate copy in
Sea-Haven-Industries/.github/oidc-deploy-roles.yaml and has 26 LIVE
boundary-carrying roles, where tightening is a production change with a silent,
deploy-time-invisible failure mode; it needs its own validated rollout and is
explicitly out of scope. The header's parity rule is therefore now SCOPED, not
global: SamCfnIamManagementPolicy and SamCfnExecutionRole stay byte-identical
and must still change together, while LambdaExecutionBoundary must NOT be
reconciled in either direction. A DELIBERATE DIVERGENCE block records this so a
future mechanical drift check does not "fix" it away, following the same pattern
terraform-substrate.template.yaml uses for its divergences.
Content-only change: ManagedPolicyName, the policy ARN and the logical id
LambdaExecutionBoundary are unchanged. Eight StringEquals iam:PermissionsBoundary
conditions across this file and terraform-substrate.template.yaml pin the
boundary by literal name, and a rename fails SILENTLY -- an IAM condition naming
a non-existent policy simply never matches.
Verification:
- npx tsc --noEmit: clean
- npx cdk synth deploy-substrate-prod deploy-substrate-dev: succeeds
- synthesized resource diff vs main: LambdaExecutionBoundary is the ONLY
changed resource; GitHubOIDCProvider, SamCfnExecutionRole and
SamCfnIamManagementPolicy are byte-identical
- policy document 4,060 chars / 6,144 cap (2,084 headroom), 13 statements,
identical in both accounts
- iam simulate-custom-policy against live prod, every deny re-checked against
an Allow */* positive control: 11/11 cross-tenant denies are real (Config
and flow-log buckets, proposal-system-uploads, proposal-system/db-credentials,
workorder-ingest/shoc-webhook-hmac, proposal-system-api, proposal-system-jobs,
WorkOrders, /seahaven/dynamodb/cmk-arn, the flow-log group, seahavenind.com)
and 23/23 enumerated workload resources still allow
Checkov suppressions re-keyed: CKV_AWS_111 still fires on the boundary because
three statements legitimately retain Resource:"*", so the suppression is still
required. All three line-keyed ids shifted (139->291, 329->713, 805->1189); new
ids added, superseded ids retained, and the boundary justification's stale "OPEN
follow-up: tighten to per-workload prefixes" sentence rewritten to CLOSED since
this commit is what closes it. Scanners: RESULT PASS.
Refs: INFRA-186
Security review (6 detectors + proof-or-kill verifier) confirmed 1 critical and
1 high in the first revision, both inherited by mirroring the SAM copy's
Resource "*" role grants:
- C1 (critical): iam:UpdateAssumeRolePolicy on "*" with DenySelfMutation
covering only three name patterns lets the principal repoint the
AdministratorAccess CDK bootstrap role's trust policy to an external account.
- C2 (high): the SAM justification for role/* (SAM auto-roles land at path /
with no settable RolePath) does not transfer -- Terraform's aws_iam_role
supports path.
Fixes, closing the class at the root rather than by denylist:
- All role writes, boundary sets and PassRole confined to role/tf-managed/*;
reads split into a separate statement that keeps Resource "*".
- DenySelfMutation extended to cdk-hnb659fds-*, OrganizationAccountAccessRole
and seahaven-* as defense in depth.
- OIDC provider made conditional (CreateOIDCProvider), mirroring the sibling
substrate, so a first-create rollback is recoverable rather than wedging the
stack in ROLLBACK_COMPLETE against a Retained orphan.
- README corrected: the guardrail policy is NOT Retain (only the provider is),
so the Deny backstops do not survive a stack delete.
checkov CKV_AWS_109 no longer fires on this template, so no suppression is
needed. The template header records every divergence from the SAM copy.
New stack seahaven-terraform-substrate (instances terraform-substrate-prod +
terraform-substrate-dev): app.terraform.io OIDC provider and the shared
boundary-gated guardrail policy seahaven-hcptf-iam-management that
per-workspace Terraform apply roles attach at migration time. No roles are
pre-provisioned (accumulator pattern, parallel to githubdeploy-*).
Guardrail statements mirror seahaven-cfn-exec-iam-management byte-identically
except DenySelfMutation, whose scope extends to hcptf-* alongside the
GitHub-substrate principals. Explicit stack dependency on the same-account
deploy-substrate stack (boundary ARN appears only in Condition strings, so
CFN infers no edge).
Review of the DenySelfMutation port found the header's 'reconciled' claim
was not yet true: mgmt Phase A also added the CloudWatch Logs
metric-filter actions (afterhours-shift-manager creates an
AWS::Logs::MetricFilter through this role), and without them a migrating
SAM stack fails mid-deploy with AccessDenied. Ports those three actions
and corrects two stale header notes. Every IAM statement in the three
shared resources is now byte-identical across both files, verified
programmatically; the only delta left is the DependsOn ordering line.
The seahaven-cfn-exec-iam-management policy in prod and dev carried only
DenyBoundaryTampering + DenyBoundaryPolicyEdit: the mgmt Phase A review
later showed a Deny-in-a-managed-policy control is self-detachable
(iam:DetachRolePolicy on * is unconditioned), so without DenySelfMutation
the exec role can detach the very policy carrying the Denies and
reinstate the boundary-removal escalation. Latent today (no PassRole
grants, zero SAM stacks in prod/dev) but must be closed before the first
SAM workload migrates.
Ports verbatim from .github/oidc-deploy-roles.yaml (mgmt, PRs #95/#98):
- DenySelfMutation over role/github-cfn-execution-role + githubdeploy-*
- DenyBoundaryPolicyEdit widened to policy/seahaven-*
Statement set verified byte-identical to the mgmt copy (9 sids);
provenance header updated - the two copies are reconciled.
The first deploy of seahaven-deploy-substrate failed in both prod and dev
with ServiceLimitExceeded: 'Maximum policy size of 10240 bytes exceeded
for role github-cfn-execution-role'. The role's inline policies already
sat ~94 bytes under IAM's hard 10,240-byte per-role limit, so the two
Deny statements added to close the boundary-removal escalation did not
fit (10,656 total).
Moves the whole boundary-gated IAM block (6 Allow + 2 Deny statements)
into an attached managed policy, which carries its own separate
6,144-byte budget. Inline drops to 8,285 with ~1.9 KB of headroom;
the managed policy sits at 2,371.
Effective permissions are unchanged: the union of role statements
(inline + attached) is byte-identical as a sorted set before and after
the move (27 statements both sides), identity policies are unioned, and
an explicit Deny still wins. Boundary and trust policy untouched.
Both failed stacks rolled back cleanly with zero orphaned resources and
were deleted before this retry.
SAM repos migrating off the frozen management account need the shared
deploy plumbing (permissions boundary + github-cfn-execution-role) in
their target account; none of it existed outside mgmt, so there was no
OIDC SAM deploy path into seahaven-prod or seahaven-dev at all.
Adds a templated, per-account substrate stack so onboarding a future
account is one bin/app.ts instance plus one CD job, not a hand-rolled
copy. Per-repo githubdeploy-* roles stay out by design: they are
provisioned per repo at migration time so an account never accumulates
trust for repos that do not deploy to it.
The template is a verbatim extraction of the reviewed mgmt substrate,
with deliberate, documented divergences — notably the removal of
iam:DeleteRolePermissionsBoundary plus explicit Deny backstops, which
closes a confirmed privilege-escalation path (see PR body).
AWS Chatbot's control-plane API is homed in us-east-2, so every chatbot call
carries aws:RequestedRegion=us-east-2 and is denied by the workloads-region-lock
region deny (approved set = us-east-1/us-west-2). This blocked Slack
workspace/channel setup in seahaven-prod (chatbot:GetSlackOauthParameters
denied), which the prod site-alerts topic needs for Slack delivery. Adds
chatbot:* to the SCP's global-service NotAction exemption list alongside
iam/organizations/cloudfront/route53 — a region-agnostic full-prefix exemption,
the same shape as the other global services. targetIds unchanged (workloads OU);
regional services (s3/kms/logs) and the Bedrock carve-out untouched.
GPT-4.1 cross-review: SAFE TO MERGE. /sh-security-review: block=false (0 confirmed
critical/high). Security-OU region-lock deliberately NOT changed (runs no such
workloads, same asymmetry as its missing Bedrock carve-out).
* feat(prod): add seahaven-prod DynamoDB CMK and site-alerts alarm-topic stacks
Provisions the two shared dependencies procurement-ingest imports by name,
ahead of its migration from mgmt to seahaven-prod:
- dynamodb-cmk-prod: second DynamoDbCmkStack instance (same stack name,
prod account) creating alias/seahaven-dynamodb + the
/seahaven/dynamodb/cmk-arn SSM param. Adds a cross-account key-policy
statement so the mgmt seahaven-slack-bot roles can keep reading the
CMK-encrypted purchase-orders table after it moves (ViaService +
PrincipalArn-wildcard scoped; identity-policy half lands in the
slack-bot repo's cutover PR).
- alarm-topic-prod: codified site-alerts SNS topic + seahaven-alarm-topics
CMK with the cloudwatch.amazonaws.com publish grant (mirrors the working
mgmt pattern; mgmt's topic remains CLI-managed debt).
- deploy.yaml: both appended to the deploy-prod job's explicit stack list
(SH-ORG-005 rule: unlisted stacks silently never deploy).
* fix(scripts): account-id assertion in cfn-stack-decommission; complete the aws-cdk-lib 2.262.0 bump (patched brace-expansion); document CMK cutover trap
- cfn-stack-decommission.sh: --account-id is now REQUIRED and asserted
against sts get-caller-identity before anything runs. Stack names are no
longer org-unique (seahaven-dynamodb-cmk now exists in mgmt AND prod), so
a name-only lookup under the wrong ambient profile could report or delete
the wrong account's stack (security-review LOGIC-001).
- package.json/lock: PR #56's bump-for-patched-brace-expansion landed the
commit title but not the pin; package.json still said 2.261.0 and the
lockfile still resolved brace-expansion 5.0.6 (GHSA-3jxr-9vmj-r5cp HIGH,
blocking the pre-commit scanner). Pin 2.262.0 and regenerate; npm audit
now clean.
- bin/app.ts comments: slack-bot cutover MUST grant the PROD key ARN, never
the account-local mgmt SSM param (LOGIC-005); failed-first-create orphan
CMK recovery note (LOGIC-004).
* refactor(prod): drop cross-account CMK grant (slack-bot decommissioned 2026-07-23)
The AllowMgmtSlackBotReadViaDynamoDb key-policy statement targeted the
seahaven-slack-bot roles, which were decommissioned 2026-07-23. Its successor
sh-mcp is undeployed and uses same-account DynamoDB access, so no cross-account
reader of the CMK-encrypted purchase-orders table exists. The prod CMK + SSM
param + alarm-topic stacks remain (procurement-ingest still imports them). Add
a scoped cross-account grant if/when a real cross-account consumer deploys.
* docs: update aws profile specified in script (local renaming)
* ci: add least-privilege permissions blocks to workflow callers
Resolves code scanning alerts #3 and #4 (actions/missing-workflow-permissions). Both callable workflows only need contents: read; the dependency-review callable already declares it internally, this caps the caller token to match."
* chore(deps): bump aws-cdk-lib to 2.262.0 for patched brace-expansion
Resolves Dependabot alert #4 (CVE-2026-13149, exponential-time DoS in brace-expansion expand()). The vulnerable 5.0.6 is a bundled dependency inside the aws-cdk-lib tarball, so it cannot be updated independently; 2.262.0 bundles the patched 5.0.7.
Also migrates Stack#addDependency to addStackDependency (deprecated in this release) in bin/app.ts.
* fix(scp): carve out Bedrock InvokeModel/Converse to us-east-2 for cross-region inference
Add bedrock:InvokeModel, bedrock:InvokeModelWithResponseStream,
bedrock:Converse, and bedrock:ConverseStream to the existing
DenyRegionsOutsideApproved NotAction list so the us-east-1/us-west-2
region condition no longer denies them. Add a companion
DenyBedrockInvokeOutsideInference statement that re-denies those same
four actions outside {us-east-1, us-west-2, us-east-2}, bounding the
carve-out to us-east-2 only.
Without this, Anthropic cross-region inference profiles (us.anthropic.*)
that route InvokeModel to us-east-2 are denied, blocking all Claude
generation in workload accounts.
Refs: #53
* fix: add ACCEPTED RISK disposition, hoist Bedrock actions to shared const, mark security-asymmetry
- ACCEPTED RISK: Bedrock carve-out is resource-unscoped (NotAction
can't be resource-scoped); us-east-2 window admits four actions
against any Bedrock resource. Per-account IAM and model-access
enablement gate actual access.
- Hoist the four Bedrock invoke actions into BEDROCK_INVOKE_ACTIONS
shared const referenced by both NotAction and DenyBedrockInvoke
statements to prevent future one-sided edit divergence.
- Mark asymmetry in security-guardrails DenyRegionsOutsideApproved:
no Bedrock carve-out by design — security account runs no Bedrock
workloads.
---------
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Extdev root credentials were deleted 2026-07-14 via centralized root
management (four-surface verified), closing the deferred root-hardening
blocker. Root recovery is central (assume-root, drill-proven) plus a
temporary gated detach, so the OU now gets the same root lockout as the
other six.
Gates: GPT-4.1 cross-review APPROVE; /sh-security-review PASS (0 confirmed
critical/high). Note: this attach puts the extdev OU at the 5-SCP hard
quota - future guardrails attach at the account or consolidate.
* Add seahaven-prod member baseline (Phase 5)
Account 011934824531 is the target for all new production stacks; the
management account is frozen for new workloads. First proven exercise
of the automatic enrollment sweep (Enabled in 124s, no manual
create-members) and of AutoEnableStandards=NONE (no pre-enabled
standards, so CFN owns FSBP + CIS v3.0 cleanly). Default VPC deleted;
budget starts at $100 and resizes as tenants land.
* Apply Phase-5 review findings
Fleet gap closed: EBS encryption-by-default + IAM password policy were
management-account-only (the runbook's unscoped 'applied' claim hid
it); now applied and verified in all three member accounts, runbook
scoped per account. README stack inventory corrected (eleven stacks,
org-governance rows restored). Sweep comments reconciled: the
automatic enrollment sweep is proven (seahaven-prod, ~2min).
* Add seahaven-dev member baseline with org-managed detection
Account 710827005802 (internal dev/staging) is the first account born
after delegation: GuardDuty/Security Hub enroll it via the org admin,
so DetectiveControls gains a localDetectiveServices flag (default true
— zero diff on the three deployed consumers, verified) and the dev
instance sets orgManagedDetection to skip the colliding local
detector/hub/analyzer. Default VPC kept and flow-logged (dev runs real
workloads). Enrollment verified Enabled in both services before this
commit.
* Fix Phase-4 review findings: standards + analyzer stay CFN-owned
SH-DEV-001: org AutoEnableStandards DEFAULT gave dev legacy CIS v1.2.0
and nothing owned CIS v3.0 — org config set to NONE, standards are now
unconditional in DetectiveControls (attach fine to an org-enabled hub),
legacy ruleset disabled in dev. SH-DEVBASE-002: the ORGANIZATION
analyzer treats the whole org as trusted so it cannot flag intra-org
exposure — account analyzer restored unconditionally (coexistence
verified live). Enrollment comments corrected: manual create-members,
the automatic sweep is still unexercised. Zero diff re-verified on all
three deployed baseline stacks.
* Add seahaven-security member baseline (Phase 3)
Account 001520130573 is the org's delegated security administrator.
Same member-baseline construct set as external-dev; own CD job under
its own OIDC role. Created at org root pending manual root hardening
before the OU move (deny-root-user invariant).
* Document delegated security administration runbook
Delegation to seahaven-security has no CloudFormation types; the CLI
sequence is the record, same pattern as the other account toggles.
* Apply Phase-3 security-review findings
Delegation runbook marked PENDING with hard preconditions (baseline
deployed, root MFA verified, account inside the security OU) — it had
read as applied before execution, the org's known claimed-done-but-NOT
failure mode (SEC-BASE-A/B). New security-guardrails SCP on the
security OU: region lock, IAM user/key lockout, privileged-role
protection, delegated-admin membership protection (SEC-BASE-C,
cross-reviewed APPROVE). deploy-security gains stack-name pre-flight
(SEC-BASE-D). Default VPC in 001520130573 deleted; empty flow-log list
and aws@ alert routing documented as deliberate (SEC-BASE-F/H).
Budget alerts and CIS alarm subscriptions now go to aws@seahaven.com
(management) and aws-external-dev@seahaven.com (external-dev) instead
of personal addresses (Adam, 2026-07-14; resolves security-review flag
SH-ORG-007). Owner tags are informational and stay decoupled.
cdk import by Id (non-mutating), content byte-exact from
describe-policy, targetIds = exact live attachments. Post-import drift
detection: IN_SYNC, 0 drifted. All four Retain — the full org guardrail
set is now drift-checked IaC.
* Add org-governance stack: OU skeleton + generalized SCPs
Phase 2 of the multi-account segregation plan: codifies the OU tree
(workloads/prod/nonprod, security, sandbox, graveyard) and three
org-wide SCPs (workloads-region-lock, protect-security-baseline,
deny-root-user) generalized from the proven external-dev guardrails.
All resources Retain — CFN must never detach a live guardrail. New SCPs
attach only to the new empty OUs; extending to external-dev is a
separate gated targetIds change after live verification.
* Record SCP cross-review dispositions in org-governance
Root hardening must precede the OU move (deny-root-user blocks root MFA
enrollment), delegated-admin flows ride service-linked roles that SCPs
never evaluate, and the cdk exec-role exemption is accepted risk
mirroring the external-dev guardrails.
* Add org-governance to the management deploy job
Explicit stack selectors require every new stack to join exactly one
CD job (SH-ORG-005 discipline documented in this file).
* Parameterize baseline constructs for multi-account reuse
DetectiveControls, FlowLogs, and GovernanceToggles were forked into
seahaven-external-dev-baseline with only physical-name and VPC-sourcing
differences. Prefix/name props let one implementation serve both
accounts; synthesized templates are unchanged (verified: empty cdk diff
against all deployed stacks).
* Absorb external-dev member baseline stack
Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack,
construct ids and physical names byte-identical to the deployed stack
(logical IDs are path-derived; empty cdk diff verified via change set
against 396287094661). Retires the forked repo so member-account
baselines share one drift surface and one dependency pin.
* Rename package to seahaven-org-baseline
Prepares the repo rename: the app now spans the management account and
org member accounts, so 'account-baseline' undersells the scope. README
documents the two-account deploy topology and logical-ID constraints.
* Commit extdev flow-log VPC ids in code, not -c context
Security review SH-ORG-004 (confirmed high): with the ids sourced from
ephemeral cdk context, any context-less deploy silently removes every
flow log in the isolated account. A committed list makes the attachment
set reviewable and immune to a forgotten -c flag. Empty list matches
the deployed stack (zero diff).
* Split CD into per-account deploy jobs
The app now spans two AWS accounts; cdk deploy --all under one role
fails on the other account's stacks (security review IAC-01). Each job
passes explicit stack selectors and its own account's OIDC role via the
new cd-cdk stacks input.
Switch UnauthorizedApiCalls alarm from 3/3 consecutive to 3/6 M-of-N
so a single quiet 5-min window can't reset detection. The 3/3 setting
pre-dates #36 and was sized to suppress CFN/Config noise that #36 now
removes at the filter level, making a wider M-of-N evaluation window
safe from flap risk.
Enable CloudTrail Insights (ApiCallRateInsight + ApiErrorRateInsight)
on seahaven-org-trail as a compensating control for the residual risk
accepted in #36 — the CFN/Config-proxied denials intentionally excluded
from CIS 4.1 — and as a backstop for low-and-slow patterns the 5-min
alarm may miss. Cost ≈$35–$53/month at current org trail volume.
Refs: #37
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
* Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm
CIS 4.1 (cis-UnauthorizedAPICalls) flapped OK<->ALARM 15 times in 30 days,
all from benign AWS-service AccessDenied noise (CloudFormation deploy/drift
describe-scans, AWS Config recorder). A single CFN run on 2026-07-07 emitted
100+ such denials in 15 min, tripping the alarm and burying the real CIS 4.1
security signal in email noise (alert fatigue).
- Group both error codes so the exclusions apply to the whole filter (the old
pattern leaked the UnauthorizedOperation branch past the exclusions due to
&& binding tighter than ||).
- Exclude denials whose sourceIPAddress is an AWS service host (*.amazonaws.com)
— AWS acting on our behalf, not a principal of concern. Real unauthorized
calls from a console/CLI/attacker present a routable IP and are still counted.
Validated against the trail log group: spike window 107 -> 4 matches, the 4
remaining all from a routable admin IP (genuine activity CIS should retain).
- Keep the 3/3 evaluation as a backstop against one-off human fat-fingers.
Billing: deleted the manually-created AWS-MonthlyBilling CloudWatch alarm
($50 threshold on EstimatedCharges, routed to site-alerts). It was unmanaged
drift, permanently in ALARM, and fully redundant with the managed M-10 budget
(seahaven-monthly-cost). README updated with rationale + restore command.
* Address sh-security-review: scope CIS 4.1 exclusion to named benign sources
The high-recall security review (detector fan-out + proof-or-kill verifier)
confirmed a MEDIUM detection blind spot in the first revision: excluding all
`*.amazonaws.com` source hosts would hide denials driven through ANY AWS
service (SSM Automation, Step Functions, Lambda, etc.), which CloudTrail
records with that service's host as sourceIPAddress — i.e. service-proxied
privesc/recon attempts would evade CIS 4.1.
Remediation: scope the exclusion to the specific benign sources that actually
flap this account — `*cloudformation.amazonaws.com` (covers both
cloudformation. and hooks.cloudformation.) and `config.amazonaws.com` — plus
the pre-existing delivery.logs exclusion. Every other service-proxied denial
is now retained. Residual (accepted, documented inline): CloudFormation/Config-
proxied denials are still excluded — that path needs near-admin privilege
(CreateStack + PassRole), successful changes still trip the other CIS 4.x
alarms, and GuardDuty backstops.
Validated on the live trail log group: spike window still 107 -> 4 matches
(identical noise suppression), the 4 from a routable admin IP. tsc + synth clean.
The cis-UnauthorizedAPICalls alarm fired on a single breaching 5-min period
(1/1), so any CloudFormation/CDK deploy burst of benign describe-API denials or
a one-off console fat-finger paged and self-recovered, producing notification
storms (~17 flips on 2026-06-10). Require 3 consecutive breaching periods so
only sustained unauthorized activity alarms; CIS detection of a real persistent
problem is preserved (detection window up to ~15 min). Per-control override so
the other 14 CIS alarms keep their 1/1 sensitivity (templates untouched).
GPT-4.1 cross-review: no blockers (documentation FIX noted re: detection latency).
Previously the L2 cloudtrail.Trail auto-created a log group with a
CDK-generated hash suffix in its name. The 15 CIS Section 4 metric
filters in CisMonitoring imported that group by the hardcoded generated
name. If the Trail or log group was ever recreated the suffix changes
and all 15 filters would silently detach with no error, leaving the
account unmonitored.
Replace with an explicit logs.LogGroup named
seahaven-account-baseline-trail-logs (stable, no hash suffix) with
RemovalPolicy.RETAIN. Pass the CDK object — not a name constant — to
the Trail via cloudWatchLogGroup and forward it to CisMonitoring via a
new trailLogGroup prop. All 15 filters now reference the CDK object so
they can never drift from the group the Trail actually delivers to.
The old auto-named log group is orphaned by this deploy (CloudFormation
loses track of it and does not delete it). Historical audit logs in the
old group remain accessible in CloudWatch under the old name; no audit
history is destroyed.
Refs: INFRA-19
The L1 AWS::Config::ConfigurationRecorder deadlocks the CDK stack
(recorder can't complete without a delivery channel; channel can't be
created without a recorder — observed 2026-06-01).
Fix: three AwsCustomResource nodes call PutConfigurationRecorder →
PutDeliveryChannel → StartConfigurationRecorder in sequence. Put* is
an idempotent upsert, so the deploy adopts the existing CLI-created
recorder and channel without destroying or interrupting them. onDelete
stops recording rather than deleting the per-account singleton.
New IAM permissions on the custom-resource role (cross-reviewed,
GPT-4.1 APPROVE — no BLOCK):
config:PutConfigurationRecorder
config:PutDeliveryChannel
config:StartConfigurationRecorder
config:StopConfigurationRecorder
iam:PassRole → seahaven-config-recorder-role (service=config)
cdk diff shows [+] adds only — no existing resources destroyed or
replaced. Removes README note that recorder/channel are CLI-only.
Refs: INFRA-17
* [INFRA-96] CMK-encrypt sensitive CloudWatch log groups (M-24)
Add a dedicated customer-managed CMK (alias/seahaven-logs) for encrypting
the sensitive CloudWatch Logs groups (CloudTrail + finance/PII Lambdas).
- lib/logs-key.ts: LogsKey construct. Key policy grants the CloudWatch Logs
service principal (logs.us-east-1.amazonaws.com) Encrypt*/Decrypt*/
ReEncrypt*/GenerateDataKey*/DescribeKey, scoped by the
kms:EncryptionContext:aws:logs:arn condition (REQUIRED per AWS docs or log
delivery breaks). Cross-reviewed (GPT-4.1): tightened Describe* -> DescribeKey;
CreateGrant omitted (not needed for plain log-group encryption).
- account-baseline-stack.ts: instantiate LogsKey and set KmsKeyId on the L2
Trail's CloudWatch log group in place (escape hatch on the existing
AWS::Logs::LogGroup) so it keeps the same logical id + physical name -
additive, no replacement, CIS Section-4 metric filters (which import the
group by name) keep working, live audit trail not disrupted. Gated by
context `encryptTrailLogGroup` so the CMK can be smoke-tested on a low-risk
Lambda group before the most-sensitive CloudTrail group.
Finance/PII Lambda log groups (exec-aide-*, payments-*, po-email-processor,
vendor-reply-processor) are owned by other stacks and associated to this CMK
via the CLI for now; codifying KmsKeyId in those repos is tracked as drift.
* [INFRA-96] Document sensitive-logs CMK (M-24) in README
Reintroduce the scoped vault access policy that was split out of INFRA-89
after two lockout-class bugs. Adds a Deny on the destructive recovery-point
and vault-lifecycle actions (DeleteRecoveryPoint, UpdateRecoveryPointLifecycle,
DeleteBackupVault, DeleteBackupVaultAccessPolicy,
DeleteBackupVaultLockConfiguration, PutBackupVaultLockConfiguration) for every
principal except three exempted operational identities via StringNotLike on
aws:PrincipalArn:
1. SSO AdministratorAccess role (break-glass human admin)
2. seahaven-backup-service-role (AWS Backup lifecycle)
3. cdk-hnb659fds-cfn-exec-role-* (CloudFormation manages the vault)
The CFN-exec-role exemption is the fix for the 2026-06-08 strand failure: without
it CloudFormation cannot re-assert the vault lock config and the deploy strands
the policy. Uses Deny + AnyPrincipal + StringNotLike (not NotPrincipal, which
rejects wildcard ARNs). aws:PrincipalArn normalizes assumed-role sessions to the
IAM role ARN, so the iam::role/ ARN forms are correct (AWS docs: "Do not specify
the assumed role session ARN as a value for this condition key").
Deployed and verified: deploy succeeded (proves exec role not locked out),
access policy present with all three exemptions, vault still Locked
(min1/max2555, LockDate null, 168 RPs), follow-up cdk diff clean (no drift).
* Codify primary vault lock + add backups (INFRA-89, INFRA-88)
INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the
seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no
changeableFor = admin-removable) so it lives in IaC. Values match the
live lock exactly, so the deploy is a no-op adoption.
Add a scoped vault access policy that denies manual recovery-point
deletion and lock/policy tampering to all principals except the AWS
Backup service role and the break-glass SSO AdministratorAccess role,
so automatic lifecycle expiry still works but humans cannot prune
recovery points by hand.
Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard
ARN matching, so the SSO exemption is expressed as Effect DENY with
Principal * and a StringNotLike condition on aws:PrincipalArn, which
does support wildcards. This avoids an unrecoverable vault lockout.
INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po,
extracted-amazon-po, proposal-system uploads + generated) to the
phase2-offsite-everything selection. Versioning verified enabled on
all 6 against the live account (S3 backup requires versioning).
Refs: INFRA-89, INFRA-88
* Promote account trail to organization trail (INFRA-73)
INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId
(o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket
PutObject statement for member-account delivery. CloudTrail org
trusted-access is already enabled on the management account.
Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey
statement for member-account trail delivery, guarded by
aws:PrincipalOrgID. The existing single-account statements are
preserved so management-account delivery is unaffected.
Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption
context must be wildcarded across accounts (org-trail shadow trails
present the member account id), not pinned to the management account,
or member delivery silently fails. Fixed before checkpoint.
CHECKPOINT: delicate org-trail KMS/bucket-policy change — code +
diff captured for review, NOT deployed.
Refs: INFRA-73
* Add secondary-region baseline stacks (INFRA-91, INFRA-16)
INFRA-91: codify the Bedrock model-invocation logging applied
out-of-band in us-west-2 and us-east-2 (per-region delivery role
seahaven-bedrock-invocation-logging-<region> + log group
/aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level
logging config itself has no CFN resource type and is applied via CLI
(already live), same as us-east-1.
INFRA-16: add the still-missing us-east-2 detective controls — AWS
Config recorder role + delivery bucket (recorder/channel via CLI to
avoid the CFN stabilization deadlock seen in us-east-1) and Security
Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in
us-east-2 and are left for a follow-up adoption to keep this change
non-destructive.
The us-east-1 baseline stays region-pinned; these are separate
RegionalBaselineStack instances composed opt-in per region.
CHECKPOINT: new multi-region stacks. The live Bedrock role + log group
already exist (CLI-created), so a plain deploy would collide — these
need cdk import / changeset adoption, not cdk deploy. Code + diff
captured for review, NOT deployed.
Refs: INFRA-91, INFRA-16
* Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
Audit L-14, plus a latent Day-2 bug: seahaven-cis-alarms was
encrypted with the AWS-managed alias/aws/sns key, whose policy cannot
grant cloudwatch.amazonaws.com - CloudWatch alarms silently fail to
publish to topics it encrypts. All 15 CIS alarms would have fired
into the void.
New customer-managed key (rotation on) grants CloudWatch
GenerateDataKey*/Decrypt/DescribeKey scoped by SourceAccount. The
unmanaged site-alerts topic now uses the same key (set via CLI).
Cross-reviewed: no BLOCKs. Verified: forced ALARM on the payroll DLQ
alarm published successfully through the encrypted site-alerts.
Audit finding H-20: no audit trail of model I/O for seahaven-alex,
which returns payments, invoices, WO/PO, and HR/SA8000 data.
S3 bucket (Glacier at 90d, expire 365d) + CloudWatch log group (90d)
+ delivery role assumable only by bedrock.amazonaws.com scoped by
SourceAccount/SourceArn. The account-level logging configuration has
no CloudFormation resource type, so it is applied via CLI post-deploy
(documented in the construct header) - same pattern as the Config
recorder (INFRA-17).
Cross-reviewed: no BLOCKs. Verified live: converse invocation logged
to /aws/bedrock/model-invocations.
database-1 (audit H-19) and the LedgerFlow stack (incl. ledgerflow-pos) were
decommissioned 2026-06-03. Drop database-1 from the critical-data selection and
ledgerflow-pos from phase2-offsite-everything so daily jobs don't target missing
resources. database-1's final recovery point is retained encrypted in the
seahaven-offsite vault (7yr); ledgerflow-pos has a final on-demand DynamoDB
backup. Deployed before merge (seahaven-backup UPDATE_COMPLETE).
Add a second BackupSelection 'phase2-offsite-everything' on the existing
seahaven-critical-daily plan covering the 15 remaining DynamoDB tables and
all 9 in-use EBS volumes, with the same daily backup + cross-region copy to
the GOVERNANCE-locked seahaven-offsite vault ('offsite for everything').
Reuses seahaven-backup-service-role (AWSBackupServiceRolePolicyForBackup
already grants DDB/RDS/EBS) - no IAM change. Explicit-ARN (not tag-based) to
avoid drifting the standalone file-share volumes and stack-owned tables, same
as phase-1. The 4 deprecated ledgerflow delete-targets are excluded.
Cross-reviewed (no BLOCK). Follow-up: migrate EBS to tag-based selection with
tags codified in owning stacks for resilience to volume replacement.
Deployed to seahaven-backup before merge; selection verified live (24 resources).
seahaven-app-waf (CLOUDFRONT scope, us-east-1): AWS managed Common + Known Bad
Inputs rule groups + per-IP rate limit (2000/5min). ARN published to SSM
/seahaven/waf/app-web-acl-arn for app stacks (meal-order/orders) to consume.
seahaven.com already has its own WAF; ledgerflow is being decommissioned
(INFRA-26); proposal-system-web skipped (not live).
* Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10)
Adds to the seahaven-account-baseline stack:
- AWS Config recorder (all + global resources) + delivery channel + role +
hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed.
- GuardDuty detector, us-east-1 (H-3)
- Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4)
- IAM Access Analyzer, account scope (M-5)
- Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to
adam@seahavenind.com (M-10)
Scope us-east-1 only (all workloads here); multi-region is a follow-up.
The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented
separately in the README runbook.
* Document Day 1 detective layer + CLI governance toggles in README
* Move Config recorder+channel to CLI (L1 stabilization deadlock)
The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches
CREATE_COMPLETE until recording is active (needs a delivery channel), and the
delivery channel cannot be created until the recorder completes — a deadlock
that hung the deploy ~27 min before manual cancel (2026-06-01).
Keep the cross-reviewed recorder role + delivery bucket in IaC; create the
recorder, delivery channel, and start recording via CLI (documented in README).
Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls
evaluate once Config is recording. Verified live: recording=true, SUCCESS.
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.