seahaven-org-baseline/README.md

1011 lines
60 KiB
Markdown
Raw Normal View History

Merge external-dev member baseline; rename to seahaven-org-baseline (#43) * Parameterize baseline constructs for multi-account reuse DetectiveControls, FlowLogs, and GovernanceToggles were forked into seahaven-external-dev-baseline with only physical-name and VPC-sourcing differences. Prefix/name props let one implementation serve both accounts; synthesized templates are unchanged (verified: empty cdk diff against all deployed stacks). * Absorb external-dev member baseline stack Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack, construct ids and physical names byte-identical to the deployed stack (logical IDs are path-derived; empty cdk diff verified via change set against 396287094661). Retires the forked repo so member-account baselines share one drift surface and one dependency pin. * Rename package to seahaven-org-baseline Prepares the repo rename: the app now spans the management account and org member accounts, so 'account-baseline' undersells the scope. README documents the two-account deploy topology and logical-ID constraints. * Commit extdev flow-log VPC ids in code, not -c context Security review SH-ORG-004 (confirmed high): with the ids sourced from ephemeral cdk context, any context-less deploy silently removes every flow log in the isolated account. A committed list makes the attachment set reviewable and immune to a forgotten -c flag. Empty list matches the deployed stack (zero diff). * Split CD into per-account deploy jobs The app now spans two AWS accounts; cdk deploy --all under one role fails on the other account's stacks (security review IAC-01). Each job passes explicit stack selectors and its own account's OIDC role via the new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
# seahaven-org-baseline
![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
![AWS CDK](https://img.shields.io/badge/AWS-CDK-FF9900?logo=amazonaws&logoColor=white)
Merge external-dev member baseline; rename to seahaven-org-baseline (#43) * Parameterize baseline constructs for multi-account reuse DetectiveControls, FlowLogs, and GovernanceToggles were forked into seahaven-external-dev-baseline with only physical-name and VPC-sourcing differences. Prefix/name props let one implementation serve both accounts; synthesized templates are unchanged (verified: empty cdk diff against all deployed stacks). * Absorb external-dev member baseline stack Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack, construct ids and physical names byte-identical to the deployed stack (logical IDs are path-derived; empty cdk diff verified via change set against 396287094661). Retires the forked repo so member-account baselines share one drift surface and one dependency pin. * Rename package to seahaven-org-baseline Prepares the repo rename: the app now spans the management account and org member accounts, so 'account-baseline' undersells the scope. README documents the two-account deploy topology and logical-ID constraints. * Commit extdev flow-log VPC ids in code, not -c context Security review SH-ORG-004 (confirmed high): with the ids sourced from ephemeral cdk context, any context-less deploy silently removes every flow log in the isolated account. A committed list makes the attachment set reviewable and immune to a forgotten -c flag. Empty list matches the deployed stack (zero diff). * Split CD into per-account deploy jobs The app now spans two AWS accounts; cdk deploy --all under one role fails on the other account's stacks (security review IAC-01). Each job passes explicit stack selectors and its own account's OIDC role via the new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
![CI](https://github.com/Sea-Haven-Industries/seahaven-org-baseline/actions/workflows/ci.yaml/badge.svg)
Organization-wide security and governance baseline for Sea Haven Industries,
managed as a single CDK TypeScript app. Covers the management account
(**328440206208**: primary baseline in **us-east-1**, secondary-region
baselines in **us-east-2**/**us-west-2**, offsite backup vault in
**us-west-2**) and org **member accounts** (first tenant:
`seahaven-external-dev` **396287094661**, absorbed from the retired
`seahaven-external-dev-baseline` repo). This is where account-wide detective
and recovery controls live, so they are versioned, reviewed, and drift-checked
like any other stack.
> **History:** this repo was `seahaven-account-baseline` (management account
> only) until 2026-07-14, when the external-dev member baseline was merged in
> and the repo renamed. Deployed CloudFormation stack names are unchanged.
Stacks (normally deployed by one CD job per target account; staged exceptions
are noted):
Merge external-dev member baseline; rename to seahaven-org-baseline (#43) * Parameterize baseline constructs for multi-account reuse DetectiveControls, FlowLogs, and GovernanceToggles were forked into seahaven-external-dev-baseline with only physical-name and VPC-sourcing differences. Prefix/name props let one implementation serve both accounts; synthesized templates are unchanged (verified: empty cdk diff against all deployed stacks). * Absorb external-dev member baseline stack Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack, construct ids and physical names byte-identical to the deployed stack (logical IDs are path-derived; empty cdk diff verified via change set against 396287094661). Retires the forked repo so member-account baselines share one drift surface and one dependency pin. * Rename package to seahaven-org-baseline Prepares the repo rename: the app now spans the management account and org member accounts, so 'account-baseline' undersells the scope. README documents the two-account deploy topology and logical-ID constraints. * Commit extdev flow-log VPC ids in code, not -c context Security review SH-ORG-004 (confirmed high): with the ids sourced from ephemeral cdk context, any context-less deploy silently removes every flow log in the isolated account. A committed list makes the attachment set reviewable and immune to a forgotten -c flag. Empty list matches the deployed stack (zero diff). * Split CD into per-account deploy jobs The app now spans two AWS accounts; cdk deploy --all under one role fails on the other account's stacks (security review IAC-01). Each job passes explicit stack selectors and its own account's OIDC role via the new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
| Stack | Account | Region | Purpose |
|---|---|---|---|
| `seahaven-account-baseline` | 328440206208 | us-east-1 | CloudTrail + detective controls (C-1) |
| `seahaven-dynamodb-cmk` | 328440206208 | us-east-1 | Shared customer-managed KMS key for finance/PII DynamoDB tables; ARN published to SSM `/seahaven/dynamodb/cmk-arn` (INFRA-95 / M-3) |
| `seahaven-regional-baseline-us-west-2` | 328440206208 | us-west-2 | Bedrock invocation logging (INFRA-91) |
| `seahaven-regional-baseline-us-east-2` | 328440206208 | us-east-2 | Bedrock invocation logging + AWS Config recorder + Security Hub (INFRA-91 / INFRA-16) |
| `seahaven-backup` | 328440206208 | us-east-1 | Primary AWS Backup vault + plan + role (C-7) |
| `seahaven-backup-offsite` | 328440206208 | us-west-2 | Governance-locked offsite copy vault (C-7) |
| `seahaven-org-governance` | 328440206208 | us-east-1 | AWS Organizations OU tree + SCPs (incl. the cdk-imported external-dev guardrails) |
Merge external-dev member baseline; rename to seahaven-org-baseline (#43) * Parameterize baseline constructs for multi-account reuse DetectiveControls, FlowLogs, and GovernanceToggles were forked into seahaven-external-dev-baseline with only physical-name and VPC-sourcing differences. Prefix/name props let one implementation serve both accounts; synthesized templates are unchanged (verified: empty cdk diff against all deployed stacks). * Absorb external-dev member baseline stack Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack, construct ids and physical names byte-identical to the deployed stack (logical IDs are path-derived; empty cdk diff verified via change set against 396287094661). Retires the forked repo so member-account baselines share one drift surface and one dependency pin. * Rename package to seahaven-org-baseline Prepares the repo rename: the app now spans the management account and org member accounts, so 'account-baseline' undersells the scope. README documents the two-account deploy topology and logical-ID constraints. * Commit extdev flow-log VPC ids in code, not -c context Security review SH-ORG-004 (confirmed high): with the ids sourced from ephemeral cdk context, any context-less deploy silently removes every flow log in the isolated account. A committed list makes the attachment set reviewable and immune to a forgotten -c flag. Empty list matches the deployed stack (zero diff). * Split CD into per-account deploy jobs The app now spans two AWS accounts; cdk deploy --all under one role fails on the other account's stacks (security review IAC-01). Each job passes explicit stack selectors and its own account's OIDC role via the new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
| `seahaven-external-dev-baseline` | 396287094661 | us-east-1 | Member-account baseline: Config, GuardDuty, Security Hub (FSBP + CIS v3.0), Access Analyzer, flow logs, budget |
| `seahaven-terraform-substrate` | 396287094661 | us-east-1 | Staged manually for SHOC backend/frontend adoption; exact HCP roles and deploy boundaries referencing the existing OIDC provider |
| `seahaven-security-baseline` | 001520130573 | us-east-1 | Member-account baseline for the delegated security-admin account (same construct set) |
| `seahaven-dev-baseline` | 710827005802 | us-east-1 | Member-account baseline for internal dev/staging (org-managed detection — no local GuardDuty/SecurityHub) |
| `seahaven-prod-baseline` | 011934824531 | us-east-1 | Member-account baseline for production workloads (org-managed detection; mgmt account frozen for new workloads) |
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
## CDK app
The repo is a single AWS CDK app written in TypeScript. `cdk.json` is the
project config the `cdk` CLI reads on every command: its `app` key
(`npx ts-node bin/app.ts`) tells CDK how to synthesize the app straight from
the TypeScript source — no separate compile step needed for `cdk synth` /
`diff` / `deploy` — and its `context` block carries the AWS CDK feature flags.
| Path | Role |
|---|---|
| `cdk.json` | CDK config: `app` synth command, `watch` includes/excludes, `context` feature flags |
| `bin/app.ts` | App entry point — instantiates every stack with an explicit kebab-case `stackName` and its target `env` (five accounts, per-account/per-region) |
| `lib/*-stack.ts` | Stack definitions (one class per stack; larger stacks compose the constructs in `lib/*.ts`) |
| `tsconfig.json` | TypeScript compiler options (`outDir: cdk.out`) |
| `package.json` | Pinned `aws-cdk-lib`, CDK CLI, and the `build` / `synth` / `diff` / `deploy` npm scripts |
`bin/app.ts` synthesizes fifteen stacks across three regions and five accounts:
Merge external-dev member baseline; rename to seahaven-org-baseline (#43) * Parameterize baseline constructs for multi-account reuse DetectiveControls, FlowLogs, and GovernanceToggles were forked into seahaven-external-dev-baseline with only physical-name and VPC-sourcing differences. Prefix/name props let one implementation serve both accounts; synthesized templates are unchanged (verified: empty cdk diff against all deployed stacks). * Absorb external-dev member baseline stack Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack, construct ids and physical names byte-identical to the deployed stack (logical IDs are path-derived; empty cdk diff verified via change set against 396287094661). Retires the forked repo so member-account baselines share one drift surface and one dependency pin. * Rename package to seahaven-org-baseline Prepares the repo rename: the app now spans the management account and org member accounts, so 'account-baseline' undersells the scope. README documents the two-account deploy topology and logical-ID constraints. * Commit extdev flow-log VPC ids in code, not -c context Security review SH-ORG-004 (confirmed high): with the ids sourced from ephemeral cdk context, any context-less deploy silently removes every flow log in the isolated account. A committed list makes the attachment set reviewable and immune to a forgotten -c flag. Empty list matches the deployed stack (zero diff). * Split CD into per-account deploy jobs The app now spans two AWS accounts; cdk deploy --all under one role fails on the other account's stacks (security review IAC-01). Each job passes explicit stack selectors and its own account's OIDC role via the new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
| Construct id | Stack name | Account | Region | Source |
|---|---|---|---|---|
| `account-baseline` | `seahaven-account-baseline` | 328440206208 | us-east-1 | `lib/account-baseline-stack.ts` |
| `dynamodb-cmk` | `seahaven-dynamodb-cmk` | 328440206208 | us-east-1 | `lib/dynamodb-cmk-stack.ts` |
| `regional-baseline-us-west-2` | `seahaven-regional-baseline-us-west-2` | 328440206208 | us-west-2 | `lib/regional-baseline-stack.ts` |
| `regional-baseline-us-east-2` | `seahaven-regional-baseline-us-east-2` | 328440206208 | us-east-2 | `lib/regional-baseline-stack.ts` |
| `backup-offsite` | `seahaven-backup-offsite` | 328440206208 | us-west-2 | `lib/backup-offsite-stack.ts` |
| `backup` | `seahaven-backup` | 328440206208 | us-east-1 | `lib/backup-stack.ts` |
| `org-governance` | `seahaven-org-governance` | 328440206208 | us-east-1 | `lib/org-governance-stack.ts` |
Merge external-dev member baseline; rename to seahaven-org-baseline (#43) * Parameterize baseline constructs for multi-account reuse DetectiveControls, FlowLogs, and GovernanceToggles were forked into seahaven-external-dev-baseline with only physical-name and VPC-sourcing differences. Prefix/name props let one implementation serve both accounts; synthesized templates are unchanged (verified: empty cdk diff against all deployed stacks). * Absorb external-dev member baseline stack Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack, construct ids and physical names byte-identical to the deployed stack (logical IDs are path-derived; empty cdk diff verified via change set against 396287094661). Retires the forked repo so member-account baselines share one drift surface and one dependency pin. * Rename package to seahaven-org-baseline Prepares the repo rename: the app now spans the management account and org member accounts, so 'account-baseline' undersells the scope. README documents the two-account deploy topology and logical-ID constraints. * Commit extdev flow-log VPC ids in code, not -c context Security review SH-ORG-004 (confirmed high): with the ids sourced from ephemeral cdk context, any context-less deploy silently removes every flow log in the isolated account. A committed list makes the attachment set reviewable and immune to a forgotten -c flag. Empty list matches the deployed stack (zero diff). * Split CD into per-account deploy jobs The app now spans two AWS accounts; cdk deploy --all under one role fails on the other account's stacks (security review IAC-01). Each job passes explicit stack selectors and its own account's OIDC role via the new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
| `external-dev-baseline` | `seahaven-external-dev-baseline` | 396287094661 | us-east-1 | `lib/member-baseline-stack.ts` |
| `security-baseline` | `seahaven-security-baseline` | 001520130573 | us-east-1 | `lib/member-baseline-stack.ts` |
| `dev-baseline` | `seahaven-dev-baseline` | 710827005802 | us-east-1 | `lib/member-baseline-stack.ts` (orgManagedDetection) |
| `prod-baseline` | `seahaven-prod-baseline` | 011934824531 | us-east-1 | `lib/member-baseline-stack.ts` (orgManagedDetection) |
| `deploy-substrate-prod` | `seahaven-deploy-substrate` | 011934824531 | us-east-1 | `lib/deploy-substrate-stack.ts` |
| `deploy-substrate-dev` | `seahaven-deploy-substrate` | 710827005802 | us-east-1 | `lib/deploy-substrate-stack.ts` |
| `terraform-substrate-external-dev` | `seahaven-terraform-substrate` | 396287094661 | us-east-1 | `lib/terraform-substrate-stack.ts` |
| `dynamodb-cmk-prod` | `seahaven-dynamodb-cmk` | 011934824531 | us-east-1 | `lib/dynamodb-cmk-stack.ts` |
| `alarm-topic-prod` | `seahaven-alarm-topic` | 011934824531 | us-east-1 | `lib/alarm-topic-stack.ts` |
| `app-web-acl-prod` | `seahaven-app-web-acl` | 011934824531 | us-east-1 | `lib/app-web-acl-stack.ts` |
Member-account stacks deploy with per-account credentials — the CD workflow
runs one job per account, each assuming that account's OIDC deploy role. Local
deploys/diffs assume `OrganizationAccountAccessRole` in the target account.
| Account | OIDC deploy role | Repo secret |
|---|---|---|
| 396287094661 (external-dev) | `githubdeploy-seahaven-external-dev-baseline` | `AWS_DEPLOY_ROLE_ARN_EXTDEV` |
| 001520130573 (security) | `githubdeploy-seahaven-org-baseline` | `AWS_DEPLOY_ROLE_ARN_SECURITY` |
| 710827005802 (dev) | `githubdeploy-seahaven-org-baseline` | `AWS_DEPLOY_ROLE_ARN_DEV` |
| 011934824531 (prod) | `githubdeploy-seahaven-org-baseline` | `AWS_DEPLOY_ROLE_ARN_PROD` |
Shared constructs (`DetectiveControls`, `FlowLogs`, `GovernanceToggles`) are
prefix-parameterized — construct ids and physical names must stay
byte-identical to the deployed stacks (logical IDs are path-derived).
Accounts enrolled by the org delegated admin (post 2026-07-14) set
`orgManagedDetection: true`: the GuardDuty detector + Security Hub hub come
from the org, while STANDARDS (FSBP + CIS v3.0) and the account analyzer stay
CFN-owned (org `AutoEnableStandards` is `NONE` — the DEFAULT setting enrolls
legacy CIS v1.2.0). Enrollment (member `Enabled` in GuardDuty + Security Hub)
is a hard precondition for such a stack's first deploy.
`backup` declares an explicit dependency on `backup-offsite` so the offsite copy
vault exists before the primary plan that copies into it. Stack names are set
explicitly to enforce kebab-case (CDK defaults to PascalCase). Synthesized
CloudFormation templates land in `cdk.out/` (git-ignored).
Common commands:
```
npm ci # install pinned deps
npm run build # tsc type-check (compiles to cdk.out/)
npx cdk synth # synthesize CloudFormation for all stacks
npx cdk diff # diff synthesized stacks against deployed state
npx cdk deploy --all # deploy every stack
npx cdk deploy <stack-name> # deploy a single stack
```
The `--context <key>=<value>` flag overrides `cdk.json` context at the command
line (e.g. the `encryptTrailLogGroup` toggle under *Monitoring + logging*).
## Documentation
The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This project's `seahaven-account-baseline`, `seahaven-backup`, and `seahaven-backup-offsite` stacks are represented there as Mermaid subgraphs.
- **[AWS Architecture Map](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098)** (Confluence, IT space, page 1540098)
## What it deploys
### GitHub Actions deploy substrate (per account)
`lib/deploy-substrate-stack.ts` + `lib/deploy-substrate/deploy-substrate.template.yaml`
deploy `seahaven-deploy-substrate` into each member account that hosts SAM
workloads (currently seahaven-prod and seahaven-dev). It contains the shared
account-level deploy plumbing:
- the `seahaven-lambda-execution-boundary` permissions boundary (legacy
shared ceiling for roles not yet retargeted) plus per-workload policies
`seahaven-lambda-execution-boundary-<workload>` (PLAT-52),
- the `github-cfn-execution-role` CloudFormation execution role that `cd-sam`
callers pass as `cfn-role-arn`,
- optionally the GitHub OIDC identity provider (`createOidcProvider: true`,
only for an account that does not already have one — one provider per URL
per account).
fix(iam): correct two boundary-scoping defects found in review Post-implementation verification of the INFRA-186 prod/dev scoping found two functional defects that would have denied permissions the migrating stacks actually need. Neither is live today (prod/dev boundary usage is 0), but both would have surfaced as AccessDenied at first migration. - KMS: the ViaService list omitted ssm., while SSMParameterRead in the same policy grants ssm:GetParameter*. A SecureString read decrypts via the SSM service principal, so the boundary denied reads it also granted. - S3: payments-dashboard was classified read-only from the template's own permission-source comment, but that enumeration is incomplete -- the real stack grants s3:PutObject on BoaRawBucket (template.yaml:272, 1098-1099). Write is now allowed on seahaven-payments-boa-raw-* only; payroll-emails and payments-csv stay read-only, preserving the evidence-deletion protection. The seahaven-payments-* wildcard is replaced by the three literal bucket names, verified against payments-dashboard/template.yaml. Not changed: SES configuration-set/*. Review claimed dropping it rested on a false premise; verified live -- prod and dev both have ZERO configuration sets and member-baseline-stack.ts:44 excludes SES monitoring. The drop is correct. README: the 'substrate changes must edit both files' rule is now false for the boundary specifically, and said so uniformly. Corrected to distinguish the deliberately divergent boundary from the still-at-parity substrate resources.
2026-07-30 18:03:40 -04:00
The template began as a verbatim extraction of the substrate section of
`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml`, which remains the source
of truth for mgmt (328440206208) until its stacks migrate out.
**The two copies are no longer at parity, and the old "edit both files" rule no
longer applies uniformly.** Under INFRA-186, `seahaven-lambda-execution-boundary`
in *this* copy was reduced to a fleet-wide floor for prod and dev; later
migrations packed per-workload data plane back into that shared document until
it hit the 6,144-character cap (PLAT-93 / PLAT-100). PLAT-52 adds per-workload
policies `seahaven-lambda-execution-boundary-<workload>` (floor plus that
stack's data plane) and switches both prod/dev guardrails to a StringEquals
allow-list of the shared ARN plus each per-workload ARN. The shared document
is left unchanged until live roles retarget. Mgmt's copy keeps the
account-wide wildcards and a **single-ARN** pin pending its own separately
validated rollout across 26 live boundary-carrying roles (PLAT-51). So: **the
boundary resource is deliberately divergent**, and the guardrail
`iam:PermissionsBoundary` condition **cardinality** is also divergent
(enumerated list here, scalar on mgmt). Do not weaken mgmt to ArnLike. Every
*other* substrate resource (`github-cfn-execution-role`,
`seahaven-cfn-exec-iam-management` Sid/Action/Resource sets) is still expected
to change in both files together. The template's provenance header records
which is which — read it before assuming either parity or divergence.
fix(iam): correct two boundary-scoping defects found in review Post-implementation verification of the INFRA-186 prod/dev scoping found two functional defects that would have denied permissions the migrating stacks actually need. Neither is live today (prod/dev boundary usage is 0), but both would have surfaced as AccessDenied at first migration. - KMS: the ViaService list omitted ssm., while SSMParameterRead in the same policy grants ssm:GetParameter*. A SecureString read decrypts via the SSM service principal, so the boundary denied reads it also granted. - S3: payments-dashboard was classified read-only from the template's own permission-source comment, but that enumeration is incomplete -- the real stack grants s3:PutObject on BoaRawBucket (template.yaml:272, 1098-1099). Write is now allowed on seahaven-payments-boa-raw-* only; payroll-emails and payments-csv stay read-only, preserving the evidence-deletion protection. The seahaven-payments-* wildcard is replaced by the three literal bucket names, verified against payments-dashboard/template.yaml. Not changed: SES configuration-set/*. Review claimed dropping it rested on a false premise; verified live -- prod and dev both have ZERO configuration sets and member-baseline-stack.ts:44 excludes SES monitoring. The drop is correct. README: the 'substrate changes must edit both files' rule is now false for the boundary specifically, and said so uniformly. Corrected to distinguish the deliberately divergent boundary from the still-at-parity substrate resources.
2026-07-30 18:03:40 -04:00
Per-repo `githubdeploy-*` deploy roles are deliberately NOT part
of the substrate — they are provisioned per repo at migration/onboarding time
so an account never carries trust relationships for repos that do not deploy
to it.
**Escalation controls on `github-cfn-execution-role`.** Every `iam:CreateRole`,
`AttachRolePolicy` and `PutRolePolicy` is conditioned on the target carrying
one of the enumerated `seahaven-lambda-execution-boundary` ARNs (the shared
policy plus each `seahaven-lambda-execution-boundary-<workload>`). That
condition alone is not sufficient,
so the attached `seahaven-cfn-exec-iam-management` managed policy also carries
three explicit Deny statements:
- `DenyBoundaryTampering` — no removing a boundary from any role or user.
Granting the delete under the same `StringEquals` condition self-defeats the
gate, because for a delete the condition key resolves to the boundary already
on the target.
- `DenyBoundaryPolicyEdit` — no rewriting any `seahaven-*` managed policy.
- `DenySelfMutation` — the role cannot modify or delete itself or any
`githubdeploy-*` role. Without it the control is one API call from being
undone: `IAMRoleReadAndDelete` grants `iam:DetachRolePolicy` on `Resource:
"*"` unconditioned, so the role could detach the very policy carrying these
Denies.
Verify a change to these with `aws iam simulate-principal-policy` against the
role's own ARN (expect `explicitDeny`) and against a `<stack>-<Function>Role-`
name (expect `allowed`, no regression for normal SAM deploys). Note that
simulation currently does **not** see this role's *inline* policies in
seahaven-prod or seahaven-dev — read those back with `get-role-policy` instead.
Known consequence of `DenyBoundaryTampering`: a CloudFormation rollback of an
update that *adds* a boundary to an existing role wedges in
`UPDATE_ROLLBACK_FAILED`. Recovery is an administrator action, not a pipeline
retry — `aws cloudformation continue-update-rollback --stack-name <stack>
--resources-to-skip <RoleLogicalId>`. Unreachable while every SAM role is
created with the boundary already attached.
**Onboarding a future account as a deploy target:**
1. CDK-bootstrap the account (`npx cdk bootstrap aws://<account>/us-east-1`
via `OrganizationAccountAccessRole`).
2. Create `githubdeploy-seahaven-org-baseline` in the account (same trust and
policy as the dev/prod copies) and add the repo secret
`AWS_DEPLOY_ROLE_ARN_<ACCT>`.
3. Add a `DeploySubstrateStack` instance in `bin/app.ts`
(`createOidcProvider: true` if the account has no GitHub OIDC provider)
and append its construct id to a new per-account job in
`.github/workflows/deploy.yaml` (explicit `stacks` selector, one job per
account).
4. Merge; the substrate deploys via CD. Per-repo deploy roles and app stacks
follow the cross-account migration playbook from there.
### Terraform deploy substrate (per account)
`lib/terraform-substrate-stack.ts` + `lib/terraform-substrate/terraform-substrate.template.yaml`
deploy `seahaven-terraform-substrate` into each member account that hosts
Terraform-managed workloads (currently seahaven-prod, seahaven-dev, and
external-dev; never mgmt — mgmt stays SAM until its stacks migrate out).
Prod/dev use the shared IAM-management policy. External-dev references its
existing `app.terraform.io` provider and carries only exact SHOC
import/adoption roles:
- the `app.terraform.io` OIDC identity provider (audience
`aws.workload.identity`; Retain — it is the federation anchor for every
future `hcptf-*` role),
- the `seahaven-hcptf-iam-management` guardrail policy: the boundary-gated
IAM role lifecycle (conditioned on the enumerated
`seahaven-lambda-execution-boundary` allow-list owned by the
deploy-substrate stack — hence the explicit stack dependency in
`bin/app.ts`) plus the `DenyBoundaryTampering` / `DenyBoundaryPolicyEdit`
/ `DenySelfMutation` backstops,
- external-dev-only deploy boundaries
`shoc-backend-{tf-poc,dev,staging}-deploy-boundary`. Each is the maximum
current policy for one exact `githubdeploy-shoc-backend-*` role. Dev
temporarily retains its live broad `elasticbeanstalk-*` S3 grants so
boundary attachment cannot regress deployment before the separately
reviewed policy narrowing,
- external-dev-only runtime boundaries
`shoc-backend-{tf-poc,dev,staging}-runtime-boundary`. These retain only the
account-scoped S3, environment health/log, and X-Ray portions of
`AWSElasticBeanstalkWebTier`, plus each environment's exact secrets/KMS/STS
data plane. They deliberately exclude the managed policy's 2026
Bedrock/Marketplace additions.
**This policy derives from `seahaven-cfn-exec-iam-management` but is
deliberately stricter — it is not a mirror.** The 2026-07-30 security review
confirmed the SAM copy's `Resource: "*"` role grants as a critical escalation
primitive (`iam:UpdateAssumeRolePolicy` on `*` repoints the AdministratorAccess
CDK bootstrap role's trust policy to an external account), and its justification
for the wildcard — SAM auto-generates execution roles at path `/` with no
settable `RolePath` — does not transfer, because Terraform's `aws_iam_role`
supports `path`. So here:
- every role **write** (create, delete, detach, `UpdateAssumeRolePolicy`,
boundary set) and `iam:PassRole` is confined to the Terraform-owned path
`role/tf-managed/*`; reads stay on `*` for data sources,
- **Terraform configs must set `path = "/tf-managed/"` on every
`aws_iam_role`** — a role created anywhere else is denied,
- `DenySelfMutation` additionally covers `cdk-hnb659fds-*`,
`OrganizationAccountAccessRole` and `seahaven-*` (detective-control roles,
which no prod/nonprod SCP shields from `iam:DeleteRole`).
Do not "reconcile" the two files by copying statements between them. The
durable org-level fix for the same class is extending the existing
`ProtectPrivilegedRoles` SCP (currently security-OU only) to prod and nonprod.
Per-workspace roles (`hcptf-<stack>` apply + `hcptf-<stack>-plan`) are
deliberately NOT pre-provisioned — they are appended to the template at each
stack's migration time so an account never carries trust for workspaces that
do not deploy to it.
**External-dev SHOC role adoption is a staged CloudFormation import, not a
normal first deploy.** Six roles exist today:
`hcptf-shoc-backend-{dev,staging}` and their `-plan` partners, plus the
`hcptf-shoc-backend-tf-poc` pair. Two independent CDK contexts make each
transition explicit:
`enableShocBackendPocRoles` and `enableShocBackendLiveRoles`. Both began as
`false` for the initial rollout and remain version-controlled as `true` after
their completed ownership transitions.
`terraform-substrate-external-dev` is deliberately absent from the automatic
external-dev deploy job during this sequence; `external-dev-baseline` remains
automatic and unchanged.
1. Create the role-free base stack:
```bash
npx cdk deploy terraform-substrate-external-dev \
-c enableShocBackendPocRoles=false \
-c enableShocBackendLiveRoles=false
```
`CreateOIDCProvider=false` is fixed in `bin/app.ts`; the external-dev
account therefore creates neither the existing provider, SHOC roles, nor
the prod/dev-only shared IAM policy. The base stack does create all six
retained external-dev deploy/runtime boundary policies.
2. Set `enableShocBackendPocRoles` to `true` in `cdk.json`, leave the live gate
`false`, review the synthesized two-role addition, then run the normal
external-dev stack update. This creates only the new tf-poc HCP plan/apply
pair. The retained POC CDK stack references
`shoc-backend-tf-poc-deploy-boundary` when it creates
`githubdeploy-shoc-backend-tf-poc` and
`shoc-backend-tf-poc-runtime-boundary` when it creates the POC runtime
role; do not attach the generic account execution boundary to either role.
3. Prove both tf-poc HCP assumptions and the retained POC import rehearsal
before touching the live-role ownership boundary.
4. In a separately approved administrator/CDK migration, tag the existing HCP
apply roles first:
`hcptf-shoc-backend-dev` gets
`HcpTerraformWorkspace=shoc-backend-dev`, and
`hcptf-shoc-backend-staging` gets
`HcpTerraformWorkspace=shoc-backend-staging`. Next attach
`shoc-backend-dev-deploy-boundary` and
`shoc-backend-staging-deploy-boundary` to the exact `githubdeploy-*` roles,
and attach `shoc-backend-dev-runtime-boundary` /
`shoc-backend-staging-runtime-boundary` to the exact runtime roles. Verify
the boundary ceilings before adding the matching manager tag to either
target `githubdeploy-*` role. The POC CDK
creates its deploy role with `HcpTerraformWorkspace=shoc-backend-tf-poc`;
the substrate-created POC apply role already carries the same principal
tag. Verify each effective deployment action before continuing. HCP remains
blocked while a target tag is missing/different or the target lacks its
exact dedicated boundary, so a partial migration cannot authorize policy
writes. Complete both runtime/deploy boundary attachments before workload
imports. HCP apply roles deliberately have no
`iam:PutRolePermissionsBoundary` or boundary-policy mutation permissions.
5. In the backend bootstrap, add Terraform `removed` blocks with
`destroy = false` for only the four dev/staging HCP roles and their inline
policies. Apply and verify Terraform state no longer owns them while all
four physical roles and ARNs remain unchanged.
6. Set both contexts to `true`, synthesize with
`npx cdk synth terraform-substrate-external-dev`, and create a
CloudFormation **IMPORT** change set for the four existing
`AWS::IAM::Role` resources by exact role name. Do not run a normal
CREATE/UPDATE change set for this ownership transition. The POC gate must
remain true so the already-managed pair stays in the template. Import
records ownership; it does not update existing role properties or inline
policies.
7. Run a separate, reviewed CloudFormation reconcile update after import and
before switching workspace credentials. Each current live role has one
inline policy: `shoc-backend-dev-import-plan`,
`shoc-backend-dev-import-apply`,
`shoc-backend-staging-import-plan`, or
`shoc-backend-staging-import-apply`. Existing descriptions and tags are
inventoried in the backend handoff. Reconcile those explicit differences
to the final baseline shape without replacing a role.
8. After reconcile and HCP assumption proof, keep both context values committed
as `true`. Every subsequent normal deployment must synthesize all six
roles. Never return either gate to false as a rollback mechanism; Retain
protects the physical role but removing it from the stack abandons
CloudFormation ownership.
9. Only after all imports/reconciliation complete and both context defaults
are permanently `true`, add `terraform-substrate-external-dev` back to the
external-dev workflow stack selector. Until then all substrate operations
are deliberate manual deploy/import actions.
10. Retire the backend bootstrap only after the POC pair and all four imported
live roles are proven under this stack. Role deletion/recreation is never a
migration step.
The external-dev apply roles intentionally omit role create/delete,
managed-policy attach/detach, trust or boundary mutation, `iam:PassRole`, and
secret-value APIs. IAM writes are limited to exact-role inline-policy and
ordinary tag updates plus exact-profile tags; role descriptions remain stable
and HCP receives no `UpdateRole` or `UpdateRoleDescription`. The SCP permits
only the three enumerated HCP apply roles to mutate a `githubdeploy-*` role
whose locked `HcpTerraformWorkspace` resource tag equals the caller's immutable
principal tag. Adding or changing that manager tag remains administrator/CDK
only. POC DNS and certificate access is tag/name constrained because their
physical IDs are allocated by the temporary retained CDK stack before
Terraform imports them. Dev and staging DNS writes are pinned to their existing
hosted-zone IDs and API record names.
Current compact policy-document sizes are 1,387 / 1,873 / 1,844 characters for
the POC/dev/staging deploy boundaries and 1,176 / 1,779 / 1,656 for their
runtime boundaries, each below IAM's 6,144-character managed-policy limit. The
external-dev IAM guardrail SCP is 5,095 compact characters against its
5,120-character Organizations limit; keep size assertions in every change.
**External-dev SHOC frontend adoption uses separate gates and creates its
boundaries first.** The three retained boundaries are
`shoc-frontend-new-{tf-poc,dev,staging}-deploy-boundary`. Each permits only
bucket location/list/version reads, object get/put/current and version delete,
and invalidation create/read for one exact distribution. Dev is pinned to
`E2CWLM1AFB964P`; staging is pinned to `E2JDVEZ6EGD49J`. The tf-poc
distribution, OAC, function, hosted-zone, and certificate identifiers are
intentionally empty in `cdk.json`. They must come from the frontend shared
creator outputs; this substrate does not reuse the backend tf-poc zone or
certificate. Its site name is `frontend-tf-poc.seahaven.com`. While the
identifier set is empty, its boundary omits invalidation access and
`ShouldManageShocFrontendPocRoles` remains false even if its role gate is
mistakenly enabled.
The frontend role transition is manual and is not part of the external-dev CD
job:
1. Deploy `terraform-substrate-external-dev` with both frontend role gates
false. Verify the three boundary documents before touching any GitHub role.
2. For dev and staging, an administrator attaches the matching dedicated
boundary to `githubdeploy-shoc-frontend-new-<env>`, then adds
`HcpTerraformWorkspace=shoc-frontend-new-<env>`. Verify the current GitHub
deployment still uploads to only the exact bucket and invalidates only the
exact distribution. The SCP blocks HCP from adding or changing this manager
tag itself.
3. Set `enableShocFrontendLiveRoles=true` and run a reviewed reconcile update.
This creates the exact plan/apply pairs with 3,600-second sessions and
phase-specific HCP `StringEquals` trust. It does not import or replace the
existing GitHub roles.
4. Reconcile the frontend Terraform roots with `removed { destroy = false }`
ownership handoff for HCP roles/boundaries where applicable. Import the
existing site resources only after a no-replacement plan. Apply-role writes
are limited to ordinary tags, `PutRolePolicy` on the exact deploy role,
`PutBucketPolicy` on the exact bucket, and A/AAAA changes for the exact site
name with CREATE/DELETE/UPSERT conditions.
5. For a future tf-poc, first provision and inventory the site outside these
adoption roles. Set all five `shocFrontendPoc*` identifiers from the
frontend shared creator outputs while its role gate remains false, deploy
and reconcile the boundary, attach it and the matching manager tag to the
exact GitHub role, then set
`enableShocFrontendPocRoles=true`.
6. Keep `terraform-substrate-external-dev` out of automatic deployment until
all enabled frontend roles and target-role guardrails are proven. Do not use
a false gate as rollback after CloudFormation owns a role.
The apply roles explicitly deny role lifecycle/trust/boundary/managed-policy
changes, `PassRole`, secret and parameter reads, CloudFront/S3 infrastructure
mutation, and deletion of inline role or bucket policies. IAM does not expose a
condition key for an inline policy name, so `PutRolePolicy` is constrained to
the exact target-role ARN and requires the exact dedicated deploy boundary to
already be attached. The boundary limits effective permissions, and the SCP
requires the target role's locked `HcpTerraformWorkspace` tag to equal the
apply role's principal tag. The Terraform resource must retain the inventoried
inline policy name.
**HCP Terraform layout (org-level setup, console):** one org `seahaven`
(free tier: 500 managed resources, 1 concurrent run); one HCP **project per
AWS account** (`seahaven-prod`, `seahaven-dev`); one **workspace per stack**
(`<stack>-<env>`, one state file = one blast radius). Default execution mode
Remote. Never use HCP's "Quick setup AWS dynamic credentials" button — it
writes the single `TFC_AWS_RUN_ROLE_ARN`, which collapses the plan/apply role
split this substrate exists to enforce.
**Reference implementation:** first workload was `afi-backup-monitor` in
seahaven-prod (PLAT-56). Copy
`Sea-Haven-Industries/afi-backup-monitor` `terraform/` and the live
`hcptf-afi-backup-monitor*` / `hcptf-afi-backup-monitor-plan` statements in
this template rather than inventing new IAM shapes.
**Migration checklist (per stack, in order):**
0. **Freeze the app's SAM/CDK CD** (remove or disable the deploy workflow) so
HCP Terraform becomes the sole deploy path before the first apply. Leave
the source-account stack frozen until cutover.
1. **Secrets first.** Create exact secret shells in the target account; strip
trailing newlines/whitespace before `put-secret-value` (a trailing `\n`
breaks HTTP headers at runtime). Capture ARNs. Never put secret *values*
in Terraform state (ARN references only).
2. **HCP workspace** in the target account's project (`<stack>-<env>`). Apply
method **Manual**; automatic speculative plans on if VCS-connected;
working directory `terraform/`. (CLI `terraform plan` runs are inherently
speculative.)
3. **Substrate PR** to this repo appending `hcptf-<stack>-plan` and
`hcptf-<stack>` (see 3a/3b). Trust: this account's `app.terraform.io`
provider; `StringEquals` on `app.terraform.io:aud` =
`aws.workload.identity` and on `app.terraform.io:sub` =
`organization:seahaven:project:seahaven-<env>:workspace:<workspace>:run_phase:plan`
(or `:apply`). Exact `StringEquals` only — never `StringLike`, never a
wildcarded `run_phase` (a speculative PR plan must never hold write
credentials). **If the stack creates Lambda execution roles, this same PR
must also add `seahaven-lambda-execution-boundary-<stack>`** per the
WIDENING PATH in `lib/deploy-substrate/deploy-substrate.template.yaml`
(floor plus that stack's data plane, **exact** secret ARNs from step 1,
no `secret:afi-*` patterns) **and** append that policy's ARN to both
guardrail StringEquals allow-lists. Do not add data-plane to the shared
`seahaven-lambda-execution-boundary` document. The guardrail forces every
Terraform-created role to carry a listed boundary; an unlisted or
floor-only boundary deploys green, then every data-plane call is denied
at first invoke and async/DLQ writes are discarded silently. IAM roles and
boundary policies = mandatory cross-family review +
`/sh-security-review` on the diff.
3a. **Plan role (required for every stack):** attach
`arn:aws:iam::aws:policy/job-function/ViewOnlyAccess` (never
`ReadOnlyAccess`, which grants `secretsmanager:GetSecretValue`,
`s3:GetObject` and `kms:Decrypt` and would let any PR-triggered speculative
plan render secret values into HCP run output) **plus** a scoped
plan-refresh sidecar inline policy. ViewOnly alone is insufficient for
Terraform refresh after partial apply — it lacks `iam:GetRole`,
`events:DescribeRule`, and several Lambda/S3 reads. Sidecar minimum:
`iam:GetRole` / related reads on `role/tf-managed/<prefix>-*`;
`events:DescribeRule` (and list-targets/tags as needed) on
`rule/<prefix>-*`; `lambda:*` (or at least the Get*/List* the provider
uses) on `function:<prefix>-*` / `layer:<prefix>-*`; `s3:Get*` /
`s3:ListBucket` on the stack artifact bucket. **No** IAM writes, **no**
guardrail-policy attach on the plan role. Copy
`afi-backup-monitor-plan-refresh` on `hcptf-afi-backup-monitor-plan`.
3b. **Apply role (Lambda/EventBridge stacks):** attach
`seahaven-hcptf-iam-management` plus stack-scoped service statements.
Prefer prefix-scoped `lambda:*` on `function:<prefix>-*` /
`layer:<prefix>-*`, `events:*` on `rule/<prefix>-*`, and bucket-scoped
`s3:*` on the artifact bucket — do **not** enumerate individual provider
Get* APIs (`GetFunctionCodeSigningConfig`, `GetBucketAcl`, …); that list
lags and fails first apply. Keep list/describe-on-`*` only where the
service requires it (e.g. `lambda:ListFunctions`). Copy
`afi-backup-monitor-services` on `hcptf-afi-backup-monitor`.
4. **Deploy substrate** to `UPDATE_COMPLETE`. Verify: both roles exist;
`hcptf-<stack>` lists `seahaven-hcptf-iam-management` in
`list-attached-role-policies`; trust subs match the live
org/project/workspace names byte-for-byte; simulate the apply role against
a `hcptf-*` ARN (expect `explicitDeny` from `DenySelfMutation`) and against
a normal stack role name (expect `allowed`); and if step 3 added a
per-workload boundary, confirm the deployed default version of
`seahaven-lambda-execution-boundary-<stack>` carries the stack's
data-plane statements (`aws iam get-policy-version`) — role verification
alone never checks boundary content. Mechanical template↔deployed policy
reconcile as for other substrate policies.
5. Set **workspace-level** variables `TFC_AWS_PLAN_ROLE_ARN` +
`TFC_AWS_APPLY_ROLE_ARN` (category env) to the verified role ARNs, plus
`TFC_AWS_PROVIDER_AUTH=true`. Never project-scoped variable sets — the
trust is pinned per workspace, so a shared set breaks every other
workspace. Auto-apply stays OFF until the stack is sealed.
6. **App Terraform PR:** every `aws_iam_role` sets `path = "/tf-managed/"` and
`permissions_boundary` to that stack's
`seahaven-lambda-execution-boundary-<stack>` ARN (not the shared name,
once the per-workload policy exists); package Lambda/layer zips via an account artifact S3 bucket
and `aws_s3_object` `content_base64` (HCP plan and apply run on separate
workers and do not share local `archive_file` paths — see
`afi-backup-monitor/terraform/artifacts.tf`); functions `depends_on` their
IAM policies before create; commit `.terraform.lock.hcl` with
multi-platform hashes.
7. **First Manual apply** from the HCP workspace (not local apply against
prod). Tolerate partial state on permission misses; widen the apply/plan
roles and retry. Confirm all expected resources exist in the target
account.
8. **Live-path proof:** real invoke of every critical function must hit real
external APIs / Slack (not synth or simulate alone) before cutover.
9. **Cutover + decommission:** disable source-account schedules (e.g.
EventBridge rules); observe a clean prod path; delete the source
CloudFormation/CDK stack per the decommission playbook; sweep or retain
log groups deliberately; delete source secrets last.
10. **Docs:** update Confluence AWS Architecture Map and the stack ops page;
promote durable gotchas to the convention ledger when they are general.
**HCP-side authority is AWS authority.** AWS exposes only `aud`, `sub` and
`amr` as trust-policy condition keys for a generic OIDC provider — HCP's
immutable `terraform_workspace_id` / `terraform_project_id` claims are *not*
usable in an IAM condition (AWS's provider-specific claim validation covers
Google, GitHub, CircleCI and OCI only). The `sub` pin therefore rests on HCP
display names, so whoever can create, rename, move or delete a workspace in the
`seahaven-prod` project effectively holds prod deploy authority. Restrict that
HCP team permission to the same people, and when a workspace is retired, delete
its `hcptf-*` roles in the same change so a reused name cannot inherit them.
**Terraform state is secret-bearing.** HCP-hosted state records sensitive
attributes in full and lives outside the AWS accounts, readable by any HCP
principal with workspace read. Per the handbook's secrets-and-config rule,
secrets stay in Secrets Manager / SSM and are referenced by ARN: do not manage
secret *values* in Terraform (create the secret shell, populate out of band or
via write-only/ephemeral arguments) so no value enters state.
**Rollback (proven in mgmt 2026-07-30):** delete any `hcptf-*` roles first —
they reference the provider, and while any of them still attaches the guardrail
policy the stack delete cannot remove it. Then delete the stack. Only the
**provider** is `Retain`: it survives as an orphan and is removed with
`aws iam delete-open-id-connect-provider`. The **guardrail policy is deleted
with the stack** — do not expect it to persist, and note that every
`DenySelfMutation` / `DenyBoundaryTampering` backstop goes with it, so an
`hcptf-*` role recreated out of band afterwards is *not* gated. Workspaces
holding state must be migrated or destroyed HCP-side first; deleting the OIDC
provider strands them mid-run rather than cleaning them up.
**First-create rollback trap.** The provider is `Retain`, so if any other
resource in this stack fails on first create, CloudFormation rolls back, the
provider survives untracked, and the stack lands in `ROLLBACK_COMPLETE` — which
cannot be updated, and cannot be recreated because an account holds exactly one
provider per URL. Recovery: delete the stack, then either remove the orphaned
provider with the command above before retrying, or redeploy with
`createOidcProvider: false`. Note `cd-cdk`'s pre-flight and health check probe
only the job's single `stack-name` input (the account baseline), so a wedged
substrate stack does not show up there — check it directly.
**Verification of record for the guardrail policy** is mechanical
reconciliation — tag-preserving YAML load of the template vs
`get-policy-version` readback, sorted `json.dumps` compare per statement —
same discipline as the deploy-substrate reconciliation (2026-07-27), not
header-reading. The managed-policy document budget is 6,144 characters;
measure before appending statements.
### CloudTrail (audit finding C-1)
| Resource | Logical ID | Notes |
|---|---|---|
| Multi-region trail | `Trail` (`seahaven-org-trail`) | Management events read+write, global service events, **log-file validation on**, **CloudTrail Insights on** (ApiCallRate + ApiErrorRate, §37) |
| Log bucket | `TrailLogBucket` (`seahaven-cloudtrail-logs-328440206208`) | Private (Block Public Access all), SSE-KMS, versioned, **TLS-only**, **Object Lock GOVERNANCE 365d**, lifecycle (Glacier @90d, expire @365d), server access logging → `seahaven-s3-access-logs` |
| KMS CMK | `TrailKey` (`alias/cloudtrail-logs`) | Encrypts log files; **automatic rotation enabled** |
| CloudWatch Logs group | created by the L2 `Trail` | 365-day retention; this is the group the CIS Section 4 metric filters (H-1) attach to |
**Data flow:** API activity across all regions → CloudTrail → (a) KMS-encrypted,
Object-Locked S3 bucket for durable/tamper-resistant storage and (b) CloudWatch
Logs for real-time querying and metric-filter alarms.
**Compliance impact:** closes CIS 3.1 (multi-region trail), 3.2 (log-file
validation), 3.4 (CloudWatch Logs integration), 3.6 (bucket access logging),
3.7 (KMS CMK encryption), and 3.8 (CMK rotation). Unblocks CIS Section 4 /
finding H-1 (metric filters + alarms now have a log group to target).
### Design decisions
- **Management events only.** Object-level S3/Lambda data events (CIS 3.10/3.11)
are deferred to control cost; revisit with targeted S3 *write* data events on
sensitive buckets (payments / accounting / kb) if needed.
- **Object Lock GOVERNANCE, not COMPLIANCE.** Tamper-resistant but still
deletable by a principal holding `s3:BypassGovernanceRetention` — avoids the
irreversibility of COMPLIANCE mode. Revisit if a stricter posture is required.
- **RETAIN** on the bucket and KMS key so a stack teardown never destroys the
audit trail.
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
### AWS Backup (audit finding C-7)
Phase 1 ("critical data first") of fixing the account's complete lack of AWS
Backup. Protects the data stores with no offsite leg today and copies each
recovery point cross-region into a governance-locked vault.
| Resource | Logical ID | Notes |
|---|---|---|
| Primary vault | `seahaven-primary` (us-east-1) | KMS-CMK encrypted, unlocked (working copy), RETAIN |
| Offsite vault | `seahaven-offsite` (us-west-2) | KMS-CMK encrypted, **Vault Lock GOVERNANCE** (min-retention 30d, no cooling-off window), RETAIN |
| Backup plan | `seahaven-critical-daily` | Daily 06:00 UTC, delete-after 35d, **cross-region CopyAction → offsite** (retain 90d) |
| Service role | `seahaven-backup-service-role` | **Backup-only** (Backup + S3-Backup managed policies); restore perms intentionally deferred |
**Phase-1 scope** (selected by explicit ARN, not tags, to avoid drifting other
stacks): RDS `proposal-system-db`, DynamoDB `PaymentsDashboard`,
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
DynamoDB `purchase-orders`, S3 `accounting.seahaven.com`,
`seahaven-payments-csv-328440206208`, `google-workspace-seahavenind.com`.
*(RDS `database-1` was originally in this set but was retired 2026-06-03 —
audit H-19, idle 0 conn/60d — and removed from the selection; its final
encrypted recovery point is retained in `seahaven-offsite` for 7 years.)*
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
**Coexists with** existing EBS DLM snapshots and DynamoDB PITR — it supplements
them with the missing offsite + immutable leg; it does not replace them.
**Design decisions:**
- **Governance lock first, not compliance.** Recovery points can't be silently
deleted, but a principal with explicit permission can still intervene while
we validate. Graduate to COMPLIANCE (irreversible) later by adding
`changeableFor` to the offsite vault lock + redeploy.
- **Backup-only role.** Restore policies and `allowRestores` are not granted;
restores get a separate audited path once a restore-test process exists.
**Pre-deploy gates** (must clear before the first scheduled run):
1. Enable S3 versioning on `seahaven-payments-csv-328440206208` and
`google-workspace-seahavenind.com` (`accounting.seahaven.com` already has it,
audit C-9), or their jobs fail silently (folds in H-21).
2. `database-1` is unencrypted (H-19): smoke-test an on-demand backup + copy of
it to us-west-2 first; if the copy fails, encrypt it or drop it from the copy.
3. Enable DynamoDB PITR (H-7) on the two tables for between-window recovery.
### AWS Backup phase 2 (audit Day 4)
Expands the same `seahaven-critical-daily` plan to every remaining data store, so
all of DynamoDB + EBS get the offsite + immutable leg ("offsite for everything").
| Resource | Logical ID | Notes |
|---|---|---|
| Phase-2 selection | `Plan/Phase2Resources` (`phase2-offsite-everything`) | Same plan, same `seahaven-backup-service-role`, same daily + cross-region copy rule |
**Phase-2 scope:** the 15 remaining DynamoDB tables (all except the two phase-1
financial tables + the deleted ledgerflow tables) and all 9 in-use EBS volumes,
again **by explicit ARN** — tag-based selection was deliberately avoided because
the file-share volumes are standalone-managed and the tables are owned by other
stacks, so tagging here would drift them.
**No IAM change:** `AWSBackupServiceRolePolicyForBackup` already grants the
DynamoDB/RDS/EBS backup actions, so phase 2 reuses the phase-1 role unchanged
(cross-reviewed, no BLOCK).
**Known tradeoff (→ Jira INFRA-31):** explicit-ARN EBS entries go stale if a
volume is replaced (new volume id), silently dropping it from backup. Migrating
the EBS portion to tag-based selection (with the tag codified in each owning
stack) is the resilient follow-up; scheduled drift detection is the interim
backstop.
**Also enabled outside this stack (audit H-7, via CLI — codify per stack →
INFRA-30):** PITR + `DeletionProtectionEnabled` on 12 more DynamoDB tables
(account-wide PITR now 19/21).
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
### Detective controls + budget (audit Day 1)
Account-level detective layer, in `lib/detective-controls.ts`, plus the cost
budget in `lib/governance-toggles.ts`. **Scope is us-east-1 only** (all workloads
live here); multi-region coverage is a follow-up.
| Resource | Logical ID | Finding | Notes |
|---|---|---|---|
| Config delivery bucket | `seahaven-config-328440206208` | H-2 | Private (BPA all), SSE-S3, versioned, TLS-only, 365d lifecycle |
| Config recorder role | `seahaven-config-recorder-role` | H-2 | `AWS_ConfigRole` + scoped S3 delivery; **IAM cross-reviewed** |
| Config recorder + channel | `DetectiveControls/ConfigPutRecorder`, `ConfigPutChannel`, `ConfigStartRecorder` | H-2 | `AwsCustomResource` calls `PutConfigurationRecorder` → `PutDeliveryChannel` → `StartConfigurationRecorder` in sequence; idempotent upsert avoids the L1 CFN deadlock (INFRA-17). Custom-resource role cross-reviewed. |
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
| GuardDuty detector | `DetectiveControls/GuardDutyDetector` | H-3 | Findings every 15 min |
| Security Hub | `DetectiveControls/SecurityHub` | H-4 | FSBP v1.0.0 + CIS v3.0.0; controls evaluate once Config is recording |
| Access Analyzer | `seahaven-account-analyzer` | M-5 | ACCOUNT external-access analyzer (free) |
| Monthly budget | `GovernanceToggles/MonthlyCostBudget` (`seahaven-monthly-cost`) | M-10 | $1,200/mo, 80%/100% actual + 100% forecast → adam@seahavenind.com |
**Config recorder + delivery channel are managed by `AwsCustomResource` (INFRA-17).**
The L1 `AWS::Config::ConfigurationRecorder` deadlocks the stack (recorder never
reaches `CREATE_COMPLETE` without a delivery channel; channel can't be created
without a recorder — hit 2026-06-01). The custom resource sidesteps this by
calling the Config SDK directly: `Put*` is an upsert, so the deploy adopts the
existing CLI-created recorder and channel without destroying them. Active
recording is never interrupted.
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
### CLI-applied governance toggles (no CloudFormation resource)
These account toggles have no native CloudFormation resource, so they are applied
via CLI and recorded here — **per account** (they are account-scoped; a new
member account has NONE of them until applied). Applied: 328440206208
(2026-06-01); 001520130573, 710827005802, 011934824531 (2026-07-14 — EBS
encryption-by-default + password policy, verified per account; Inspector2 via
delegated admin; cost-allocation tags are org-level).
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
```bash
# M-3 EBS encryption-by-default (new volumes; existing 5 plaintext volumes are H-19-adjacent)
aws ec2 enable-ebs-encryption-by-default --region us-east-1
# M-6 Inspector2 (EC2 + Lambda + ECR)
aws inspector2 enable --resource-types EC2 LAMBDA ECR --region us-east-1
# M-7 IAM password policy (CIS 1.8/1.9): >=14 chars, full complexity, no reuse of last 24
aws iam update-account-password-policy \
--minimum-password-length 14 \
--require-symbols --require-numbers \
--require-uppercase-characters --require-lowercase-characters \
--allow-users-to-change-password --password-reuse-prevention 24
# M-11 Activate cost-allocation tags (only activates keys already seen on resources)
aws ce update-cost-allocation-tags-status --cost-allocation-tags-status \
'TagKey=Project,Status=Active' 'TagKey=Owner,Status=Active' 'TagKey=Environment,Status=Active'
```
### Centralized root access management (org-level, no CloudFormation resource)
**STATUS: ENABLED 2026-07-14, all member root credentials DELETED** (evidence:
`~/Documents/repositories/_audits/centralized-root-access-evidence-2026-07-14.md`).
Member accounts have NO root credentials; the only root path is a privileged
session from the management account. The management account's own root is NOT
centrally manageable and stays password+MFA hardened.
```bash
# Enable (mgmt account). ORDER MATTERS: trusted access must be enabled
# explicitly first — enable-organizations-root-credentials-management does
# NOT auto-enable it (fails ServiceAccessNotEnabledException).
aws organizations enable-aws-service-access --service-principal iam.amazonaws.com
aws iam enable-organizations-root-credentials-management
aws iam enable-organizations-root-sessions
# Periodic verification (add to governance checks): expect BOTH features
aws iam list-organizations-features
```
**Audit / delete member root credentials** (task-scoped root sessions, 15-min):
```bash
aws sts assume-root --target-principal <acct> \
--task-policy-arn arn=arn:aws:iam::aws:policy/root-task/IAMAuditRootUserCredentials
# then, with the session creds (no --user-name; root has none):
# get-login-profile / list-mfa-devices / list-access-keys / list-signing-certificates
aws sts assume-root --target-principal <acct> \
--task-policy-arn arn=arn:aws:iam::aws:policy/root-task/IAMDeleteRootUserCredentials
# delete-login-profile; deactivate-mfa-device --serial-number <arn>
# GOTCHA: the delete task policy explicitly DENIES iam:DeleteVirtualMFADevice —
# remove the orphaned virtual-device OBJECT via OrganizationAccountAccessRole.
# DONE = four surfaces clear: login profile NoSuchEntity, MFA/keys/certs all empty.
```
**Root recovery runbook** (proven by drill on prod 2026-07-14):
1. `deny-root-user` (p-2idoxozz) DENIES root sessions in every covered OU
(SCPs evaluate `sts:AssumeRoot` sessions — the principal is the member
root ARN). Recovery therefore starts with a **manual, temporary detach**
(`aws organizations detach-policy` — NOT a CDK deploy), timeboxed minutes.
2. GOTCHA (inheritance): p-2idoxozz is attached to `workloads` AND its child
OUs — for an account under workloads/, detach from BOTH the child OU and
workloads, or the inherited deny still applies. Allow ~10s propagation.
3. Freeze deploys of `seahaven-org-governance` for the window (a concurrent
deploy would re-attach mid-recovery); verify no CD run in flight first.
4. `aws sts assume-root --target-principal <acct> --task-policy-arn
arn=arn:aws:iam::aws:policy/root-task/IAMCreateRootUserPassword` →
`create-login-profile` (no args) restores a login profile.
5. Do the root-only task, DELETE the credentials again (four-surface verify),
reattach the SCP(s), confirm `list-targets-for-policy` matches the
pre-detach capture and stack drift is IN_SYNC.
6. extdev extra: `external-dev-iam-guardrails` also denies
`iam:CreateLoginProfile` — recovery there needs that SCP temporarily
detached too. The extdev OU sits at the **5-SCP hard quota**: any new
guardrail for extdev must attach at the ACCOUNT (396287094661) or
consolidate into an existing policy.
**New-account flow (supersedes root-harden-before-OU-move):** create the
account at the org ROOT → it has no root credentials from birth (verify with
the audit session) → bootstrap + deploy role + baseline → `move-account` into
the target OU → verify SCP inheritance + region-lock canary. No mailbox or
MFA enrollment step. Root-usage monitoring: GuardDuty
`Policy:IAMUser/RootCredentialUsage` + CIS 4.3 alarm remain active.
### Delegated security administration (Phase 3, no CloudFormation resource)
Account **seahaven-security (001520130573)** is the org's delegated
administrator for the detective services. **STATUS: APPLIED 2026-07-14,
verified** (see evidence below). The hard preconditions were enforced before
the first delegation call (security review SEC-BASE-B/D — never delegate to an
account with unhardened root or before its baseline stack exists):
- Baseline stack `UPDATE_COMPLETE`; root `AccountMFAEnabled: 1`; account
parent `ou-nbuj-v0s9630u` with SCPs deny-root-user +
protect-security-baseline + security-guardrails inherited.
Verification evidence (2026-07-14):
- `organizations list-delegated-administrators` → `001520130573` (all five
service principals registered).
- GuardDuty: `AutoEnableOrganizationMembers: ALL`; members 328440206208 +
396287094661 both `Enabled`.
- Security Hub: org auto-enable on; both members `Enabled`.
- Org Access Analyzer `seahaven-org-analyzer` created; Config org aggregator
`seahaven-org-aggregator` (AllAwsRegions) on the Config SLR; Inspector2
auto-enable ec2/ecr/lambda + both members associated.
- **Findings flow verified end-to-end:** GuardDuty sample findings created in
member 396287094661 were listed and fully readable from the admin detector
in 001520130573 (`Recon:EC2/PortProbeUnprotectedPort`, AccountId
396287094661, sample=true), then archived.
The delegation runbook (all calls idempotent, run from the **management
account**):
```bash
# GuardDuty: delegate + auto-enable all org members (adopts existing detectors)
aws guardduty enable-organization-admin-account --admin-account-id 001520130573
# then AS 001520130573: update-organization-configuration --auto-enable-organization-members ALL
# + create-members for pre-existing accounts (mgmt, external-dev)
# Security Hub: delegate + auto-enable new members
aws securityhub enable-organization-admin-account --admin-account-id 001520130573
# then AS 001520130573: update-organization-configuration --auto-enable
# IAM Access Analyzer: delegate + ORGANIZATION-scoped analyzer
aws organizations register-delegated-administrator \
--account-id 001520130573 --service-principal access-analyzer.amazonaws.com
# then AS 001520130573: create-analyzer --type ORGANIZATION
# Config: delegate the aggregator (recorders stay per-account in the baselines;
# the aggregator's recorder-status view is the drift detector)
aws organizations register-delegated-administrator \
--account-id 001520130573 --service-principal config.amazonaws.com
# then AS 001520130573: put-configuration-aggregator --organization-aggregation-source
# Inspector2: delegate + associate members
aws inspector2 enable-delegated-admin-account --delegated-admin-account-id 001520130573
```
ONLY once delegation is live AND auto-enrollment is verified (a new member
shows enrolled in the security account's GuardDuty/Security Hub consoles):
new member accounts are then detected/enrolled automatically, and future
member baselines can drop per-account GuardDuty/SecurityHub resources.
Until then, every member baseline MUST keep them (slimming the existing
member stacks is a separate, verification-gated change; note the DA
account's own CFN-owned detector/hub become co-managed after delegation —
never rename/remove them via CFN while the account is delegated admin).
Accepted read-surface note (SEC-BASE-I): the org Config aggregator +
ORGANIZATION Access Analyzer give principals in 001520130573 org-wide READ of
resource configurations (including recorded Lambda env vars) and IAM policies.
Main-branch write access to this repo therefore implies that read surface —
verify no prod Lambda keeps secrets in env vars before creating the
aggregator, and keep branch protection tight.
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
**L-8 (billing-metrics preference) is OUTSTANDING — console only.** Enabling the
CloudWatch `EstimatedCharges` metric in us-east-1 requires turning on *Receive
Billing Alerts* under Billing → Billing preferences; there is no public API/CLI.
fix: Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm (#36) * Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm CIS 4.1 (cis-UnauthorizedAPICalls) flapped OK<->ALARM 15 times in 30 days, all from benign AWS-service AccessDenied noise (CloudFormation deploy/drift describe-scans, AWS Config recorder). A single CFN run on 2026-07-07 emitted 100+ such denials in 15 min, tripping the alarm and burying the real CIS 4.1 security signal in email noise (alert fatigue). - Group both error codes so the exclusions apply to the whole filter (the old pattern leaked the UnauthorizedOperation branch past the exclusions due to && binding tighter than ||). - Exclude denials whose sourceIPAddress is an AWS service host (*.amazonaws.com) — AWS acting on our behalf, not a principal of concern. Real unauthorized calls from a console/CLI/attacker present a routable IP and are still counted. Validated against the trail log group: spike window 107 -> 4 matches, the 4 remaining all from a routable admin IP (genuine activity CIS should retain). - Keep the 3/3 evaluation as a backstop against one-off human fat-fingers. Billing: deleted the manually-created AWS-MonthlyBilling CloudWatch alarm ($50 threshold on EstimatedCharges, routed to site-alerts). It was unmanaged drift, permanently in ALARM, and fully redundant with the managed M-10 budget (seahaven-monthly-cost). README updated with rationale + restore command. * Address sh-security-review: scope CIS 4.1 exclusion to named benign sources The high-recall security review (detector fan-out + proof-or-kill verifier) confirmed a MEDIUM detection blind spot in the first revision: excluding all `*.amazonaws.com` source hosts would hide denials driven through ANY AWS service (SSM Automation, Step Functions, Lambda, etc.), which CloudTrail records with that service's host as sourceIPAddress — i.e. service-proxied privesc/recon attempts would evade CIS 4.1. Remediation: scope the exclusion to the specific benign sources that actually flap this account — `*cloudformation.amazonaws.com` (covers both cloudformation. and hooks.cloudformation.) and `config.amazonaws.com` — plus the pre-existing delivery.logs exclusion. Every other service-proxied denial is now retained. Residual (accepted, documented inline): CloudFormation/Config- proxied denials are still excluded — that path needs near-admin privilege (CreateStack + PassRole), successful changes still trip the other CIS 4.x alarms, and GuardDuty backstops. Validated on the live trail log group: spike window still 107 -> 4 matches (identical noise suppression), the 4 from a routable admin IP. tsc + synth clean.
2026-07-07 15:32:10 -04:00
The M-10 budget (`seahaven-monthly-cost`, 80%/100% actual + 100% forecast)
provides cost alerting independent of that metric.
> The legacy, manually-created `AWS-MonthlyBilling` CloudWatch alarm ($50
> threshold on `EstimatedCharges`, routed to `site-alerts`) was **deleted
> 2026-07-07** as unmanaged drift: it was fully redundant with the M-10 budget,
> sat permanently in ALARM (spend has far exceeded $50/mo), and was never in
> IaC. Billing alerting is now solely the managed M-10 budget. To restore the
> old alarm if ever needed: `aws cloudwatch put-metric-alarm --alarm-name
> AWS-MonthlyBilling --namespace AWS/Billing --metric-name EstimatedCharges
> --dimensions Name=Currency,Value=USD --statistic Maximum --period 86400
> --evaluation-periods 1 --threshold 50 --comparison-operator GreaterThanThreshold
> --alarm-actions arn:aws:sns:us-east-1:328440206208:site-alerts`.
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
### Monitoring + logging (audit Day 2)
| Resource | Logical ID | Finding | Notes |
|---|---|---|---|
| CIS metric filters + alarms | `CisMonitoring/*` | H-1 | 15 filters (CIS 4.1–4.15) on the CloudTrail log group, each with an alarm → `seahaven-cis-alarms`. ALARM-only actions (no OK). 4.16 = Security Hub (Day 1) |
| CIS alarm topic | `seahaven-cis-alarms` | H-1 | SNS, SSE (`alias/aws/sns`), email sub to adam@seahavenind.com |
| VPC flow logs | `FlowLogs/FlowLog0..4` | H-14 | ALL traffic on all 5 VPCs → S3 |
| Flow-logs bucket | `seahaven-vpc-flow-logs-328440206208` | H-14 | Private, SSE-S3, TLS-only, Glacier @90d / expire @365d; delivery bucket policy cross-reviewed |
| SES config set | `seahaven-email-events` | M-13 | Bounce/complaint/reject → CloudWatch metrics for reputation visibility |
2026-06-08 19:04:36 -04:00
| Sensitive-logs CMK | `LogsKey/Key` (`alias/seahaven-logs`) | M-24 | Encrypts sensitive CloudWatch Logs groups. Key policy grants `logs.us-east-1.amazonaws.com` Encrypt*/Decrypt*/ReEncrypt*/GenerateDataKey*/DescribeKey scoped by `kms:EncryptionContext:aws:logs:arn` (required or log delivery breaks). Rotation on, RETAIN. Applied in place to `TrailLogGroup` via escape hatch (same logical id/name). Cross-reviewed |
**M-24 sensitive log groups:** `alias/seahaven-logs` encrypts the CloudTrail CW
log group (codified here) plus the finance/PII Lambda groups owned by other
stacks — `exec-aide-*`, `payments-*`, `po-email-processor`, `vendor-reply-processor`
— which are associated via `aws logs associate-kms-key` and tracked as drift to
codify in their owning repos. The CloudTrail group is encrypted in place (escape
hatch on the existing `AWS::Logs::LogGroup`) so it is additive: same logical id +
physical name, no replacement, CIS Section-4 metric filters keep working. A
context flag `encryptTrailLogGroup` (default `true`) allows rolling the CMK out
and smoke-testing it on a low-risk Lambda group before the CloudTrail group:
```bash
# Phase 1: deploy CMK only, validate on a low-risk group
cdk deploy account-baseline --context encryptTrailLogGroup=false
# Phase 2: encrypt the CloudTrail group (default)
cdk deploy account-baseline
```
**H-1 log group:** the metric filters attach to the existing CloudTrail
CloudWatch Logs group by name (`seahaven-account-baseline-TrailLogGroup4CBE3AF5-…`),
imported read-only so the live audit trail is never replaced. Stable unless the
Trail is recreated.
**H-14 bucket policy note:** the flow-logs delivery policy keeps
`s3:x-amz-acl=bucket-owner-full-control` and the `arn:aws:logs:…:*` source-ARN
wildcard — both are required by AWS's documented flow-logs-to-S3 policy
(`flow-logs-s3-permissions.html`). A cross-review suggested dropping them; that
was rejected as it would break delivery. `s3:ListBucket` was dropped (not needed).
**M-13 follow-up:** associate `seahaven-email-events` as the default config set
on the live sending identities to capture events from existing senders:
```bash
aws sesv2 put-email-identity-configuration-set-attributes \
--email-identity int.seahaven.com --configuration-set-name seahaven-email-events
```
### Log-group retention + alarm wiring (audit L-4, L-5)
Applied via CLI (auto-created groups spread across stacks; one alarm in another
stack). Applied 2026-06-02.
```bash
# L-4 90-day retention on the 13 never-expire log groups (CodeBuild + CDK helpers)
for lg in <the 13 groups>; do aws logs put-retention-policy --log-group-name "$lg" --retention-in-days 90; done
# L-5 wire the actionless forgejo backup-verification alarm to site-alerts
aws cloudwatch put-metric-alarm --alarm-name forgejo-backup-verification-errors \
--alarm-actions arn:aws:sns:us-east-1:328440206208:site-alerts # (preserve existing alarm config)
```
## Roadmap (same stack)
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
Detective layer multi-region expansion (GuardDuty/Config/Security Hub beyond
us-east-1). Backup: phase 2 is deployed (see above); remaining is migrating the
phase-2 EBS entries to tag-based selection (INFRA-31) and graduating the offsite
vault to compliance mode.
## Deploy
CI/CD via the org reusable workflows (`ci-typescript-cdk.yaml`,
`cd-cdk.yaml`); pushes to `main` deploy through the OIDC role in
`secrets.AWS_DEPLOY_ROLE_ARN`. Local: `npm ci && npm run build && npx cdk diff`.
```
npx cdk deploy seahaven-account-baseline
```
## Verify
```
aws cloudtrail get-trail-status --name seahaven-org-trail # IsLogging: true
aws cloudtrail describe-trails --trail-name-list seahaven-org-trail
aws cloudtrail validate-logs --trail-arn <arn> --start-time <t> # digest integrity
```
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
AWS Backup (C-7):
```
aws backup list-backup-vaults # seahaven-primary
aws backup list-backup-vaults --region us-west-2 # seahaven-offsite
aws backup describe-backup-vault --backup-vault-name seahaven-offsite --region us-west-2 # Locked, MinRetentionDays
aws backup get-backup-plan --backup-plan-id <id> # daily rule + CopyAction
# Smoke test: on-demand backup of one resource, then confirm the cross-region copy lands
aws backup start-backup-job --backup-vault-name seahaven-primary \
--resource-arn arn:aws:rds:us-east-1:328440206208:db:proposal-system-db \
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
--iam-role-arn arn:aws:iam::328440206208:role/seahaven-backup-service-role
aws backup list-copy-jobs --region us-west-2 # copy to offsite present + COMPLETED
# Phase-2 selections live on the plan:
aws backup list-backup-selections --backup-plan-id <id> --query 'BackupSelectionsList[].SelectionName' # critical-data + phase2-offsite-everything
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
```
Account detective layer + budget (audit Day 1) (#5) * Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10) Adds to the seahaven-account-baseline stack: - AWS Config recorder (all + global resources) + delivery channel + role + hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed. - GuardDuty detector, us-east-1 (H-3) - Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4) - IAM Access Analyzer, account scope (M-5) - Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to adam@seahavenind.com (M-10) Scope us-east-1 only (all workloads here); multi-region is a follow-up. The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented separately in the README runbook. * Document Day 1 detective layer + CLI governance toggles in README * Move Config recorder+channel to CLI (L1 stabilization deadlock) The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches CREATE_COMPLETE until recording is active (needs a delivery channel), and the delivery channel cannot be created until the recorder completes — a deadlock that hung the deploy ~27 min before manual cancel (2026-06-01). Keep the cross-reviewed recorder role + delivery bucket in IaC; create the recorder, delivery channel, and start recording via CLI (documented in README). Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
Detective layer + governance (Day 1):
```
aws configservice describe-configuration-recorder-status # recording: true
aws guardduty list-detectors # one detector id
aws securityhub get-enabled-standards # FSBP + CIS v3.0.0
aws accessanalyzer list-analyzers # seahaven-account-analyzer ACTIVE
aws inspector2 batch-get-account-status --region us-east-1 # ec2/ecr/lambda ENABLED
aws iam get-account-password-policy # length 14, reuse 24
aws ec2 get-ebs-encryption-by-default --region us-east-1 # EbsEncryptionByDefault: true
aws budgets describe-budgets --account-id 328440206208 # seahaven-monthly-cost $1,200
aws ce list-cost-allocation-tags --status Active # Project/Owner/Environment Active
```