2026-07-14 13:53:07 -04:00
# seahaven-org-baseline
2026-05-29 17:44:55 -04:00
2026-06-11 14:42:34 -04:00


2026-07-14 13:53:07 -04:00

Organization-wide security and governance baseline for Sea Haven Industries,
managed as a single CDK TypeScript app. Covers the management account
(**328440206208**: primary baseline in **us-east-1** , secondary-region
baselines in **us-east-2** /**us-west-2**, offsite backup vault in
**us-west-2**) and org **member accounts** (first tenant:
`seahaven-external-dev` **396287094661** , absorbed from the retired
`seahaven-external-dev-baseline` repo). This is where account-wide detective
and recovery controls live, so they are versioned, reviewed, and drift-checked
like any other stack.
> **History:** this repo was `seahaven-account-baseline` (management account
> only) until 2026-07-14, when the external-dev member baseline was merged in
> and the repo renamed. Deployed CloudFormation stack names are unchanged.
2026-08-29 21:04:40 +00:00
Stacks (normally deployed by one CD job per target account; staged exceptions
are noted):
2026-07-14 13:53:07 -04:00
| Stack | Account | Region | Purpose |
|---|---|---|---|
| `seahaven-account-baseline` | 328440206208 | us-east-1 | CloudTrail + detective controls (C-1) |
| `seahaven-dynamodb-cmk` | 328440206208 | us-east-1 | Shared customer-managed KMS key for finance/PII DynamoDB tables; ARN published to SSM `/seahaven/dynamodb/cmk-arn` (INFRA-95 / M-3) |
| `seahaven-regional-baseline-us-west-2` | 328440206208 | us-west-2 | Bedrock invocation logging (INFRA-91) |
| `seahaven-regional-baseline-us-east-2` | 328440206208 | us-east-2 | Bedrock invocation logging + AWS Config recorder + Security Hub (INFRA-91 / INFRA-16) |
| `seahaven-backup` | 328440206208 | us-east-1 | Primary AWS Backup vault + plan + role (C-7) |
| `seahaven-backup-offsite` | 328440206208 | us-west-2 | Governance-locked offsite copy vault (C-7) |
2026-07-14 17:17:55 -04:00
| `seahaven-org-governance` | 328440206208 | us-east-1 | AWS Organizations OU tree + SCPs (incl. the cdk-imported external-dev guardrails) |
2026-07-14 13:53:07 -04:00
| `seahaven-external-dev-baseline` | 396287094661 | us-east-1 | Member-account baseline: Config, GuardDuty, Security Hub (FSBP + CIS v3.0), Access Analyzer, flow logs, budget |
2026-08-29 21:04:40 +00:00
| `seahaven-terraform-substrate` | 396287094661 | us-east-1 | Staged manually until SHOC role imports complete; HCP roles/deploy boundaries for the backend rehearsal, referencing the existing OIDC provider |
2026-07-14 15:32:50 -04:00
| `seahaven-security-baseline` | 001520130573 | us-east-1 | Member-account baseline for the delegated security-admin account (same construct set) |
2026-07-14 16:41:36 -04:00
| `seahaven-dev-baseline` | 710827005802 | us-east-1 | Member-account baseline for internal dev/staging (org-managed detection — no local GuardDuty/SecurityHub) |
2026-07-14 17:17:55 -04:00
| `seahaven-prod-baseline` | 011934824531 | us-east-1 | Member-account baseline for production workloads (org-managed detection; mgmt account frozen for new workloads) |
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
2026-07-10 16:07:20 -04:00
## CDK app
The repo is a single AWS CDK app written in TypeScript. `cdk.json` is the
project config the `cdk` CLI reads on every command: its `app` key
(`npx ts-node bin/app.ts` ) tells CDK how to synthesize the app straight from
the TypeScript source — no separate compile step needed for `cdk synth` /
`diff` / `deploy` — and its `context` block carries the AWS CDK feature flags.
| Path | Role |
|---|---|
| `cdk.json` | CDK config: `app` synth command, `watch` includes/excludes, `context` feature flags |
2026-07-14 17:17:55 -04:00
| `bin/app.ts` | App entry point — instantiates every stack with an explicit kebab-case `stackName` and its target `env` (five accounts, per-account/per-region) |
2026-07-10 16:07:20 -04:00
| `lib/*-stack.ts` | Stack definitions (one class per stack; larger stacks compose the constructs in `lib/*.ts` ) |
| `tsconfig.json` | TypeScript compiler options (`outDir: cdk.out` ) |
| `package.json` | Pinned `aws-cdk-lib` , CDK CLI, and the `build` / `synth` / `diff` / `deploy` npm scripts |
2026-07-27 16:24:09 -04:00
`bin/app.ts` synthesizes fifteen stacks across three regions and five accounts:
2026-07-10 16:07:20 -04:00
2026-07-14 13:53:07 -04:00
| Construct id | Stack name | Account | Region | Source |
|---|---|---|---|---|
| `account-baseline` | `seahaven-account-baseline` | 328440206208 | us-east-1 | `lib/account-baseline-stack.ts` |
| `dynamodb-cmk` | `seahaven-dynamodb-cmk` | 328440206208 | us-east-1 | `lib/dynamodb-cmk-stack.ts` |
| `regional-baseline-us-west-2` | `seahaven-regional-baseline-us-west-2` | 328440206208 | us-west-2 | `lib/regional-baseline-stack.ts` |
| `regional-baseline-us-east-2` | `seahaven-regional-baseline-us-east-2` | 328440206208 | us-east-2 | `lib/regional-baseline-stack.ts` |
| `backup-offsite` | `seahaven-backup-offsite` | 328440206208 | us-west-2 | `lib/backup-offsite-stack.ts` |
| `backup` | `seahaven-backup` | 328440206208 | us-east-1 | `lib/backup-stack.ts` |
2026-07-14 17:17:55 -04:00
| `org-governance` | `seahaven-org-governance` | 328440206208 | us-east-1 | `lib/org-governance-stack.ts` |
2026-07-14 13:53:07 -04:00
| `external-dev-baseline` | `seahaven-external-dev-baseline` | 396287094661 | us-east-1 | `lib/member-baseline-stack.ts` |
2026-07-14 15:32:50 -04:00
| `security-baseline` | `seahaven-security-baseline` | 001520130573 | us-east-1 | `lib/member-baseline-stack.ts` |
2026-07-14 16:41:36 -04:00
| `dev-baseline` | `seahaven-dev-baseline` | 710827005802 | us-east-1 | `lib/member-baseline-stack.ts` (orgManagedDetection) |
2026-07-14 17:17:55 -04:00
| `prod-baseline` | `seahaven-prod-baseline` | 011934824531 | us-east-1 | `lib/member-baseline-stack.ts` (orgManagedDetection) |
2026-07-27 16:24:09 -04:00
| `deploy-substrate-prod` | `seahaven-deploy-substrate` | 011934824531 | us-east-1 | `lib/deploy-substrate-stack.ts` |
| `deploy-substrate-dev` | `seahaven-deploy-substrate` | 710827005802 | us-east-1 | `lib/deploy-substrate-stack.ts` |
2026-08-29 21:04:40 +00:00
| `terraform-substrate-external-dev` | `seahaven-terraform-substrate` | 396287094661 | us-east-1 | `lib/terraform-substrate-stack.ts` |
2026-07-27 16:24:09 -04:00
| `dynamodb-cmk-prod` | `seahaven-dynamodb-cmk` | 011934824531 | us-east-1 | `lib/dynamodb-cmk-stack.ts` |
| `alarm-topic-prod` | `seahaven-alarm-topic` | 011934824531 | us-east-1 | `lib/alarm-topic-stack.ts` |
2026-08-07 17:07:04 -04:00
| `app-web-acl-prod` | `seahaven-app-web-acl` | 011934824531 | us-east-1 | `lib/app-web-acl-stack.ts` |
2026-07-14 16:41:36 -04:00
Member-account stacks deploy with per-account credentials — the CD workflow
runs one job per account, each assuming that account's OIDC deploy role. Local
deploys/diffs assume `OrganizationAccountAccessRole` in the target account.
| Account | OIDC deploy role | Repo secret |
|---|---|---|
| 396287094661 (external-dev) | `githubdeploy-seahaven-external-dev-baseline` | `AWS_DEPLOY_ROLE_ARN_EXTDEV` |
| 001520130573 (security) | `githubdeploy-seahaven-org-baseline` | `AWS_DEPLOY_ROLE_ARN_SECURITY` |
| 710827005802 (dev) | `githubdeploy-seahaven-org-baseline` | `AWS_DEPLOY_ROLE_ARN_DEV` |
2026-07-14 17:17:55 -04:00
| 011934824531 (prod) | `githubdeploy-seahaven-org-baseline` | `AWS_DEPLOY_ROLE_ARN_PROD` |
2026-07-14 16:41:36 -04:00
Shared constructs (`DetectiveControls` , `FlowLogs` , `GovernanceToggles` ) are
prefix-parameterized — construct ids and physical names must stay
byte-identical to the deployed stacks (logical IDs are path-derived).
Accounts enrolled by the org delegated admin (post 2026-07-14) set
`orgManagedDetection: true` : the GuardDuty detector + Security Hub hub come
from the org, while STANDARDS (FSBP + CIS v3.0) and the account analyzer stay
CFN-owned (org `AutoEnableStandards` is `NONE` — the DEFAULT setting enrolls
legacy CIS v1.2.0). Enrollment (member `Enabled` in GuardDuty + Security Hub)
is a hard precondition for such a stack's first deploy.
2026-07-10 16:07:20 -04:00
`backup` declares an explicit dependency on `backup-offsite` so the offsite copy
vault exists before the primary plan that copies into it. Stack names are set
explicitly to enforce kebab-case (CDK defaults to PascalCase). Synthesized
CloudFormation templates land in `cdk.out/` (git-ignored).
Common commands:
```
npm ci # install pinned deps
npm run build # tsc type-check (compiles to cdk.out/)
npx cdk synth # synthesize CloudFormation for all stacks
npx cdk diff # diff synthesized stacks against deployed state
npx cdk deploy --all # deploy every stack
npx cdk deploy < stack-name > # deploy a single stack
```
The `--context <key>=<value>` flag overrides `cdk.json` context at the command
line (e.g. the `encryptTrailLogGroup` toggle under *Monitoring + logging* ).
2026-07-06 17:44:22 -04:00
## Documentation
The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This project's `seahaven-account-baseline` , `seahaven-backup` , and `seahaven-backup-offsite` stacks are represented there as Mermaid subgraphs.
- **[AWS Architecture Map ](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098 )** (Confluence, IT space, page 1540098)
2026-05-29 17:44:55 -04:00
## What it deploys
2026-07-27 16:24:09 -04:00
### GitHub Actions deploy substrate (per account)
`lib/deploy-substrate-stack.ts` + `lib/deploy-substrate/deploy-substrate.template.yaml`
deploy `seahaven-deploy-substrate` into each member account that hosts SAM
workloads (currently seahaven-prod and seahaven-dev). It contains the shared
account-level deploy plumbing:
2026-08-13 16:47:53 -04:00
- the `seahaven-lambda-execution-boundary` permissions boundary (legacy
shared ceiling for roles not yet retargeted) plus per-workload policies
`seahaven-lambda-execution-boundary-<workload>` (PLAT-52),
2026-07-27 16:24:09 -04:00
- the `github-cfn-execution-role` CloudFormation execution role that `cd-sam`
callers pass as `cfn-role-arn` ,
- optionally the GitHub OIDC identity provider (`createOidcProvider: true` ,
only for an account that does not already have one — one provider per URL
per account).
2026-07-30 18:03:40 -04:00
The template began as a verbatim extraction of the substrate section of
`Sea-Haven-Industries/.github/oidc-deploy-roles.yaml` , which remains the source
of truth for mgmt (328440206208) until its stacks migrate out.
**The two copies are no longer at parity, and the old "edit both files" rule no
longer applies uniformly.** Under INFRA-186, `seahaven-lambda-execution-boundary`
2026-08-13 16:47:53 -04:00
in *this* copy was reduced to a fleet-wide floor for prod and dev; later
migrations packed per-workload data plane back into that shared document until
it hit the 6,144-character cap (PLAT-93 / PLAT-100). PLAT-52 adds per-workload
policies `seahaven-lambda-execution-boundary-<workload>` (floor plus that
stack's data plane) and switches both prod/dev guardrails to a StringEquals
allow-list of the shared ARN plus each per-workload ARN. The shared document
is left unchanged until live roles retarget. Mgmt's copy keeps the
account-wide wildcards and a **single-ARN** pin pending its own separately
validated rollout across 26 live boundary-carrying roles (PLAT-51). So: **the
boundary resource is deliberately divergent**, and the guardrail
`iam:PermissionsBoundary` condition **cardinality** is also divergent
(enumerated list here, scalar on mgmt). Do not weaken mgmt to ArnLike. Every
*other* substrate resource (`github-cfn-execution-role` ,
`seahaven-cfn-exec-iam-management` Sid/Action/Resource sets) is still expected
to change in both files together. The template's provenance header records
which is which — read it before assuming either parity or divergence.
2026-07-30 18:03:40 -04:00
Per-repo `githubdeploy-*` deploy roles are deliberately NOT part
2026-07-27 16:24:09 -04:00
of the substrate — they are provisioned per repo at migration/onboarding time
so an account never carries trust relationships for repos that do not deploy
to it.
2026-07-27 19:15:08 -04:00
**Escalation controls on `github-cfn-execution-role` .** Every `iam:CreateRole` ,
`AttachRolePolicy` and `PutRolePolicy` is conditioned on the target carrying
2026-08-13 16:47:53 -04:00
one of the enumerated `seahaven-lambda-execution-boundary` ARNs (the shared
policy plus each `seahaven-lambda-execution-boundary-<workload>` ). That
condition alone is not sufficient,
2026-07-27 19:15:08 -04:00
so the attached `seahaven-cfn-exec-iam-management` managed policy also carries
three explicit Deny statements:
- `DenyBoundaryTampering` — no removing a boundary from any role or user.
Granting the delete under the same `StringEquals` condition self-defeats the
gate, because for a delete the condition key resolves to the boundary already
on the target.
- `DenyBoundaryPolicyEdit` — no rewriting any `seahaven-*` managed policy.
- `DenySelfMutation` — the role cannot modify or delete itself or any
`githubdeploy-*` role. Without it the control is one API call from being
undone: `IAMRoleReadAndDelete` grants `iam:DetachRolePolicy` on `Resource:
"*"` unconditioned, so the role could detach the very policy carrying these
Denies.
Verify a change to these with `aws iam simulate-principal-policy` against the
role's own ARN (expect `explicitDeny` ) and against a `<stack>-<Function>Role-`
name (expect `allowed` , no regression for normal SAM deploys). Note that
simulation currently does **not** see this role's *inline* policies in
seahaven-prod or seahaven-dev — read those back with `get-role-policy` instead.
Known consequence of `DenyBoundaryTampering` : a CloudFormation rollback of an
update that *adds* a boundary to an existing role wedges in
`UPDATE_ROLLBACK_FAILED` . Recovery is an administrator action, not a pipeline
retry — `aws cloudformation continue-update-rollback --stack-name < stack >
--resources-to-skip < RoleLogicalId > `. Unreachable while every SAM role is
created with the boundary already attached.
2026-07-27 16:24:09 -04:00
**Onboarding a future account as a deploy target:**
1. CDK-bootstrap the account (`npx cdk bootstrap aws://<account>/us-east-1`
via `OrganizationAccountAccessRole` ).
2. Create `githubdeploy-seahaven-org-baseline` in the account (same trust and
policy as the dev/prod copies) and add the repo secret
`AWS_DEPLOY_ROLE_ARN_<ACCT>` .
3. Add a `DeploySubstrateStack` instance in `bin/app.ts`
(`createOidcProvider: true` if the account has no GitHub OIDC provider)
and append its construct id to a new per-account job in
`.github/workflows/deploy.yaml` (explicit `stacks` selector, one job per
account).
4. Merge; the substrate deploys via CD. Per-repo deploy roles and app stacks
follow the cross-account migration playbook from there.
2026-07-30 16:31:34 -04:00
### Terraform deploy substrate (per account)
`lib/terraform-substrate-stack.ts` + `lib/terraform-substrate/terraform-substrate.template.yaml`
deploy `seahaven-terraform-substrate` into each member account that hosts
2026-08-29 21:04:40 +00:00
Terraform-managed workloads (currently seahaven-prod, seahaven-dev, and
external-dev; never mgmt — mgmt stays SAM until its stacks migrate out).
Prod/dev use the shared IAM-management policy. External-dev references its
existing `app.terraform.io` provider and carries only exact SHOC
import/adoption roles:
2026-07-30 16:31:34 -04:00
- the `app.terraform.io` OIDC identity provider (audience
`aws.workload.identity` ; Retain — it is the federation anchor for every
future `hcptf-*` role),
- the `seahaven-hcptf-iam-management` guardrail policy: the boundary-gated
2026-08-13 16:47:53 -04:00
IAM role lifecycle (conditioned on the enumerated
`seahaven-lambda-execution-boundary` allow-list owned by the
deploy-substrate stack — hence the explicit stack dependency in
`bin/app.ts` ) plus the `DenyBoundaryTampering` / `DenyBoundaryPolicyEdit`
2026-08-29 21:04:40 +00:00
/ `DenySelfMutation` backstops,
- external-dev-only deploy boundaries
`shoc-backend-{tf-poc,dev,staging}-deploy-boundary` . Each is the maximum
current policy for one exact `githubdeploy-shoc-backend-*` role. Dev
temporarily retains its live broad `elasticbeanstalk-*` S3 grants so
boundary attachment cannot regress deployment before the separately
reviewed policy narrowing,
- external-dev-only runtime boundaries
`shoc-backend-{tf-poc,dev,staging}-runtime-boundary` . These retain only the
account-scoped S3, environment health/log, and X-Ray portions of
`AWSElasticBeanstalkWebTier` , plus each environment's exact secrets/KMS/STS
data plane. They deliberately exclude the managed policy's 2026
Bedrock/Marketplace additions.
2026-07-30 16:55:45 -04:00
**This policy derives from `seahaven-cfn-exec-iam-management` but is
deliberately stricter — it is not a mirror.** The 2026-07-30 security review
confirmed the SAM copy's `Resource: "*"` role grants as a critical escalation
primitive (`iam:UpdateAssumeRolePolicy` on `*` repoints the AdministratorAccess
CDK bootstrap role's trust policy to an external account), and its justification
for the wildcard — SAM auto-generates execution roles at path `/` with no
settable `RolePath` — does not transfer, because Terraform's `aws_iam_role`
supports `path` . So here:
- every role **write** (create, delete, detach, `UpdateAssumeRolePolicy` ,
boundary set) and `iam:PassRole` is confined to the Terraform-owned path
`role/tf-managed/*` ; reads stay on `*` for data sources,
- **Terraform configs must set `path = "/tf-managed/"` on every
`aws_iam_role` ** — a role created anywhere else is denied,
- `DenySelfMutation` additionally covers `cdk-hnb659fds-*` ,
`OrganizationAccountAccessRole` and `seahaven-*` (detective-control roles,
which no prod/nonprod SCP shields from `iam:DeleteRole` ).
Do not "reconcile" the two files by copying statements between them. The
durable org-level fix for the same class is extending the existing
`ProtectPrivilegedRoles` SCP (currently security-OU only) to prod and nonprod.
2026-07-30 16:31:34 -04:00
Per-workspace roles (`hcptf-<stack>` apply + `hcptf-<stack>-plan` ) are
deliberately NOT pre-provisioned — they are appended to the template at each
stack's migration time so an account never carries trust for workspaces that
do not deploy to it.
2026-08-29 21:04:40 +00:00
**External-dev SHOC role adoption is a staged CloudFormation import, not a
2026-08-30 20:12:54 +00:00
normal first deploy.** Six roles exist today:
`hcptf-shoc-backend-{dev,staging}` and their `-plan` partners, plus the
`hcptf-shoc-backend-tf-poc` pair. Two independent CDK contexts make each
transition explicit:
`enableShocBackendPocRoles` and `enableShocBackendLiveRoles` . Both began as
`false` for the initial rollout and remain version-controlled as `true` after
their completed ownership transitions.
2026-08-29 21:04:40 +00:00
`terraform-substrate-external-dev` is deliberately absent from the automatic
external-dev deploy job during this sequence; `external-dev-baseline` remains
automatic and unchanged.
1. Create the role-free base stack:
```bash
npx cdk deploy terraform-substrate-external-dev \
-c enableShocBackendPocRoles=false \
-c enableShocBackendLiveRoles=false
```
`CreateOIDCProvider=false` is fixed in `bin/app.ts` ; the external-dev
account therefore creates neither the existing provider, SHOC roles, nor
the prod/dev-only shared IAM policy. The base stack does create all six
retained external-dev deploy/runtime boundary policies.
2. Set `enableShocBackendPocRoles` to `true` in `cdk.json` , leave the live gate
`false` , review the synthesized two-role addition, then run the normal
external-dev stack update. This creates only the new tf-poc HCP plan/apply
pair. The retained POC CDK stack references
`shoc-backend-tf-poc-deploy-boundary` when it creates
`githubdeploy-shoc-backend-tf-poc` and
`shoc-backend-tf-poc-runtime-boundary` when it creates the POC runtime
role; do not attach the generic account execution boundary to either role.
3. Prove both tf-poc HCP assumptions and the retained POC import rehearsal
before touching the live-role ownership boundary.
4. In a separately approved administrator/CDK migration, tag the existing HCP
apply roles first:
`hcptf-shoc-backend-dev` gets
`HcpTerraformWorkspace=shoc-backend-dev` , and
`hcptf-shoc-backend-staging` gets
`HcpTerraformWorkspace=shoc-backend-staging` . Next attach
`shoc-backend-dev-deploy-boundary` and
`shoc-backend-staging-deploy-boundary` to the exact `githubdeploy-*` roles,
and attach `shoc-backend-dev-runtime-boundary` /
`shoc-backend-staging-runtime-boundary` to the exact runtime roles. Verify
the boundary ceilings before adding the matching manager tag to either
target `githubdeploy-*` role. The POC CDK
creates its deploy role with `HcpTerraformWorkspace=shoc-backend-tf-poc` ;
the substrate-created POC apply role already carries the same principal
tag. Verify each effective deployment action before continuing. HCP remains
blocked while a target tag is missing/different or the target lacks its
exact dedicated boundary, so a partial migration cannot authorize policy
writes. Complete both runtime/deploy boundary attachments before workload
imports. HCP apply roles deliberately have no
`iam:PutRolePermissionsBoundary` or boundary-policy mutation permissions.
5. In the backend bootstrap, add Terraform `removed` blocks with
`destroy = false` for only the four dev/staging HCP roles and their inline
policies. Apply and verify Terraform state no longer owns them while all
four physical roles and ARNs remain unchanged.
6. Set both contexts to `true` , synthesize with
`npx cdk synth terraform-substrate-external-dev` , and create a
CloudFormation **IMPORT** change set for the four existing
`AWS::IAM::Role` resources by exact role name. Do not run a normal
CREATE/UPDATE change set for this ownership transition. The POC gate must
remain true so the already-managed pair stays in the template. Import
records ownership; it does not update existing role properties or inline
policies.
7. Run a separate, reviewed CloudFormation reconcile update after import and
before switching workspace credentials. Each current live role has one
inline policy: `shoc-backend-dev-import-plan` ,
`shoc-backend-dev-import-apply` ,
`shoc-backend-staging-import-plan` , or
`shoc-backend-staging-import-apply` . Existing descriptions and tags are
inventoried in the backend handoff. Reconcile those explicit differences
to the final baseline shape without replacing a role.
8. After reconcile and HCP assumption proof, keep both context values committed
as `true` . Every subsequent normal deployment must synthesize all six
roles. Never return either gate to false as a rollback mechanism; Retain
protects the physical role but removing it from the stack abandons
CloudFormation ownership.
9. Only after all imports/reconciliation complete and both context defaults
are permanently `true` , add `terraform-substrate-external-dev` back to the
external-dev workflow stack selector. Until then all substrate operations
are deliberate manual deploy/import actions.
10. Retire the backend bootstrap only after the POC pair and all four imported
live roles are proven under this stack. Role deletion/recreation is never a
migration step.
The external-dev apply roles intentionally omit role create/delete,
managed-policy attach/detach, trust or boundary mutation, `iam:PassRole` , and
secret-value APIs. IAM writes are limited to exact-role inline-policy and
ordinary tag updates plus exact-profile tags; role descriptions remain stable
and HCP receives no `UpdateRole` or `UpdateRoleDescription` . The SCP permits
only the three enumerated HCP apply roles to mutate a `githubdeploy-*` role
whose locked `HcpTerraformWorkspace` resource tag equals the caller's immutable
principal tag. Adding or changing that manager tag remains administrator/CDK
only. POC DNS and certificate access is tag/name constrained because their
physical IDs are allocated by the temporary retained CDK stack before
Terraform imports them. Dev and staging DNS writes are pinned to their existing
hosted-zone IDs and API record names.
Current compact policy-document sizes are 1,387 / 1,873 / 1,844 characters for
the POC/dev/staging deploy boundaries and 1,176 / 1,779 / 1,656 for their
runtime boundaries, each below IAM's 6,144-character managed-policy limit. The
external-dev IAM guardrail SCP is 4,922 compact characters against its
5,120-character Organizations limit; keep size assertions in every change.
2026-07-30 16:31:34 -04:00
**HCP Terraform layout (org-level setup, console):** one org `seahaven`
(free tier: 500 managed resources, 1 concurrent run); one HCP **project per
AWS account** (`seahaven-prod` , `seahaven-dev` ); one **workspace per stack**
(`<stack>-<env>` , one state file = one blast radius). Default execution mode
Remote. Never use HCP's "Quick setup AWS dynamic credentials" button — it
writes the single `TFC_AWS_RUN_ROLE_ARN` , which collapses the plan/apply role
split this substrate exists to enforce.
2026-08-05 16:14:16 -04:00
**Reference implementation:** first workload was `afi-backup-monitor` in
seahaven-prod (PLAT-56). Copy
`Sea-Haven-Industries/afi-backup-monitor` `terraform/` and the live
`hcptf-afi-backup-monitor*` / `hcptf-afi-backup-monitor-plan` statements in
this template rather than inventing new IAM shapes.
2026-07-30 16:31:34 -04:00
**Migration checklist (per stack, in order):**
2026-08-05 16:14:16 -04:00
0. **Freeze the app's SAM/CDK CD** (remove or disable the deploy workflow) so
HCP Terraform becomes the sole deploy path before the first apply. Leave
the source-account stack frozen until cutover.
1. **Secrets first.** Create exact secret shells in the target account; strip
trailing newlines/whitespace before `put-secret-value` (a trailing `\n`
breaks HTTP headers at runtime). Capture ARNs. Never put secret *values*
in Terraform state (ARN references only).
2. **HCP workspace** in the target account's project (`<stack>-<env>` ). Apply
method **Manual** ; automatic speculative plans on if VCS-connected;
working directory `terraform/` . (CLI `terraform plan` runs are inherently
speculative.)
3. **Substrate PR** to this repo appending `hcptf-<stack>-plan` and
`hcptf-<stack>` (see 3a/3b). Trust: this account's `app.terraform.io`
provider; `StringEquals` on `app.terraform.io:aud` =
`aws.workload.identity` and on `app.terraform.io:sub` =
2026-07-30 16:31:34 -04:00
`organization:seahaven:project:seahaven-<env>:workspace:<workspace>:run_phase:plan`
(or `:apply` ). Exact `StringEquals` only — never `StringLike` , never a
wildcarded `run_phase` (a speculative PR plan must never hold write
docs(iam): resolve confirmed review findings from both INFRA-186 gates
Cross-family round 1 plus the /sh-security-review verifier confirmed 11
findings on the floor reduction, all documentation defects; no policy
statement changes. The one HIGH: the Terraform migration checklist never
widened the boundary, so a Lambda-bearing Terraform migration would deploy
green and lose every data-plane call at first invoke. Checklist step 2 now
carries the widening requirement, step 3 verifies deployed boundary content,
and the terraform-substrate header no longer reads as 'Terraform path
unaffected'. Also corrected: Description is a REPLACEMENT property (a
Description edit wedges the custom-named policy and CFN's remedy is the
forbidden rename), the sanctioned-source contradiction, the false
AWSLambdaVPCAccessExecutionRole parity claim, the KMS log-group category
error, stale size numbers (691/5,453), the same-PR widening contradiction,
per-workload residue text, a LoggingConfig silent-log-loss note, the
us-east-1 region pin rationale, and ENI DoS deferral now tracked as
INFRA-200.
2026-07-31 13:44:21 -04:00
credentials). **If the stack creates Lambda execution roles, this same PR
2026-08-13 16:47:53 -04:00
must also add `seahaven-lambda-execution-boundary-<stack>` ** per the
WIDENING PATH in `lib/deploy-substrate/deploy-substrate.template.yaml`
(floor plus that stack's data plane, **exact** secret ARNs from step 1,
no `secret:afi-*` patterns) **and** append that policy's ARN to both
guardrail StringEquals allow-lists. Do not add data-plane to the shared
`seahaven-lambda-execution-boundary` document. The guardrail forces every
Terraform-created role to carry a listed boundary; an unlisted or
floor-only boundary deploys green, then every data-plane call is denied
docs(iam): resolve confirmed review findings from both INFRA-186 gates
Cross-family round 1 plus the /sh-security-review verifier confirmed 11
findings on the floor reduction, all documentation defects; no policy
statement changes. The one HIGH: the Terraform migration checklist never
widened the boundary, so a Lambda-bearing Terraform migration would deploy
green and lose every data-plane call at first invoke. Checklist step 2 now
carries the widening requirement, step 3 verifies deployed boundary content,
and the terraform-substrate header no longer reads as 'Terraform path
unaffected'. Also corrected: Description is a REPLACEMENT property (a
Description edit wedges the custom-named policy and CFN's remedy is the
forbidden rename), the sanctioned-source contradiction, the false
AWSLambdaVPCAccessExecutionRole parity claim, the KMS log-group category
error, stale size numbers (691/5,453), the same-PR widening contradiction,
per-workload residue text, a LoggingConfig silent-log-loss note, the
us-east-1 region pin rationale, and ENI DoS deferral now tracked as
INFRA-200.
2026-07-31 13:44:21 -04:00
at first invoke and async/DLQ writes are discarded silently. IAM roles and
2026-08-13 16:47:53 -04:00
boundary policies = mandatory cross-family review +
2026-07-30 16:31:34 -04:00
`/sh-security-review` on the diff.
2026-08-05 16:14:16 -04:00
3a. **Plan role (required for every stack):** attach
`arn:aws:iam::aws:policy/job-function/ViewOnlyAccess` (never
`ReadOnlyAccess` , which grants `secretsmanager:GetSecretValue` ,
`s3:GetObject` and `kms:Decrypt` and would let any PR-triggered speculative
plan render secret values into HCP run output) **plus** a scoped
plan-refresh sidecar inline policy. ViewOnly alone is insufficient for
Terraform refresh after partial apply — it lacks `iam:GetRole` ,
`events:DescribeRule` , and several Lambda/S3 reads. Sidecar minimum:
`iam:GetRole` / related reads on `role/tf-managed/<prefix>-*` ;
`events:DescribeRule` (and list-targets/tags as needed) on
`rule/<prefix>-*` ; `lambda:*` (or at least the Get*/List* the provider
uses) on `function:<prefix>-*` / `layer:<prefix>-*` ; `s3:Get*` /
`s3:ListBucket` on the stack artifact bucket. **No** IAM writes, **no**
guardrail-policy attach on the plan role. Copy
`afi-backup-monitor-plan-refresh` on `hcptf-afi-backup-monitor-plan` .
3b. **Apply role (Lambda/EventBridge stacks):** attach
`seahaven-hcptf-iam-management` plus stack-scoped service statements.
Prefer prefix-scoped `lambda:*` on `function:<prefix>-*` /
`layer:<prefix>-*` , `events:*` on `rule/<prefix>-*` , and bucket-scoped
`s3:*` on the artifact bucket — do **not** enumerate individual provider
Get* APIs (`GetFunctionCodeSigningConfig` , `GetBucketAcl` , …); that list
lags and fails first apply. Keep list/describe-on-`*` only where the
service requires it (e.g. `lambda:ListFunctions` ). Copy
`afi-backup-monitor-services` on `hcptf-afi-backup-monitor` .
4. **Deploy substrate** to `UPDATE_COMPLETE` . Verify: both roles exist;
`hcptf-<stack>` lists `seahaven-hcptf-iam-management` in
`list-attached-role-policies` ; trust subs match the live
org/project/workspace names byte-for-byte; simulate the apply role against
a `hcptf-*` ARN (expect `explicitDeny` from `DenySelfMutation` ) and against
2026-08-13 16:47:53 -04:00
a normal stack role name (expect `allowed` ); and if step 3 added a
per-workload boundary, confirm the deployed default version of
`seahaven-lambda-execution-boundary-<stack>` carries the stack's
2026-08-05 16:14:16 -04:00
data-plane statements (`aws iam get-policy-version` ) — role verification
alone never checks boundary content. Mechanical template↔deployed policy
reconcile as for other substrate policies.
5. Set **workspace-level** variables `TFC_AWS_PLAN_ROLE_ARN` +
2026-07-30 16:31:34 -04:00
`TFC_AWS_APPLY_ROLE_ARN` (category env) to the verified role ARNs, plus
`TFC_AWS_PROVIDER_AUTH=true` . Never project-scoped variable sets — the
trust is pinned per workspace, so a shared set breaks every other
2026-08-05 16:14:16 -04:00
workspace. Auto-apply stays OFF until the stack is sealed.
6. **App Terraform PR:** every `aws_iam_role` sets `path = "/tf-managed/"` and
2026-08-13 16:47:53 -04:00
`permissions_boundary` to that stack's
`seahaven-lambda-execution-boundary-<stack>` ARN (not the shared name,
once the per-workload policy exists); package Lambda/layer zips via an account artifact S3 bucket
2026-08-05 16:14:16 -04:00
and `aws_s3_object` `content_base64` (HCP plan and apply run on separate
workers and do not share local `archive_file` paths — see
`afi-backup-monitor/terraform/artifacts.tf` ); functions `depends_on` their
IAM policies before create; commit `.terraform.lock.hcl` with
multi-platform hashes.
7. **First Manual apply** from the HCP workspace (not local apply against
prod). Tolerate partial state on permission misses; widen the apply/plan
roles and retry. Confirm all expected resources exist in the target
account.
8. **Live-path proof:** real invoke of every critical function must hit real
external APIs / Slack (not synth or simulate alone) before cutover.
9. **Cutover + decommission:** disable source-account schedules (e.g.
EventBridge rules); observe a clean prod path; delete the source
CloudFormation/CDK stack per the decommission playbook; sweep or retain
log groups deliberately; delete source secrets last.
10. **Docs:** update Confluence AWS Architecture Map and the stack ops page;
promote durable gotchas to the convention ledger when they are general.
2026-07-30 16:31:34 -04:00
2026-07-30 16:55:45 -04:00
**HCP-side authority is AWS authority.** AWS exposes only `aud` , `sub` and
`amr` as trust-policy condition keys for a generic OIDC provider — HCP's
immutable `terraform_workspace_id` / `terraform_project_id` claims are *not*
usable in an IAM condition (AWS's provider-specific claim validation covers
Google, GitHub, CircleCI and OCI only). The `sub` pin therefore rests on HCP
display names, so whoever can create, rename, move or delete a workspace in the
`seahaven-prod` project effectively holds prod deploy authority. Restrict that
HCP team permission to the same people, and when a workspace is retired, delete
its `hcptf-*` roles in the same change so a reused name cannot inherit them.
**Terraform state is secret-bearing.** HCP-hosted state records sensitive
attributes in full and lives outside the AWS accounts, readable by any HCP
principal with workspace read. Per the handbook's secrets-and-config rule,
secrets stay in Secrets Manager / SSM and are referenced by ARN: do not manage
secret *values* in Terraform (create the secret shell, populate out of band or
via write-only/ephemeral arguments) so no value enters state.
**Rollback (proven in mgmt 2026-07-30):** delete any `hcptf-*` roles first —
they reference the provider, and while any of them still attaches the guardrail
policy the stack delete cannot remove it. Then delete the stack. Only the
**provider** is `Retain` : it survives as an orphan and is removed with
`aws iam delete-open-id-connect-provider` . The **guardrail policy is deleted
with the stack** — do not expect it to persist, and note that every
`DenySelfMutation` / `DenyBoundaryTampering` backstop goes with it, so an
`hcptf-*` role recreated out of band afterwards is *not* gated. Workspaces
holding state must be migrated or destroyed HCP-side first; deleting the OIDC
provider strands them mid-run rather than cleaning them up.
**First-create rollback trap.** The provider is `Retain` , so if any other
resource in this stack fails on first create, CloudFormation rolls back, the
provider survives untracked, and the stack lands in `ROLLBACK_COMPLETE` — which
cannot be updated, and cannot be recreated because an account holds exactly one
provider per URL. Recovery: delete the stack, then either remove the orphaned
provider with the command above before retrying, or redeploy with
`createOidcProvider: false` . Note `cd-cdk` 's pre-flight and health check probe
only the job's single `stack-name` input (the account baseline), so a wedged
substrate stack does not show up there — check it directly.
2026-07-30 16:31:34 -04:00
**Verification of record for the guardrail policy** is mechanical
reconciliation — tag-preserving YAML load of the template vs
`get-policy-version` readback, sorted `json.dumps` compare per statement —
same discipline as the deploy-substrate reconciliation (2026-07-27), not
header-reading. The managed-policy document budget is 6,144 characters;
measure before appending statements.
2026-05-29 17:44:55 -04:00
### CloudTrail (audit finding C-1)
| Resource | Logical ID | Notes |
|---|---|---|
2026-07-07 15:47:41 -04:00
| Multi-region trail | `Trail` (`seahaven-org-trail` ) | Management events read+write, global service events, **log-file validation on** , **CloudTrail Insights on** (ApiCallRate + ApiErrorRate, §37) |
2026-05-29 17:44:55 -04:00
| Log bucket | `TrailLogBucket` (`seahaven-cloudtrail-logs-328440206208` ) | Private (Block Public Access all), SSE-KMS, versioned, **TLS-only** , **Object Lock GOVERNANCE 365d** , lifecycle (Glacier @90d , expire @365d ), server access logging → `seahaven-s3-access-logs` |
| KMS CMK | `TrailKey` (`alias/cloudtrail-logs` ) | Encrypts log files; **automatic rotation enabled** |
| CloudWatch Logs group | created by the L2 `Trail` | 365-day retention; this is the group the CIS Section 4 metric filters (H-1) attach to |
**Data flow:** API activity across all regions → CloudTrail → (a) KMS-encrypted,
Object-Locked S3 bucket for durable/tamper-resistant storage and (b) CloudWatch
Logs for real-time querying and metric-filter alarms.
**Compliance impact:** closes CIS 3.1 (multi-region trail), 3.2 (log-file
validation), 3.4 (CloudWatch Logs integration), 3.6 (bucket access logging),
3.7 (KMS CMK encryption), and 3.8 (CMK rotation). Unblocks CIS Section 4 /
finding H-1 (metric filters + alarms now have a log group to target).
### Design decisions
- **Management events only.** Object-level S3/Lambda data events (CIS 3.10/3.11)
are deferred to control cost; revisit with targeted S3 *write* data events on
sensitive buckets (payments / accounting / kb) if needed.
- **Object Lock GOVERNANCE, not COMPLIANCE.** Tamper-resistant but still
deletable by a principal holding `s3:BypassGovernanceRetention` — avoids the
irreversibility of COMPLIANCE mode. Revisit if a stricter posture is required.
- **RETAIN** on the bucket and KMS key so a stack teardown never destroys the
audit trail.
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
### AWS Backup (audit finding C-7)
Phase 1 ("critical data first") of fixing the account's complete lack of AWS
Backup. Protects the data stores with no offsite leg today and copies each
recovery point cross-region into a governance-locked vault.
| Resource | Logical ID | Notes |
|---|---|---|
| Primary vault | `seahaven-primary` (us-east-1) | KMS-CMK encrypted, unlocked (working copy), RETAIN |
| Offsite vault | `seahaven-offsite` (us-west-2) | KMS-CMK encrypted, **Vault Lock GOVERNANCE** (min-retention 30d, no cooling-off window), RETAIN |
| Backup plan | `seahaven-critical-daily` | Daily 06:00 UTC, delete-after 35d, **cross-region CopyAction → offsite** (retain 90d) |
| Service role | `seahaven-backup-service-role` | **Backup-only** (Backup + S3-Backup managed policies); restore perms intentionally deferred |
**Phase-1 scope** (selected by explicit ARN, not tags, to avoid drifting other
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.
scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
stacks): RDS `proposal-system-db` , DynamoDB `PaymentsDashboard` ,
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
DynamoDB `purchase-orders` , S3 `accounting.seahaven.com` ,
`seahaven-payments-csv-328440206208` , `google-workspace-seahavenind.com` .
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.
scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
*(RDS `database-1` was originally in this set but was retired 2026-06-03 —
audit H-19, idle 0 conn/60d — and removed from the selection; its final
encrypted recovery point is retained in `seahaven-offsite` for 7 years.)*
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
**Coexists with** existing EBS DLM snapshots and DynamoDB PITR — it supplements
them with the missing offsite + immutable leg; it does not replace them.
**Design decisions:**
- **Governance lock first, not compliance.** Recovery points can't be silently
deleted, but a principal with explicit permission can still intervene while
we validate. Graduate to COMPLIANCE (irreversible) later by adding
`changeableFor` to the offsite vault lock + redeploy.
- **Backup-only role.** Restore policies and `allowRestores` are not granted;
restores get a separate audited path once a restore-test process exists.
**Pre-deploy gates** (must clear before the first scheduled run):
1. Enable S3 versioning on `seahaven-payments-csv-328440206208` and
`google-workspace-seahavenind.com` (`accounting.seahaven.com` already has it,
audit C-9), or their jobs fail silently (folds in H-21).
2. `database-1` is unencrypted (H-19): smoke-test an on-demand backup + copy of
it to us-west-2 first; if the copy fails, encrypt it or drop it from the copy.
3. Enable DynamoDB PITR (H-7) on the two tables for between-window recovery.
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.
scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
### AWS Backup phase 2 (audit Day 4)
Expands the same `seahaven-critical-daily` plan to every remaining data store, so
all of DynamoDB + EBS get the offsite + immutable leg ("offsite for everything").
| Resource | Logical ID | Notes |
|---|---|---|
| Phase-2 selection | `Plan/Phase2Resources` (`phase2-offsite-everything` ) | Same plan, same `seahaven-backup-service-role` , same daily + cross-region copy rule |
**Phase-2 scope:** the 15 remaining DynamoDB tables (all except the two phase-1
financial tables + the deleted ledgerflow tables) and all 9 in-use EBS volumes,
again **by explicit ARN** — tag-based selection was deliberately avoided because
the file-share volumes are standalone-managed and the tables are owned by other
stacks, so tagging here would drift them.
**No IAM change:** `AWSBackupServiceRolePolicyForBackup` already grants the
DynamoDB/RDS/EBS backup actions, so phase 2 reuses the phase-1 role unchanged
(cross-reviewed, no BLOCK).
**Known tradeoff (→ Jira INFRA-31):** explicit-ARN EBS entries go stale if a
volume is replaced (new volume id), silently dropping it from backup. Migrating
the EBS portion to tag-based selection (with the tag codified in each owning
stack) is the resilient follow-up; scheduled drift detection is the interim
backstop.
**Also enabled outside this stack (audit H-7, via CLI — codify per stack →
INFRA-30):** PITR + `DeletionProtectionEnabled` on 12 more DynamoDB tables
(account-wide PITR now 19/21).
2026-06-01 17:56:12 -04:00
### Detective controls + budget (audit Day 1)
Account-level detective layer, in `lib/detective-controls.ts` , plus the cost
budget in `lib/governance-toggles.ts` . **Scope is us-east-1 only** (all workloads
live here); multi-region coverage is a follow-up.
| Resource | Logical ID | Finding | Notes |
|---|---|---|---|
| Config delivery bucket | `seahaven-config-328440206208` | H-2 | Private (BPA all), SSE-S3, versioned, TLS-only, 365d lifecycle |
| Config recorder role | `seahaven-config-recorder-role` | H-2 | `AWS_ConfigRole` + scoped S3 delivery; **IAM cross-reviewed** |
2026-06-10 14:38:23 -04:00
| Config recorder + channel | `DetectiveControls/ConfigPutRecorder` , `ConfigPutChannel` , `ConfigStartRecorder` | H-2 | `AwsCustomResource` calls `PutConfigurationRecorder` → `PutDeliveryChannel` → `StartConfigurationRecorder` in sequence; idempotent upsert avoids the L1 CFN deadlock (INFRA-17). Custom-resource role cross-reviewed. |
2026-06-01 17:56:12 -04:00
| GuardDuty detector | `DetectiveControls/GuardDutyDetector` | H-3 | Findings every 15 min |
| Security Hub | `DetectiveControls/SecurityHub` | H-4 | FSBP v1.0.0 + CIS v3.0.0; controls evaluate once Config is recording |
| Access Analyzer | `seahaven-account-analyzer` | M-5 | ACCOUNT external-access analyzer (free) |
| Monthly budget | `GovernanceToggles/MonthlyCostBudget` (`seahaven-monthly-cost` ) | M-10 | $1,200/mo, 80%/100% actual + 100% forecast → adam@seahavenind .com |
2026-06-10 14:38:23 -04:00
**Config recorder + delivery channel are managed by `AwsCustomResource` (INFRA-17).**
The L1 `AWS::Config::ConfigurationRecorder` deadlocks the stack (recorder never
reaches `CREATE_COMPLETE` without a delivery channel; channel can't be created
without a recorder — hit 2026-06-01). The custom resource sidesteps this by
calling the Config SDK directly: `Put*` is an upsert, so the deploy adopts the
existing CLI-created recorder and channel without destroying them. Active
recording is never interrupted.
2026-06-01 17:56:12 -04:00
### CLI-applied governance toggles (no CloudFormation resource)
These account toggles have no native CloudFormation resource, so they are applied
2026-07-14 17:17:55 -04:00
via CLI and recorded here — **per account** (they are account-scoped; a new
member account has NONE of them until applied). Applied: 328440206208
(2026-06-01); 001520130573, 710827005802, 011934824531 (2026-07-14 — EBS
encryption-by-default + password policy, verified per account; Inspector2 via
delegated admin; cost-allocation tags are org-level).
2026-06-01 17:56:12 -04:00
```bash
# M-3 EBS encryption-by-default (new volumes; existing 5 plaintext volumes are H-19-adjacent)
aws ec2 enable-ebs-encryption-by-default --region us-east-1
# M-6 Inspector2 (EC2 + Lambda + ECR)
aws inspector2 enable --resource-types EC2 LAMBDA ECR --region us-east-1
# M-7 IAM password policy (CIS 1.8/1.9): >=14 chars, full complexity, no reuse of last 24
aws iam update-account-password-policy \
--minimum-password-length 14 \
--require-symbols --require-numbers \
--require-uppercase-characters --require-lowercase-characters \
--allow-users-to-change-password --password-reuse-prevention 24
# M-11 Activate cost-allocation tags (only activates keys already seen on resources)
aws ce update-cost-allocation-tags-status --cost-allocation-tags-status \
'TagKey=Project,Status=Active' 'TagKey=Owner,Status=Active' 'TagKey=Environment,Status=Active'
```
2026-07-14 18:32:40 -04:00
### Centralized root access management (org-level, no CloudFormation resource)
**STATUS: ENABLED 2026-07-14, all member root credentials DELETED** (evidence:
`~/Documents/repositories/_audits/centralized-root-access-evidence-2026-07-14.md` ).
Member accounts have NO root credentials; the only root path is a privileged
session from the management account. The management account's own root is NOT
centrally manageable and stays password+MFA hardened.
```bash
# Enable (mgmt account). ORDER MATTERS: trusted access must be enabled
# explicitly first — enable-organizations-root-credentials-management does
# NOT auto-enable it (fails ServiceAccessNotEnabledException).
aws organizations enable-aws-service-access --service-principal iam.amazonaws.com
aws iam enable-organizations-root-credentials-management
aws iam enable-organizations-root-sessions
# Periodic verification (add to governance checks): expect BOTH features
aws iam list-organizations-features
```
**Audit / delete member root credentials** (task-scoped root sessions, 15-min):
```bash
aws sts assume-root --target-principal < acct > \
--task-policy-arn arn=arn:aws:iam::aws:policy/root-task/IAMAuditRootUserCredentials
# then, with the session creds (no --user-name; root has none):
# get-login-profile / list-mfa-devices / list-access-keys / list-signing-certificates
aws sts assume-root --target-principal < acct > \
--task-policy-arn arn=arn:aws:iam::aws:policy/root-task/IAMDeleteRootUserCredentials
# delete-login-profile; deactivate-mfa-device --serial-number <arn>
# GOTCHA: the delete task policy explicitly DENIES iam:DeleteVirtualMFADevice —
# remove the orphaned virtual-device OBJECT via OrganizationAccountAccessRole.
# DONE = four surfaces clear: login profile NoSuchEntity, MFA/keys/certs all empty.
```
**Root recovery runbook** (proven by drill on prod 2026-07-14):
1. `deny-root-user` (p-2idoxozz) DENIES root sessions in every covered OU
(SCPs evaluate `sts:AssumeRoot` sessions — the principal is the member
root ARN). Recovery therefore starts with a **manual, temporary detach**
(`aws organizations detach-policy` — NOT a CDK deploy), timeboxed minutes.
2. GOTCHA (inheritance): p-2idoxozz is attached to `workloads` AND its child
OUs — for an account under workloads/, detach from BOTH the child OU and
workloads, or the inherited deny still applies. Allow ~10s propagation.
3. Freeze deploys of `seahaven-org-governance` for the window (a concurrent
deploy would re-attach mid-recovery); verify no CD run in flight first.
4. `aws sts assume-root --target-principal < acct > --task-policy-arn
arn=arn:aws:iam::aws:policy/root-task/IAMCreateRootUserPassword` →
`create-login-profile` (no args) restores a login profile.
5. Do the root-only task, DELETE the credentials again (four-surface verify),
reattach the SCP(s), confirm `list-targets-for-policy` matches the
pre-detach capture and stack drift is IN_SYNC.
6. extdev extra: `external-dev-iam-guardrails` also denies
`iam:CreateLoginProfile` — recovery there needs that SCP temporarily
detached too. The extdev OU sits at the **5-SCP hard quota** : any new
guardrail for extdev must attach at the ACCOUNT (396287094661) or
consolidate into an existing policy.
**New-account flow (supersedes root-harden-before-OU-move):** create the
account at the org ROOT → it has no root credentials from birth (verify with
the audit session) → bootstrap + deploy role + baseline → `move-account` into
the target OU → verify SCP inheritance + region-lock canary. No mailbox or
MFA enrollment step. Root-usage monitoring: GuardDuty
`Policy:IAMUser/RootCredentialUsage` + CIS 4.3 alarm remain active.
2026-07-14 15:32:50 -04:00
### Delegated security administration (Phase 3, no CloudFormation resource)
2026-07-14 15:53:42 -04:00
Account **seahaven-security (001520130573)** is the org's delegated
administrator for the detective services. **STATUS: APPLIED 2026-07-14,
verified** (see evidence below). The hard preconditions were enforced before
the first delegation call (security review SEC-BASE-B/D — never delegate to an
account with unhardened root or before its baseline stack exists):
- Baseline stack `UPDATE_COMPLETE` ; root `AccountMFAEnabled: 1` ; account
parent `ou-nbuj-v0s9630u` with SCPs deny-root-user +
protect-security-baseline + security-guardrails inherited.
Verification evidence (2026-07-14):
- `organizations list-delegated-administrators` → `001520130573` (all five
service principals registered).
- GuardDuty: `AutoEnableOrganizationMembers: ALL` ; members 328440206208 +
396287094661 both `Enabled` .
- Security Hub: org auto-enable on; both members `Enabled` .
- Org Access Analyzer `seahaven-org-analyzer` created; Config org aggregator
`seahaven-org-aggregator` (AllAwsRegions) on the Config SLR; Inspector2
auto-enable ec2/ecr/lambda + both members associated.
- **Findings flow verified end-to-end:** GuardDuty sample findings created in
member 396287094661 were listed and fully readable from the admin detector
in 001520130573 (`Recon:EC2/PortProbeUnprotectedPort` , AccountId
396287094661, sample=true), then archived.
The delegation runbook (all calls idempotent, run from the **management
account**):
2026-07-14 15:32:50 -04:00
```bash
# GuardDuty: delegate + auto-enable all org members (adopts existing detectors)
aws guardduty enable-organization-admin-account --admin-account-id 001520130573
# then AS 001520130573: update-organization-configuration --auto-enable-organization-members ALL
# + create-members for pre-existing accounts (mgmt, external-dev)
# Security Hub: delegate + auto-enable new members
aws securityhub enable-organization-admin-account --admin-account-id 001520130573
# then AS 001520130573: update-organization-configuration --auto-enable
# IAM Access Analyzer: delegate + ORGANIZATION-scoped analyzer
aws organizations register-delegated-administrator \
--account-id 001520130573 --service-principal access-analyzer.amazonaws.com
# then AS 001520130573: create-analyzer --type ORGANIZATION
# Config: delegate the aggregator (recorders stay per-account in the baselines;
# the aggregator's recorder-status view is the drift detector)
aws organizations register-delegated-administrator \
--account-id 001520130573 --service-principal config.amazonaws.com
# then AS 001520130573: put-configuration-aggregator --organization-aggregation-source
# Inspector2: delegate + associate members
aws inspector2 enable-delegated-admin-account --delegated-admin-account-id 001520130573
```
ONLY once delegation is live AND auto-enrollment is verified (a new member
shows enrolled in the security account's GuardDuty/Security Hub consoles):
new member accounts are then detected/enrolled automatically, and future
member baselines can drop per-account GuardDuty/SecurityHub resources.
Until then, every member baseline MUST keep them (slimming the existing
member stacks is a separate, verification-gated change; note the DA
account's own CFN-owned detector/hub become co-managed after delegation —
never rename/remove them via CFN while the account is delegated admin).
Accepted read-surface note (SEC-BASE-I): the org Config aggregator +
ORGANIZATION Access Analyzer give principals in 001520130573 org-wide READ of
resource configurations (including recorded Lambda env vars) and IAM policies.
Main-branch write access to this repo therefore implies that read surface —
verify no prod Lambda keeps secrets in env vars before creating the
aggregator, and keep branch protection tight.
2026-06-01 17:56:12 -04:00
**L-8 (billing-metrics preference) is OUTSTANDING — console only.** Enabling the
CloudWatch `EstimatedCharges` metric in us-east-1 requires turning on *Receive
Billing Alerts* under Billing → Billing preferences; there is no public API/CLI.
fix: Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm (#36)
* Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm
CIS 4.1 (cis-UnauthorizedAPICalls) flapped OK<->ALARM 15 times in 30 days,
all from benign AWS-service AccessDenied noise (CloudFormation deploy/drift
describe-scans, AWS Config recorder). A single CFN run on 2026-07-07 emitted
100+ such denials in 15 min, tripping the alarm and burying the real CIS 4.1
security signal in email noise (alert fatigue).
- Group both error codes so the exclusions apply to the whole filter (the old
pattern leaked the UnauthorizedOperation branch past the exclusions due to
&& binding tighter than ||).
- Exclude denials whose sourceIPAddress is an AWS service host (*.amazonaws.com)
— AWS acting on our behalf, not a principal of concern. Real unauthorized
calls from a console/CLI/attacker present a routable IP and are still counted.
Validated against the trail log group: spike window 107 -> 4 matches, the 4
remaining all from a routable admin IP (genuine activity CIS should retain).
- Keep the 3/3 evaluation as a backstop against one-off human fat-fingers.
Billing: deleted the manually-created AWS-MonthlyBilling CloudWatch alarm
($50 threshold on EstimatedCharges, routed to site-alerts). It was unmanaged
drift, permanently in ALARM, and fully redundant with the managed M-10 budget
(seahaven-monthly-cost). README updated with rationale + restore command.
* Address sh-security-review: scope CIS 4.1 exclusion to named benign sources
The high-recall security review (detector fan-out + proof-or-kill verifier)
confirmed a MEDIUM detection blind spot in the first revision: excluding all
`*.amazonaws.com` source hosts would hide denials driven through ANY AWS
service (SSM Automation, Step Functions, Lambda, etc.), which CloudTrail
records with that service's host as sourceIPAddress — i.e. service-proxied
privesc/recon attempts would evade CIS 4.1.
Remediation: scope the exclusion to the specific benign sources that actually
flap this account — `*cloudformation.amazonaws.com` (covers both
cloudformation. and hooks.cloudformation.) and `config.amazonaws.com` — plus
the pre-existing delivery.logs exclusion. Every other service-proxied denial
is now retained. Residual (accepted, documented inline): CloudFormation/Config-
proxied denials are still excluded — that path needs near-admin privilege
(CreateStack + PassRole), successful changes still trip the other CIS 4.x
alarms, and GuardDuty backstops.
Validated on the live trail log group: spike window still 107 -> 4 matches
(identical noise suppression), the 4 from a routable admin IP. tsc + synth clean.
2026-07-07 15:32:10 -04:00
The M-10 budget (`seahaven-monthly-cost` , 80%/100% actual + 100% forecast)
provides cost alerting independent of that metric.
> The legacy, manually-created `AWS-MonthlyBilling` CloudWatch alarm ($50
> threshold on `EstimatedCharges`, routed to `site-alerts`) was **deleted
> 2026-07-07** as unmanaged drift: it was fully redundant with the M-10 budget,
> sat permanently in ALARM (spend has far exceeded $50/mo), and was never in
> IaC. Billing alerting is now solely the managed M-10 budget. To restore the
> old alarm if ever needed: `aws cloudwatch put-metric-alarm --alarm-name
> AWS-MonthlyBilling --namespace AWS/Billing --metric-name EstimatedCharges
> --dimensions Name=Currency,Value=USD --statistic Maximum --period 86400
> --evaluation-periods 1 --threshold 50 --comparison-operator GreaterThanThreshold
> --alarm-actions arn:aws:sns:us-east-1:328440206208:site-alerts`.
2026-06-01 17:56:12 -04:00
2026-06-02 15:16:24 -04:00
### Monitoring + logging (audit Day 2)
| Resource | Logical ID | Finding | Notes |
|---|---|---|---|
| CIS metric filters + alarms | `CisMonitoring/*` | H-1 | 15 filters (CIS 4.1– 4.15) on the CloudTrail log group, each with an alarm → `seahaven-cis-alarms` . ALARM-only actions (no OK). 4.16 = Security Hub (Day 1) |
| CIS alarm topic | `seahaven-cis-alarms` | H-1 | SNS, SSE (`alias/aws/sns` ), email sub to adam@seahavenind .com |
| VPC flow logs | `FlowLogs/FlowLog0..4` | H-14 | ALL traffic on all 5 VPCs → S3 |
| Flow-logs bucket | `seahaven-vpc-flow-logs-328440206208` | H-14 | Private, SSE-S3, TLS-only, Glacier @90d / expire @365d ; delivery bucket policy cross-reviewed |
| SES config set | `seahaven-email-events` | M-13 | Bounce/complaint/reject → CloudWatch metrics for reputation visibility |
2026-06-08 19:04:36 -04:00
| Sensitive-logs CMK | `LogsKey/Key` (`alias/seahaven-logs` ) | M-24 | Encrypts sensitive CloudWatch Logs groups. Key policy grants `logs.us-east-1.amazonaws.com` Encrypt*/Decrypt*/ReEncrypt*/GenerateDataKey*/DescribeKey scoped by `kms:EncryptionContext:aws:logs:arn` (required or log delivery breaks). Rotation on, RETAIN. Applied in place to `TrailLogGroup` via escape hatch (same logical id/name). Cross-reviewed |
**M-24 sensitive log groups:** `alias/seahaven-logs` encrypts the CloudTrail CW
log group (codified here) plus the finance/PII Lambda groups owned by other
stacks — `exec-aide-*` , `payments-*` , `po-email-processor` , `vendor-reply-processor`
— which are associated via `aws logs associate-kms-key` and tracked as drift to
codify in their owning repos. The CloudTrail group is encrypted in place (escape
hatch on the existing `AWS::Logs::LogGroup` ) so it is additive: same logical id +
physical name, no replacement, CIS Section-4 metric filters keep working. A
context flag `encryptTrailLogGroup` (default `true` ) allows rolling the CMK out
and smoke-testing it on a low-risk Lambda group before the CloudTrail group:
```bash
# Phase 1: deploy CMK only, validate on a low-risk group
cdk deploy account-baseline --context encryptTrailLogGroup=false
# Phase 2: encrypt the CloudTrail group (default)
cdk deploy account-baseline
```
2026-06-02 15:16:24 -04:00
**H-1 log group:** the metric filters attach to the existing CloudTrail
CloudWatch Logs group by name (`seahaven-account-baseline-TrailLogGroup4CBE3AF5-…` ),
imported read-only so the live audit trail is never replaced. Stable unless the
Trail is recreated.
**H-14 bucket policy note:** the flow-logs delivery policy keeps
`s3:x-amz-acl=bucket-owner-full-control` and the `arn:aws:logs:…:*` source-ARN
wildcard — both are required by AWS's documented flow-logs-to-S3 policy
(`flow-logs-s3-permissions.html` ). A cross-review suggested dropping them; that
was rejected as it would break delivery. `s3:ListBucket` was dropped (not needed).
**M-13 follow-up:** associate `seahaven-email-events` as the default config set
on the live sending identities to capture events from existing senders:
```bash
aws sesv2 put-email-identity-configuration-set-attributes \
--email-identity int.seahaven.com --configuration-set-name seahaven-email-events
```
### Log-group retention + alarm wiring (audit L-4, L-5)
Applied via CLI (auto-created groups spread across stacks; one alarm in another
stack). Applied 2026-06-02.
```bash
# L-4 90-day retention on the 13 never-expire log groups (CodeBuild + CDK helpers)
for lg in < the 13 groups > ; do aws logs put-retention-policy --log-group-name "$lg" --retention-in-days 90; done
# L-5 wire the actionless forgejo backup-verification alarm to site-alerts
aws cloudwatch put-metric-alarm --alarm-name forgejo-backup-verification-errors \
--alarm-actions arn:aws:sns:us-east-1:328440206208:site-alerts # (preserve existing alarm config)
```
2026-05-29 17:44:55 -04:00
## Roadmap (same stack)
2026-06-01 17:56:12 -04:00
Detective layer multi-region expansion (GuardDuty/Config/Security Hub beyond
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.
scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
us-east-1). Backup: phase 2 is deployed (see above); remaining is migrating the
phase-2 EBS entries to tag-based selection (INFRA-31) and graduating the offsite
vault to compliance mode.
2026-05-29 17:44:55 -04:00
## Deploy
CI/CD via the org reusable workflows (`ci-typescript-cdk.yaml` ,
`cd-cdk.yaml` ); pushes to `main` deploy through the OIDC role in
`secrets.AWS_DEPLOY_ROLE_ARN` . Local: `npm ci && npm run build && npx cdk diff` .
```
npx cdk deploy seahaven-account-baseline
```
## Verify
```
aws cloudtrail get-trail-status --name seahaven-org-trail # IsLogging: true
aws cloudtrail describe-trails --trail-name-list seahaven-org-trail
aws cloudtrail validate-logs --trail-arn < arn > --start-time < t > # digest integrity
```
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
AWS Backup (C-7):
```
aws backup list-backup-vaults # seahaven-primary
aws backup list-backup-vaults --region us-west-2 # seahaven-offsite
aws backup describe-backup-vault --backup-vault-name seahaven-offsite --region us-west-2 # Locked, MinRetentionDays
aws backup get-backup-plan --backup-plan-id < id > # daily rule + CopyAction
# Smoke test: on-demand backup of one resource, then confirm the cross-region copy lands
aws backup start-backup-job --backup-vault-name seahaven-primary \
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.
scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
--resource-arn arn:aws:rds:us-east-1:328440206208:db:proposal-system-db \
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
--iam-role-arn arn:aws:iam::328440206208:role/seahaven-backup-service-role
aws backup list-copy-jobs --region us-west-2 # copy to offsite present + COMPLETED
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.
scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
# Phase-2 selections live on the plan:
aws backup list-backup-selections --backup-plan-id < id > --query 'BackupSelectionsList[].SelectionName' # critical-data + phase2-offsite-everything
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
```
2026-06-01 17:56:12 -04:00
Detective layer + governance (Day 1):
```
aws configservice describe-configuration-recorder-status # recording: true
aws guardduty list-detectors # one detector id
aws securityhub get-enabled-standards # FSBP + CIS v3.0.0
aws accessanalyzer list-analyzers # seahaven-account-analyzer ACTIVE
aws inspector2 batch-get-account-status --region us-east-1 # ec2/ecr/lambda ENABLED
aws iam get-account-password-policy # length 14, reuse 24
aws ec2 get-ebs-encryption-by-default --region us-east-1 # EbsEncryptionByDefault: true
aws budgets describe-budgets --account-id 328440206208 # seahaven-monthly-cost $1,200
aws ce list-cost-allocation-tags --status Active # Project/Owner/Environment Active
```