Reintroduce the scoped vault access policy that was split out of INFRA-89
after two lockout-class bugs. Adds a Deny on the destructive recovery-point
and vault-lifecycle actions (DeleteRecoveryPoint, UpdateRecoveryPointLifecycle,
DeleteBackupVault, DeleteBackupVaultAccessPolicy,
DeleteBackupVaultLockConfiguration, PutBackupVaultLockConfiguration) for every
principal except three exempted operational identities via StringNotLike on
aws:PrincipalArn:
1. SSO AdministratorAccess role (break-glass human admin)
2. seahaven-backup-service-role (AWS Backup lifecycle)
3. cdk-hnb659fds-cfn-exec-role-* (CloudFormation manages the vault)
The CFN-exec-role exemption is the fix for the 2026-06-08 strand failure: without
it CloudFormation cannot re-assert the vault lock config and the deploy strands
the policy. Uses Deny + AnyPrincipal + StringNotLike (not NotPrincipal, which
rejects wildcard ARNs). aws:PrincipalArn normalizes assumed-role sessions to the
IAM role ARN, so the iam::role/ ARN forms are correct (AWS docs: "Do not specify
the assumed role session ARN as a value for this condition key").
Deployed and verified: deploy succeeded (proves exec role not locked out),
access policy present with all three exemptions, vault still Locked
(min1/max2555, LockDate null, 168 RPs), follow-up cdk diff clean (no drift).
* Codify primary vault lock + add backups (INFRA-89, INFRA-88)
INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the
seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no
changeableFor = admin-removable) so it lives in IaC. Values match the
live lock exactly, so the deploy is a no-op adoption.
Add a scoped vault access policy that denies manual recovery-point
deletion and lock/policy tampering to all principals except the AWS
Backup service role and the break-glass SSO AdministratorAccess role,
so automatic lifecycle expiry still works but humans cannot prune
recovery points by hand.
Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard
ARN matching, so the SSO exemption is expressed as Effect DENY with
Principal * and a StringNotLike condition on aws:PrincipalArn, which
does support wildcards. This avoids an unrecoverable vault lockout.
INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po,
extracted-amazon-po, proposal-system uploads + generated) to the
phase2-offsite-everything selection. Versioning verified enabled on
all 6 against the live account (S3 backup requires versioning).
Refs: INFRA-89, INFRA-88
* Promote account trail to organization trail (INFRA-73)
INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId
(o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket
PutObject statement for member-account delivery. CloudTrail org
trusted-access is already enabled on the management account.
Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey
statement for member-account trail delivery, guarded by
aws:PrincipalOrgID. The existing single-account statements are
preserved so management-account delivery is unaffected.
Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption
context must be wildcarded across accounts (org-trail shadow trails
present the member account id), not pinned to the management account,
or member delivery silently fails. Fixed before checkpoint.
CHECKPOINT: delicate org-trail KMS/bucket-policy change — code +
diff captured for review, NOT deployed.
Refs: INFRA-73
* Add secondary-region baseline stacks (INFRA-91, INFRA-16)
INFRA-91: codify the Bedrock model-invocation logging applied
out-of-band in us-west-2 and us-east-2 (per-region delivery role
seahaven-bedrock-invocation-logging-<region> + log group
/aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level
logging config itself has no CFN resource type and is applied via CLI
(already live), same as us-east-1.
INFRA-16: add the still-missing us-east-2 detective controls — AWS
Config recorder role + delivery bucket (recorder/channel via CLI to
avoid the CFN stabilization deadlock seen in us-east-1) and Security
Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in
us-east-2 and are left for a follow-up adoption to keep this change
non-destructive.
The us-east-1 baseline stays region-pinned; these are separate
RegionalBaselineStack instances composed opt-in per region.
CHECKPOINT: new multi-region stacks. The live Bedrock role + log group
already exist (CLI-created), so a plain deploy would collide — these
need cdk import / changeset adoption, not cdk deploy. Code + diff
captured for review, NOT deployed.
Refs: INFRA-91, INFRA-16
* Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
database-1 (audit H-19) and the LedgerFlow stack (incl. ledgerflow-pos) were
decommissioned 2026-06-03. Drop database-1 from the critical-data selection and
ledgerflow-pos from phase2-offsite-everything so daily jobs don't target missing
resources. database-1's final recovery point is retained encrypted in the
seahaven-offsite vault (7yr); ledgerflow-pos has a final on-demand DynamoDB
backup. Deployed before merge (seahaven-backup UPDATE_COMPLETE).
Add a second BackupSelection 'phase2-offsite-everything' on the existing
seahaven-critical-daily plan covering the 15 remaining DynamoDB tables and
all 9 in-use EBS volumes, with the same daily backup + cross-region copy to
the GOVERNANCE-locked seahaven-offsite vault ('offsite for everything').
Reuses seahaven-backup-service-role (AWSBackupServiceRolePolicyForBackup
already grants DDB/RDS/EBS) - no IAM change. Explicit-ARN (not tag-based) to
avoid drifting the standalone file-share volumes and stack-owned tables, same
as phase-1. The 4 deprecated ledgerflow delete-targets are excluded.
Cross-reviewed (no BLOCK). Follow-up: migrate EBS to tag-based selection with
tags codified in owning stacks for resilience to volume replacement.
Deployed to seahaven-backup before merge; selection verified live (24 resources).
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.