Commit graph

11 commits

Author SHA1 Message Date
Adam Moussa
42781f743e
fix(baseline): drop departed mgmt resources and retain drifted web acl (#171)
Nightly backup jobs fail on resources that have left the management account.
The CloudFront WebACL is already gone while CloudFormation still owns it, so
the deletion policy has to be Retain before a later change can remove it.
2026-10-01 23:57:42 +00:00
Adam Moussa
1e10c40e52
chore(backup): drop the file-share rollback volume from phase 2 (PLAT-77) (#168)
Some checks failed
Deploy / deploy-management (push) Has been cancelled
Deploy / deploy-external-dev (push) Has been cancelled
Deploy / deploy-security (push) Has been cancelled
Deploy / deploy-dev (push) Has been cancelled
Deploy / deploy-prod (push) Has been cancelled
* chore(backup): drop the file-share rollback volume from phase 2 (PLAT-77)

The rollback hold is being deleted now, so the selection must not keep an ARN that will not exist.

* docs(backup): state why the file-share volume leaves phase 2 (PLAT-77)

The comment now matches the cleanup: the missing volume is already gone, and the rollback volume leaves because the hold ended.
2026-09-29 20:23:47 -04:00
Adam Moussa
78e4bcf128
chore(backup): drop the missing file-share volume from phase 2 (PLAT-77) (#167)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* chore(backup): drop retired file-share volumes from phase 2 (PLAT-77)

The 20 GiB volume is already gone and the 500 GiB volume is only a rollback hold, so the backup selection should not include either one.

* fix(backup): keep the file-share rollback volume in phase 2 (PLAT-77)

The 500 GiB disk is the rollback hold until 2026-10-06, so it stays in the offsite selection. Only the missing 20 GiB volume comes out.
2026-09-29 23:51:26 +00:00
Adam Moussa
7e706aef2f
chore(backup): drop retired apm-wo grafana EBS volume (PLAT-75) (#166)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
The mgmt Grafana instance is gone after the HCP cutover, so the org backup
selection no longer needs vol-0488e0bad1f9afbfb.
2026-09-29 22:16:07 +00:00
Adam Moussa
7cce026f4a
chore: drop deleted tables from Phase2 backup selection (#59)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
seahaven-conversations + seahaven-unanswered-questions (seahaven-slack-bot
teardown 2026-07-23), exec-aide (decommissioned 2026-07-10), and
internal-portal-data (internal-portal decommission) no longer exist;
their backup selections would fail nightly.
2026-07-23 16:21:04 -04:00
Adam Moussa
cc54b1e28b
chore(security): add explicit workflow permissions and bump aws-cdk-lib to 2.262.0 (#56)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* docs: update aws profile specified in script (local renaming)

* ci: add least-privilege permissions blocks to workflow callers

Resolves code scanning alerts #3 and #4 (actions/missing-workflow-permissions). Both callable workflows only need contents: read; the dependency-review callable already declares it internally, this caps the caller token to match."

* chore(deps): bump aws-cdk-lib to 2.262.0 for patched brace-expansion

Resolves Dependabot alert #4 (CVE-2026-13149, exponential-time DoS in brace-expansion expand()). The vulnerable 5.0.6 is a bundled dependency inside the aws-cdk-lib tarball, so it cannot be updated independently; 2.262.0 bundles the patched 5.0.7.

Also migrates Stack#addDependency to addStackDependency (deprecated in this release) in bin/app.ts.
2026-07-23 17:38:17 +00:00
Adam Moussa
5002ed86d8
[INFRA-94] Add Backup vault access policy on seahaven-primary (#19)
Reintroduce the scoped vault access policy that was split out of INFRA-89
after two lockout-class bugs. Adds a Deny on the destructive recovery-point
and vault-lifecycle actions (DeleteRecoveryPoint, UpdateRecoveryPointLifecycle,
DeleteBackupVault, DeleteBackupVaultAccessPolicy,
DeleteBackupVaultLockConfiguration, PutBackupVaultLockConfiguration) for every
principal except three exempted operational identities via StringNotLike on
aws:PrincipalArn:

  1. SSO AdministratorAccess role (break-glass human admin)
  2. seahaven-backup-service-role (AWS Backup lifecycle)
  3. cdk-hnb659fds-cfn-exec-role-* (CloudFormation manages the vault)

The CFN-exec-role exemption is the fix for the 2026-06-08 strand failure: without
it CloudFormation cannot re-assert the vault lock config and the deploy strands
the policy. Uses Deny + AnyPrincipal + StringNotLike (not NotPrincipal, which
rejects wildcard ARNs). aws:PrincipalArn normalizes assumed-role sessions to the
IAM role ARN, so the iam::role/ ARN forms are correct (AWS docs: "Do not specify
the assumed role session ARN as a value for this condition key").

Deployed and verified: deploy succeeded (proves exec role not locked out),
access policy present with all three exemptions, vault still Locked
(min1/max2555, LockDate null, 168 RPs), follow-up cdk diff clean (no drift).
2026-06-08 17:34:01 -04:00
Adam Moussa
e9a184cf96
[INFRA-91/89/16/88/73] Reconcile out-of-band baseline changes + add missing detective controls (#18)
* Codify primary vault lock + add backups (INFRA-89, INFRA-88)

INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the
seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no
changeableFor = admin-removable) so it lives in IaC. Values match the
live lock exactly, so the deploy is a no-op adoption.

Add a scoped vault access policy that denies manual recovery-point
deletion and lock/policy tampering to all principals except the AWS
Backup service role and the break-glass SSO AdministratorAccess role,
so automatic lifecycle expiry still works but humans cannot prune
recovery points by hand.

Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard
ARN matching, so the SSO exemption is expressed as Effect DENY with
Principal * and a StringNotLike condition on aws:PrincipalArn, which
does support wildcards. This avoids an unrecoverable vault lockout.

INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po,
extracted-amazon-po, proposal-system uploads + generated) to the
phase2-offsite-everything selection. Versioning verified enabled on
all 6 against the live account (S3 backup requires versioning).

Refs: INFRA-89, INFRA-88

* Promote account trail to organization trail (INFRA-73)

INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId
(o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket
PutObject statement for member-account delivery. CloudTrail org
trusted-access is already enabled on the management account.

Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey
statement for member-account trail delivery, guarded by
aws:PrincipalOrgID. The existing single-account statements are
preserved so management-account delivery is unaffected.

Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption
context must be wildcarded across accounts (org-trail shadow trails
present the member account id), not pinned to the management account,
or member delivery silently fails. Fixed before checkpoint.

CHECKPOINT: delicate org-trail KMS/bucket-policy change — code +
diff captured for review, NOT deployed.

Refs: INFRA-73

* Add secondary-region baseline stacks (INFRA-91, INFRA-16)

INFRA-91: codify the Bedrock model-invocation logging applied
out-of-band in us-west-2 and us-east-2 (per-region delivery role
seahaven-bedrock-invocation-logging-<region> + log group
/aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level
logging config itself has no CFN resource type and is applied via CLI
(already live), same as us-east-1.

INFRA-16: add the still-missing us-east-2 detective controls — AWS
Config recorder role + delivery bucket (recorder/channel via CLI to
avoid the CFN stabilization deadlock seen in us-east-1) and Security
Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in
us-east-2 and are left for a follow-up adoption to keep this change
non-destructive.

The us-east-1 baseline stays region-pinned; these are separate
RegionalBaselineStack instances composed opt-in per region.

CHECKPOINT: new multi-region stacks. The live Bedrock role + log group
already exist (CLI-created), so a plain deploy would collide — these
need cdk import / changeset adoption, not cdk deploy. Code + diff
captured for review, NOT deployed.

Refs: INFRA-91, INFRA-16

* Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
2026-06-08 17:03:18 -04:00
Adam Moussa
1d6668090c
Remove retired database-1 + ledgerflow-pos from backup selections (audit Day 4) (#9)
database-1 (audit H-19) and the LedgerFlow stack (incl. ledgerflow-pos) were
decommissioned 2026-06-03. Drop database-1 from the critical-data selection and
ledgerflow-pos from phase2-offsite-everything so daily jobs don't target missing
resources. database-1's final recovery point is retained encrypted in the
seahaven-offsite vault (7yr); ledgerflow-pos has a final on-demand DynamoDB
backup. Deployed before merge (seahaven-backup UPDATE_COMPLETE).
2026-06-03 11:45:37 -04:00
Adam Moussa
f2a0cc40d6
Expand AWS Backup to remaining DDB + EBS (audit Day 4 phase-2) (#8)
Add a second BackupSelection 'phase2-offsite-everything' on the existing
seahaven-critical-daily plan covering the 15 remaining DynamoDB tables and
all 9 in-use EBS volumes, with the same daily backup + cross-region copy to
the GOVERNANCE-locked seahaven-offsite vault ('offsite for everything').

Reuses seahaven-backup-service-role (AWSBackupServiceRolePolicyForBackup
already grants DDB/RDS/EBS) - no IAM change. Explicit-ARN (not tag-based) to
avoid drifting the standalone file-share volumes and stack-owned tables, same
as phase-1. The 4 deprecated ledgerflow delete-targets are excluded.

Cross-reviewed (no BLOCK). Follow-up: migrate EBS to tag-based selection with
tags codified in owning stacks for resilience to volume replacement.

Deployed to seahaven-backup before merge; selection verified live (24 resources).
2026-06-03 11:18:53 -04:00
Adam Moussa
64ef25dc5b
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)

The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.

Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.

The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.

Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.

* Grant AWS Backup service use of vault CMKs

The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.

Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00