Commit graph

10 commits

Author SHA1 Message Date
Adam Moussa
1e10c40e52
chore(backup): drop the file-share rollback volume from phase 2 (PLAT-77) (#168)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* chore(backup): drop the file-share rollback volume from phase 2 (PLAT-77)

The rollback hold is being deleted now, so the selection must not keep an ARN that will not exist.

* docs(backup): state why the file-share volume leaves phase 2 (PLAT-77)

The comment now matches the cleanup: the missing volume is already gone, and the rollback volume leaves because the hold ended.
2026-09-29 20:23:47 -04:00
Adam Moussa
78e4bcf128
chore(backup): drop the missing file-share volume from phase 2 (PLAT-77) (#167)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* chore(backup): drop retired file-share volumes from phase 2 (PLAT-77)

The 20 GiB volume is already gone and the 500 GiB volume is only a rollback hold, so the backup selection should not include either one.

* fix(backup): keep the file-share rollback volume in phase 2 (PLAT-77)

The 500 GiB disk is the rollback hold until 2026-10-06, so it stays in the offsite selection. Only the missing 20 GiB volume comes out.
2026-09-29 23:51:26 +00:00
Adam Moussa
7e706aef2f
chore(backup): drop retired apm-wo grafana EBS volume (PLAT-75) (#166)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
The mgmt Grafana instance is gone after the HCP cutover, so the org backup
selection no longer needs vol-0488e0bad1f9afbfb.
2026-09-29 22:16:07 +00:00
Adam Moussa
7cce026f4a
chore: drop deleted tables from Phase2 backup selection (#59)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
seahaven-conversations + seahaven-unanswered-questions (seahaven-slack-bot
teardown 2026-07-23), exec-aide (decommissioned 2026-07-10), and
internal-portal-data (internal-portal decommission) no longer exist;
their backup selections would fail nightly.
2026-07-23 16:21:04 -04:00
Adam Moussa
cc54b1e28b
chore(security): add explicit workflow permissions and bump aws-cdk-lib to 2.262.0 (#56)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* docs: update aws profile specified in script (local renaming)

* ci: add least-privilege permissions blocks to workflow callers

Resolves code scanning alerts #3 and #4 (actions/missing-workflow-permissions). Both callable workflows only need contents: read; the dependency-review callable already declares it internally, this caps the caller token to match."

* chore(deps): bump aws-cdk-lib to 2.262.0 for patched brace-expansion

Resolves Dependabot alert #4 (CVE-2026-13149, exponential-time DoS in brace-expansion expand()). The vulnerable 5.0.6 is a bundled dependency inside the aws-cdk-lib tarball, so it cannot be updated independently; 2.262.0 bundles the patched 5.0.7.

Also migrates Stack#addDependency to addStackDependency (deprecated in this release) in bin/app.ts.
2026-07-23 17:38:17 +00:00
Adam Moussa
5002ed86d8
[INFRA-94] Add Backup vault access policy on seahaven-primary (#19)
Reintroduce the scoped vault access policy that was split out of INFRA-89
after two lockout-class bugs. Adds a Deny on the destructive recovery-point
and vault-lifecycle actions (DeleteRecoveryPoint, UpdateRecoveryPointLifecycle,
DeleteBackupVault, DeleteBackupVaultAccessPolicy,
DeleteBackupVaultLockConfiguration, PutBackupVaultLockConfiguration) for every
principal except three exempted operational identities via StringNotLike on
aws:PrincipalArn:

  1. SSO AdministratorAccess role (break-glass human admin)
  2. seahaven-backup-service-role (AWS Backup lifecycle)
  3. cdk-hnb659fds-cfn-exec-role-* (CloudFormation manages the vault)

The CFN-exec-role exemption is the fix for the 2026-06-08 strand failure: without
it CloudFormation cannot re-assert the vault lock config and the deploy strands
the policy. Uses Deny + AnyPrincipal + StringNotLike (not NotPrincipal, which
rejects wildcard ARNs). aws:PrincipalArn normalizes assumed-role sessions to the
IAM role ARN, so the iam::role/ ARN forms are correct (AWS docs: "Do not specify
the assumed role session ARN as a value for this condition key").

Deployed and verified: deploy succeeded (proves exec role not locked out),
access policy present with all three exemptions, vault still Locked
(min1/max2555, LockDate null, 168 RPs), follow-up cdk diff clean (no drift).
2026-06-08 17:34:01 -04:00
Adam Moussa
e9a184cf96
[INFRA-91/89/16/88/73] Reconcile out-of-band baseline changes + add missing detective controls (#18)
* Codify primary vault lock + add backups (INFRA-89, INFRA-88)

INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the
seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no
changeableFor = admin-removable) so it lives in IaC. Values match the
live lock exactly, so the deploy is a no-op adoption.

Add a scoped vault access policy that denies manual recovery-point
deletion and lock/policy tampering to all principals except the AWS
Backup service role and the break-glass SSO AdministratorAccess role,
so automatic lifecycle expiry still works but humans cannot prune
recovery points by hand.

Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard
ARN matching, so the SSO exemption is expressed as Effect DENY with
Principal * and a StringNotLike condition on aws:PrincipalArn, which
does support wildcards. This avoids an unrecoverable vault lockout.

INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po,
extracted-amazon-po, proposal-system uploads + generated) to the
phase2-offsite-everything selection. Versioning verified enabled on
all 6 against the live account (S3 backup requires versioning).

Refs: INFRA-89, INFRA-88

* Promote account trail to organization trail (INFRA-73)

INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId
(o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket
PutObject statement for member-account delivery. CloudTrail org
trusted-access is already enabled on the management account.

Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey
statement for member-account trail delivery, guarded by
aws:PrincipalOrgID. The existing single-account statements are
preserved so management-account delivery is unaffected.

Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption
context must be wildcarded across accounts (org-trail shadow trails
present the member account id), not pinned to the management account,
or member delivery silently fails. Fixed before checkpoint.

CHECKPOINT: delicate org-trail KMS/bucket-policy change — code +
diff captured for review, NOT deployed.

Refs: INFRA-73

* Add secondary-region baseline stacks (INFRA-91, INFRA-16)

INFRA-91: codify the Bedrock model-invocation logging applied
out-of-band in us-west-2 and us-east-2 (per-region delivery role
seahaven-bedrock-invocation-logging-<region> + log group
/aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level
logging config itself has no CFN resource type and is applied via CLI
(already live), same as us-east-1.

INFRA-16: add the still-missing us-east-2 detective controls — AWS
Config recorder role + delivery bucket (recorder/channel via CLI to
avoid the CFN stabilization deadlock seen in us-east-1) and Security
Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in
us-east-2 and are left for a follow-up adoption to keep this change
non-destructive.

The us-east-1 baseline stays region-pinned; these are separate
RegionalBaselineStack instances composed opt-in per region.

CHECKPOINT: new multi-region stacks. The live Bedrock role + log group
already exist (CLI-created), so a plain deploy would collide — these
need cdk import / changeset adoption, not cdk deploy. Code + diff
captured for review, NOT deployed.

Refs: INFRA-91, INFRA-16

* Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
2026-06-08 17:03:18 -04:00
Adam Moussa
1d6668090c
Remove retired database-1 + ledgerflow-pos from backup selections (audit Day 4) (#9)
database-1 (audit H-19) and the LedgerFlow stack (incl. ledgerflow-pos) were
decommissioned 2026-06-03. Drop database-1 from the critical-data selection and
ledgerflow-pos from phase2-offsite-everything so daily jobs don't target missing
resources. database-1's final recovery point is retained encrypted in the
seahaven-offsite vault (7yr); ledgerflow-pos has a final on-demand DynamoDB
backup. Deployed before merge (seahaven-backup UPDATE_COMPLETE).
2026-06-03 11:45:37 -04:00
Adam Moussa
f2a0cc40d6
Expand AWS Backup to remaining DDB + EBS (audit Day 4 phase-2) (#8)
Add a second BackupSelection 'phase2-offsite-everything' on the existing
seahaven-critical-daily plan covering the 15 remaining DynamoDB tables and
all 9 in-use EBS volumes, with the same daily backup + cross-region copy to
the GOVERNANCE-locked seahaven-offsite vault ('offsite for everything').

Reuses seahaven-backup-service-role (AWSBackupServiceRolePolicyForBackup
already grants DDB/RDS/EBS) - no IAM change. Explicit-ARN (not tag-based) to
avoid drifting the standalone file-share volumes and stack-owned tables, same
as phase-1. The 4 deprecated ledgerflow delete-targets are excluded.

Cross-reviewed (no BLOCK). Follow-up: migrate EBS to tag-based selection with
tags codified in owning stacks for resilience to volume replacement.

Deployed to seahaven-backup before merge; selection verified live (24 resources).
2026-06-03 11:18:53 -04:00
Adam Moussa
64ef25dc5b
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)

The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.

Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.

The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.

Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.

* Grant AWS Backup service use of vault CMKs

The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.

Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00