seahaven-org-baseline/lib/backup-stack.ts

332 lines
15 KiB
TypeScript
Raw Normal View History

Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
import * as cdk from "aws-cdk-lib";
import * as kms from "aws-cdk-lib/aws-kms";
import * as iam from "aws-cdk-lib/aws-iam";
import * as events from "aws-cdk-lib/aws-events";
import * as backup from "aws-cdk-lib/aws-backup";
import { Construct } from "constructs";
/**
* Primary AWS Backup vault + plan for Sea Haven (account 328440206208), us-east-1.
*
* Closes audit finding C-7 (AWS Backup entirely unused) together with
* backup-offsite-stack. Phase 1 ("critical data first"): protect the data
* stores with no offsite leg today and copy each recovery point cross-region
* to the GOVERNANCE-locked `seahaven-offsite` vault (us-west-2).
*
* Coexistence: this SUPPLEMENTS the existing EBS DLM snapshots and DynamoDB
* PITR — it does not replace them. It adds the missing Copy3 (offsite) +
* immutability leg. The DLM/PITR overlap is rationalized in a later phase.
*
* Selection is by explicit ARN (not tag-based) so we don't have to tag — and
* drift — resources owned by other stacks (proposal-system, payments-dashboard).
* Switch to tag-based selection when expanding past the phase-1 set.
*
* PRE-DEPLOY GATES (validate before the first scheduled run):
* - S3 backup requires bucket versioning. `accounting.seahaven.com` already
* has it (audit C-9); `seahaven-payments-csv-328440206208` and
* `google-workspace-seahavenind.com` must have versioning enabled first or
* their jobs fail silently (folds in audit H-21).
* - `database-1` is unencrypted (audit H-19). Cross-region copy of an
* unencrypted RDS recovery point may fail or land unencrypted. Smoke-test
* an on-demand backup of `database-1` FIRST and confirm the copy job to
* us-west-2 succeeds; if not, encrypt database-1 (H-19) or drop it from the
* copy until then.
* - DynamoDB PITR (H-7) is independent of this plan; enable it on the two
* tables for between-window point-in-time recovery.
*/
export class BackupStack extends cdk.Stack {
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
super(scope, id, props);
// CMK encrypting the primary (operational) vault. RETAIN + rotation.
const vaultKey = new kms.Key(this, "PrimaryVaultKey", {
alias: "backup-primary-vault",
description: "Encrypts primary AWS Backup recovery points (us-east-1)",
enableKeyRotation: true,
removalPolicy: cdk.RemovalPolicy.RETAIN,
});
// The L2 BackupVault does NOT grant the backup service use of a customer
// CMK; the default key policy only delegates to account IAM. Grant
// backup.amazonaws.com the minimum KMS actions (incl. CreateGrant for
// RDS/EBS recovery points) so backup jobs can write to this vault.
vaultKey.addToResourcePolicy(
new iam.PolicyStatement({
sid: "AllowAwsBackupUseOfTheKey",
principals: [new iam.ServicePrincipal("backup.amazonaws.com")],
// Action set matches AWS's documented Backup vault-key policy; scoped
// to this account so only this account's Backup service can use it.
actions: [
"kms:Decrypt",
"kms:GenerateDataKey",
"kms:GenerateDataKeyWithoutPlaintext",
"kms:ReEncrypt*",
"kms:DescribeKey",
],
resources: ["*"],
conditions: { StringEquals: { "aws:SourceAccount": this.account } },
})
);
vaultKey.addToResourcePolicy(
new iam.PolicyStatement({
sid: "AllowAwsBackupCreateGrant",
principals: [new iam.ServicePrincipal("backup.amazonaws.com")],
actions: ["kms:CreateGrant"],
resources: ["*"],
conditions: {
Bool: { "kms:GrantIsForAWSResource": "true" },
StringEquals: { "aws:SourceAccount": this.account },
},
})
);
[INFRA-91/89/16/88/73] Reconcile out-of-band baseline changes + add missing detective controls (#18) * Codify primary vault lock + add backups (INFRA-89, INFRA-88) INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no changeableFor = admin-removable) so it lives in IaC. Values match the live lock exactly, so the deploy is a no-op adoption. Add a scoped vault access policy that denies manual recovery-point deletion and lock/policy tampering to all principals except the AWS Backup service role and the break-glass SSO AdministratorAccess role, so automatic lifecycle expiry still works but humans cannot prune recovery points by hand. Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard ARN matching, so the SSO exemption is expressed as Effect DENY with Principal * and a StringNotLike condition on aws:PrincipalArn, which does support wildcards. This avoids an unrecoverable vault lockout. INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po, extracted-amazon-po, proposal-system uploads + generated) to the phase2-offsite-everything selection. Versioning verified enabled on all 6 against the live account (S3 backup requires versioning). Refs: INFRA-89, INFRA-88 * Promote account trail to organization trail (INFRA-73) INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId (o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket PutObject statement for member-account delivery. CloudTrail org trusted-access is already enabled on the management account. Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey statement for member-account trail delivery, guarded by aws:PrincipalOrgID. The existing single-account statements are preserved so management-account delivery is unaffected. Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption context must be wildcarded across accounts (org-trail shadow trails present the member account id), not pinned to the management account, or member delivery silently fails. Fixed before checkpoint. CHECKPOINT: delicate org-trail KMS/bucket-policy change — code + diff captured for review, NOT deployed. Refs: INFRA-73 * Add secondary-region baseline stacks (INFRA-91, INFRA-16) INFRA-91: codify the Bedrock model-invocation logging applied out-of-band in us-west-2 and us-east-2 (per-region delivery role seahaven-bedrock-invocation-logging-<region> + log group /aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level logging config itself has no CFN resource type and is applied via CLI (already live), same as us-east-1. INFRA-16: add the still-missing us-east-2 detective controls — AWS Config recorder role + delivery bucket (recorder/channel via CLI to avoid the CFN stabilization deadlock seen in us-east-1) and Security Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in us-east-2 and are left for a follow-up adoption to keep this change non-destructive. The us-east-1 baseline stays region-pinned; these are separate RegionalBaselineStack instances composed opt-in per region. CHECKPOINT: new multi-region stacks. The live Bedrock role + log group already exist (CLI-created), so a plain deploy would collide — these need cdk import / changeset adoption, not cdk deploy. Code + diff captured for review, NOT deployed. Refs: INFRA-91, INFRA-16 * Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
2026-06-08 17:03:18 -04:00
// Primary vault — GOVERNANCE Vault Lock (INFRA-89). Codifies the lock applied
// out-of-band 2026-06: MinRetention 1d, MaxRetention 2555d (~7y), no
// `changeableFor` (LockDate null = admin-removable, GOVERNANCE not
// COMPLIANCE) so it stays adjustable while the plan is validated. This block
// is written to MATCH the live lock exactly, so the deploy diff is a no-op
// adoption — it does not change the live vault.
//
// NOTE: the scoped vault access policy is (re)introduced below (INFRA-94)
// with the fix for the two lockout-class bugs that got it split out of
// INFRA-89: the deny statement now exempts THREE principals via
// StringNotLike on aws:PrincipalArn — the SSO AdministratorAccess role (break
// glass), the backup service role, AND the CDK CFN execution role. The
// CFN-exec-role exemption is MANDATORY: without it CloudFormation cannot
// re-assert the vault lock config / manage the vault and the deploy strands
// the policy (this happened 2026-06-08). NotPrincipal is deliberately NOT
// used (it rejects wildcard ARNs). Governance lock codify + backup
// selections land here.
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
const primaryVault = new backup.BackupVault(this, "PrimaryVault", {
backupVaultName: "seahaven-primary",
encryptionKey: vaultKey,
removalPolicy: cdk.RemovalPolicy.RETAIN,
[INFRA-91/89/16/88/73] Reconcile out-of-band baseline changes + add missing detective controls (#18) * Codify primary vault lock + add backups (INFRA-89, INFRA-88) INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no changeableFor = admin-removable) so it lives in IaC. Values match the live lock exactly, so the deploy is a no-op adoption. Add a scoped vault access policy that denies manual recovery-point deletion and lock/policy tampering to all principals except the AWS Backup service role and the break-glass SSO AdministratorAccess role, so automatic lifecycle expiry still works but humans cannot prune recovery points by hand. Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard ARN matching, so the SSO exemption is expressed as Effect DENY with Principal * and a StringNotLike condition on aws:PrincipalArn, which does support wildcards. This avoids an unrecoverable vault lockout. INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po, extracted-amazon-po, proposal-system uploads + generated) to the phase2-offsite-everything selection. Versioning verified enabled on all 6 against the live account (S3 backup requires versioning). Refs: INFRA-89, INFRA-88 * Promote account trail to organization trail (INFRA-73) INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId (o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket PutObject statement for member-account delivery. CloudTrail org trusted-access is already enabled on the management account. Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey statement for member-account trail delivery, guarded by aws:PrincipalOrgID. The existing single-account statements are preserved so management-account delivery is unaffected. Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption context must be wildcarded across accounts (org-trail shadow trails present the member account id), not pinned to the management account, or member delivery silently fails. Fixed before checkpoint. CHECKPOINT: delicate org-trail KMS/bucket-policy change — code + diff captured for review, NOT deployed. Refs: INFRA-73 * Add secondary-region baseline stacks (INFRA-91, INFRA-16) INFRA-91: codify the Bedrock model-invocation logging applied out-of-band in us-west-2 and us-east-2 (per-region delivery role seahaven-bedrock-invocation-logging-<region> + log group /aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level logging config itself has no CFN resource type and is applied via CLI (already live), same as us-east-1. INFRA-16: add the still-missing us-east-2 detective controls — AWS Config recorder role + delivery bucket (recorder/channel via CLI to avoid the CFN stabilization deadlock seen in us-east-1) and Security Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in us-east-2 and are left for a follow-up adoption to keep this change non-destructive. The us-east-1 baseline stays region-pinned; these are separate RegionalBaselineStack instances composed opt-in per region. CHECKPOINT: new multi-region stacks. The live Bedrock role + log group already exist (CLI-created), so a plain deploy would collide — these need cdk import / changeset adoption, not cdk deploy. Code + diff captured for review, NOT deployed. Refs: INFRA-91, INFRA-16 * Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
2026-06-08 17:03:18 -04:00
lockConfiguration: {
minRetention: cdk.Duration.days(1),
maxRetention: cdk.Duration.days(2555),
// No `changeableFor` → GOVERNANCE mode, admin-removable (matches live).
},
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
});
// Vault access policy (INFRA-94): Deny the destructive recovery-point and
// vault-lifecycle actions to EVERY principal EXCEPT the three operational
// identities below. This is defense-in-depth on top of the GOVERNANCE lock —
// it blocks manual deletion / lifecycle tampering even from accounts that
// hold the equivalent IAM permissions.
//
// Deny (not Allow): a resource-policy Deny overrides any identity-based
// Allow, which is exactly what we want for a guardrail. The StringNotLike
// condition means "this Deny applies UNLESS the caller's ARN matches one of
// the exempted patterns" — i.e. the three exempt principals are NOT denied.
//
// Exemptions (all THREE required):
// 1. SSO AdministratorAccess role — break-glass human admin path. Matched by
// wildcard because the AWSReservedSSO role name carries a permission-set
// hash suffix.
// 2. seahaven-backup-service-role — AWS Backup uses it for lifecycle
// expiry of recovery points; denying it would break the plan's
// deleteAfter cleanup.
// 3. cdk-hnb659fds-cfn-exec-role — the CloudFormation execution role. CFN
// re-asserts the vault lock config and manages the vault on every deploy;
// omitting it strands the policy and fails the deploy (INFRA-94,
// 2026-06-08). Wildcard-suffixed to cover the region-qualified name.
//
// NotPrincipal is intentionally avoided — it does not accept wildcard ARNs.
primaryVault.addToAccessPolicy(
new iam.PolicyStatement({
sid: "DenyDestructiveActionsExceptOperationalRoles",
effect: iam.Effect.DENY,
principals: [new iam.AnyPrincipal()],
actions: [
"backup:DeleteRecoveryPoint",
"backup:UpdateRecoveryPointLifecycle",
"backup:DeleteBackupVault",
"backup:DeleteBackupVaultAccessPolicy",
"backup:DeleteBackupVaultLockConfiguration",
"backup:PutBackupVaultLockConfiguration",
],
resources: ["*"],
conditions: {
StringNotLike: {
"aws:PrincipalArn": [
// 1. SSO AdministratorAccess (break-glass human admin)
`arn:aws:iam::${this.account}:role/aws-reserved/sso.amazonaws.com/*AWSReservedSSO_AdministratorAccess*`,
// 2. AWS Backup service role (lifecycle expiry of recovery points)
`arn:aws:iam::${this.account}:role/seahaven-backup-service-role`,
// 3. CDK CloudFormation execution role (MANDATORY — manages vault)
`arn:aws:iam::${this.account}:role/cdk-hnb659fds-cfn-exec-role-*`,
],
},
},
})
);
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
// Cross-region copy destination, referenced by literal ARN (the offsite
// stack is in another region; a literal ARN avoids crossRegionReferences /
// SSM exports). Stack ordering is enforced via addStackDependency in bin/app.ts.
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
const offsiteVault = backup.BackupVault.fromBackupVaultArn(
this,
"OffsiteVaultRef",
`arn:aws:backup:us-west-2:${this.account}:backup-vault:seahaven-offsite`
);
// AWS Backup service role. Explicit (not auto-generated) because S3 backup
// needs the S3-specific managed policy on top of the standard backup one.
// Least-privilege: BACKUP + S3-backup only. Restore policies
// (AWSBackupServiceRolePolicyForRestores / ...ForS3Restore) and
// BackupSelection allowRestores are intentionally NOT granted — restores
// are a deliberate, audited action and will get their own scoped role/path
// once a restore-test process exists (cross-review F-1/F-2). A known role
// name lets the deploy role's iam:PassRole be scoped to this exact ARN.
// NOTE: creating this role is an IAM change → Sea Haven cross-review gate.
const backupRole = new iam.Role(this, "BackupRole", {
roleName: "seahaven-backup-service-role",
assumedBy: new iam.ServicePrincipal("backup.amazonaws.com"),
description: "AWS Backup service role (backup-only) for seahaven-primary",
managedPolicies: [
iam.ManagedPolicy.fromAwsManagedPolicyName(
"service-role/AWSBackupServiceRolePolicyForBackup"
),
iam.ManagedPolicy.fromAwsManagedPolicyName(
"AWSBackupServiceRolePolicyForS3Backup"
),
],
});
// Daily backup → primary vault (35d), cross-region copy → offsite (90d).
const plan = new backup.BackupPlan(this, "Plan", {
backupPlanName: "seahaven-critical-daily",
backupVault: primaryVault,
backupPlanRules: [
new backup.BackupPlanRule({
ruleName: "daily-crr-offsite",
backupVault: primaryVault,
// 06:00 UTC — offset from the file-share DLM run.
scheduleExpression: events.Schedule.cron({ hour: "6", minute: "0" }),
startWindow: cdk.Duration.hours(1),
completionWindow: cdk.Duration.hours(6),
deleteAfter: cdk.Duration.days(35),
copyActions: [
{
destinationBackupVault: offsiteVault,
deleteAfter: cdk.Duration.days(90),
},
],
}),
],
});
// Phase-1 critical set, by explicit ARN (identifiers verified against the
// live account 2026-05-29).
plan.addSelection("CriticalResources", {
backupSelectionName: "critical-data",
role: backupRole,
// allowRestores omitted (defaults false) — backup-only, see role comment.
resources: [
// RDS. database-1 was retired 2026-06-03 (audit H-19: idle SQL Server
// Express, snapshot-and-delete) — its final recovery point lives in the
// offsite vault; removed from the selection so backup jobs don't fail on
// a missing resource.
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
backup.BackupResource.fromArn(
`arn:aws:rds:us-east-1:${this.account}:db:proposal-system-db`
),
// DynamoDB (financial)
backup.BackupResource.fromArn(
`arn:aws:dynamodb:us-east-1:${this.account}:table/PaymentsDashboard`
),
backup.BackupResource.fromArn(
`arn:aws:dynamodb:us-east-1:${this.account}:table/purchase-orders`
),
// S3 (single-copy critical buckets) — versioning required (see header)
backup.BackupResource.fromArn("arn:aws:s3:::accounting.seahaven.com"),
backup.BackupResource.fromArn(
"arn:aws:s3:::seahaven-payments-csv-328440206208"
),
backup.BackupResource.fromArn(
"arn:aws:s3:::google-workspace-seahavenind.com"
),
],
});
// Phase-2 expansion (audit Day 4): bring the remaining DynamoDB tables and
// all in-use EBS volumes under the same daily plan + cross-region copy to
// the locked offsite vault ("offsite for everything"). Same role and rule
// as phase-1; a separate selection keeps the phase-1 critical set readable.
//
// Still EXPLICIT-ARN (not tag-based) on purpose: the file-share volumes are
// standalone CDK-managed (RETAIN) and the DynamoDB tables are owned by other
// stacks, so tagging them here would drift those stacks — the same reason
// phase-1 avoided tags. Tradeoff: if a volume is replaced (new vol-id) it
// silently drops from this selection; scheduled drift detection + the audit
// re-run are the backstop. Identifiers verified against the live account
// 2026-06-03.
//
// Excluded by intent: the ledgerflow tables — the whole LedgerFlow stack
// was decommissioned 2026-06-03 (audit Day 4), so they no longer exist.
// database-1 was retired the same day and removed from the phase-1 selection
// above (audit H-19).
plan.addSelection("Phase2Resources", {
backupSelectionName: "phase2-offsite-everything",
role: backupRole,
resources: [
// DynamoDB — all remaining tables (10)
...[
"SiteAssignments",
"VendorReplies",
"WorkOrderComments",
"WorkOrders",
"afterhours-shifts",
"front-sla-alerts",
"last-war-bot",
"meal-order-manager-orders",
"pending-site-review",
"verified-sites",
].map((t) =>
backup.BackupResource.fromArn(
`arn:aws:dynamodb:us-east-1:${this.account}:table/${t}`
)
),
// EBS — all 9 in-use volumes (unencrypted sources land encrypted at the
// vault CMK, as the C-7 database-1 smoke-test confirmed)
...[
"vol-05cb0eb5c145d799b", // SeaHavenIndustries-dev
"vol-054cf918f227d88f6", // file-share (20 GiB)
"vol-00f05a5a809697ce5", // forgejo
"vol-04d951cccacc435b5", // file-share NAS (500 GiB)
"vol-07094902194638fff", // syslog-server
"vol-0488e0bad1f9afbfb", // apm-wo-analysis grafana
"vol-0c2cbe9e71517a517", // Mutual Aid Data
"vol-0fe224f13812f47e7", // jump box
"vol-0f0c167f3d7f85542", // last-war-rankings
].map((v) =>
backup.BackupResource.fromArn(
`arn:aws:ec2:us-east-1:${this.account}:volume/${v}`
)
),
[INFRA-91/89/16/88/73] Reconcile out-of-band baseline changes + add missing detective controls (#18) * Codify primary vault lock + add backups (INFRA-89, INFRA-88) INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no changeableFor = admin-removable) so it lives in IaC. Values match the live lock exactly, so the deploy is a no-op adoption. Add a scoped vault access policy that denies manual recovery-point deletion and lock/policy tampering to all principals except the AWS Backup service role and the break-glass SSO AdministratorAccess role, so automatic lifecycle expiry still works but humans cannot prune recovery points by hand. Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard ARN matching, so the SSO exemption is expressed as Effect DENY with Principal * and a StringNotLike condition on aws:PrincipalArn, which does support wildcards. This avoids an unrecoverable vault lockout. INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po, extracted-amazon-po, proposal-system uploads + generated) to the phase2-offsite-everything selection. Versioning verified enabled on all 6 against the live account (S3 backup requires versioning). Refs: INFRA-89, INFRA-88 * Promote account trail to organization trail (INFRA-73) INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId (o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket PutObject statement for member-account delivery. CloudTrail org trusted-access is already enabled on the management account. Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey statement for member-account trail delivery, guarded by aws:PrincipalOrgID. The existing single-account statements are preserved so management-account delivery is unaffected. Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption context must be wildcarded across accounts (org-trail shadow trails present the member account id), not pinned to the management account, or member delivery silently fails. Fixed before checkpoint. CHECKPOINT: delicate org-trail KMS/bucket-policy change — code + diff captured for review, NOT deployed. Refs: INFRA-73 * Add secondary-region baseline stacks (INFRA-91, INFRA-16) INFRA-91: codify the Bedrock model-invocation logging applied out-of-band in us-west-2 and us-east-2 (per-region delivery role seahaven-bedrock-invocation-logging-<region> + log group /aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level logging config itself has no CFN resource type and is applied via CLI (already live), same as us-east-1. INFRA-16: add the still-missing us-east-2 detective controls — AWS Config recorder role + delivery bucket (recorder/channel via CLI to avoid the CFN stabilization deadlock seen in us-east-1) and Security Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in us-east-2 and are left for a follow-up adoption to keep this change non-destructive. The us-east-1 baseline stays region-pinned; these are separate RegionalBaselineStack instances composed opt-in per region. CHECKPOINT: new multi-region stacks. The live Bedrock role + log group already exist (CLI-created), so a plain deploy would collide — these need cdk import / changeset adoption, not cdk deploy. Code + diff captured for review, NOT deployed. Refs: INFRA-91, INFRA-16 * Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
2026-06-08 17:03:18 -04:00
// S3 — additional critical/PII buckets (INFRA-88). Explicit-ARN, same as
// the phase-1 set. Versioning verified enabled on all 6 against the live
// account 2026-06-08 (S3 backup requires versioning). payroll-emails is
// PII → immutable offsite copy is the point of including it.
...[
"seahaven-kb-docs-328440206208",
"seahaven-payroll-emails-328440206208",
"amazon-po",
"extracted-amazon-po",
"proposal-system-uploads-328440206208",
"proposal-system-generated-328440206208",
].map((b) => backup.BackupResource.fromArn(`arn:aws:s3:::${b}`)),
],
});
Add AWS Backup with offsite vault (audit C-7) (#3) * Add AWS Backup with offsite vault (audit C-7) The account had zero AWS Backup vaults/plans, so 22 of 23 data stores had no immutable, cross-region recovery path (audit finding C-7). One ransomware event or rogue delete would erase primary plus same-region snapshots/PITR. Phase 1 ("critical data first") protects the seven highest-risk stores with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in a new us-east-1 vault, copied cross-region into a governance-locked us-west-2 vault. Governance (not compliance) mode first so the plan can be validated before committing to irreversible immutability. The backup service role is backup-only (no restore policies) to stay least-privilege; restores get a separate audited path later. Resources are selected by explicit ARN to avoid drifting the stacks that own them. Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail stack. See the README pre-deploy gates (S3 versioning, database-1 unencrypted copy smoke-test, DynamoDB PITR) before the first run. * Grant AWS Backup service use of vault CMKs The L2 BackupVault does not grant the backup service principal use of a customer-managed key; the synthesized key policy only delegated to account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses KMS grants on the destination key, so without an explicit grant those copy jobs fail — and silently, since the account has no CloudTrail yet. Add backup.amazonaws.com crypto + CreateGrant statements to both vault keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the discipline used on the C-1 CloudTrail key). Same class of bug the C-1 cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
cdk.Tags.of(this).add("Project", "account-baseline");
cdk.Tags.of(this).add("Owner", "adam@seahavenind.com");
cdk.Tags.of(this).add("Environment", "prod");
cdk.Tags.of(this).add("ManagedBy", "cdk");
new cdk.CfnOutput(this, "PrimaryVaultName", { value: "seahaven-primary" });
new cdk.CfnOutput(this, "PrimaryVaultKmsKeyArn", { value: vaultKey.keyArn });
new cdk.CfnOutput(this, "BackupPlanId", { value: plan.backupPlanId });
new cdk.CfnOutput(this, "BackupRoleArn", { value: backupRole.roleArn });
}
}