fix(baseline): drop departed mgmt resources and retain drifted web acl (#171)

Nightly backup jobs fail on resources that have left the management account.
The CloudFront WebACL is already gone while CloudFormation still owns it, so
the deletion policy has to be Retain before a later change can remove it.
This commit is contained in:
Adam Moussa 2026-10-01 23:57:42 +00:00 • committed by GitHub
parent 7ab7346f78
commit 42781f743e
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
6 changed files with 95 additions and 122 deletions

View file

@ -710,12 +710,12 @@ recovery point cross-region into a governance-locked vault.
| Service role | `seahaven-backup-service-role` | **Backup-only** (Backup + S3-Backup managed policies); restore perms intentionally deferred |
**Phase-1 scope** (selected by explicit ARN, not tags, to avoid drifting other
stacks): RDS `proposal-system-db`, DynamoDB `PaymentsDashboard`,
DynamoDB `purchase-orders`, S3 `accounting.seahaven.com`,
`seahaven-payments-csv-328440206208`, `google-workspace-seahavenind.com`.
*(RDS `database-1` was originally in this set but was retired 2026-06-03 —
audit H-19, idle 0 conn/60d — and removed from the selection; its final
encrypted recovery point is retained in `seahaven-offsite` for 7 years.)*
stacks): DynamoDB `purchase-orders`, S3 `accounting.seahaven.com`,
`google-workspace-seahavenind.com`.
*(Removed once the resource was gone: RDS `database-1` on 2026-06-03, audit
H-19, final recovery point retained in `seahaven-offsite`; RDS
`proposal-system-db`, DynamoDB `PaymentsDashboard`, and S3
`seahaven-payments-csv-328440206208` on 2026-10-01.)*
**Coexists with** existing EBS DLM snapshots and DynamoDB PITR — it supplements
them with the missing offsite + immutable leg; it does not replace them.
@ -729,39 +729,43 @@ them with the missing offsite + immutable leg; it does not replace them.
- **Backup-only role.** Restore policies and `allowRestores` are not granted;
restores get a separate audited path once a restore-test process exists.
**Pre-deploy gates** (must clear before the first scheduled run):
**Pre-deploy gates** (cleared before the first scheduled run; historical):
1. Enable S3 versioning on `seahaven-payments-csv-328440206208` and
`google-workspace-seahavenind.com` (`accounting.seahaven.com` already has it,
audit C-9), or their jobs fail silently (folds in H-21).
2. `database-1` is unencrypted (H-19): smoke-test an on-demand backup + copy of
it to us-west-2 first; if the copy fails, encrypt it or drop it from the copy.
3. Enable DynamoDB PITR (H-7) on the two tables for between-window recovery.
1. S3 versioning on the original phase-1 buckets (`accounting.seahaven.com`
already had it, audit C-9). `seahaven-payments-csv-328440206208` has since
been deleted and dropped from the selection.
2. `database-1` was retired 2026-06-03 (H-19) and is not in the selection.
3. DynamoDB PITR (H-7) is independent of this plan.
### AWS Backup phase 2 (audit Day 4)
Expands the same `seahaven-critical-daily` plan to every remaining data store, so
all of DynamoDB + EBS get the offsite + immutable leg ("offsite for everything").
Expands the same `seahaven-critical-daily` plan to the remaining DynamoDB tables
and S3 buckets that still exist in the management account.
| Resource | Logical ID | Notes |
|---|---|---|
| Phase-2 selection | `Plan/Phase2Resources` (`phase2-offsite-everything`) | Same plan, same `seahaven-backup-service-role`, same daily + cross-region copy rule |
**Phase-2 scope:** the 15 remaining DynamoDB tables (all except the two phase-1
financial tables + the deleted ledgerflow tables) and all 9 in-use EBS volumes,
again **by explicit ARN** — tag-based selection was deliberately avoided because
the file-share volumes are standalone-managed and the tables are owned by other
stacks, so tagging here would drift them.
**Phase-2 scope** (explicit ARN): DynamoDB `SiteAssignments`, `VendorReplies`,
`WorkOrderComments`, `WorkOrders`, `last-war-bot`, `pending-site-review`,
`verified-sites`; S3 `amazon-po`, `extracted-amazon-po`.
`amazon-po` stays selected even though its job fails: the bucket exists and
versioning is on. Dropped 2026-10-01 because the resources are gone and the
nightly jobs were failing: DynamoDB `afterhours-shifts`, `front-sla-alerts`,
`meal-order-manager-orders`; every EBS volume; S3
`seahaven-kb-docs-328440206208`, `seahaven-payroll-emails-328440206208`,
`proposal-system-uploads-328440206208`,
`proposal-system-generated-328440206208`. Tag-based selection stays avoided
because the tables are owned by other stacks.
**No IAM change:** `AWSBackupServiceRolePolicyForBackup` already grants the
DynamoDB/RDS/EBS backup actions, so phase 2 reuses the phase-1 role unchanged
(cross-reviewed, no BLOCK).
**Known tradeoff (→ Jira INFRA-31):** explicit-ARN EBS entries go stale if a
volume is replaced (new volume id), silently dropping it from backup. Migrating
the EBS portion to tag-based selection (with the tag codified in each owning
stack) is the resilient follow-up; scheduled drift detection is the interim
backstop.
**Known tradeoff (→ Jira INFRA-31):** explicit-ARN entries go stale when a
resource is replaced or deleted, and the nightly job then fails until the ARN
is removed. The EBS portion of this selection was removed 2026-10-01 after the
instances left the account.
**Also enabled outside this stack (audit H-7, via CLI — codify per stack →
INFRA-30):** PITR + `DeletionProtectionEnabled` on 12 more DynamoDB tables
@ -979,7 +983,7 @@ provides cost alerting independent of that metric.
|---|---|---|---|
| CIS metric filters + alarms | `CisMonitoring/*` | H-1 | 15 filters (CIS 4.1–4.15) on the CloudTrail log group, each with an alarm → `seahaven-cis-alarms`. ALARM-only actions (no OK). 4.16 = Security Hub (Day 1) |
| CIS alarm topic | `seahaven-cis-alarms` | H-1 | SNS, SSE (`alias/aws/sns`), email sub to adam@seahavenind.com |
| VPC flow logs | `FlowLogs/FlowLog0..4` | H-14 | ALL traffic on all 5 VPCs → S3 |
| VPC flow logs | `FlowLogs/FlowLog3`, `FlowLogs/FlowLog4` | H-14 | ALL traffic on `seahaven-vpc` and the default VPC → S3. Slots 0–2 are empty so those logical ids are not reused |
| Flow-logs bucket | `seahaven-vpc-flow-logs-328440206208` | H-14 | Private, SSE-S3, TLS-only, Glacier @90d / expire @365d; delivery bucket policy cross-reviewed |
| SES config set | `seahaven-email-events` | M-13 | Bounce/complaint/reject → CloudWatch metrics for reputation visibility |
| Sensitive-logs CMK | `LogsKey/Key` (`alias/seahaven-logs`) | M-24 | Encrypts sensitive CloudWatch Logs groups. Key policy grants `logs.us-east-1.amazonaws.com` Encrypt*/Decrypt*/ReEncrypt*/GenerateDataKey*/DescribeKey scoped by `kms:EncryptionContext:aws:logs:arn` (required or log delivery breaks). Rotation on, RETAIN. Applied in place to `TrailLogGroup` via escape hatch (same logical id/name). Cross-reviewed |
@ -1037,9 +1041,9 @@ aws cloudwatch put-metric-alarm --alarm-name forgejo-backup-verification-errors
## Roadmap (same stack)
Detective layer multi-region expansion (GuardDuty/Config/Security Hub beyond
us-east-1). Backup: phase 2 is deployed (see above); remaining is migrating the
phase-2 EBS entries to tag-based selection (INFRA-31) and graduating the offsite
vault to compliance mode.
us-east-1). Backup phase 2 is deployed (see above). The EBS entries were
removed 2026-10-01 (INFRA-31's tag-migration follow-up no longer applies).
Remaining backup work is graduating the offsite vault to compliance mode.
## Deploy
@ -1068,7 +1072,7 @@ aws backup describe-backup-vault --backup-vault-name seahaven-offsite --region u
aws backup get-backup-plan --backup-plan-id <id> # daily rule + CopyAction
# Smoke test: on-demand backup of one resource, then confirm the cross-region copy lands
aws backup start-backup-job --backup-vault-name seahaven-primary \
--resource-arn arn:aws:rds:us-east-1:328440206208:db:proposal-system-db \
--resource-arn arn:aws:dynamodb:us-east-1:328440206208:table/purchase-orders \
--iam-role-arn arn:aws:iam::328440206208:role/seahaven-backup-service-role
aws backup list-copy-jobs --region us-west-2 # copy to offsite present + COMPLETED
# Phase-2 selections live on the plan:

View file

@ -22,12 +22,13 @@ const SECURITY_ACCOUNT = "001520130573";
const DEV_ACCOUNT = "710827005802";
const PROD_ACCOUNT = "011934824531";
// All 5 VPCs in 328440206208 / us-east-1 (4 custom + default), audit Agent 7.
// Index-derived logical IDs — append only, never reorder.
const PROD_VPC_IDS = [
"vpc-061d66990b6a4d1fb",
"vpc-0542a9e934b417d23",
"vpc-062d200c68bd4ca0e",
// Index-derived logical IDs — append only, never reorder, never close a hole.
// Slots 0-2 are retired VPCs (FlowLog0, FlowLog1, FlowLog2). Clearing them
// deletes those flow logs. Slots 3 and 4 stay: seahaven-vpc and the default VPC.
const PROD_VPC_IDS: readonly (string | undefined)[] = [
undefined, // vpc-061d66990b6a4d1fb
undefined, // vpc-0542a9e934b417d23
undefined, // vpc-062d200c68bd4ca0e (flow log was still ACTIVE; VPC is gone)
"vpc-0d3d4b67bd0cf8a68",
"vpc-02c10a89d66f6f9b8",
];

View file

@ -35,7 +35,7 @@ export interface AccountBaselineStackProps extends cdk.StackProps {
* VPC ids to attach ALL-traffic flow logs to (H-14). Logical IDs are
* index-derived — only append, never reorder (see lib/flow-logs.ts).
*/
readonly flowLogVpcIds: string[];
readonly flowLogVpcIds: readonly (string | undefined)[];
}
export class AccountBaselineStack extends cdk.Stack {
@ -253,7 +253,14 @@ export class AccountBaselineStack extends cdk.Stack {
});
new SesMonitoring(this, "SesMonitoring");
// Shared CloudFront WAF WebACL (M-17); ARN published to SSM for app stacks.
new AppWebAcl(this, "AppWebAcl");
// RETAIN on the WebACL only. The physical ACL is already gone
// (WAFNonexistentItemException) while this stack still owns it. The
// deletion policy has to be in the live template before the construct is
// removed, or CloudFormation Deletes, fails, and rollback tries to
// recreate it. The SSM parameter stays on Delete.
new AppWebAcl(this, "AppWebAcl", {
webAclRemovalPolicy: cdk.RemovalPolicy.RETAIN,
});
// ── Day 5 AI governance ──
// Bedrock model invocation logging destinations + delivery role (H-20).

View file

@ -21,18 +21,10 @@ import { Construct } from "constructs";
* drift — resources owned by other stacks (proposal-system, payments-dashboard).
* Switch to tag-based selection when expanding past the phase-1 set.
*
* PRE-DEPLOY GATES (validate before the first scheduled run):
* - S3 backup requires bucket versioning. `accounting.seahaven.com` already
* has it (audit C-9); `seahaven-payments-csv-328440206208` and
* `google-workspace-seahavenind.com` must have versioning enabled first or
* their jobs fail silently (folds in audit H-21).
* - `database-1` is unencrypted (audit H-19). Cross-region copy of an
* unencrypted RDS recovery point may fail or land unencrypted. Smoke-test
* an on-demand backup of `database-1` FIRST and confirm the copy job to
* us-west-2 succeeds; if not, encrypt database-1 (H-19) or drop it from the
* copy until then.
* - DynamoDB PITR (H-7) is independent of this plan; enable it on the two
* tables for between-window point-in-time recovery.
* Departed resources were removed from the selections on 2026-10-01 so the
* daily job stops failing on missing ARNs. The vaults, plan, role, and lock
* stay. `amazon-po` stays selected: the bucket exists and is versioned, and
* its job fails for a reason that is not "bucket gone".
*/
export class BackupStack extends cdk.Stack {
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
@ -216,71 +208,48 @@ export class BackupStack extends cdk.Stack {
],
});
// Phase-1 critical set, by explicit ARN (identifiers verified against the
// live account 2026-05-29).
// Phase-1 critical set, by explicit ARN.
// Removed 2026-10-01 (resources gone; jobs were failing or about to):
// RDS proposal-system-db, DynamoDB PaymentsDashboard, S3
// seahaven-payments-csv-328440206208. database-1 was removed 2026-06-03
// (audit H-19); its final recovery point stays in the offsite vault.
plan.addSelection("CriticalResources", {
backupSelectionName: "critical-data",
role: backupRole,
// allowRestores omitted (defaults false) — backup-only, see role comment.
resources: [
// RDS. database-1 was retired 2026-06-03 (audit H-19: idle SQL Server
// Express, snapshot-and-delete) — its final recovery point lives in the
// offsite vault; removed from the selection so backup jobs don't fail on
// a missing resource.
backup.BackupResource.fromArn(
`arn:aws:rds:us-east-1:${this.account}:db:proposal-system-db`
),
// DynamoDB (financial)
backup.BackupResource.fromArn(
`arn:aws:dynamodb:us-east-1:${this.account}:table/PaymentsDashboard`
),
backup.BackupResource.fromArn(
`arn:aws:dynamodb:us-east-1:${this.account}:table/purchase-orders`
),
// S3 (single-copy critical buckets) — versioning required (see header)
backup.BackupResource.fromArn("arn:aws:s3:::accounting.seahaven.com"),
backup.BackupResource.fromArn(
"arn:aws:s3:::seahaven-payments-csv-328440206208"
),
backup.BackupResource.fromArn(
"arn:aws:s3:::google-workspace-seahavenind.com"
),
],
});
// Phase-2 expansion (audit Day 4): bring the remaining DynamoDB tables and
// all in-use EBS volumes under the same daily plan + cross-region copy to
// the locked offsite vault ("offsite for everything"). Same role and rule
// as phase-1; a separate selection keeps the phase-1 critical set readable.
// Phase-2 selection, same plan and role as phase-1. Explicit ARN, not tags:
// the tables are owned by other stacks.
//
// Still EXPLICIT-ARN (not tag-based) on purpose: the DynamoDB tables are
// owned by other stacks, so tagging them here would drift those stacks.
// File-share left this account (PLAT-77). vol-054cf918f227d88f6 does not
// exist. vol-04d951cccacc435b5 is out of this selection because the
// rollback hold ended and that volume is deleted. Tradeoff: if a volume
// is replaced (new vol-id) it
// silently drops from this selection; scheduled drift detection and the
// audit re-run are the backstop. Identifiers verified against the live
// account 2026-06-03.
//
// Excluded by intent: the ledgerflow tables — the whole LedgerFlow stack
// was decommissioned 2026-06-03 (audit Day 4), so they no longer exist.
// database-1 was retired the same day and removed from the phase-1 selection
// above (audit H-19).
// Removed 2026-10-01 (resources gone; jobs failing nightly): DynamoDB
// afterhours-shifts, front-sla-alerts, meal-order-manager-orders; every
// EBS volume (no instances remain); S3 seahaven-kb-docs-328440206208,
// seahaven-payroll-emails-328440206208,
// proposal-system-uploads-328440206208,
// proposal-system-generated-328440206208.
// amazon-po stays. The bucket exists and versioning is enabled, but the
// backup job fails with a generic message.
// LedgerFlow tables were excluded 2026-06-03 when that stack was deleted.
plan.addSelection("Phase2Resources", {
backupSelectionName: "phase2-offsite-everything",
role: backupRole,
resources: [
// DynamoDB — all remaining tables (10)
...[
"SiteAssignments",
"VendorReplies",
"WorkOrderComments",
"WorkOrders",
"afterhours-shifts",
"front-sla-alerts",
"last-war-bot",
"meal-order-manager-orders",
"pending-site-review",
"verified-sites",
].map((t) =>
@ -288,32 +257,9 @@ export class BackupStack extends cdk.Stack {
`arn:aws:dynamodb:us-east-1:${this.account}:table/${t}`
)
),
// EBS — remaining in-use volumes (unencrypted sources land encrypted at
// the vault CMK, as the C-7 database-1 smoke-test confirmed)
...[
"vol-05cb0eb5c145d799b", // SeaHavenIndustries-dev
"vol-00f05a5a809697ce5", // forgejo
"vol-07094902194638fff", // syslog-server
"vol-0c2cbe9e71517a517", // Mutual Aid Data
"vol-0fe224f13812f47e7", // jump box
"vol-0f0c167f3d7f85542", // last-war-rankings
].map((v) =>
backup.BackupResource.fromArn(
`arn:aws:ec2:us-east-1:${this.account}:volume/${v}`
)
...["amazon-po", "extracted-amazon-po"].map((b) =>
backup.BackupResource.fromArn(`arn:aws:s3:::${b}`)
),
// S3 — additional critical/PII buckets (INFRA-88). Explicit-ARN, same as
// the phase-1 set. Versioning verified enabled on all 6 against the live
// account 2026-06-08 (S3 backup requires versioning). payroll-emails is
// PII → immutable offsite copy is the point of including it.
...[
"seahaven-kb-docs-328440206208",
"seahaven-payroll-emails-328440206208",
"amazon-po",
"extracted-amazon-po",
"proposal-system-uploads-328440206208",
"proposal-system-generated-328440206208",
].map((b) => backup.BackupResource.fromArn(`arn:aws:s3:::${b}`)),
],
});

View file

@ -10,10 +10,10 @@ import { Construct } from "constructs";
* Athena. ALL traffic (accept + reject).
*
* Serves both the management-account baseline and member-account baselines:
* VPC ids are passed via props (the management account pins its 5 audited
* VPCs in bin/app.ts; member accounts source theirs from cdk context because
* their VPCs change over time). Pass an empty list to create the hardened
* destination bucket without any flow logs attached yet.
* VPC ids are passed via props (the management account pins a stable-index
* list in bin/app.ts; member accounts pass a dense list). Pass an empty list
* to create the hardened destination bucket without any flow logs attached
* yet. An empty slot keeps its index so later logical ids do not shift.
*
* S3 delivery needs no IAM role; instead the bucket policy grants the
* `delivery.logs.amazonaws.com` service principal write access, scoped to this
@ -23,11 +23,12 @@ export interface FlowLogsProps {
/** Physical-name prefix for the destination bucket (e.g. "seahaven" or "seahaven-extdev"). */
readonly namePrefix: string;
/**
* VPC ids to attach ALL-traffic flow logs to. May be empty. Logical IDs are
* VPC ids to attach ALL-traffic flow logs to. May be empty. A hole
* (`undefined`) reserves that index and creates no flow log. Logical IDs are
* index-derived (FlowLog0, FlowLog1, ...) — REORDERING this list replaces
* deployed flow logs; only append.
* deployed flow logs; only append, and never close a hole.
*/
readonly vpcIds: string[];
readonly vpcIds: readonly (string | undefined)[];
}
export class FlowLogs extends Construct {
@ -98,6 +99,7 @@ export class FlowLogs extends Construct {
);
props.vpcIds.forEach((vpcId, i) => {
if (!vpcId) return;
const flowLog = new ec2.CfnFlowLog(this, `FlowLog${i}`, {
resourceId: vpcId,
resourceType: "VPC",

View file

@ -13,8 +13,18 @@ import { Construct } from "constructs";
* The ARN is published to SSM (`/seahaven/waf/app-web-acl-arn`) so app stacks in
* other repos can consume it via `{{resolve:ssm:...}}` without a hard CFN export.
*/
export interface AppWebAclProps {
/**
* Deletion policy for the WebACL only. The SSM parameter keeps the default
* Delete policy so a later stack update can remove the parameter. Prod does
* not set this. Management sets RETAIN first because the live WebACL is
* already gone and a Delete call would fail and roll back into a recreate.
*/
readonly webAclRemovalPolicy?: cdk.RemovalPolicy;
}
export class AppWebAcl extends Construct {
constructor(scope: Construct, id: string) {
constructor(scope: Construct, id: string, props?: AppWebAclProps) {
super(scope, id);
const vis = (metric: string): wafv2.CfnWebACL.VisibilityConfigProperty => ({
@ -64,6 +74,9 @@ export class AppWebAcl extends Construct {
},
],
});
if (props?.webAclRemovalPolicy) {
webAcl.applyRemovalPolicy(props.webAclRemovalPolicy);
}
new ssm.StringParameter(this, "AppWebAclArnParam", {
parameterName: "/seahaven/waf/app-web-acl-arn",