CI and deploy both run on Node 24; @types/node was stale at ^22. Dependabot
proposed jumping to ^25 (#2), but Node 25 is a non-LTS, non-Lambda release
ahead of the runtime. Pinning to ^24 to match. Build + cdk synth verified clean.
Previously the L2 cloudtrail.Trail auto-created a log group with a
CDK-generated hash suffix in its name. The 15 CIS Section 4 metric
filters in CisMonitoring imported that group by the hardcoded generated
name. If the Trail or log group was ever recreated the suffix changes
and all 15 filters would silently detach with no error, leaving the
account unmonitored.
Replace with an explicit logs.LogGroup named
seahaven-account-baseline-trail-logs (stable, no hash suffix) with
RemovalPolicy.RETAIN. Pass the CDK object — not a name constant — to
the Trail via cloudWatchLogGroup and forward it to CisMonitoring via a
new trailLogGroup prop. All 15 filters now reference the CDK object so
they can never drift from the group the Trail actually delivers to.
The old auto-named log group is orphaned by this deploy (CloudFormation
loses track of it and does not delete it). Historical audit logs in the
old group remain accessible in CloudWatch under the old name; no audit
history is destroyed.
Refs: INFRA-19
The L1 AWS::Config::ConfigurationRecorder deadlocks the CDK stack
(recorder can't complete without a delivery channel; channel can't be
created without a recorder — observed 2026-06-01).
Fix: three AwsCustomResource nodes call PutConfigurationRecorder →
PutDeliveryChannel → StartConfigurationRecorder in sequence. Put* is
an idempotent upsert, so the deploy adopts the existing CLI-created
recorder and channel without destroying or interrupting them. onDelete
stops recording rather than deleting the per-account singleton.
New IAM permissions on the custom-resource role (cross-reviewed,
GPT-4.1 APPROVE — no BLOCK):
config:PutConfigurationRecorder
config:PutDeliveryChannel
config:StartConfigurationRecorder
config:StopConfigurationRecorder
iam:PassRole → seahaven-config-recorder-role (service=config)
cdk diff shows [+] adds only — no existing resources destroyed or
replaced. Removes README note that recorder/channel are CLI-only.
Refs: INFRA-17
* [INFRA-96] CMK-encrypt sensitive CloudWatch log groups (M-24)
Add a dedicated customer-managed CMK (alias/seahaven-logs) for encrypting
the sensitive CloudWatch Logs groups (CloudTrail + finance/PII Lambdas).
- lib/logs-key.ts: LogsKey construct. Key policy grants the CloudWatch Logs
service principal (logs.us-east-1.amazonaws.com) Encrypt*/Decrypt*/
ReEncrypt*/GenerateDataKey*/DescribeKey, scoped by the
kms:EncryptionContext:aws:logs:arn condition (REQUIRED per AWS docs or log
delivery breaks). Cross-reviewed (GPT-4.1): tightened Describe* -> DescribeKey;
CreateGrant omitted (not needed for plain log-group encryption).
- account-baseline-stack.ts: instantiate LogsKey and set KmsKeyId on the L2
Trail's CloudWatch log group in place (escape hatch on the existing
AWS::Logs::LogGroup) so it keeps the same logical id + physical name -
additive, no replacement, CIS Section-4 metric filters (which import the
group by name) keep working, live audit trail not disrupted. Gated by
context `encryptTrailLogGroup` so the CMK can be smoke-tested on a low-risk
Lambda group before the most-sensitive CloudTrail group.
Finance/PII Lambda log groups (exec-aide-*, payments-*, po-email-processor,
vendor-reply-processor) are owned by other stacks and associated to this CMK
via the CLI for now; codifying KmsKeyId in those repos is tracked as drift.
* [INFRA-96] Document sensitive-logs CMK (M-24) in README
Reintroduce the scoped vault access policy that was split out of INFRA-89
after two lockout-class bugs. Adds a Deny on the destructive recovery-point
and vault-lifecycle actions (DeleteRecoveryPoint, UpdateRecoveryPointLifecycle,
DeleteBackupVault, DeleteBackupVaultAccessPolicy,
DeleteBackupVaultLockConfiguration, PutBackupVaultLockConfiguration) for every
principal except three exempted operational identities via StringNotLike on
aws:PrincipalArn:
1. SSO AdministratorAccess role (break-glass human admin)
2. seahaven-backup-service-role (AWS Backup lifecycle)
3. cdk-hnb659fds-cfn-exec-role-* (CloudFormation manages the vault)
The CFN-exec-role exemption is the fix for the 2026-06-08 strand failure: without
it CloudFormation cannot re-assert the vault lock config and the deploy strands
the policy. Uses Deny + AnyPrincipal + StringNotLike (not NotPrincipal, which
rejects wildcard ARNs). aws:PrincipalArn normalizes assumed-role sessions to the
IAM role ARN, so the iam::role/ ARN forms are correct (AWS docs: "Do not specify
the assumed role session ARN as a value for this condition key").
Deployed and verified: deploy succeeded (proves exec role not locked out),
access policy present with all three exemptions, vault still Locked
(min1/max2555, LockDate null, 168 RPs), follow-up cdk diff clean (no drift).
* Codify primary vault lock + add backups (INFRA-89, INFRA-88)
INFRA-89: codify the GOVERNANCE Vault Lock applied out-of-band on the
seahaven-primary vault (MinRetention 1d, MaxRetention 2555d, no
changeableFor = admin-removable) so it lives in IaC. Values match the
live lock exactly, so the deploy is a no-op adoption.
Add a scoped vault access policy that denies manual recovery-point
deletion and lock/policy tampering to all principals except the AWS
Backup service role and the break-glass SSO AdministratorAccess role,
so automatic lifecycle expiry still works but humans cannot prune
recovery points by hand.
Cross-review (GPT-4.1) BLOCK: NotPrincipal does not support wildcard
ARN matching, so the SSO exemption is expressed as Effect DENY with
Principal * and a StringNotLike condition on aws:PrincipalArn, which
does support wildcards. This avoids an unrecoverable vault lockout.
INFRA-88: add 6 S3 buckets (kb-docs, payroll-emails [PII], amazon-po,
extracted-amazon-po, proposal-system uploads + generated) to the
phase2-offsite-everything selection. Versioning verified enabled on
all 6 against the live account (S3 backup requires versioning).
Refs: INFRA-89, INFRA-88
* Promote account trail to organization trail (INFRA-73)
INFRA-73: set isOrganizationTrail on seahaven-org-trail and pass orgId
(o-9kufuzz6b4) so the L2 Trail attaches the AWSLogs/<org-id>/* bucket
PutObject statement for member-account delivery. CloudTrail org
trusted-access is already enabled on the management account.
Broaden the KMS key policy with an org-scoped GenerateDataKey/DescribeKey
statement for member-account trail delivery, guarded by
aws:PrincipalOrgID. The existing single-account statements are
preserved so management-account delivery is unaffected.
Cross-review (GPT-4.1) BLOCK: the member KMS SourceArn and encryption
context must be wildcarded across accounts (org-trail shadow trails
present the member account id), not pinned to the management account,
or member delivery silently fails. Fixed before checkpoint.
CHECKPOINT: delicate org-trail KMS/bucket-policy change — code +
diff captured for review, NOT deployed.
Refs: INFRA-73
* Add secondary-region baseline stacks (INFRA-91, INFRA-16)
INFRA-91: codify the Bedrock model-invocation logging applied
out-of-band in us-west-2 and us-east-2 (per-region delivery role
seahaven-bedrock-invocation-logging-<region> + log group
/aws/bedrock/model-invocations 90d, CloudWatch-only). The account-level
logging config itself has no CFN resource type and is applied via CLI
(already live), same as us-east-1.
INFRA-16: add the still-missing us-east-2 detective controls — AWS
Config recorder role + delivery bucket (recorder/channel via CLI to
avoid the CFN stabilization deadlock seen in us-east-1) and Security
Hub with FSBP + CIS v3.0. GuardDuty + flow logs already live in
us-east-2 and are left for a follow-up adoption to keep this change
non-destructive.
The us-east-1 baseline stays region-pinned; these are separate
RegionalBaselineStack instances composed opt-in per region.
CHECKPOINT: new multi-region stacks. The live Bedrock role + log group
already exist (CLI-created), so a plain deploy would collide — these
need cdk import / changeset adoption, not cdk deploy. Code + diff
captured for review, NOT deployed.
Refs: INFRA-91, INFRA-16
* Drop vault access policy from this deploy; tracked in INFRA-94 (kept governance lock codify + selection)
* Add dependency-review caller workflow
Add a pull_request-triggered caller that invokes the org-level
callable-dependency-review workflow to scan dependency changes and
fail on high-severity advisories.
* chore: retrigger checks
* chore: retrigger dep review (post-fix)
Audit L-14, plus a latent Day-2 bug: seahaven-cis-alarms was
encrypted with the AWS-managed alias/aws/sns key, whose policy cannot
grant cloudwatch.amazonaws.com - CloudWatch alarms silently fail to
publish to topics it encrypts. All 15 CIS alarms would have fired
into the void.
New customer-managed key (rotation on) grants CloudWatch
GenerateDataKey*/Decrypt/DescribeKey scoped by SourceAccount. The
unmanaged site-alerts topic now uses the same key (set via CLI).
Cross-reviewed: no BLOCKs. Verified: forced ALARM on the payroll DLQ
alarm published successfully through the encrypted site-alerts.
Audit finding H-20: no audit trail of model I/O for seahaven-alex,
which returns payments, invoices, WO/PO, and HR/SA8000 data.
S3 bucket (Glacier at 90d, expire 365d) + CloudWatch log group (90d)
+ delivery role assumable only by bedrock.amazonaws.com scoped by
SourceAccount/SourceArn. The account-level logging configuration has
no CloudFormation resource type, so it is applied via CLI post-deploy
(documented in the construct header) - same pattern as the Config
recorder (INFRA-17).
Cross-reviewed: no BLOCKs. Verified live: converse invocation logged
to /aws/bedrock/model-invocations.
cfn-stack-decommission.sh: report-by-default stack retirement; pre-flight
predicts DeletionPolicy:Retain orphans + consumed-export blocks before delete
(distilled from the LedgerFlow decommission). --execute to act.
resource-usage-probe.sh: is-it-used probe (RDS connections/Lambda invocations/
DDB capacity/EBS attachment) to choose retire-vs-harden before acting on an
encrypt/migrate finding (the database-1 H-19 lesson).
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.
scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
database-1 (audit H-19) and the LedgerFlow stack (incl. ledgerflow-pos) were
decommissioned 2026-06-03. Drop database-1 from the critical-data selection and
ledgerflow-pos from phase2-offsite-everything so daily jobs don't target missing
resources. database-1's final recovery point is retained encrypted in the
seahaven-offsite vault (7yr); ledgerflow-pos has a final on-demand DynamoDB
backup. Deployed before merge (seahaven-backup UPDATE_COMPLETE).
Add a second BackupSelection 'phase2-offsite-everything' on the existing
seahaven-critical-daily plan covering the 15 remaining DynamoDB tables and
all 9 in-use EBS volumes, with the same daily backup + cross-region copy to
the GOVERNANCE-locked seahaven-offsite vault ('offsite for everything').
Reuses seahaven-backup-service-role (AWSBackupServiceRolePolicyForBackup
already grants DDB/RDS/EBS) - no IAM change. Explicit-ARN (not tag-based) to
avoid drifting the standalone file-share volumes and stack-owned tables, same
as phase-1. The 4 deprecated ledgerflow delete-targets are excluded.
Cross-reviewed (no BLOCK). Follow-up: migrate EBS to tag-based selection with
tags codified in owning stacks for resilience to volume replacement.
Deployed to seahaven-backup before merge; selection verified live (24 resources).
seahaven-app-waf (CLOUDFRONT scope, us-east-1): AWS managed Common + Known Bad
Inputs rule groups + per-IP rate limit (2000/5min). ARN published to SSM
/seahaven/waf/app-web-acl-arn for app stacks (meal-order/orders) to consume.
seahaven.com already has its own WAF; ledgerflow is being decommissioned
(INFRA-26); proposal-system-web skipped (not live).
* Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10)
Adds to the seahaven-account-baseline stack:
- AWS Config recorder (all + global resources) + delivery channel + role +
hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed.
- GuardDuty detector, us-east-1 (H-3)
- Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4)
- IAM Access Analyzer, account scope (M-5)
- Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to
adam@seahavenind.com (M-10)
Scope us-east-1 only (all workloads here); multi-region is a follow-up.
The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented
separately in the README runbook.
* Document Day 1 detective layer + CLI governance toggles in README
* Move Config recorder+channel to CLI (L1 stabilization deadlock)
The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches
CREATE_COMPLETE until recording is active (needs a delivery channel), and the
delivery channel cannot be created until the recorder completes — a deadlock
that hung the deploy ~27 min before manual cancel (2026-06-01).
Keep the cross-reviewed recorder role + delivery bucket in IaC; create the
recorder, delivery channel, and start recording via CLI (documented in README).
Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls
evaluate once Config is recording. Verified live: recording=true, SUCCESS.
* Add AWS Backup with offsite vault (audit C-7)
The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.
Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.
The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.
Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.
* Grant AWS Backup service use of vault CMKs
The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.
Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.