Commit graph

47 commits

Author SHA1 Message Date
Adam Moussa
ca179bbdf6
chore(terraform-substrate): forget imported prod hcptf pairs (PLAT-147) (#164)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
The six pairs are already in HCP state and DeletionPolicy is Retain, so CloudFormation drops the logical IDs without deleting the roles.
2026-09-28 15:56:04 -04:00
Adam Moussa
2f2858f885
fix(iam): let the site plan role describe SSM parameters (PLAT-225) (#154)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* fix(iam): let the site plan role describe SSM parameters

* fix(iam): address review feedback
2026-09-25 16:40:34 +00:00
Adam Moussa
38ed1bfe00
feat(iam): move seahaven-site exec roles into their own stack (PLAT-225) (#153)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): move seahaven-site exec roles into their own stack (PLAT-225)

Drop the retained roles from the substrate template so the new stack can import them without a second owner.

* fix(iam): address review feedback
2026-09-24 23:37:39 +00:00
Adam Moussa
7e41625e4b
feat(iam): allow PassRole to ECS tasks and EventBridge Scheduler (#151)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
A first apply of Fargate and Scheduler targets cannot create those service-linked attachments while PassRole is Lambda-only.
2026-09-21 18:59:08 +00:00
Adam Moussa
960e4619b4
fix(iam): allow frontend HCP apply to write deploy SSM and githubdeploy trust (PLAT-212) (#150)
Some checks failed
Deploy / deploy-management (push) Has been cancelled
Deploy / deploy-external-dev (push) Has been cancelled
Deploy / deploy-security (push) Has been cancelled
Deploy / deploy-dev (push) Has been cancelled
Deploy / deploy-prod (push) Has been cancelled
* fix(iam): allow frontend HCP apply to write deploy SSM and githubdeploy trust (PLAT-212)

* fix(iam): grant frontend HCP plan named SSM describe and tag reads (PLAT-212)

* fix(iam): allow frontend githubdeploy to read deploy SSM (PLAT-212)

HCP apply already writes /shoc-frontend-new/<env>/deploy/*, but the
githubdeploy ceiling omitted GetParameter so Deploy Web cannot resolve
bucket and distribution after origin moves to the bucket root.

* fix(iam): allow staging HCP apply to update the SHOC backend EB stack (PLAT-213)

* fix(iam): allow staging HCP apply to use the Elastic Beanstalk bucket (PLAT-213)

* fix(iam): allow staging HCP apply to copy the current release zip (PLAT-213)

* fix(iam): allow staging HCP apply versioned ACLs on EB env objects (PLAT-213)

* fix(iam): give staging HCP apply the proven Elastic Beanstalk bucket grants (PLAT-213)

* fix(iam): allow staging HCP apply to write CloudFormation template buckets (PLAT-213)

* fix(iam): let staging HCP apply read Elastic Beanstalk service templates (PLAT-213)
2026-09-18 21:37:05 +00:00
Adam Moussa
db9465deda
fix(iam): allow shoc-backend HCP apply to write deploy SSM and matching githubdeploy trust (PLAT-148) (#148)
* fix(iam): allow shoc-backend HCP apply to write deploy SSM and matching githubdeploy trust (PLAT-148)

* fix(iam): grant shoc-backend staging plan named inventory reads (PLAT-148)

* fix(iam): allow staging githubdeploy to GetObject release zips (PLAT-148)

* fix(iam): allow staging githubdeploy to write EB processed extensions (PLAT-148)

* fix(iam): allow staging githubdeploy GetObjectAcl on release zips (PLAT-148)

* fix(iam): grant staging githubdeploy named S3 reads on EB resources prefix (PLAT-148)

* fix(iam): allow staging githubdeploy to delete EB version cache objects (PLAT-148)

* fix(iam): scope staging githubdeploy S3 object access to the EB bucket (PLAT-148)

* fix(iam): allow staging githubdeploy PutObjectVersionAcl on EB artifacts (PLAT-148)

* fix(iam): allow staging githubdeploy GetBucketPolicy on the EB bucket (PLAT-148)

* fix(iam): scope staging githubdeploy S3 objects to SHOC and staging EB prefixes (PLAT-148)
2026-09-18 18:13:01 +00:00
Adam Moussa
a492a45e07
chore(iam): remove frontend tf-poc substrate after teardown (PLAT-194) (#146)
Some checks failed
Deploy / deploy-management (push) Has been cancelled
Deploy / deploy-external-dev (push) Has been cancelled
Deploy / deploy-security (push) Has been cancelled
Deploy / deploy-dev (push) Has been cancelled
Deploy / deploy-prod (push) Has been cancelled
2026-09-11 22:17:00 +00:00
Adam Moussa
5d613c73bc
feat(iam): allow frontend tf-poc HCP apply destroy (PLAT-193) (#145) 2026-09-11 21:46:59 +00:00
Adam Moussa
3c54df6341
fix(iam): allow GitHub frontend deploy roles to GetDistribution (PLAT-192) (#144)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
Verify and live-state summary call get-distribution; the identity policy already granted it, but the permissions boundary denied the action.
2026-09-11 19:37:24 +00:00
Adam Moussa
60b978aa2a
feat(iam): allow frontend HCP apply to own release pointer and invalidation (PLAT-188) (#143)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): allow frontend HCP apply to own release pointer and invalidation (PLAT-188)

Plan and apply roles can read .release/current; apply can PutObject that key and CreateInvalidation on the exact distribution.

* fix(iam): allow frontend HCP roles to tag the release pointer (PLAT-188)

Terraform aws_s3_object lists object tags on every refresh, so plan and apply need GetObjectTagging and apply needs PutObjectTagging on the exact .release/current key.
2026-09-11 17:37:42 +00:00
Adam Moussa
0c6f307b61
feat(iam): allow frontend HCP apply to update exact CloudFront resources (PLAT-187) (#142)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): allow frontend HCP apply to update exact CloudFront resources (PLAT-187)

Phase 2 ownership tags cannot apply while UpdateDistribution and UpdateFunction are denied on *. Allow those two actions only on the pinned distribution and function ARNs.

* fix(iam): allow PublishFunction on exact frontend CloudFront functions (PLAT-187)

The AWS provider publishes after UpdateFunction, including tag-only applies, so denying PublishFunction on * still blocked Phase 2 function updates.
2026-09-11 15:08:21 +00:00
Adam Moussa
a829854cd0
feat(hcp): flag workspaces that skip source-path file triggers (PLAT-183) (#141)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
2026-09-10 21:00:33 +00:00
Adam Moussa
6bc4f6e095
chore(iam): remove backend tf-poc boundaries (#139)
Some checks failed
Deploy / deploy-management (push) Has been cancelled
Deploy / deploy-external-dev (push) Has been cancelled
Deploy / deploy-security (push) Has been cancelled
Deploy / deploy-dev (push) Has been cancelled
Deploy / deploy-prod (push) Has been cancelled
2026-09-03 15:04:58 +00:00
Adam Moussa
4f0d84cddb
fix(iam): use unique HCP bootstrap workspace names (PLAT-143) (#138)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
HCP workspace names are org-unique, so prod and dev cannot both be iam-bootstrap. Pin trust to iam-bootstrap-prod and iam-bootstrap-dev.
2026-09-02 15:51:39 +00:00
Adam Moussa
b02f52b805
feat(iam): lock app-owned HCP IAM and add bootstrap SCP (PLAT-143) (#137)
* feat(iam): lock app-owned HCP IAM and add bootstrap SCP (PLAT-143)

* fix(iam): pin HCP boundary ARNs and bootstrap trust window (PLAT-143)

Null on iam:PermissionsBoundary accepted any ceiling, including AdministratorAccess. Import apply cannot self-mutate hcptf-* while bootstrap trust is iam-bootstrap only; add a time-boxed exact StringEquals workspace grant instead of StringLike.
2026-09-02 15:22:48 +00:00
Adam Moussa
6a0713f49d
feat(iam): add frontend Terraform substrate (#133)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): add frontend Terraform substrate

* fix(iam): align frontend Terraform substrate

* feat(iam): enable frontend live Terraform roles
2026-08-31 02:25:47 +00:00
Adam Moussa
08191ded4c
chore(iam): finalize backend role ownership (#132)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* chore(iam): finalize backend role ownership

* fix(iam): complete backend import permissions
2026-08-30 20:12:54 +00:00
Adam Moussa
dba0871587
feat(iam): add external-dev backend Terraform substrate (#131)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(iam): add external-dev backend terraform substrate

* fix(iam): require boundaries for SHOC policy writes
2026-08-29 21:04:40 +00:00
0f84d7808b
feat(iam): add per-workload lambda execution boundaries
Shared seahaven-lambda-execution-boundary stays unchanged for live roles.
New named policies plus an enumerated StringEquals allow-list unblock the
next PLAT-71 widen without growing the 6144-character shared document.
2026-08-13 16:47:53 -04:00
Adam Moussa
a6f22880db
feat(waf): add seahaven-prod shared CloudFront WebACL (PLAT-92) (#96)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* feat(waf): add seahaven-prod shared CloudFront WebACL stack

Stand up AppWebAcl in a thin prod stack and widen seahaven-site HCP
roles to read the SSM ARN so CloudFront can associate the ACL in-account.

* fix(deploy): add app-web-acl-prod to deploy.yaml
2026-08-07 17:07:04 -04:00
Adam Moussa
9ee4d4a3d7
docs(iam): codify hcp terraform migration checklist from PLAT-56 (#79)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
Expand the README playbook to steps 0–10 and document the required
plan-refresh sidecar plus prefix-scoped apply-role wildcards so the
next workload copies afi patterns instead of relearning first-apply misses.
2026-08-05 16:14:16 -04:00
08a1d41b05
docs(iam): resolve confirmed review findings from both INFRA-186 gates
Cross-family round 1 plus the /sh-security-review verifier confirmed 11
findings on the floor reduction, all documentation defects; no policy
statement changes. The one HIGH: the Terraform migration checklist never
widened the boundary, so a Lambda-bearing Terraform migration would deploy
green and lose every data-plane call at first invoke. Checklist step 2 now
carries the widening requirement, step 3 verifies deployed boundary content,
and the terraform-substrate header no longer reads as 'Terraform path
unaffected'. Also corrected: Description is a REPLACEMENT property (a
Description edit wedges the custom-named policy and CFN's remedy is the
forbidden rename), the sanctioned-source contradiction, the false
AWSLambdaVPCAccessExecutionRole parity claim, the KMS log-group category
error, stale size numbers (691/5,453), the same-PR widening contradiction,
per-workload residue text, a LoggingConfig silent-log-loss note, the
us-east-1 region pin rationale, and ENI DoS deferral now tracked as
INFRA-200.
2026-07-31 13:44:21 -04:00
a1086e04fb
docs(readme): describe the boundary floor, not the superseded prefix design 2026-07-31 13:23:20 -04:00
b264f74f01
fix(iam): correct two boundary-scoping defects found in review
Post-implementation verification of the INFRA-186 prod/dev scoping found two
functional defects that would have denied permissions the migrating stacks
actually need. Neither is live today (prod/dev boundary usage is 0), but both
would have surfaced as AccessDenied at first migration.

- KMS: the ViaService list omitted ssm., while SSMParameterRead in the same
  policy grants ssm:GetParameter*. A SecureString read decrypts via the SSM
  service principal, so the boundary denied reads it also granted.
- S3: payments-dashboard was classified read-only from the template's own
  permission-source comment, but that enumeration is incomplete -- the real
  stack grants s3:PutObject on BoaRawBucket (template.yaml:272, 1098-1099).
  Write is now allowed on seahaven-payments-boa-raw-* only; payroll-emails and
  payments-csv stay read-only, preserving the evidence-deletion protection.
  The seahaven-payments-* wildcard is replaced by the three literal bucket
  names, verified against payments-dashboard/template.yaml.

Not changed: SES configuration-set/*. Review claimed dropping it rested on a
false premise; verified live -- prod and dev both have ZERO configuration sets
and member-baseline-stack.ts:44 excludes SES monitoring. The drop is correct.

README: the 'substrate changes must edit both files' rule is now false for the
boundary specifically, and said so uniformly. Corrected to distinguish the
deliberately divergent boundary from the still-at-parity substrate resources.
2026-07-30 18:03:40 -04:00
981960433f
fix(iam): scope Terraform guardrail role writes to a Terraform-owned path
Security review (6 detectors + proof-or-kill verifier) confirmed 1 critical and
1 high in the first revision, both inherited by mirroring the SAM copy's
Resource "*" role grants:

- C1 (critical): iam:UpdateAssumeRolePolicy on "*" with DenySelfMutation
  covering only three name patterns lets the principal repoint the
  AdministratorAccess CDK bootstrap role's trust policy to an external account.
- C2 (high): the SAM justification for role/* (SAM auto-roles land at path /
  with no settable RolePath) does not transfer -- Terraform's aws_iam_role
  supports path.

Fixes, closing the class at the root rather than by denylist:
- All role writes, boundary sets and PassRole confined to role/tf-managed/*;
  reads split into a separate statement that keeps Resource "*".
- DenySelfMutation extended to cdk-hnb659fds-*, OrganizationAccountAccessRole
  and seahaven-* as defense in depth.
- OIDC provider made conditional (CreateOIDCProvider), mirroring the sibling
  substrate, so a first-create rollback is recoverable rather than wedging the
  stack in ROLLBACK_COMPLETE against a Retained orphan.
- README corrected: the guardrail policy is NOT Retain (only the provider is),
  so the Deny backstops do not survive a stack delete.

checkov CKV_AWS_109 no longer fires on this template, so no suppression is
needed. The template header records every divergence from the SAM copy.
2026-07-30 16:55:45 -04:00
ea27635ef2
feat(iac): add per-account HCP Terraform deploy substrate for prod and dev
New stack seahaven-terraform-substrate (instances terraform-substrate-prod +
terraform-substrate-dev): app.terraform.io OIDC provider and the shared
boundary-gated guardrail policy seahaven-hcptf-iam-management that
per-workspace Terraform apply roles attach at migration time. No roles are
pre-provisioned (accumulator pattern, parallel to githubdeploy-*).

Guardrail statements mirror seahaven-cfn-exec-iam-management byte-identically
except DenySelfMutation, whose scope extends to hcptf-* alongside the
GitHub-substrate principals. Explicit stack dependency on the same-account
deploy-substrate stack (boundary ARN appears only in Condition strings, so
CFN infers no edge).
2026-07-30 16:31:34 -04:00
83d0d0eb0d
docs(readme): document the deploy substrate's escalation controls
The substrate section described what the stack contains but not the three
Deny statements that make the boundary gate hold, so a future editor could
remove or weaken them without knowing what they defend. Records why
DenySelfMutation is required (the role holds unconditioned DetachRolePolicy
on * and could detach the Deny-carrying policy from itself), how to verify a
change by simulation, the prod/dev caveat that simulation cannot see this
role's inline policies, and the rollback-wedge recovery the mgmt README
already carried.
2026-07-27 19:15:08 -04:00
d6bea33436
feat(deploy-substrate): per-account GitHub Actions deploy substrate for prod/dev
SAM repos migrating off the frozen management account need the shared
deploy plumbing (permissions boundary + github-cfn-execution-role) in
their target account; none of it existed outside mgmt, so there was no
OIDC SAM deploy path into seahaven-prod or seahaven-dev at all.

Adds a templated, per-account substrate stack so onboarding a future
account is one bin/app.ts instance plus one CD job, not a hand-rolled
copy. Per-repo githubdeploy-* roles stay out by design: they are
provisioned per repo at migration time so an account never accumulates
trust for repos that do not deploy to it.

The template is a verbatim extraction of the reviewed mgmt substrate,
with deliberate, documented divergences — notably the removal of
iam:DeleteRolePermissionsBoundary plus explicit Deny backstops, which
closes a confirmed privilege-escalation path (see PR body).
2026-07-27 16:24:09 -04:00
Adam Moussa
ed26ff937d
Document centralized root access: README runbook + stack comment updates (#52)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
Comment/docs only, no template change (synth verified). Records the
2026-07-14 rollout: features enabled, member root credentials deleted,
recovery runbook (manual detach, inheritance + propagation gotchas),
new-account flow superseding root-harden-before-OU-move, extdev 5-SCP
quota saturation.
2026-07-14 18:32:40 -04:00
Adam Moussa
2303a54ebc
seahaven-prod account baseline (Phase 5) (#50)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* Add seahaven-prod member baseline (Phase 5)

Account 011934824531 is the target for all new production stacks; the
management account is frozen for new workloads. First proven exercise
of the automatic enrollment sweep (Enabled in 124s, no manual
create-members) and of AutoEnableStandards=NONE (no pre-enabled
standards, so CFN owns FSBP + CIS v3.0 cleanly). Default VPC deleted;
budget starts at $100 and resizes as tenants land.

* Apply Phase-5 review findings

Fleet gap closed: EBS encryption-by-default + IAM password policy were
management-account-only (the runbook's unscoped 'applied' claim hid
it); now applied and verified in all three member accounts, runbook
scoped per account. README stack inventory corrected (eleven stacks,
org-governance rows restored). Sweep comments reconciled: the
automatic enrollment sweep is proven (seahaven-prod, ~2min).
2026-07-14 17:17:55 -04:00
Adam Moussa
0a7c1bc450
seahaven-dev account baseline with org-managed detection (Phase 4) (#49)
* Add seahaven-dev member baseline with org-managed detection

Account 710827005802 (internal dev/staging) is the first account born
after delegation: GuardDuty/Security Hub enroll it via the org admin,
so DetectiveControls gains a localDetectiveServices flag (default true
— zero diff on the three deployed consumers, verified) and the dev
instance sets orgManagedDetection to skip the colliding local
detector/hub/analyzer. Default VPC kept and flow-logged (dev runs real
workloads). Enrollment verified Enabled in both services before this
commit.

* Fix Phase-4 review findings: standards + analyzer stay CFN-owned

SH-DEV-001: org AutoEnableStandards DEFAULT gave dev legacy CIS v1.2.0
and nothing owned CIS v3.0 — org config set to NONE, standards are now
unconditional in DetectiveControls (attach fine to an org-enabled hub),
legacy ruleset disabled in dev. SH-DEVBASE-002: the ORGANIZATION
analyzer treats the whole org as trusted so it cannot flag intra-org
exposure — account analyzer restored unconditionally (coexistence
verified live). Enrollment comments corrected: manual create-members,
the automatic sweep is still unexercised. Zero diff re-verified on all
three deployed baseline stacks.
2026-07-14 16:41:36 -04:00
Adam Moussa
2d3ba94140
Flip delegation runbook to applied with verification evidence (#48)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
All five services delegated to seahaven-security 2026-07-14 after the
hard preconditions verified (root MFA, OU placement, baseline live).
Member adoption and central findings flow verified end-to-end with a
sample finding; evidence recorded inline per SEC-BASE-A.
2026-07-14 15:53:42 -04:00
Adam Moussa
18f0f40e74
seahaven-security account baseline + security-OU guardrails (Phase 3) (#47)
* Add seahaven-security member baseline (Phase 3)

Account 001520130573 is the org's delegated security administrator.
Same member-baseline construct set as external-dev; own CD job under
its own OIDC role. Created at org root pending manual root hardening
before the OU move (deny-root-user invariant).

* Document delegated security administration runbook

Delegation to seahaven-security has no CloudFormation types; the CLI
sequence is the record, same pattern as the other account toggles.

* Apply Phase-3 security-review findings

Delegation runbook marked PENDING with hard preconditions (baseline
deployed, root MFA verified, account inside the security OU) — it had
read as applied before execution, the org's known claimed-done-but-NOT
failure mode (SEC-BASE-A/B). New security-guardrails SCP on the
security OU: region lock, IAM user/key lockout, privileged-role
protection, delegated-admin membership protection (SEC-BASE-C,
cross-reviewed APPROVE). deploy-security gains stack-name pre-flight
(SEC-BASE-D). Default VPC in 001520130573 deleted; empty flow-log list
and aws@ alert routing documented as deliberate (SEC-BASE-F/H).
2026-07-14 15:32:50 -04:00
Adam Moussa
3ab3bc773a
Merge external-dev member baseline; rename to seahaven-org-baseline (#43)
* Parameterize baseline constructs for multi-account reuse

DetectiveControls, FlowLogs, and GovernanceToggles were forked into
seahaven-external-dev-baseline with only physical-name and VPC-sourcing
differences. Prefix/name props let one implementation serve both
accounts; synthesized templates are unchanged (verified: empty cdk diff
against all deployed stacks).

* Absorb external-dev member baseline stack

Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack,
construct ids and physical names byte-identical to the deployed stack
(logical IDs are path-derived; empty cdk diff verified via change set
against 396287094661). Retires the forked repo so member-account
baselines share one drift surface and one dependency pin.

* Rename package to seahaven-org-baseline

Prepares the repo rename: the app now spans the management account and
org member accounts, so 'account-baseline' undersells the scope. README
documents the two-account deploy topology and logical-ID constraints.

* Commit extdev flow-log VPC ids in code, not -c context

Security review SH-ORG-004 (confirmed high): with the ids sourced from
ephemeral cdk context, any context-less deploy silently removes every
flow log in the isolated account. A committed list makes the attachment
set reviewable and immune to a forgotten -c flag. Empty list matches
the deployed stack (zero diff).

* Split CD into per-account deploy jobs

The app now spans two AWS accounts; cdk deploy --all under one role
fails on the other account's stacks (security review IAC-01). Each job
passes explicit stack selectors and its own account's OIDC role via the
new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
Adam Moussa
fd6fa5518b
Complete stack table in README (#42)
The intro summary table listed only 3 of the 6 stacks that bin/app.ts
synthesizes, omitting seahaven-dynamodb-cmk and the two secondary-region
baselines. Bring it in line with the detailed CDK-app table and fix the
region summary sentence.
2026-07-10 16:22:58 -04:00
Adam Moussa
3ee66ef63b
Document CDK app structure and cdk.json in README (#41)
The README covered what the stacks deploy but never documented the
CDK app itself — the cdk.json config, the bin/app.ts entry point, or
the full set of stacks it synthesizes (only three of six were listed).
Add a CDK app section mapping cdk.json, bin/, and lib/ to their roles,
listing all six stacks with their names/regions/source files, and the
common build/synth/deploy commands.
2026-07-10 16:07:20 -04:00
seahaven-openswe[bot]
142e221c47
feat: harden CIS 4.1 detection depth with M-of-N alarm tuning and CloudTrail Insights (#38)
Switch UnauthorizedApiCalls alarm from 3/3 consecutive to 3/6 M-of-N
so a single quiet 5-min window can't reset detection. The 3/3 setting
pre-dates #36 and was sized to suppress CFN/Config noise that #36 now
removes at the filter level, making a wider M-of-N evaluation window
safe from flap risk.

Enable CloudTrail Insights (ApiCallRateInsight + ApiErrorRateInsight)
on seahaven-org-trail as a compensating control for the residual risk
accepted in #36 — the CFN/Config-proxied denials intentionally excluded
from CIS 4.1 — and as a backstop for low-and-slow patterns the 5-min
alarm may miss. Cost ≈$35–$53/month at current org trail volume.

Refs: #37

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-07 15:47:41 -04:00
Adam Moussa
0d654edb35
fix: Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm (#36)
* Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm

CIS 4.1 (cis-UnauthorizedAPICalls) flapped OK<->ALARM 15 times in 30 days,
all from benign AWS-service AccessDenied noise (CloudFormation deploy/drift
describe-scans, AWS Config recorder). A single CFN run on 2026-07-07 emitted
100+ such denials in 15 min, tripping the alarm and burying the real CIS 4.1
security signal in email noise (alert fatigue).

- Group both error codes so the exclusions apply to the whole filter (the old
  pattern leaked the UnauthorizedOperation branch past the exclusions due to
  && binding tighter than ||).
- Exclude denials whose sourceIPAddress is an AWS service host (*.amazonaws.com)
  — AWS acting on our behalf, not a principal of concern. Real unauthorized
  calls from a console/CLI/attacker present a routable IP and are still counted.
  Validated against the trail log group: spike window 107 -> 4 matches, the 4
  remaining all from a routable admin IP (genuine activity CIS should retain).
- Keep the 3/3 evaluation as a backstop against one-off human fat-fingers.

Billing: deleted the manually-created AWS-MonthlyBilling CloudWatch alarm
($50 threshold on EstimatedCharges, routed to site-alerts). It was unmanaged
drift, permanently in ALARM, and fully redundant with the managed M-10 budget
(seahaven-monthly-cost). README updated with rationale + restore command.

* Address sh-security-review: scope CIS 4.1 exclusion to named benign sources

The high-recall security review (detector fan-out + proof-or-kill verifier)
confirmed a MEDIUM detection blind spot in the first revision: excluding all
`*.amazonaws.com` source hosts would hide denials driven through ANY AWS
service (SSM Automation, Step Functions, Lambda, etc.), which CloudTrail
records with that service's host as sourceIPAddress — i.e. service-proxied
privesc/recon attempts would evade CIS 4.1.

Remediation: scope the exclusion to the specific benign sources that actually
flap this account — `*cloudformation.amazonaws.com` (covers both
cloudformation. and hooks.cloudformation.) and `config.amazonaws.com` — plus
the pre-existing delivery.logs exclusion. Every other service-proxied denial
is now retained. Residual (accepted, documented inline): CloudFormation/Config-
proxied denials are still excluded — that path needs near-admin privilege
(CreateStack + PassRole), successful changes still trip the other CIS 4.x
alarms, and GuardDuty backstops.

Validated on the live trail log group: spike window still 107 -> 4 matches
(identical noise suppression), the 4 from a routable admin IP. tsc + synth clean.
2026-07-07 19:32:10 +00:00
Adam Moussa
4713172af1
docs: link Confluence AWS Architecture Map (INFRA-53) (#34) 2026-07-06 17:44:22 -04:00
Adam Moussa
6a63a4f9b0
Repo hygiene: PR labeler + README badges (INFRA-56/57) (#26) 2026-06-11 14:42:34 -04:00
Adam Moussa
e22c4ac005
Bring Config recorder + channel under IaC via AwsCustomResource (#22)
The L1 AWS::Config::ConfigurationRecorder deadlocks the CDK stack
(recorder can't complete without a delivery channel; channel can't be
created without a recorder — observed 2026-06-01).

Fix: three AwsCustomResource nodes call PutConfigurationRecorder →
PutDeliveryChannel → StartConfigurationRecorder in sequence. Put* is
an idempotent upsert, so the deploy adopts the existing CLI-created
recorder and channel without destroying or interrupting them. onDelete
stops recording rather than deleting the per-account singleton.

New IAM permissions on the custom-resource role (cross-reviewed,
GPT-4.1 APPROVE — no BLOCK):
  config:PutConfigurationRecorder
  config:PutDeliveryChannel
  config:StartConfigurationRecorder
  config:StopConfigurationRecorder
  iam:PassRole → seahaven-config-recorder-role (service=config)

cdk diff shows [+] adds only — no existing resources destroyed or
replaced. Removes README note that recorder/channel are CLI-only.

Refs: INFRA-17
2026-06-10 14:38:23 -04:00
Adam Moussa
77d9d6c574
[INFRA-96] CMK-encrypt sensitive CloudWatch log groups (M-24) (#20)
* [INFRA-96] CMK-encrypt sensitive CloudWatch log groups (M-24)

Add a dedicated customer-managed CMK (alias/seahaven-logs) for encrypting
the sensitive CloudWatch Logs groups (CloudTrail + finance/PII Lambdas).

- lib/logs-key.ts: LogsKey construct. Key policy grants the CloudWatch Logs
  service principal (logs.us-east-1.amazonaws.com) Encrypt*/Decrypt*/
  ReEncrypt*/GenerateDataKey*/DescribeKey, scoped by the
  kms:EncryptionContext:aws:logs:arn condition (REQUIRED per AWS docs or log
  delivery breaks). Cross-reviewed (GPT-4.1): tightened Describe* -> DescribeKey;
  CreateGrant omitted (not needed for plain log-group encryption).
- account-baseline-stack.ts: instantiate LogsKey and set KmsKeyId on the L2
  Trail's CloudWatch log group in place (escape hatch on the existing
  AWS::Logs::LogGroup) so it keeps the same logical id + physical name -
  additive, no replacement, CIS Section-4 metric filters (which import the
  group by name) keep working, live audit trail not disrupted. Gated by
  context `encryptTrailLogGroup` so the CMK can be smoke-tested on a low-risk
  Lambda group before the most-sensitive CloudTrail group.

Finance/PII Lambda log groups (exec-aide-*, payments-*, po-email-processor,
vendor-reply-processor) are owned by other stacks and associated to this CMK
via the CLI for now; codifying KmsKeyId in those repos is tracked as drift.

* [INFRA-96] Document sensitive-logs CMK (M-24) in README
2026-06-08 19:04:36 -04:00
Adam Moussa
0dd8d2a7af
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.

scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
Adam Moussa
3ba90ddc40
Add monitoring + logging layer (audit Day 2: H-1/H-14/M-13) (#6)
- H-1: 15 CIS Section 4 metric filters (4.1-4.15) on the CloudTrail log group,
  each alarming to a new SSE SNS topic seahaven-cis-alarms (email to adam).
  ALARM-only actions per Sea Haven preference. 4.16 = Security Hub (Day 1).
- H-14: VPC flow logs (ALL traffic) on all 5 VPCs → hardened S3 bucket. Delivery
  bucket policy cross-reviewed; kept the AWS-required s3:x-amz-acl condition +
  logs:*:* source-ARN (cross-reviewer wrongly flagged these; verified against
  AWS flow-logs-s3-permissions docs), dropped the unneeded s3:ListBucket.
- M-13: SES configuration set seahaven-email-events capturing bounce/complaint/
  reject to CloudWatch for reputation visibility.

L-4 (log retention) and L-5 (alarm action) applied via CLI, documented in README.
2026-06-02 15:16:24 -04:00
Adam Moussa
38d4a5753a
Account detective layer + budget (audit Day 1) (#5)
* Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10)

Adds to the seahaven-account-baseline stack:
- AWS Config recorder (all + global resources) + delivery channel + role +
  hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed.
- GuardDuty detector, us-east-1 (H-3)
- Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4)
- IAM Access Analyzer, account scope (M-5)
- Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to
  adam@seahavenind.com (M-10)

Scope us-east-1 only (all workloads here); multi-region is a follow-up.
The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented
separately in the README runbook.

* Document Day 1 detective layer + CLI governance toggles in README

* Move Config recorder+channel to CLI (L1 stabilization deadlock)

The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches
CREATE_COMPLETE until recording is active (needs a delivery channel), and the
delivery channel cannot be created until the recorder completes — a deadlock
that hung the deploy ~27 min before manual cancel (2026-06-01).

Keep the cross-reviewed recorder role + delivery bucket in IaC; create the
recorder, delivery channel, and start recording via CLI (documented in README).
Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls
evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
Adam Moussa
64ef25dc5b
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)

The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.

Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.

The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.

Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.

* Grant AWS Backup service use of vault CMKs

The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.

Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
Adam Moussa
dc079edb93 Initial account-baseline stack with CloudTrail (audit C-1)
Multi-region CloudTrail with log-file validation, a rotating KMS CMK, an
Object-Lock'd S3 log bucket, and CloudWatch Logs delivery. First resident of
the account-level security baseline; AWS Backup / 3-2-1 (C-7) lands alongside.

IAM/KMS/S3 policies cross-reviewed; review caught a missing CloudTrail KMS
grant, now added (SourceArn + encryption-context scoped).
2026-05-29 17:44:55 -04:00