Commit graph

20 commits

Author SHA1 Message Date
d6bea33436
feat(deploy-substrate): per-account GitHub Actions deploy substrate for prod/dev
SAM repos migrating off the frozen management account need the shared
deploy plumbing (permissions boundary + github-cfn-execution-role) in
their target account; none of it existed outside mgmt, so there was no
OIDC SAM deploy path into seahaven-prod or seahaven-dev at all.

Adds a templated, per-account substrate stack so onboarding a future
account is one bin/app.ts instance plus one CD job, not a hand-rolled
copy. Per-repo githubdeploy-* roles stay out by design: they are
provisioned per repo at migration time so an account never accumulates
trust for repos that do not deploy to it.

The template is a verbatim extraction of the reviewed mgmt substrate,
with deliberate, documented divergences — notably the removal of
iam:DeleteRolePermissionsBoundary plus explicit Deny backstops, which
closes a confirmed privilege-escalation path (see PR body).
2026-07-27 16:24:09 -04:00
Adam Moussa
ed26ff937d
Document centralized root access: README runbook + stack comment updates (#52)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
Comment/docs only, no template change (synth verified). Records the
2026-07-14 rollout: features enabled, member root credentials deleted,
recovery runbook (manual detach, inheritance + propagation gotchas),
new-account flow superseding root-harden-before-OU-move, extdev 5-SCP
quota saturation.
2026-07-14 18:32:40 -04:00
Adam Moussa
2303a54ebc
seahaven-prod account baseline (Phase 5) (#50)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run
* Add seahaven-prod member baseline (Phase 5)

Account 011934824531 is the target for all new production stacks; the
management account is frozen for new workloads. First proven exercise
of the automatic enrollment sweep (Enabled in 124s, no manual
create-members) and of AutoEnableStandards=NONE (no pre-enabled
standards, so CFN owns FSBP + CIS v3.0 cleanly). Default VPC deleted;
budget starts at $100 and resizes as tenants land.

* Apply Phase-5 review findings

Fleet gap closed: EBS encryption-by-default + IAM password policy were
management-account-only (the runbook's unscoped 'applied' claim hid
it); now applied and verified in all three member accounts, runbook
scoped per account. README stack inventory corrected (eleven stacks,
org-governance rows restored). Sweep comments reconciled: the
automatic enrollment sweep is proven (seahaven-prod, ~2min).
2026-07-14 17:17:55 -04:00
Adam Moussa
0a7c1bc450
seahaven-dev account baseline with org-managed detection (Phase 4) (#49)
* Add seahaven-dev member baseline with org-managed detection

Account 710827005802 (internal dev/staging) is the first account born
after delegation: GuardDuty/Security Hub enroll it via the org admin,
so DetectiveControls gains a localDetectiveServices flag (default true
— zero diff on the three deployed consumers, verified) and the dev
instance sets orgManagedDetection to skip the colliding local
detector/hub/analyzer. Default VPC kept and flow-logged (dev runs real
workloads). Enrollment verified Enabled in both services before this
commit.

* Fix Phase-4 review findings: standards + analyzer stay CFN-owned

SH-DEV-001: org AutoEnableStandards DEFAULT gave dev legacy CIS v1.2.0
and nothing owned CIS v3.0 — org config set to NONE, standards are now
unconditional in DetectiveControls (attach fine to an org-enabled hub),
legacy ruleset disabled in dev. SH-DEVBASE-002: the ORGANIZATION
analyzer treats the whole org as trusted so it cannot flag intra-org
exposure — account analyzer restored unconditionally (coexistence
verified live). Enrollment comments corrected: manual create-members,
the automatic sweep is still unexercised. Zero diff re-verified on all
three deployed baseline stacks.
2026-07-14 16:41:36 -04:00
Adam Moussa
2d3ba94140
Flip delegation runbook to applied with verification evidence (#48)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
All five services delegated to seahaven-security 2026-07-14 after the
hard preconditions verified (root MFA, OU placement, baseline live).
Member adoption and central findings flow verified end-to-end with a
sample finding; evidence recorded inline per SEC-BASE-A.
2026-07-14 15:53:42 -04:00
Adam Moussa
18f0f40e74
seahaven-security account baseline + security-OU guardrails (Phase 3) (#47)
* Add seahaven-security member baseline (Phase 3)

Account 001520130573 is the org's delegated security administrator.
Same member-baseline construct set as external-dev; own CD job under
its own OIDC role. Created at org root pending manual root hardening
before the OU move (deny-root-user invariant).

* Document delegated security administration runbook

Delegation to seahaven-security has no CloudFormation types; the CLI
sequence is the record, same pattern as the other account toggles.

* Apply Phase-3 security-review findings

Delegation runbook marked PENDING with hard preconditions (baseline
deployed, root MFA verified, account inside the security OU) — it had
read as applied before execution, the org's known claimed-done-but-NOT
failure mode (SEC-BASE-A/B). New security-guardrails SCP on the
security OU: region lock, IAM user/key lockout, privileged-role
protection, delegated-admin membership protection (SEC-BASE-C,
cross-reviewed APPROVE). deploy-security gains stack-name pre-flight
(SEC-BASE-D). Default VPC in 001520130573 deleted; empty flow-log list
and aws@ alert routing documented as deliberate (SEC-BASE-F/H).
2026-07-14 15:32:50 -04:00
Adam Moussa
3ab3bc773a
Merge external-dev member baseline; rename to seahaven-org-baseline (#43)
* Parameterize baseline constructs for multi-account reuse

DetectiveControls, FlowLogs, and GovernanceToggles were forked into
seahaven-external-dev-baseline with only physical-name and VPC-sourcing
differences. Prefix/name props let one implementation serve both
accounts; synthesized templates are unchanged (verified: empty cdk diff
against all deployed stacks).

* Absorb external-dev member baseline stack

Moves seahaven-external-dev-baseline's stack in as MemberBaselineStack,
construct ids and physical names byte-identical to the deployed stack
(logical IDs are path-derived; empty cdk diff verified via change set
against 396287094661). Retires the forked repo so member-account
baselines share one drift surface and one dependency pin.

* Rename package to seahaven-org-baseline

Prepares the repo rename: the app now spans the management account and
org member accounts, so 'account-baseline' undersells the scope. README
documents the two-account deploy topology and logical-ID constraints.

* Commit extdev flow-log VPC ids in code, not -c context

Security review SH-ORG-004 (confirmed high): with the ids sourced from
ephemeral cdk context, any context-less deploy silently removes every
flow log in the isolated account. A committed list makes the attachment
set reviewable and immune to a forgotten -c flag. Empty list matches
the deployed stack (zero diff).

* Split CD into per-account deploy jobs

The app now spans two AWS accounts; cdk deploy --all under one role
fails on the other account's stacks (security review IAC-01). Each job
passes explicit stack selectors and its own account's OIDC role via the
new cd-cdk stacks input.
2026-07-14 13:53:07 -04:00
Adam Moussa
fd6fa5518b
Complete stack table in README (#42)
The intro summary table listed only 3 of the 6 stacks that bin/app.ts
synthesizes, omitting seahaven-dynamodb-cmk and the two secondary-region
baselines. Bring it in line with the detailed CDK-app table and fix the
region summary sentence.
2026-07-10 16:22:58 -04:00
Adam Moussa
3ee66ef63b
Document CDK app structure and cdk.json in README (#41)
The README covered what the stacks deploy but never documented the
CDK app itself — the cdk.json config, the bin/app.ts entry point, or
the full set of stacks it synthesizes (only three of six were listed).
Add a CDK app section mapping cdk.json, bin/, and lib/ to their roles,
listing all six stacks with their names/regions/source files, and the
common build/synth/deploy commands.
2026-07-10 16:07:20 -04:00
seahaven-openswe[bot]
142e221c47
feat: harden CIS 4.1 detection depth with M-of-N alarm tuning and CloudTrail Insights (#38)
Switch UnauthorizedApiCalls alarm from 3/3 consecutive to 3/6 M-of-N
so a single quiet 5-min window can't reset detection. The 3/3 setting
pre-dates #36 and was sized to suppress CFN/Config noise that #36 now
removes at the filter level, making a wider M-of-N evaluation window
safe from flap risk.

Enable CloudTrail Insights (ApiCallRateInsight + ApiErrorRateInsight)
on seahaven-org-trail as a compensating control for the residual risk
accepted in #36 — the CFN/Config-proxied denials intentionally excluded
from CIS 4.1 — and as a backstop for low-and-slow patterns the 5-min
alarm may miss. Cost ≈$35–$53/month at current org trail volume.

Refs: #37

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-07 15:47:41 -04:00
Adam Moussa
0d654edb35
fix: Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm (#36)
* Fix noisy CIS 4.1 unauthorized-API alarm; drop redundant billing alarm

CIS 4.1 (cis-UnauthorizedAPICalls) flapped OK<->ALARM 15 times in 30 days,
all from benign AWS-service AccessDenied noise (CloudFormation deploy/drift
describe-scans, AWS Config recorder). A single CFN run on 2026-07-07 emitted
100+ such denials in 15 min, tripping the alarm and burying the real CIS 4.1
security signal in email noise (alert fatigue).

- Group both error codes so the exclusions apply to the whole filter (the old
  pattern leaked the UnauthorizedOperation branch past the exclusions due to
  && binding tighter than ||).
- Exclude denials whose sourceIPAddress is an AWS service host (*.amazonaws.com)
  — AWS acting on our behalf, not a principal of concern. Real unauthorized
  calls from a console/CLI/attacker present a routable IP and are still counted.
  Validated against the trail log group: spike window 107 -> 4 matches, the 4
  remaining all from a routable admin IP (genuine activity CIS should retain).
- Keep the 3/3 evaluation as a backstop against one-off human fat-fingers.

Billing: deleted the manually-created AWS-MonthlyBilling CloudWatch alarm
($50 threshold on EstimatedCharges, routed to site-alerts). It was unmanaged
drift, permanently in ALARM, and fully redundant with the managed M-10 budget
(seahaven-monthly-cost). README updated with rationale + restore command.

* Address sh-security-review: scope CIS 4.1 exclusion to named benign sources

The high-recall security review (detector fan-out + proof-or-kill verifier)
confirmed a MEDIUM detection blind spot in the first revision: excluding all
`*.amazonaws.com` source hosts would hide denials driven through ANY AWS
service (SSM Automation, Step Functions, Lambda, etc.), which CloudTrail
records with that service's host as sourceIPAddress — i.e. service-proxied
privesc/recon attempts would evade CIS 4.1.

Remediation: scope the exclusion to the specific benign sources that actually
flap this account — `*cloudformation.amazonaws.com` (covers both
cloudformation. and hooks.cloudformation.) and `config.amazonaws.com` — plus
the pre-existing delivery.logs exclusion. Every other service-proxied denial
is now retained. Residual (accepted, documented inline): CloudFormation/Config-
proxied denials are still excluded — that path needs near-admin privilege
(CreateStack + PassRole), successful changes still trip the other CIS 4.x
alarms, and GuardDuty backstops.

Validated on the live trail log group: spike window still 107 -> 4 matches
(identical noise suppression), the 4 from a routable admin IP. tsc + synth clean.
2026-07-07 19:32:10 +00:00
Adam Moussa
4713172af1
docs: link Confluence AWS Architecture Map (INFRA-53) (#34) 2026-07-06 17:44:22 -04:00
Adam Moussa
6a63a4f9b0
Repo hygiene: PR labeler + README badges (INFRA-56/57) (#26) 2026-06-11 14:42:34 -04:00
Adam Moussa
e22c4ac005
Bring Config recorder + channel under IaC via AwsCustomResource (#22)
The L1 AWS::Config::ConfigurationRecorder deadlocks the CDK stack
(recorder can't complete without a delivery channel; channel can't be
created without a recorder — observed 2026-06-01).

Fix: three AwsCustomResource nodes call PutConfigurationRecorder →
PutDeliveryChannel → StartConfigurationRecorder in sequence. Put* is
an idempotent upsert, so the deploy adopts the existing CLI-created
recorder and channel without destroying or interrupting them. onDelete
stops recording rather than deleting the per-account singleton.

New IAM permissions on the custom-resource role (cross-reviewed,
GPT-4.1 APPROVE — no BLOCK):
  config:PutConfigurationRecorder
  config:PutDeliveryChannel
  config:StartConfigurationRecorder
  config:StopConfigurationRecorder
  iam:PassRole → seahaven-config-recorder-role (service=config)

cdk diff shows [+] adds only — no existing resources destroyed or
replaced. Removes README note that recorder/channel are CLI-only.

Refs: INFRA-17
2026-06-10 14:38:23 -04:00
Adam Moussa
77d9d6c574
[INFRA-96] CMK-encrypt sensitive CloudWatch log groups (M-24) (#20)
* [INFRA-96] CMK-encrypt sensitive CloudWatch log groups (M-24)

Add a dedicated customer-managed CMK (alias/seahaven-logs) for encrypting
the sensitive CloudWatch Logs groups (CloudTrail + finance/PII Lambdas).

- lib/logs-key.ts: LogsKey construct. Key policy grants the CloudWatch Logs
  service principal (logs.us-east-1.amazonaws.com) Encrypt*/Decrypt*/
  ReEncrypt*/GenerateDataKey*/DescribeKey, scoped by the
  kms:EncryptionContext:aws:logs:arn condition (REQUIRED per AWS docs or log
  delivery breaks). Cross-reviewed (GPT-4.1): tightened Describe* -> DescribeKey;
  CreateGrant omitted (not needed for plain log-group encryption).
- account-baseline-stack.ts: instantiate LogsKey and set KmsKeyId on the L2
  Trail's CloudWatch log group in place (escape hatch on the existing
  AWS::Logs::LogGroup) so it keeps the same logical id + physical name -
  additive, no replacement, CIS Section-4 metric filters (which import the
  group by name) keep working, live audit trail not disrupted. Gated by
  context `encryptTrailLogGroup` so the CMK can be smoke-tested on a low-risk
  Lambda group before the most-sensitive CloudTrail group.

Finance/PII Lambda log groups (exec-aide-*, payments-*, po-email-processor,
vendor-reply-processor) are owned by other stacks and associated to this CMK
via the CLI for now; codifying KmsKeyId in those repos is tracked as drift.

* [INFRA-96] Document sensitive-logs CMK (M-24) in README
2026-06-08 19:04:36 -04:00
Adam Moussa
0dd8d2a7af
Docs + tooling: README phase-2 backup, iam-user-delete script (#10)
README: document AWS Backup phase-2 (phase2-offsite-everything selection),
remove retired database-1 from phase-1 scope (audit H-19), update roadmap +
verify smoke-test to a live resource.

scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies,
groups, MFA, login profile, certs, SSH keys, service creds, then user) with
--profile/--yes and a guard against deleting the caller's own identity. Built
from the Day 4 audit IAM cleanup.
2026-06-03 13:15:18 -04:00
Adam Moussa
3ba90ddc40
Add monitoring + logging layer (audit Day 2: H-1/H-14/M-13) (#6)
- H-1: 15 CIS Section 4 metric filters (4.1-4.15) on the CloudTrail log group,
  each alarming to a new SSE SNS topic seahaven-cis-alarms (email to adam).
  ALARM-only actions per Sea Haven preference. 4.16 = Security Hub (Day 1).
- H-14: VPC flow logs (ALL traffic) on all 5 VPCs → hardened S3 bucket. Delivery
  bucket policy cross-reviewed; kept the AWS-required s3:x-amz-acl condition +
  logs:*:* source-ARN (cross-reviewer wrongly flagged these; verified against
  AWS flow-logs-s3-permissions docs), dropped the unneeded s3:ListBucket.
- M-13: SES configuration set seahaven-email-events capturing bounce/complaint/
  reject to CloudWatch for reputation visibility.

L-4 (log retention) and L-5 (alarm action) applied via CLI, documented in README.
2026-06-02 15:16:24 -04:00
Adam Moussa
38d4a5753a
Account detective layer + budget (audit Day 1) (#5)
* Add account detective layer + budget (audit Day 1: H-2/H-3/H-4/M-5/M-10)

Adds to the seahaven-account-baseline stack:
- AWS Config recorder (all + global resources) + delivery channel + role +
  hardened delivery bucket (H-2, CIS 3.3/3.5). Recorder role IAM cross-reviewed.
- GuardDuty detector, us-east-1 (H-3)
- Security Hub with AWS FSBP v1.0.0 + CIS v3.0.0 standards, depends on Config (H-4)
- IAM Access Analyzer, account scope (M-5)
- Monthly cost budget $1,200 with 80/100% actual + 100% forecast alerts to
  adam@seahavenind.com (M-10)

Scope us-east-1 only (all workloads here); multi-region is a follow-up.
The CLI-applied governance toggles (M-6/M-3/M-7/L-8/M-11) are documented
separately in the README runbook.

* Document Day 1 detective layer + CLI governance toggles in README

* Move Config recorder+channel to CLI (L1 stabilization deadlock)

The L1 AWS::Config::ConfigurationRecorder hangs the stack: it never reaches
CREATE_COMPLETE until recording is active (needs a delivery channel), and the
delivery channel cannot be created until the recorder completes — a deadlock
that hung the deploy ~27 min before manual cancel (2026-06-01).

Keep the cross-reviewed recorder role + delivery bucket in IaC; create the
recorder, delivery channel, and start recording via CLI (documented in README).
Security Hub no longer takes a CFN dependency on the recorder; CIS/FSBP controls
evaluate once Config is recording. Verified live: recording=true, SUCCESS.
2026-06-01 17:56:12 -04:00
Adam Moussa
64ef25dc5b
Add AWS Backup with offsite vault (audit C-7) (#3)
* Add AWS Backup with offsite vault (audit C-7)

The account had zero AWS Backup vaults/plans, so 22 of 23 data stores
had no immutable, cross-region recovery path (audit finding C-7). One
ransomware event or rogue delete would erase primary plus same-region
snapshots/PITR.

Phase 1 ("critical data first") protects the seven highest-risk stores
with no offsite leg today (2 RDS, 2 DynamoDB, 3 S3) via a daily plan in
a new us-east-1 vault, copied cross-region into a governance-locked
us-west-2 vault. Governance (not compliance) mode first so the plan can
be validated before committing to irreversible immutability.

The backup service role is backup-only (no restore policies) to stay
least-privilege; restores get a separate audited path later. Resources
are selected by explicit ARN to avoid drifting the stacks that own them.

Deploys via the shared cdk deploy --all alongside the C-1 CloudTrail
stack. See the README pre-deploy gates (S3 versioning, database-1
unencrypted copy smoke-test, DynamoDB PITR) before the first run.

* Grant AWS Backup service use of vault CMKs

The L2 BackupVault does not grant the backup service principal use of a
customer-managed key; the synthesized key policy only delegated to
account IAM. Cross-region copy of encrypted RDS/EBS recovery points uses
KMS grants on the destination key, so without an explicit grant those
copy jobs fail — and silently, since the account has no CloudTrail yet.

Add backup.amazonaws.com crypto + CreateGrant statements to both vault
keys, scoped by aws:SourceAccount (cross-review BLOCK 2; mirrors the
discipline used on the C-1 CloudTrail key). Same class of bug the C-1
cross-review caught on the CloudTrail CMK.
2026-05-29 18:06:17 -04:00
Adam Moussa
dc079edb93 Initial account-baseline stack with CloudTrail (audit C-1)
Multi-region CloudTrail with log-file validation, a rotating KMS CMK, an
Object-Lock'd S3 log bucket, and CloudWatch Logs delivery. First resident of
the account-level security baseline; AWS Backup / 3-2-1 (C-7) lands alongside.

IAM/KMS/S3 policies cross-reviewed; review caught a missing CloudTrail KMS
grant, now added (SourceArn + encryption-context scoped).
2026-05-29 17:44:55 -04:00