seahaven-org-baseline/README.md
Adam Moussa 3ee66ef63b
Document CDK app structure and cdk.json in README (#41)
The README covered what the stacks deploy but never documented the
CDK app itself — the cdk.json config, the bin/app.ts entry point, or
the full set of stacks it synthesizes (only three of six were listed).
Add a CDK app section mapping cdk.json, bin/, and lib/ to their roles,
listing all six stacks with their names/regions/source files, and the
common build/synth/deploy commands.
2026-07-10 16:07:20 -04:00

354 lines
20 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# seahaven-account-baseline
![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
![AWS CDK](https://img.shields.io/badge/AWS-CDK-FF9900?logo=amazonaws&logoColor=white)
![CI](https://github.com/Sea-Haven-Industries/seahaven-account-baseline/actions/workflows/ci.yaml/badge.svg)
Account-level security and governance baseline for Sea Haven Industries
(AWS account **328440206208**), managed as a single CDK TypeScript app. Most
resources are in **us-east-1**; the offsite backup vault is in **us-west-2**.
This is where account-wide detective and recovery controls live, so they are
versioned, reviewed, and drift-checked like any other stack.
Stacks (all deployed by `cdk deploy --all` / the CD workflow):
| Stack | Region | Purpose |
|---|---|---|
| `seahaven-account-baseline` | us-east-1 | CloudTrail + future detective controls (C-1) |
| `seahaven-backup` | us-east-1 | Primary AWS Backup vault + plan + role (C-7) |
| `seahaven-backup-offsite` | us-west-2 | Governance-locked offsite copy vault (C-7) |
## CDK app
The repo is a single AWS CDK app written in TypeScript. `cdk.json` is the
project config the `cdk` CLI reads on every command: its `app` key
(`npx ts-node bin/app.ts`) tells CDK how to synthesize the app straight from
the TypeScript source — no separate compile step needed for `cdk synth` /
`diff` / `deploy` — and its `context` block carries the AWS CDK feature flags.
| Path | Role |
|---|---|
| `cdk.json` | CDK config: `app` synth command, `watch` includes/excludes, `context` feature flags |
| `bin/app.ts` | App entry point — instantiates every stack with an explicit kebab-case `stackName` and its target `env` (account `328440206208`, per-region) |
| `lib/*-stack.ts` | Stack definitions (one class per stack; larger stacks compose the constructs in `lib/*.ts`) |
| `tsconfig.json` | TypeScript compiler options (`outDir: cdk.out`) |
| `package.json` | Pinned `aws-cdk-lib`, CDK CLI, and the `build` / `synth` / `diff` / `deploy` npm scripts |
`bin/app.ts` synthesizes six stacks across three regions:
| Construct id | Stack name | Region | Source |
|---|---|---|---|
| `account-baseline` | `seahaven-account-baseline` | us-east-1 | `lib/account-baseline-stack.ts` |
| `dynamodb-cmk` | `seahaven-dynamodb-cmk` | us-east-1 | `lib/dynamodb-cmk-stack.ts` |
| `regional-baseline-us-west-2` | `seahaven-regional-baseline-us-west-2` | us-west-2 | `lib/regional-baseline-stack.ts` |
| `regional-baseline-us-east-2` | `seahaven-regional-baseline-us-east-2` | us-east-2 | `lib/regional-baseline-stack.ts` |
| `backup-offsite` | `seahaven-backup-offsite` | us-west-2 | `lib/backup-offsite-stack.ts` |
| `backup` | `seahaven-backup` | us-east-1 | `lib/backup-stack.ts` |
`backup` declares an explicit dependency on `backup-offsite` so the offsite copy
vault exists before the primary plan that copies into it. Stack names are set
explicitly to enforce kebab-case (CDK defaults to PascalCase). Synthesized
CloudFormation templates land in `cdk.out/` (git-ignored).
Common commands:
```
npm ci # install pinned deps
npm run build # tsc type-check (compiles to cdk.out/)
npx cdk synth # synthesize CloudFormation for all stacks
npx cdk diff # diff synthesized stacks against deployed state
npx cdk deploy --all # deploy every stack
npx cdk deploy <stack-name> # deploy a single stack
```
The `--context <key>=<value>` flag overrides `cdk.json` context at the command
line (e.g. the `encryptTrailLogGroup` toggle under *Monitoring + logging*).
## Documentation
The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This project's `seahaven-account-baseline`, `seahaven-backup`, and `seahaven-backup-offsite` stacks are represented there as Mermaid subgraphs.
- **[AWS Architecture Map](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098)** (Confluence, IT space, page 1540098)
## What it deploys
### CloudTrail (audit finding C-1)
| Resource | Logical ID | Notes |
|---|---|---|
| Multi-region trail | `Trail` (`seahaven-org-trail`) | Management events read+write, global service events, **log-file validation on**, **CloudTrail Insights on** (ApiCallRate + ApiErrorRate, §37) |
| Log bucket | `TrailLogBucket` (`seahaven-cloudtrail-logs-328440206208`) | Private (Block Public Access all), SSE-KMS, versioned, **TLS-only**, **Object Lock GOVERNANCE 365d**, lifecycle (Glacier @90d, expire @365d), server access logging → `seahaven-s3-access-logs` |
| KMS CMK | `TrailKey` (`alias/cloudtrail-logs`) | Encrypts log files; **automatic rotation enabled** |
| CloudWatch Logs group | created by the L2 `Trail` | 365-day retention; this is the group the CIS Section 4 metric filters (H-1) attach to |
**Data flow:** API activity across all regions → CloudTrail → (a) KMS-encrypted,
Object-Locked S3 bucket for durable/tamper-resistant storage and (b) CloudWatch
Logs for real-time querying and metric-filter alarms.
**Compliance impact:** closes CIS 3.1 (multi-region trail), 3.2 (log-file
validation), 3.4 (CloudWatch Logs integration), 3.6 (bucket access logging),
3.7 (KMS CMK encryption), and 3.8 (CMK rotation). Unblocks CIS Section 4 /
finding H-1 (metric filters + alarms now have a log group to target).
### Design decisions
- **Management events only.** Object-level S3/Lambda data events (CIS 3.10/3.11)
are deferred to control cost; revisit with targeted S3 *write* data events on
sensitive buckets (payments / accounting / kb) if needed.
- **Object Lock GOVERNANCE, not COMPLIANCE.** Tamper-resistant but still
deletable by a principal holding `s3:BypassGovernanceRetention` — avoids the
irreversibility of COMPLIANCE mode. Revisit if a stricter posture is required.
- **RETAIN** on the bucket and KMS key so a stack teardown never destroys the
audit trail.
### AWS Backup (audit finding C-7)
Phase 1 ("critical data first") of fixing the account's complete lack of AWS
Backup. Protects the data stores with no offsite leg today and copies each
recovery point cross-region into a governance-locked vault.
| Resource | Logical ID | Notes |
|---|---|---|
| Primary vault | `seahaven-primary` (us-east-1) | KMS-CMK encrypted, unlocked (working copy), RETAIN |
| Offsite vault | `seahaven-offsite` (us-west-2) | KMS-CMK encrypted, **Vault Lock GOVERNANCE** (min-retention 30d, no cooling-off window), RETAIN |
| Backup plan | `seahaven-critical-daily` | Daily 06:00 UTC, delete-after 35d, **cross-region CopyAction → offsite** (retain 90d) |
| Service role | `seahaven-backup-service-role` | **Backup-only** (Backup + S3-Backup managed policies); restore perms intentionally deferred |
**Phase-1 scope** (selected by explicit ARN, not tags, to avoid drifting other
stacks): RDS `proposal-system-db`, DynamoDB `PaymentsDashboard`,
DynamoDB `purchase-orders`, S3 `accounting.seahaven.com`,
`seahaven-payments-csv-328440206208`, `google-workspace-seahavenind.com`.
*(RDS `database-1` was originally in this set but was retired 2026-06-03 —
audit H-19, idle 0 conn/60d — and removed from the selection; its final
encrypted recovery point is retained in `seahaven-offsite` for 7 years.)*
**Coexists with** existing EBS DLM snapshots and DynamoDB PITR — it supplements
them with the missing offsite + immutable leg; it does not replace them.
**Design decisions:**
- **Governance lock first, not compliance.** Recovery points can't be silently
deleted, but a principal with explicit permission can still intervene while
we validate. Graduate to COMPLIANCE (irreversible) later by adding
`changeableFor` to the offsite vault lock + redeploy.
- **Backup-only role.** Restore policies and `allowRestores` are not granted;
restores get a separate audited path once a restore-test process exists.
**Pre-deploy gates** (must clear before the first scheduled run):
1. Enable S3 versioning on `seahaven-payments-csv-328440206208` and
`google-workspace-seahavenind.com` (`accounting.seahaven.com` already has it,
audit C-9), or their jobs fail silently (folds in H-21).
2. `database-1` is unencrypted (H-19): smoke-test an on-demand backup + copy of
it to us-west-2 first; if the copy fails, encrypt it or drop it from the copy.
3. Enable DynamoDB PITR (H-7) on the two tables for between-window recovery.
### AWS Backup phase 2 (audit Day 4)
Expands the same `seahaven-critical-daily` plan to every remaining data store, so
all of DynamoDB + EBS get the offsite + immutable leg ("offsite for everything").
| Resource | Logical ID | Notes |
|---|---|---|
| Phase-2 selection | `Plan/Phase2Resources` (`phase2-offsite-everything`) | Same plan, same `seahaven-backup-service-role`, same daily + cross-region copy rule |
**Phase-2 scope:** the 15 remaining DynamoDB tables (all except the two phase-1
financial tables + the deleted ledgerflow tables) and all 9 in-use EBS volumes,
again **by explicit ARN** — tag-based selection was deliberately avoided because
the file-share volumes are standalone-managed and the tables are owned by other
stacks, so tagging here would drift them.
**No IAM change:** `AWSBackupServiceRolePolicyForBackup` already grants the
DynamoDB/RDS/EBS backup actions, so phase 2 reuses the phase-1 role unchanged
(cross-reviewed, no BLOCK).
**Known tradeoff (→ Jira INFRA-31):** explicit-ARN EBS entries go stale if a
volume is replaced (new volume id), silently dropping it from backup. Migrating
the EBS portion to tag-based selection (with the tag codified in each owning
stack) is the resilient follow-up; scheduled drift detection is the interim
backstop.
**Also enabled outside this stack (audit H-7, via CLI — codify per stack →
INFRA-30):** PITR + `DeletionProtectionEnabled` on 12 more DynamoDB tables
(account-wide PITR now 19/21).
### Detective controls + budget (audit Day 1)
Account-level detective layer, in `lib/detective-controls.ts`, plus the cost
budget in `lib/governance-toggles.ts`. **Scope is us-east-1 only** (all workloads
live here); multi-region coverage is a follow-up.
| Resource | Logical ID | Finding | Notes |
|---|---|---|---|
| Config delivery bucket | `seahaven-config-328440206208` | H-2 | Private (BPA all), SSE-S3, versioned, TLS-only, 365d lifecycle |
| Config recorder role | `seahaven-config-recorder-role` | H-2 | `AWS_ConfigRole` + scoped S3 delivery; **IAM cross-reviewed** |
| Config recorder + channel | `DetectiveControls/ConfigPutRecorder`, `ConfigPutChannel`, `ConfigStartRecorder` | H-2 | `AwsCustomResource` calls `PutConfigurationRecorder` → `PutDeliveryChannel` → `StartConfigurationRecorder` in sequence; idempotent upsert avoids the L1 CFN deadlock (INFRA-17). Custom-resource role cross-reviewed. |
| GuardDuty detector | `DetectiveControls/GuardDutyDetector` | H-3 | Findings every 15 min |
| Security Hub | `DetectiveControls/SecurityHub` | H-4 | FSBP v1.0.0 + CIS v3.0.0; controls evaluate once Config is recording |
| Access Analyzer | `seahaven-account-analyzer` | M-5 | ACCOUNT external-access analyzer (free) |
| Monthly budget | `GovernanceToggles/MonthlyCostBudget` (`seahaven-monthly-cost`) | M-10 | $1,200/mo, 80%/100% actual + 100% forecast → adam@seahavenind.com |
**Config recorder + delivery channel are managed by `AwsCustomResource` (INFRA-17).**
The L1 `AWS::Config::ConfigurationRecorder` deadlocks the stack (recorder never
reaches `CREATE_COMPLETE` without a delivery channel; channel can't be created
without a recorder — hit 2026-06-01). The custom resource sidesteps this by
calling the Config SDK directly: `Put*` is an upsert, so the deploy adopts the
existing CLI-created recorder and channel without destroying them. Active
recording is never interrupted.
### CLI-applied governance toggles (no CloudFormation resource)
These account toggles have no native CloudFormation resource, so they are applied
via CLI and recorded here. Applied 2026-06-01.
```bash
# M-3 EBS encryption-by-default (new volumes; existing 5 plaintext volumes are H-19-adjacent)
aws ec2 enable-ebs-encryption-by-default --region us-east-1
# M-6 Inspector2 (EC2 + Lambda + ECR)
aws inspector2 enable --resource-types EC2 LAMBDA ECR --region us-east-1
# M-7 IAM password policy (CIS 1.8/1.9): >=14 chars, full complexity, no reuse of last 24
aws iam update-account-password-policy \
--minimum-password-length 14 \
--require-symbols --require-numbers \
--require-uppercase-characters --require-lowercase-characters \
--allow-users-to-change-password --password-reuse-prevention 24
# M-11 Activate cost-allocation tags (only activates keys already seen on resources)
aws ce update-cost-allocation-tags-status --cost-allocation-tags-status \
'TagKey=Project,Status=Active' 'TagKey=Owner,Status=Active' 'TagKey=Environment,Status=Active'
```
**L-8 (billing-metrics preference) is OUTSTANDING — console only.** Enabling the
CloudWatch `EstimatedCharges` metric in us-east-1 requires turning on *Receive
Billing Alerts* under Billing → Billing preferences; there is no public API/CLI.
The M-10 budget (`seahaven-monthly-cost`, 80%/100% actual + 100% forecast)
provides cost alerting independent of that metric.
> The legacy, manually-created `AWS-MonthlyBilling` CloudWatch alarm ($50
> threshold on `EstimatedCharges`, routed to `site-alerts`) was **deleted
> 2026-07-07** as unmanaged drift: it was fully redundant with the M-10 budget,
> sat permanently in ALARM (spend has far exceeded $50/mo), and was never in
> IaC. Billing alerting is now solely the managed M-10 budget. To restore the
> old alarm if ever needed: `aws cloudwatch put-metric-alarm --alarm-name
> AWS-MonthlyBilling --namespace AWS/Billing --metric-name EstimatedCharges
> --dimensions Name=Currency,Value=USD --statistic Maximum --period 86400
> --evaluation-periods 1 --threshold 50 --comparison-operator GreaterThanThreshold
> --alarm-actions arn:aws:sns:us-east-1:328440206208:site-alerts`.
### Monitoring + logging (audit Day 2)
| Resource | Logical ID | Finding | Notes |
|---|---|---|---|
| CIS metric filters + alarms | `CisMonitoring/*` | H-1 | 15 filters (CIS 4.1–4.15) on the CloudTrail log group, each with an alarm → `seahaven-cis-alarms`. ALARM-only actions (no OK). 4.16 = Security Hub (Day 1) |
| CIS alarm topic | `seahaven-cis-alarms` | H-1 | SNS, SSE (`alias/aws/sns`), email sub to adam@seahavenind.com |
| VPC flow logs | `FlowLogs/FlowLog0..4` | H-14 | ALL traffic on all 5 VPCs → S3 |
| Flow-logs bucket | `seahaven-vpc-flow-logs-328440206208` | H-14 | Private, SSE-S3, TLS-only, Glacier @90d / expire @365d; delivery bucket policy cross-reviewed |
| SES config set | `seahaven-email-events` | M-13 | Bounce/complaint/reject → CloudWatch metrics for reputation visibility |
| Sensitive-logs CMK | `LogsKey/Key` (`alias/seahaven-logs`) | M-24 | Encrypts sensitive CloudWatch Logs groups. Key policy grants `logs.us-east-1.amazonaws.com` Encrypt*/Decrypt*/ReEncrypt*/GenerateDataKey*/DescribeKey scoped by `kms:EncryptionContext:aws:logs:arn` (required or log delivery breaks). Rotation on, RETAIN. Applied in place to `TrailLogGroup` via escape hatch (same logical id/name). Cross-reviewed |
**M-24 sensitive log groups:** `alias/seahaven-logs` encrypts the CloudTrail CW
log group (codified here) plus the finance/PII Lambda groups owned by other
stacks — `exec-aide-*`, `payments-*`, `po-email-processor`, `vendor-reply-processor`
— which are associated via `aws logs associate-kms-key` and tracked as drift to
codify in their owning repos. The CloudTrail group is encrypted in place (escape
hatch on the existing `AWS::Logs::LogGroup`) so it is additive: same logical id +
physical name, no replacement, CIS Section-4 metric filters keep working. A
context flag `encryptTrailLogGroup` (default `true`) allows rolling the CMK out
and smoke-testing it on a low-risk Lambda group before the CloudTrail group:
```bash
# Phase 1: deploy CMK only, validate on a low-risk group
cdk deploy account-baseline --context encryptTrailLogGroup=false
# Phase 2: encrypt the CloudTrail group (default)
cdk deploy account-baseline
```
**H-1 log group:** the metric filters attach to the existing CloudTrail
CloudWatch Logs group by name (`seahaven-account-baseline-TrailLogGroup4CBE3AF5-…`),
imported read-only so the live audit trail is never replaced. Stable unless the
Trail is recreated.
**H-14 bucket policy note:** the flow-logs delivery policy keeps
`s3:x-amz-acl=bucket-owner-full-control` and the `arn:aws:logs:…:*` source-ARN
wildcard — both are required by AWS's documented flow-logs-to-S3 policy
(`flow-logs-s3-permissions.html`). A cross-review suggested dropping them; that
was rejected as it would break delivery. `s3:ListBucket` was dropped (not needed).
**M-13 follow-up:** associate `seahaven-email-events` as the default config set
on the live sending identities to capture events from existing senders:
```bash
aws sesv2 put-email-identity-configuration-set-attributes \
--email-identity int.seahaven.com --configuration-set-name seahaven-email-events
```
### Log-group retention + alarm wiring (audit L-4, L-5)
Applied via CLI (auto-created groups spread across stacks; one alarm in another
stack). Applied 2026-06-02.
```bash
# L-4 90-day retention on the 13 never-expire log groups (CodeBuild + CDK helpers)
for lg in <the 13 groups>; do aws logs put-retention-policy --log-group-name "$lg" --retention-in-days 90; done
# L-5 wire the actionless forgejo backup-verification alarm to site-alerts
aws cloudwatch put-metric-alarm --alarm-name forgejo-backup-verification-errors \
--alarm-actions arn:aws:sns:us-east-1:328440206208:site-alerts # (preserve existing alarm config)
```
## Roadmap (same stack)
Detective layer multi-region expansion (GuardDuty/Config/Security Hub beyond
us-east-1). Backup: phase 2 is deployed (see above); remaining is migrating the
phase-2 EBS entries to tag-based selection (INFRA-31) and graduating the offsite
vault to compliance mode.
## Deploy
CI/CD via the org reusable workflows (`ci-typescript-cdk.yaml`,
`cd-cdk.yaml`); pushes to `main` deploy through the OIDC role in
`secrets.AWS_DEPLOY_ROLE_ARN`. Local: `npm ci && npm run build && npx cdk diff`.
```
npx cdk deploy seahaven-account-baseline
```
## Verify
```
aws cloudtrail get-trail-status --name seahaven-org-trail # IsLogging: true
aws cloudtrail describe-trails --trail-name-list seahaven-org-trail
aws cloudtrail validate-logs --trail-arn <arn> --start-time <t> # digest integrity
```
AWS Backup (C-7):
```
aws backup list-backup-vaults # seahaven-primary
aws backup list-backup-vaults --region us-west-2 # seahaven-offsite
aws backup describe-backup-vault --backup-vault-name seahaven-offsite --region us-west-2 # Locked, MinRetentionDays
aws backup get-backup-plan --backup-plan-id <id> # daily rule + CopyAction
# Smoke test: on-demand backup of one resource, then confirm the cross-region copy lands
aws backup start-backup-job --backup-vault-name seahaven-primary \
--resource-arn arn:aws:rds:us-east-1:328440206208:db:proposal-system-db \
--iam-role-arn arn:aws:iam::328440206208:role/seahaven-backup-service-role
aws backup list-copy-jobs --region us-west-2 # copy to offsite present + COMPLETED
# Phase-2 selections live on the plan:
aws backup list-backup-selections --backup-plan-id <id> --query 'BackupSelectionsList[].SelectionName' # critical-data + phase2-offsite-everything
```
Detective layer + governance (Day 1):
```
aws configservice describe-configuration-recorder-status # recording: true
aws guardduty list-detectors # one detector id
aws securityhub get-enabled-standards # FSBP + CIS v3.0.0
aws accessanalyzer list-analyzers # seahaven-account-analyzer ACTIVE
aws inspector2 batch-get-account-status --region us-east-1 # ec2/ecr/lambda ENABLED
aws iam get-account-password-policy # length 14, reuse 24
aws ec2 get-ebs-encryption-by-default --region us-east-1 # EbsEncryptionByDefault: true
aws budgets describe-budgets --account-id 328440206208 # seahaven-monthly-cost $1,200
aws ce list-cost-allocation-tags --status Active # Project/Owner/Environment Active
```