mirror of
https://github.com/Sea-Haven-Industries/seahaven-account-baseline.git
synced 2026-08-04 16:56:14 +00:00
README: document AWS Backup phase-2 (phase2-offsite-everything selection), remove retired database-1 from phase-1 scope (audit H-19), update roadmap + verify smoke-test to a live resource. scripts/iam-user-delete.sh: reusable full IAM user teardown (keys, policies, groups, MFA, login profile, certs, SSH keys, service creds, then user) with --profile/--yes and a guard against deleting the caller's own identity. Built from the Day 4 audit IAM cleanup.
269 lines
14 KiB
Markdown
269 lines
14 KiB
Markdown
# seahaven-account-baseline
|
||
|
||
Account-level security and governance baseline for Sea Haven Industries
|
||
(AWS account **328440206208**), managed as a single CDK TypeScript app. Most
|
||
resources are in **us-east-1**; the offsite backup vault is in **us-west-2**.
|
||
This is where account-wide detective and recovery controls live, so they are
|
||
versioned, reviewed, and drift-checked like any other stack.
|
||
|
||
Stacks (all deployed by `cdk deploy --all` / the CD workflow):
|
||
|
||
| Stack | Region | Purpose |
|
||
|---|---|---|
|
||
| `seahaven-account-baseline` | us-east-1 | CloudTrail + future detective controls (C-1) |
|
||
| `seahaven-backup` | us-east-1 | Primary AWS Backup vault + plan + role (C-7) |
|
||
| `seahaven-backup-offsite` | us-west-2 | Governance-locked offsite copy vault (C-7) |
|
||
|
||
## What it deploys
|
||
|
||
### CloudTrail (audit finding C-1)
|
||
|
||
| Resource | Logical ID | Notes |
|
||
|---|---|---|
|
||
| Multi-region trail | `Trail` (`seahaven-org-trail`) | Management events read+write, global service events, **log-file validation on** |
|
||
| Log bucket | `TrailLogBucket` (`seahaven-cloudtrail-logs-328440206208`) | Private (Block Public Access all), SSE-KMS, versioned, **TLS-only**, **Object Lock GOVERNANCE 365d**, lifecycle (Glacier @90d, expire @365d), server access logging → `seahaven-s3-access-logs` |
|
||
| KMS CMK | `TrailKey` (`alias/cloudtrail-logs`) | Encrypts log files; **automatic rotation enabled** |
|
||
| CloudWatch Logs group | created by the L2 `Trail` | 365-day retention; this is the group the CIS Section 4 metric filters (H-1) attach to |
|
||
|
||
**Data flow:** API activity across all regions → CloudTrail → (a) KMS-encrypted,
|
||
Object-Locked S3 bucket for durable/tamper-resistant storage and (b) CloudWatch
|
||
Logs for real-time querying and metric-filter alarms.
|
||
|
||
**Compliance impact:** closes CIS 3.1 (multi-region trail), 3.2 (log-file
|
||
validation), 3.4 (CloudWatch Logs integration), 3.6 (bucket access logging),
|
||
3.7 (KMS CMK encryption), and 3.8 (CMK rotation). Unblocks CIS Section 4 /
|
||
finding H-1 (metric filters + alarms now have a log group to target).
|
||
|
||
### Design decisions
|
||
|
||
- **Management events only.** Object-level S3/Lambda data events (CIS 3.10/3.11)
|
||
are deferred to control cost; revisit with targeted S3 *write* data events on
|
||
sensitive buckets (payments / accounting / kb) if needed.
|
||
- **Object Lock GOVERNANCE, not COMPLIANCE.** Tamper-resistant but still
|
||
deletable by a principal holding `s3:BypassGovernanceRetention` — avoids the
|
||
irreversibility of COMPLIANCE mode. Revisit if a stricter posture is required.
|
||
- **RETAIN** on the bucket and KMS key so a stack teardown never destroys the
|
||
audit trail.
|
||
|
||
### AWS Backup (audit finding C-7)
|
||
|
||
Phase 1 ("critical data first") of fixing the account's complete lack of AWS
|
||
Backup. Protects the data stores with no offsite leg today and copies each
|
||
recovery point cross-region into a governance-locked vault.
|
||
|
||
| Resource | Logical ID | Notes |
|
||
|---|---|---|
|
||
| Primary vault | `seahaven-primary` (us-east-1) | KMS-CMK encrypted, unlocked (working copy), RETAIN |
|
||
| Offsite vault | `seahaven-offsite` (us-west-2) | KMS-CMK encrypted, **Vault Lock GOVERNANCE** (min-retention 30d, no cooling-off window), RETAIN |
|
||
| Backup plan | `seahaven-critical-daily` | Daily 06:00 UTC, delete-after 35d, **cross-region CopyAction → offsite** (retain 90d) |
|
||
| Service role | `seahaven-backup-service-role` | **Backup-only** (Backup + S3-Backup managed policies); restore perms intentionally deferred |
|
||
|
||
**Phase-1 scope** (selected by explicit ARN, not tags, to avoid drifting other
|
||
stacks): RDS `proposal-system-db`, DynamoDB `PaymentsDashboard`,
|
||
DynamoDB `purchase-orders`, S3 `accounting.seahaven.com`,
|
||
`seahaven-payments-csv-328440206208`, `google-workspace-seahavenind.com`.
|
||
*(RDS `database-1` was originally in this set but was retired 2026-06-03 —
|
||
audit H-19, idle 0 conn/60d — and removed from the selection; its final
|
||
encrypted recovery point is retained in `seahaven-offsite` for 7 years.)*
|
||
|
||
**Coexists with** existing EBS DLM snapshots and DynamoDB PITR — it supplements
|
||
them with the missing offsite + immutable leg; it does not replace them.
|
||
|
||
**Design decisions:**
|
||
|
||
- **Governance lock first, not compliance.** Recovery points can't be silently
|
||
deleted, but a principal with explicit permission can still intervene while
|
||
we validate. Graduate to COMPLIANCE (irreversible) later by adding
|
||
`changeableFor` to the offsite vault lock + redeploy.
|
||
- **Backup-only role.** Restore policies and `allowRestores` are not granted;
|
||
restores get a separate audited path once a restore-test process exists.
|
||
|
||
**Pre-deploy gates** (must clear before the first scheduled run):
|
||
|
||
1. Enable S3 versioning on `seahaven-payments-csv-328440206208` and
|
||
`google-workspace-seahavenind.com` (`accounting.seahaven.com` already has it,
|
||
audit C-9), or their jobs fail silently (folds in H-21).
|
||
2. `database-1` is unencrypted (H-19): smoke-test an on-demand backup + copy of
|
||
it to us-west-2 first; if the copy fails, encrypt it or drop it from the copy.
|
||
3. Enable DynamoDB PITR (H-7) on the two tables for between-window recovery.
|
||
|
||
### AWS Backup phase 2 (audit Day 4)
|
||
|
||
Expands the same `seahaven-critical-daily` plan to every remaining data store, so
|
||
all of DynamoDB + EBS get the offsite + immutable leg ("offsite for everything").
|
||
|
||
| Resource | Logical ID | Notes |
|
||
|---|---|---|
|
||
| Phase-2 selection | `Plan/Phase2Resources` (`phase2-offsite-everything`) | Same plan, same `seahaven-backup-service-role`, same daily + cross-region copy rule |
|
||
|
||
**Phase-2 scope:** the 15 remaining DynamoDB tables (all except the two phase-1
|
||
financial tables + the deleted ledgerflow tables) and all 9 in-use EBS volumes,
|
||
again **by explicit ARN** — tag-based selection was deliberately avoided because
|
||
the file-share volumes are standalone-managed and the tables are owned by other
|
||
stacks, so tagging here would drift them.
|
||
|
||
**No IAM change:** `AWSBackupServiceRolePolicyForBackup` already grants the
|
||
DynamoDB/RDS/EBS backup actions, so phase 2 reuses the phase-1 role unchanged
|
||
(cross-reviewed, no BLOCK).
|
||
|
||
**Known tradeoff (→ Jira INFRA-31):** explicit-ARN EBS entries go stale if a
|
||
volume is replaced (new volume id), silently dropping it from backup. Migrating
|
||
the EBS portion to tag-based selection (with the tag codified in each owning
|
||
stack) is the resilient follow-up; scheduled drift detection is the interim
|
||
backstop.
|
||
|
||
**Also enabled outside this stack (audit H-7, via CLI — codify per stack →
|
||
INFRA-30):** PITR + `DeletionProtectionEnabled` on 12 more DynamoDB tables
|
||
(account-wide PITR now 19/21).
|
||
|
||
### Detective controls + budget (audit Day 1)
|
||
|
||
Account-level detective layer, in `lib/detective-controls.ts`, plus the cost
|
||
budget in `lib/governance-toggles.ts`. **Scope is us-east-1 only** (all workloads
|
||
live here); multi-region coverage is a follow-up.
|
||
|
||
| Resource | Logical ID | Finding | Notes |
|
||
|---|---|---|---|
|
||
| Config delivery bucket | `seahaven-config-328440206208` | H-2 | Private (BPA all), SSE-S3, versioned, TLS-only, 365d lifecycle |
|
||
| Config recorder role | `seahaven-config-recorder-role` | H-2 | `AWS_ConfigRole` + scoped S3 delivery; **IAM cross-reviewed** |
|
||
| GuardDuty detector | `DetectiveControls/GuardDutyDetector` | H-3 | Findings every 15 min |
|
||
| Security Hub | `DetectiveControls/SecurityHub` | H-4 | FSBP v1.0.0 + CIS v3.0.0; controls evaluate once Config is recording |
|
||
| Access Analyzer | `seahaven-account-analyzer` | M-5 | ACCOUNT external-access analyzer (free) |
|
||
| Monthly budget | `GovernanceToggles/MonthlyCostBudget` (`seahaven-monthly-cost`) | M-10 | $1,200/mo, 80%/100% actual + 100% forecast → adam@seahavenind.com |
|
||
|
||
**Config recorder + delivery channel are NOT in CloudFormation.** The L1
|
||
`AWS::Config::ConfigurationRecorder` is a stabilizing resource that hangs the
|
||
stack: it never reaches `CREATE_COMPLETE` until recording is active, which needs
|
||
a delivery channel, which can't be created until the recorder completes — a
|
||
deadlock (hit on 2026-06-01). The role + delivery bucket stay in IaC (the role
|
||
is cross-reviewed); the recorder/channel are created via CLI (below), referencing
|
||
the stack's `ConfigRecorderRoleArn` output and the `seahaven-config-328440206208`
|
||
bucket.
|
||
|
||
### CLI-applied governance toggles (no CloudFormation resource)
|
||
|
||
These account toggles have no native CloudFormation resource, so they are applied
|
||
via CLI and recorded here. Applied 2026-06-01.
|
||
|
||
```bash
|
||
# M-3 EBS encryption-by-default (new volumes; existing 5 plaintext volumes are H-19-adjacent)
|
||
aws ec2 enable-ebs-encryption-by-default --region us-east-1
|
||
|
||
# M-6 Inspector2 (EC2 + Lambda + ECR)
|
||
aws inspector2 enable --resource-types EC2 LAMBDA ECR --region us-east-1
|
||
|
||
# M-7 IAM password policy (CIS 1.8/1.9): >=14 chars, full complexity, no reuse of last 24
|
||
aws iam update-account-password-policy \
|
||
--minimum-password-length 14 \
|
||
--require-symbols --require-numbers \
|
||
--require-uppercase-characters --require-lowercase-characters \
|
||
--allow-users-to-change-password --password-reuse-prevention 24
|
||
|
||
# M-11 Activate cost-allocation tags (only activates keys already seen on resources)
|
||
aws ce update-cost-allocation-tags-status --cost-allocation-tags-status \
|
||
'TagKey=Project,Status=Active' 'TagKey=Owner,Status=Active' 'TagKey=Environment,Status=Active'
|
||
```
|
||
|
||
**L-8 (billing-metrics preference) is OUTSTANDING — console only.** Enabling the
|
||
CloudWatch `EstimatedCharges` metric in us-east-1 requires turning on *Receive
|
||
Billing Alerts* under Billing → Billing preferences; there is no public API/CLI.
|
||
The M-10 budget already provides cost alerting independent of that metric, so
|
||
this only affects the legacy `AWS-MonthlyBilling` CloudWatch alarm (L-8).
|
||
|
||
### Monitoring + logging (audit Day 2)
|
||
|
||
| Resource | Logical ID | Finding | Notes |
|
||
|---|---|---|---|
|
||
| CIS metric filters + alarms | `CisMonitoring/*` | H-1 | 15 filters (CIS 4.1–4.15) on the CloudTrail log group, each with an alarm → `seahaven-cis-alarms`. ALARM-only actions (no OK). 4.16 = Security Hub (Day 1) |
|
||
| CIS alarm topic | `seahaven-cis-alarms` | H-1 | SNS, SSE (`alias/aws/sns`), email sub to adam@seahavenind.com |
|
||
| VPC flow logs | `FlowLogs/FlowLog0..4` | H-14 | ALL traffic on all 5 VPCs → S3 |
|
||
| Flow-logs bucket | `seahaven-vpc-flow-logs-328440206208` | H-14 | Private, SSE-S3, TLS-only, Glacier @90d / expire @365d; delivery bucket policy cross-reviewed |
|
||
| SES config set | `seahaven-email-events` | M-13 | Bounce/complaint/reject → CloudWatch metrics for reputation visibility |
|
||
|
||
**H-1 log group:** the metric filters attach to the existing CloudTrail
|
||
CloudWatch Logs group by name (`seahaven-account-baseline-TrailLogGroup4CBE3AF5-…`),
|
||
imported read-only so the live audit trail is never replaced. Stable unless the
|
||
Trail is recreated.
|
||
|
||
**H-14 bucket policy note:** the flow-logs delivery policy keeps
|
||
`s3:x-amz-acl=bucket-owner-full-control` and the `arn:aws:logs:…:*` source-ARN
|
||
wildcard — both are required by AWS's documented flow-logs-to-S3 policy
|
||
(`flow-logs-s3-permissions.html`). A cross-review suggested dropping them; that
|
||
was rejected as it would break delivery. `s3:ListBucket` was dropped (not needed).
|
||
|
||
**M-13 follow-up:** associate `seahaven-email-events` as the default config set
|
||
on the live sending identities to capture events from existing senders:
|
||
|
||
```bash
|
||
aws sesv2 put-email-identity-configuration-set-attributes \
|
||
--email-identity int.seahaven.com --configuration-set-name seahaven-email-events
|
||
```
|
||
|
||
### Log-group retention + alarm wiring (audit L-4, L-5)
|
||
|
||
Applied via CLI (auto-created groups spread across stacks; one alarm in another
|
||
stack). Applied 2026-06-02.
|
||
|
||
```bash
|
||
# L-4 90-day retention on the 13 never-expire log groups (CodeBuild + CDK helpers)
|
||
for lg in <the 13 groups>; do aws logs put-retention-policy --log-group-name "$lg" --retention-in-days 90; done
|
||
|
||
# L-5 wire the actionless forgejo backup-verification alarm to site-alerts
|
||
aws cloudwatch put-metric-alarm --alarm-name forgejo-backup-verification-errors \
|
||
--alarm-actions arn:aws:sns:us-east-1:328440206208:site-alerts # (preserve existing alarm config)
|
||
```
|
||
|
||
## Roadmap (same stack)
|
||
|
||
Detective layer multi-region expansion (GuardDuty/Config/Security Hub beyond
|
||
us-east-1). Backup: phase 2 is deployed (see above); remaining is migrating the
|
||
phase-2 EBS entries to tag-based selection (INFRA-31) and graduating the offsite
|
||
vault to compliance mode.
|
||
|
||
## Deploy
|
||
|
||
CI/CD via the org reusable workflows (`ci-typescript-cdk.yaml`,
|
||
`cd-cdk.yaml`); pushes to `main` deploy through the OIDC role in
|
||
`secrets.AWS_DEPLOY_ROLE_ARN`. Local: `npm ci && npm run build && npx cdk diff`.
|
||
|
||
```
|
||
npx cdk deploy seahaven-account-baseline
|
||
```
|
||
|
||
## Verify
|
||
|
||
```
|
||
aws cloudtrail get-trail-status --name seahaven-org-trail # IsLogging: true
|
||
aws cloudtrail describe-trails --trail-name-list seahaven-org-trail
|
||
aws cloudtrail validate-logs --trail-arn <arn> --start-time <t> # digest integrity
|
||
```
|
||
|
||
AWS Backup (C-7):
|
||
|
||
```
|
||
aws backup list-backup-vaults # seahaven-primary
|
||
aws backup list-backup-vaults --region us-west-2 # seahaven-offsite
|
||
aws backup describe-backup-vault --backup-vault-name seahaven-offsite --region us-west-2 # Locked, MinRetentionDays
|
||
aws backup get-backup-plan --backup-plan-id <id> # daily rule + CopyAction
|
||
# Smoke test: on-demand backup of one resource, then confirm the cross-region copy lands
|
||
aws backup start-backup-job --backup-vault-name seahaven-primary \
|
||
--resource-arn arn:aws:rds:us-east-1:328440206208:db:proposal-system-db \
|
||
--iam-role-arn arn:aws:iam::328440206208:role/seahaven-backup-service-role
|
||
aws backup list-copy-jobs --region us-west-2 # copy to offsite present + COMPLETED
|
||
# Phase-2 selections live on the plan:
|
||
aws backup list-backup-selections --backup-plan-id <id> --query 'BackupSelectionsList[].SelectionName' # critical-data + phase2-offsite-everything
|
||
```
|
||
|
||
Detective layer + governance (Day 1):
|
||
|
||
```
|
||
aws configservice describe-configuration-recorder-status # recording: true
|
||
aws guardduty list-detectors # one detector id
|
||
aws securityhub get-enabled-standards # FSBP + CIS v3.0.0
|
||
aws accessanalyzer list-analyzers # seahaven-account-analyzer ACTIVE
|
||
aws inspector2 batch-get-account-status --region us-east-1 # ec2/ecr/lambda ENABLED
|
||
aws iam get-account-password-policy # length 14, reuse 24
|
||
aws ec2 get-ebs-encryption-by-default --region us-east-1 # EbsEncryptionByDefault: true
|
||
aws budgets describe-budgets --account-id 328440206208 # seahaven-monthly-cost $1,200
|
||
aws ce list-cost-allocation-tags --status Active # Project/Owner/Environment Active
|
||
```
|