shoc-frontend-new/infra/cdk/README.md
Adam Moussa 82361e14b5
feat(cdk): add Terraform adoption retain mode
retainForTerraformAdoption=true adds the required ManageSiteInfrastructure
parameter, conditions the 13 transferred resources and the S3 auto-delete
custom resource on it, applies Retain policies, pins the live dev origin
ID, attaches the deploy boundary and HcpTerraformWorkspace tag, and
narrows the OIDC subject to StringEquals. Normal synthesis is unchanged;
template tests cover both modes.
2026-09-10 19:15:11 -04:00

271 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Infrastructure & CI/CD — Sea Haven SHOC frontend
AWS hosting for the Vite SPA, defined as an **AWS CDK** app local to this repo.
Infrastructure deploys are administrator-run; GitHub Actions publishes content
only.
> **Dev is being adopted into HCP Terraform (SH-300).** The dev stack
> `shoc-frontend-dev` is in the retain/transfer sequence described under
> [Terraform adoption mode](#terraform-adoption-mode) and in
> [`terraform/README.md`](../terraform/README.md). Do not run a plain
> `cdk deploy` against dev while that sequence is in progress. Staging is
> unaffected and stays on this CDK path (SH-287 tracks its cutover).
- **Hosting:** private S3 bucket (origin) + CloudFront, served on the custom
domain **`dev.seahaven.com`** (ACM `*.seahaven.com`, Route 53 apex alias).
- **API:** the SPA calls the backend **directly** over HTTPS at
`https://api.dev.seahaven.com/api` (`VITE_API_URL`, cross-origin; the backend
allows CORS). CloudFront serves static content only — no `/api` proxy.
- Domain/cert/zone values live in `cdk.json` context so `cdk deploy` picks
them up with no flags. `VITE_API_URL` is baked into the build, so it's
per-environment (see the note under "Adding staging / prod").
- **Auth:** GitHub Actions → AWS via **OIDC** (no long-lived keys)
- **Content workflows:** `.github/workflows/deploy.yml` (dev,
`workflow_dispatch` only during adoption) and `deploy-staging.yml` (push to
`staging`) run `scripts/deploy-web.sh` as the environment's pinned deploy
role. Neither runs `cdk deploy`. The org reusable `cd-cdk.yaml` caller was
retired with the adoption PR.
- **Infra is local to this repo** (CDK in `infra/cdk`); the deploy role is
created by this stack, not added to the central `oidc-deploy-roles.yaml`.
```
infra/cdk/
bin/app.ts entry point (reads -c context)
lib/frontend-stack.ts S3 + CloudFront + OAC + OIDC deploy role
lib/retain-for-terraform-adoption.ts adoption-mode aspect (Retain + condition)
test/frontend-stack.test.mjs template assertions for both modes
scripts/deploy-web.sh build SPA -> s3 sync -> CloudFront invalidation
.github/workflows/
ci.yaml quality gates (lint / build / test / governance)
terraform-isolation.yaml PRs may not mix terraform/** with app code
deploy.yml dev content publish (workflow_dispatch on dev)
deploy-staging.yml standalone staging deploy (push to staging)
```
## What the stack creates
| Resource | Purpose |
| --------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| S3 bucket `seahaven-shoc-frontend-dev` | private origin (BLOCK_ALL, SSE, OAC-only reads) |
| CloudFront distribution | HTTPS, gzip/br; serves the static SPA from S3 (the app calls the API directly, cross-origin) |
| CloudFront Function (viewer request) | SPA routing: rewrites extensionless paths to `/index.html` (scoped to the S3 behavior, so it never touches `/api`) |
| IAM role `githubdeploy-shoc-frontend-new-dev` | assumed by GitHub Actions via OIDC, scoped to `repo:Sea-Haven-Industries/shoc-frontend-new:ref:refs/heads/dev` |
The dev role's inline policy still carries the legacy `cd-cdk.yaml` grants:
`sts:AssumeRole` on `cdk-hnb659fds-*`, `cloudformation:DescribeStacks`,
read/write on the bucket (`s3 sync`), and `cloudfront:CreateInvalidation`. It
is left byte-identical on purpose so the Terraform import is a no-op; the
Terraform content-CD change narrows it. The OIDC **provider** is a singleton
account resource — the stack only _imports_ it (created in step 2), so
`cdk destroy` can't delete a resource shared by other roles.
## Terraform adoption mode
`-c retainForTerraformAdoption=true` switches the stack into the safety mode
used only while HCP Terraform adopts the dev resources. It is off by default
and ordinary synthesis is unchanged (`test/frontend-stack.test.mjs` asserts
both). In adoption mode the stack:
- pins the origin ID CloudFormation generated for the live distribution
(`shocfrontenddevDistributionOrigin10CCD0EE1`) so the update is
metadata-only; environments without a verified value fail synthesis
- attaches the `seahaven-org-baseline` permissions boundary
`shoc-frontend-new-dev-deploy-boundary` and the
`HcpTerraformWorkspace=shoc-frontend-new-dev` tag to the deploy role
- narrows the OIDC subject condition from `StringLike` to `StringEquals` on the
same exact value
- applies `DeletionPolicy: Retain` and `UpdateReplacePolicy: Retain` to the 13
transferred resources (bucket, bucket policy, distribution, OAC, SPA
function, A and AAAA records, deploy role, inline policy) and to
`SiteBucket/AutoDeleteObjectsCustomResource`; the auto-delete provider
Lambda and role stay unretained
- adds the required `ManageSiteInfrastructure` parameter (`true|false`, no
default) and conditions those same resources and every output on it
- emits `TerraformImport*` outputs carrying the exact import IDs
`ManageSiteInfrastructure` has no default, so every adoption-mode deploy must
state the ownership phase:
```bash
cd infra/cdk && npm ci
# Phase 1, before the Terraform import: keep the resources in the stack and
# install Retain on them. Update-only change set.
npx cdk deploy shoc-frontend-dev \
-c retainForTerraformAdoption=true \
--parameters ManageSiteInfrastructure=true
# Phase 2, after the controlled Terraform apply and its no-op plan: relinquish
# ownership. Expect DELETE_SKIPPED on the 13 resources and the custom resource.
npx cdk deploy shoc-frontend-dev \
-c retainForTerraformAdoption=true \
--parameters ManageSiteInfrastructure=false
```
Both deploys must use the same reviewed SHA. Review the change set before
confirming: Phase 1 must show no create, delete, or replace. After the
`false` deploy succeeds, `ManageSiteInfrastructure=true` must never be used
again. If the `true` deploy rolls back, inspect the stack resources and the
live bucket before retrying; retained resources can outlive a failed update and
must not be cleaned up automatically. Never delete the auto-delete custom
resource while its handler can still empty the versioned bucket.
Local checks (`npm run test:infra` from the repo root) build the app, run the
template assertions, and synthesize both modes.
---
## One-time setup (run by a human with admin AWS creds)
### 1. Authenticate to the AWS account
```bash
aws configure # or: aws sso login --profile <admin>
aws sts get-caller-identity # confirm the right account + region (us-east-1)
```
### 2. Ensure the GitHub OIDC provider exists (once per account)
```bash
aws iam list-open-id-connect-providers
# If none ends in token.actions.githubusercontent.com, create it (thumbprint is
# no longer required — AWS validates GitHub against its own trust store):
aws iam create-open-id-connect-provider \
--url https://token.actions.githubusercontent.com \
--client-id-list sts.amazonaws.com
```
### 3. CDK bootstrap (once per account/region)
```bash
cd infra/cdk
npm ci
npx cdk bootstrap aws://<ACCOUNT_ID>/us-east-1
```
### 4. Domain, cert, and API URL (already wired for dev)
Domain/cert/zone are set in `cdk.json` context (account `396287094661`):
| Context key | Value |
| --------------------------------- | ------------------------------------------------------------ |
| `domainNames` | `dev.seahaven.com` |
| `certificateArn` | `…:certificate/2b78e74f-…` (ACM `*.seahaven.com`, us-east-1) |
| `hostedZoneId` / `hostedZoneName` | `Z07671212N75U4YLPWZR8` / `dev.seahaven.com` |
The stack creates the apex A/AAAA alias in the hosted zone (in this account,
delegated from the parent `seahaven.com` zone). The **API URL is not infra** —
it's `VITE_API_URL` in `.env.production` (`https://api.dev.seahaven.com/api`),
baked into the build. Per-environment; override for staging/prod.
### 5. First deploy (locally, with admin creds)
The deploy role doesn't exist until the first `cdk deploy`, so bootstrap it
locally. This provisions infra + the role:
```bash
cd infra/cdk
npx cdk deploy
```
Note the `DeployRoleArn` output. Then publish the first content manually:
```bash
# from repo root, optional manual first content publish:
STACK_NAME=shoc-frontend-dev AWS_REGION=us-east-1 bash scripts/deploy-web.sh
```
### 6. Content deploys
The deploy role ARN is deterministic and pinned in
`.github/workflows/deploy.yml` (no `AWS_DEPLOY_ROLE_ARN` secret). During the
Terraform adoption, dev content deploys run only through **Actions → Deploy dev
content → Run workflow** on `dev`. The workflow runs `npm run verify`, assumes
`githubdeploy-shoc-frontend-new-dev`, runs `scripts/deploy-web.sh` against the
pinned bucket and distribution, uploads source maps, and verifies the served
`index.html` matches the build. Automatic push-to-`dev` releases return with the
Terraform content-CD change.
---
## Staging environment (same account, exact OIDC subject)
Staging lives in the same AWS account (396287094661) and deploys through its
own standalone workflow, `.github/workflows/deploy-staging.yml`, on push to
`staging`:
- **Trust:** with `-c githubEnvironment=staging`, the stack's deploy role
(`githubdeploy-shoc-frontend-new-staging`) trusts ONLY the exact GitHub
environment subject
`repo:Sea-Haven-Industries/shoc-frontend-new:environment:staging`
(`StringEquals` on both `aud` and `sub`). The workflow declares
`environment: staging`, so only runs in that environment can assume the role.
Without `githubEnvironment`, the dev stack keeps its branch-ref trust
unchanged.
- **No secret:** the role ARN is static (the role name is deterministic), so
the workflow pins
`arn:aws:iam::396287094661:role/githubdeploy-shoc-frontend-new-staging`
directly — no `AWS_DEPLOY_ROLE_ARN`-style secret to set.
- **Gates first:** the workflow runs the full `npm run verify` before assuming
the staging role, then runs `scripts/deploy-web.sh` with
`STACK_NAME=shoc-frontend-staging`,
`VITE_API_URL=https://api.staging.seahaven.com/api`, and waits for the
CloudFront invalidation to complete.
- **Application-only role:** the recurring staging workflow can describe only
its exact stack, publish only to its exact bucket, and invalidate only its
exact distribution. It cannot assume the shared CDK bootstrap roles or
modify infrastructure. Staging infrastructure changes use the Administrator
command below.
- **Post-deploy checks:** bucket + distribution existence, HTTPS on
`https://staging.seahaven.com`, and the actual post-invalidation remote assets
contain the staging API URL and no dev API URL. (Not browser QA.)
### One-time setup (run by a human with admin AWS creds + GitHub Admin)
1. **GitHub Admin — create the `staging` environment** (Settings →
Environments → New environment → `staging`). Add protection rules as
appropriate (e.g. required reviewers, restrict to the `staging` branch). If
the environment does not exist, GitHub creates it unprotected on first use.
2. **AWS Admin — first deploy with admin creds** (same steps 1–3 as dev; the
OIDC provider and bootstrap already exist in this account):
```bash
cd infra/cdk
npx cdk deploy shoc-frontend-staging \
-c envName=staging \
-c deployBranch=staging \
-c githubEnvironment=staging \
-c domainNames=staging.seahaven.com \
-c certificateArn=arn:aws:acm:us-east-1:396287094661:certificate/2b78e74f-7b65-4b82-a413-7a498b102f00 \
-c hostedZoneId=Z02602739VQWBWCAGXP4 \
-c hostedZoneName=staging.seahaven.com
```
The `DeployRoleArn` output must match the ARN pinned in
`deploy-staging.yml` (it will — the role name is deterministic).
3. **Backend CORS:** the staging API (`https://api.staging.seahaven.com`) must
allow the `https://staging.seahaven.com` origin.
4. Push to `staging` — `ci.yaml` runs the quality gates and
`deploy-staging.yml` deploys.
### Adding prod later
Same pattern: a prod account/stack with its own contexts and, ideally, its own
`githubEnvironment=prod` trust + workflow. Keep in mind `VITE_API_URL` is baked
into each environment's build, and the bucket's `RemovalPolicy.DESTROY` +
`autoDeleteObjects` defaults are dev/staging-friendly but should be revisited
for prod.
## Notes
- **Teardown:** `npx cdk destroy`. The bucket uses `RemovalPolicy.DESTROY` +
`autoDeleteObjects` (dev artifacts are reproducible) — change this for prod.
Never run it against dev during or after the Terraform adoption: the
adoption-mode stack retains the transferred resources, and after Phase 2
Terraform owns them.
- **CI and staging CD both fire on push to `staging`** in parallel; the
staging CD workflow runs `npm run verify` itself before deploying. Dev has
no push-triggered deploy during the adoption.
- **npm is pinned to v11.16.0**; the committed `package-lock.json` uses
lockfileVersion 3, matching the Node 24 / npm 11 CI environment.