13 KiB
Infrastructure & CI/CD — Sea Haven SHOC frontend
AWS hosting for the Vite SPA, defined as an AWS CDK app local to this repo. Infrastructure deploys are administrator-run; GitHub Actions publishes content only.
Dev is being adopted into HCP Terraform (SH-300). The dev stack
shoc-frontend-devis in the retain/transfer sequence described under Terraform adoption mode and interraform/README.md. Do not run a plaincdk deployagainst dev while that sequence is in progress. Staging is unaffected and stays on this CDK path (SH-287 tracks its cutover).
- Hosting: private S3 bucket (origin) + CloudFront, served on the custom
domain
dev.seahaven.com(ACM*.seahaven.com, Route 53 apex alias). - API: the SPA calls the backend directly over HTTPS at
https://api.dev.seahaven.com/api(VITE_API_URL, cross-origin; the backend allows CORS). CloudFront serves static content only — no/apiproxy. - Domain/cert/zone values live in
cdk.jsoncontext socdk deploypicks them up with no flags.VITE_API_URLis baked into the build, so it's per-environment (see the note under "Adding staging / prod"). - Auth: GitHub Actions → AWS via OIDC (no long-lived keys)
- Content workflows:
.github/workflows/deploy.yml(dev,workflow_dispatchonly during adoption) anddeploy-staging.yml(push tostaging) runscripts/deploy-web.shas the environment's pinned deploy role. Neither runscdk deploy. The org reusablecd-cdk.yamlcaller was retired with the adoption PR. - Infra is local to this repo (CDK in
infra/cdk); the deploy role is created by this stack, not added to the centraloidc-deploy-roles.yaml.
infra/cdk/
bin/app.ts entry point (reads -c context)
lib/frontend-stack.ts S3 + CloudFront + OAC + OIDC deploy role
lib/retain-for-terraform-adoption.ts adoption-mode aspect (Retain + condition)
test/frontend-stack.test.mjs template assertions for both modes
scripts/deploy-web.sh build SPA -> s3 sync -> CloudFront invalidation
.github/workflows/
ci.yaml quality gates (lint / build / test / governance / terraform isolation)
deploy.yml dev content publish (workflow_dispatch on dev)
deploy-staging.yml standalone staging deploy (push to staging)
What the stack creates
| Resource | Purpose |
|---|---|
S3 bucket seahaven-shoc-frontend-dev |
private origin (BLOCK_ALL, SSE, OAC-only reads) |
| CloudFront distribution | HTTPS, gzip/br; serves the static SPA from S3 (the app calls the API directly, cross-origin) |
| CloudFront Function (viewer request) | SPA routing: rewrites extensionless paths to /index.html (scoped to the S3 behavior, so it never touches /api) |
IAM role githubdeploy-shoc-frontend-new-dev |
assumed by GitHub Actions via OIDC, scoped to repo:Sea-Haven-Industries/shoc-frontend-new:ref:refs/heads/dev |
The dev role's inline policy still carries the legacy cd-cdk.yaml grants:
sts:AssumeRole on cdk-hnb659fds-*, cloudformation:DescribeStacks,
read/write on the bucket (s3 sync), and cloudfront:CreateInvalidation. It
is left byte-identical on purpose so the Terraform import is a no-op; the
Terraform content-CD change narrows it. The OIDC provider is a singleton
account resource — the stack only imports it (created in step 2), so
cdk destroy can't delete a resource shared by other roles.
Terraform adoption mode
-c retainForTerraformAdoption=true switches the stack into the safety mode
used only while HCP Terraform adopts the dev resources. It is off by default
and ordinary synthesis is unchanged (test/frontend-stack.test.mjs asserts
both). In adoption mode the stack:
- pins the origin ID CloudFormation generated for the live distribution
(
shocfrontenddevDistributionOrigin10CCD0EE1) so the update is metadata-only; environments without a verified value fail synthesis - attaches the
seahaven-org-baselinepermissions boundaryshoc-frontend-new-dev-deploy-boundaryand theHcpTerraformWorkspace=shoc-frontend-new-devtag to the deploy role - narrows the OIDC subject condition from
StringLiketoStringEqualson the same exact value - applies
DeletionPolicy: RetainandUpdateReplacePolicy: Retainto the 13 transferred resources (bucket, bucket policy, distribution, OAC, SPA function, A and AAAA records, deploy role, inline policy) and toSiteBucket/AutoDeleteObjectsCustomResource; the auto-delete provider Lambda and role stay unretained - adds the required
ManageSiteInfrastructureparameter (true|false, no default) and conditions those same resources and every output on it - emits
TerraformImport*outputs carrying the exact import IDs
ManageSiteInfrastructure has no default, so every adoption-mode deploy must
state the ownership phase:
cd infra/cdk && npm ci
# Phase 1, before the Terraform import: keep the resources in the stack and
# install Retain on them. Update-only change set.
npx cdk deploy shoc-frontend-dev \
-c retainForTerraformAdoption=true \
--parameters ManageSiteInfrastructure=true
# Phase 2, after the controlled Terraform apply and its no-op plan: relinquish
# ownership. Expect DELETE_SKIPPED on the 13 resources and the custom resource.
npx cdk deploy shoc-frontend-dev \
-c retainForTerraformAdoption=true \
--parameters ManageSiteInfrastructure=false
Both deploys must use the same reviewed SHA. Review the change set before
confirming: Phase 1 must show no create, delete, or replace. After the
false deploy succeeds, ManageSiteInfrastructure=true must never be used
again. If the true deploy rolls back, inspect the stack resources and the
live bucket before retrying; retained resources can outlive a failed update and
must not be cleaned up automatically. Never delete the auto-delete custom
resource while its handler can still empty the versioned bucket.
Local checks (npm run test:infra from the repo root) build the app, run the
template assertions, and synthesize both modes.
One-time setup (run by a human with admin AWS creds)
1. Authenticate to the AWS account
aws configure # or: aws sso login --profile <admin>
aws sts get-caller-identity # confirm the right account + region (us-east-1)
2. Ensure the GitHub OIDC provider exists (once per account)
aws iam list-open-id-connect-providers
# If none ends in token.actions.githubusercontent.com, create it (thumbprint is
# no longer required — AWS validates GitHub against its own trust store):
aws iam create-open-id-connect-provider \
--url https://token.actions.githubusercontent.com \
--client-id-list sts.amazonaws.com
3. CDK bootstrap (once per account/region)
cd infra/cdk
npm ci
npx cdk bootstrap aws://<ACCOUNT_ID>/us-east-1
4. Domain, cert, and API URL (already wired for dev)
Domain/cert/zone are set in cdk.json context (account 396287094661):
| Context key | Value |
|---|---|
domainNames |
dev.seahaven.com |
certificateArn |
…:certificate/2b78e74f-… (ACM *.seahaven.com, us-east-1) |
hostedZoneId / hostedZoneName |
Z07671212N75U4YLPWZR8 / dev.seahaven.com |
The stack creates the apex A/AAAA alias in the hosted zone (in this account,
delegated from the parent seahaven.com zone). The API URL is not infra —
it's VITE_API_URL in .env.production (https://api.dev.seahaven.com/api),
baked into the build. Per-environment; override for staging/prod.
5. First deploy (locally, with admin creds)
The deploy role doesn't exist until the first cdk deploy, so bootstrap it
locally. This provisions infra + the role:
cd infra/cdk
npx cdk deploy
Note the DeployRoleArn output. Then publish the first content manually:
# from repo root, optional manual first content publish:
STACK_NAME=shoc-frontend-dev AWS_REGION=us-east-1 bash scripts/deploy-web.sh
6. Content deploys
The deploy role ARN is deterministic and pinned in
.github/workflows/deploy.yml (no AWS_DEPLOY_ROLE_ARN secret). During the
Terraform adoption, dev content deploys run only through Actions → Deploy dev
content → Run workflow on dev. The workflow runs npm run verify, assumes
githubdeploy-shoc-frontend-new-dev, runs scripts/deploy-web.sh against the
pinned bucket and distribution, uploads source maps, and verifies the served
index.html matches the build. Automatic push-to-dev releases return with the
Terraform content-CD change.
Staging environment (same account, exact OIDC subject)
Staging lives in the same AWS account (396287094661) and deploys through its
own standalone workflow, .github/workflows/deploy-staging.yml, on push to
staging:
- Trust: with
-c githubEnvironment=staging, the stack's deploy role (githubdeploy-shoc-frontend-new-staging) trusts ONLY the exact GitHub environment subjectrepo:Sea-Haven-Industries/shoc-frontend-new:environment:staging(StringEqualson bothaudandsub). The workflow declaresenvironment: staging, so only runs in that environment can assume the role. WithoutgithubEnvironment, the dev stack keeps its branch-ref trust unchanged. - No secret: the role ARN is static (the role name is deterministic), so
the workflow pins
arn:aws:iam::396287094661:role/githubdeploy-shoc-frontend-new-stagingdirectly — noAWS_DEPLOY_ROLE_ARN-style secret to set. - Gates first: the workflow runs the full
npm run verifybefore assuming the staging role, then runsscripts/deploy-web.shwithSTACK_NAME=shoc-frontend-staging,VITE_API_URL=https://api.staging.seahaven.com/api, and waits for the CloudFront invalidation to complete. - Application-only role: the recurring staging workflow can describe only its exact stack, publish only to its exact bucket, and invalidate only its exact distribution. It cannot assume the shared CDK bootstrap roles or modify infrastructure. Staging infrastructure changes use the Administrator command below.
- Post-deploy checks: bucket + distribution existence, HTTPS on
https://staging.seahaven.com, and the actual post-invalidation remote assets contain the staging API URL and no dev API URL. (Not browser QA.)
One-time setup (run by a human with admin AWS creds + GitHub Admin)
-
GitHub Admin — create the
stagingenvironment (Settings → Environments → New environment →staging). Add protection rules as appropriate (e.g. required reviewers, restrict to thestagingbranch). If the environment does not exist, GitHub creates it unprotected on first use. -
AWS Admin — first deploy with admin creds (same steps 1–3 as dev; the OIDC provider and bootstrap already exist in this account):
cd infra/cdk npx cdk deploy shoc-frontend-staging \ -c envName=staging \ -c deployBranch=staging \ -c githubEnvironment=staging \ -c domainNames=staging.seahaven.com \ -c certificateArn=arn:aws:acm:us-east-1:396287094661:certificate/2b78e74f-7b65-4b82-a413-7a498b102f00 \ -c hostedZoneId=Z02602739VQWBWCAGXP4 \ -c hostedZoneName=staging.seahaven.comThe
DeployRoleArnoutput must match the ARN pinned indeploy-staging.yml(it will — the role name is deterministic). -
Backend CORS: the staging API (
https://api.staging.seahaven.com) must allow thehttps://staging.seahaven.comorigin. -
Push to
staging—ci.yamlruns the quality gates anddeploy-staging.ymldeploys.
Adding prod later
Same pattern: a prod account/stack with its own contexts and, ideally, its own
githubEnvironment=prod trust + workflow. Keep in mind VITE_API_URL is baked
into each environment's build, and the bucket's RemovalPolicy.DESTROY +
autoDeleteObjects defaults are dev/staging-friendly but should be revisited
for prod.
Notes
- Teardown:
npx cdk destroy. The bucket usesRemovalPolicy.DESTROY+autoDeleteObjects(dev artifacts are reproducible) — change this for prod. Never run it against dev during or after the Terraform adoption: the adoption-mode stack retains the transferred resources, and after Phase 2 Terraform owns them. - CI and staging CD both fire on push to
stagingin parallel; the staging CD workflow runsnpm run verifyitself before deploying. Dev has no push-triggered deploy during the adoption. - npm is pinned to v11.16.0; the committed
package-lock.jsonuses lockfileVersion 3, matching the Node 24 / npm 11 CI environment.