shoc-frontend-new/terraform
2026-09-10 19:37:27 -04:00
..
live fix(terraform): lock provider hashes for linux and darwin platforms 2026-09-10 19:34:10 -04:00
README.md ci: run the Terraform isolation gate as a job in the CI workflow 2026-09-10 19:37:27 -04:00

Frontend Terraform adoption runbook (dev)

This tree adopts the existing Sea Haven SHOC frontend dev hosting resources into HCP Terraform without recreating them. It mirrors the backend adoption (shoc-backend #94, #98, #99, #102) and lands in three PRs:

PR Branch Change
A feature/frontend-terraform-adoption This PR. Dev root with adoption_complete = false, import guard, CDK retain mode, push-to-dev deploy off.
B feature/terraform-dev-adoption adoption_complete = true: ownership tags, bucket policy drops the auto-delete grant, CloudFormation detaches.
C feature/terraform-dev-content-cd Content CD through Terraform: release prefixes, pointer object, origin group, invalidation action, rollback.

Creating these files, formatting them, initializing with -backend=false, and validating them does not authorize an AWS, HCP Terraform, GitHub, CloudFormation, DNS, or deployment mutation. Every live step below is gated on an explicit go from the owner, with the production impact stated first.

Staging stays on the CDK and deploy-staging.yml path. Its cutover is tracked separately (SH-287) and adds its own root under live/staging when it starts. The staging constants in scripts/terraform_import_plan_resources.py exist only so the checker can prove a dev plan carrying a staging identifier fails.

Fixed targets

  • AWS account: 396287094661
  • AWS region: us-east-1
  • HCP organization: seahaven
  • HCP project: seahaven-external-dev
  • HCP workspace: shoc-frontend-new-dev, VCS branch dev, working directory terraform/live/dev
  • Site: dev.seahaven.com
  • API build value: https://api.dev.seahaven.com/api

Workspace invariants

Set before any Terraform lands on dev, read back after setting, and re-read before the first release after any Terraform merge:

  • Auto-apply off. GitHub or a human applies every run.
  • Automatic speculative plans on (PR plans are read-only evidence).
  • Automatic run triggering: patterns terraform/live/dev/** and terraform/live/modules/**. No trigger prefixes, no tags regex. Do not switch to tag-based triggering.
  • Execution mode remote, Terraform 1.16.x (versions.tf requires >= 1.9.0, < 2.0.0; CI validates with 1.16.0).
  • Dynamic AWS credentials only: environment variables TFC_AWS_PROVIDER_AUTH=true, TFC_AWS_PLAN_ROLE_ARN, and TFC_AWS_APPLY_ROLE_ARN pointing at the seahaven-org-baseline roles hcptf-shoc-frontend-new-dev-plan and hcptf-shoc-frontend-new-dev. No access keys.
  • No adoption_complete workspace variable. The dev root pins it in code (local.adoption_complete) so the value under review is the value that applies. scripts/test-terraform-import-plan-check.py fails if a variable block reappears in the root.

Ownership boundary

live/modules/environment-owned owns exactly these 13 addresses:

  1. module.environment_owned.aws_s3_bucket.site
  2. module.environment_owned.aws_s3_bucket_public_access_block.site
  3. module.environment_owned.aws_s3_bucket_ownership_controls.site
  4. module.environment_owned.aws_s3_bucket_server_side_encryption_configuration.site
  5. module.environment_owned.aws_s3_bucket_versioning.site
  6. module.environment_owned.aws_s3_bucket_policy.site
  7. module.environment_owned.aws_cloudfront_distribution.site
  8. module.environment_owned.aws_cloudfront_origin_access_control.site
  9. module.environment_owned.aws_cloudfront_function.spa_rewrite
  10. module.environment_owned.aws_route53_record.site_a
  11. module.environment_owned.aws_route53_record.site_aaaa
  12. module.environment_owned.aws_iam_role.github_deploy
  13. module.environment_owned.aws_iam_role_policy.github_deploy

Every managed resource has prevent_destroy = true.

live/modules/environment-inventory is data-only. It resolves and checks the caller account, provider region, public hosted zone, ACM certificate, account GitHub OIDC provider, and the AWS managed Managed-CachingOptimized cache policy against pinned values, and fails the plan on any mismatch.

The following remain outside state:

  • the dev.seahaven.com hosted zone and the *.seahaven.com certificate
  • the account-global GitHub OIDC provider
  • the AWS managed CloudFront cache policy
  • CDKToolkit resources and CDK metadata
  • the S3 auto-delete custom resource, its provider Lambda and role
  • the HCP plan/apply roles and the deploy-role permissions boundary (seahaven-org-baseline owns them)

Exact live inventory (dev)

  • Bucket and all bucket subresources: seahaven-shoc-frontend-dev
  • Distribution: E2CWLM1AFB964P
  • OAC: E30VSIK87N8H64, name shocfrontenddevDistributionOrigin1S3OriginAccessControlDFC82620, description modeled as ""
  • Distribution origin ID: shocfrontenddevDistributionOrigin10CCD0EE1
  • Function: us-east-1shocfrontenddevSpaRewrite58674DB8
  • A import ID: Z07671212N75U4YLPWZR8_dev.seahaven.com_A
  • AAAA import ID: Z07671212N75U4YLPWZR8_dev.seahaven.com_AAAA
  • Deploy role: githubdeploy-shoc-frontend-new-dev
  • Inline policy import ID: githubdeploy-shoc-frontend-new-dev:GithubDeployRoleDefaultPolicyE8F540D1
  • Hosted zone: Z07671212N75U4YLPWZR8
  • Certificate: arn:aws:acm:us-east-1:396287094661:certificate/2b78e74f-7b65-4b82-a413-7a498b102f00
  • Legacy stack: shoc-frontend-dev
  • Auto-delete helper role: arn:aws:iam::396287094661:role/shoc-frontend-dev-CustomS3AutoDeleteObjectsCustomRe-dmSDIY8EH7KV
  • Permissions boundary: arn:aws:iam::396287094661:policy/shoc-frontend-new-dev-deploy-boundary

With adoption_complete = false the root declares the configuration observed after the CDK retain deploy (Phase 1, step 2), not the configuration live today:

  • Environment=dev, ManagedBy=cdk, Project=shoc-frontend tags, plus the S3-only aws-cdk:auto-delete-objects=true tag
  • the deploy-role-only HcpTerraformWorkspace=shoc-frontend-new-dev tag
  • the permissions boundary attached to the deploy role
  • StringEquals on the OIDC subject repo:Sea-Haven-Industries/shoc-frontend-new:ref:refs/heads/dev
  • the legacy bucket policy including the auto-delete helper grant
  • the legacy deploy inline policy (AssumeCdkBootstrapRoles, DescribeStack, bucket read/write, InvalidateDistribution)

The retain deploy adds the boundary, the tag, and the StringEquals narrowing. If read-back after that deploy differs from the root in any other way, update the root to the observed value and prove a zero-change import plan. Do not approve drift through the controlled-update checker.

Phase 1: import-first adoption (this PR)

Each step is gated. State the impact, get the go, act, read back, record.

  1. Workspace invariants. Set the invariants above on shoc-frontend-new-dev. Read back the workspace and record the JSON in the PR.

  2. CDK retain deploy. From the reviewed PR head, with administrator credentials:

    cd infra/cdk && npm ci
    npx cdk deploy shoc-frontend-dev \
      -c retainForTerraformAdoption=true \
      --parameters ManageSiteInfrastructure=true
    

    Expected: an update-only change set (no create, no delete, no replace) that adds DeletionPolicy: Retain and UpdateReplacePolicy: Retain to the 13 transferred resources and the Custom::S3AutoDeleteObjects resource, attaches the boundary, adds the HcpTerraformWorkspace tag, and narrows the trust operator. Read back the role, bucket policy, and stack resources as JSON and attach it to the PR.

  3. Merge PR A. The merge triggers a VCS run on the workspace (auto-apply off). Download the plan JSON and run the guard:

    python3 scripts/check-terraform-import-plan.py plan.json --environment dev
    

    Confirm the apply only when the plan is exactly 13 imports, 0 create, 0 update, 0 delete, 0 replace and the guard exits 0. Otherwise discard the run and fix the root in a new PR.

  4. Post-import no-op. Queue a plan and require it to be no-op:

    python3 scripts/check-terraform-import-plan.py post-import.json \
      --environment dev --post-import-no-op
    

    Post the run URLs and the guard output on SH-300.

After Phase 1 CloudFormation still owns every resource. Terraform holds state for them and nothing else.

Phase 2: controlled ownership transfer (PR B)

PR B pins adoption_complete = true. The controlled apply may update only:

  • module.environment_owned.aws_s3_bucket.site (tags)
  • module.environment_owned.aws_s3_bucket_policy.site (drops only the auto-delete helper grant)
  • module.environment_owned.aws_cloudfront_distribution.site (tags)
  • module.environment_owned.aws_cloudfront_function.spa_rewrite (tags)
  • module.environment_owned.aws_iam_role.github_deploy (tags)

The OAC, both Route 53 records, and the deploy inline policy must be no-op. PR B keeps the post-adoption inline policy byte-identical to live so the policy address does not appear in the plan. Run the checker with one --allow-update-address per updating address; it rejects unused allowlist entries, unknown values, and replacements.

Dependency: hcptf-shoc-frontend-new-dev currently lacks cloudfront:UpdateDistribution and cloudfront:UpdateFunction. Codify the expansion in seahaven-org-baseline (cross-family plus security review) and deploy it before the controlled apply.

After the apply and a no-op plan, deploy the same reviewed CDK SHA with --parameters ManageSiteInfrastructure=false. Expect DELETE_SKIPPED on the 13 transferred resources and the custom resource. Never deploy with ManageSiteInfrastructure=true again after that. See infra/cdk/README.md.

Phase 3: content CD through Terraform (PR C)

Summary only; PR C carries the full design. GitHub builds and uploads to an immutable releases/<sha>-<run>-<attempt>/ prefix. Terraform owns the .release/current pointer, both origin paths of a CloudFront origin group, and the invalidation action. Rollback is one guarded Terraform run swapping the labels. Push-to-dev releases return behind the repository variable TERRAFORM_CONTENT_CD_ENABLED.

Operational rules

  • Terraform-only PRs. A PR that changes terraform/** may not change deployable application code. The terraform-isolation job in .github/workflows/ci.yaml enforces this; documentation and the scripts/*terraform* tooling are allowed alongside. A reviewer may add the terraform-isolation-override label for the rare change that must introduce Terraform variables together with the workflow that consumes them (PR A and PR C), then re-run the workflow so the job reads the label. The label is the approval record. The override is temporary: a follow-up PR after PR C removes the label path from the checker and workflow so the gate has no exception.
  • Every Terraform merge produces a VCS run. A human confirms or discards it before the next content release. Do not leave a pending run on the workspace.
  • Re-read the workspace invariants before the first release after any Terraform merge or workspace settings change.
  • A red job does not mean the site is down. Read the live state first (served index.html, distribution status, pointer body once PR C lands), then triage.
  • Exact-head evidence. Every live step records the run URL, the SHA, and a machine-readable read-back on the PR or SH-300.

Local validation

From the repository root (also run by npm run verify through scripts/governance-check.mjs):

npm run test:terraform              # fmt -check, init -backend=false, validate
npm run test:terraform-import-plan  # checker unit tests against synthetic plans
npm run test:terraform-isolation    # isolation gate unit tests
npm run test:infra                  # CDK build, template tests, synth in both modes

terraform init -backend=false -lockfile=readonly may download the provider but never contacts HCP state or plans against AWS. Only HCP runs plan against the account.

The lock file must carry h1: hashes for every platform that runs the gate (CI and HCP are linux_amd64, laptops are darwin_*). After changing the provider version, refresh them with:

terraform -chdir=terraform/live/dev providers lock \
  -platform=linux_amd64 -platform=linux_arm64 \
  -platform=darwin_amd64 -platform=darwin_arm64

Import plan safety

Import mode requires exactly the canonical 13 addresses and AWS types, valid import metadata for every resource, the exact dev import IDs (a staging ID in a dev plan fails), and zero create, update, delete, or replace actions.

Post-import mode requires all 13 resources to be no-op and rejects any remaining import metadata.

Controlled mode permits only in-place updates to the addresses explicitly listed with --allow-update-address, verifies before against the exact pre-adoption policies and tags and after against the exact adopted values, and rejects create, delete, replace, import metadata, unknown values, unapproved addresses, and unused allowlist entries.

Rollback

  • Before import apply: discard the run and correct the root.
  • After import, before the controlled update (end of Phase 1): remove only the 13 imported addresses from state under a separately reviewed state operation. CloudFormation remains authoritative; a ManageSiteInfrastructure=true stack is unchanged by this.
  • After the controlled update, before detachment: either complete the reviewed detachment or restore the exact pre-adoption policy and tags under a separate approval. Do not remove state or redeploy CloudFormation blindly.
  • After detachment: Terraform is authoritative. Restore content from the versioned bucket. Re-establishing CloudFormation ownership requires a reviewed IMPORT change set, never an ordinary update.

Any replacement, destroy, cross-environment ID, missing import, broad policy change, or failed smoke check is a hard stop.

Evidence per phase

  • HCP run URL and the workspace settings read-back
  • plan JSON and checker output
  • terraform state list showing exactly the 13 addresses
  • read-only inventory before and after each mutation
  • synthesized CloudFormation template, change set, and stack events
  • deploy, invalidation, and smoke output
  • the post-action no-op plan
  • phase close-out on SH-300: completed work, validation, risks, deviations, remaining work