engineering-handbook/hcp-terraform.md
Adam Moussa affa1a64f8
Some checks failed
ci / ci / ci (push) Has been cancelled
docs(cd): separate terraform infra from github app deploys (#46)
Make HCP Terraform plus GitHub Actions content CD the default for new
workloads, and keep SAM/CDK documented as the remaining path.
2026-09-15 22:05:33 +00:00

5.5 KiB

HCP Terraform

Default IaC for every new deployable, including a one-Lambda repo. GitHub Actions owns application content. This page is the infrastructure half of that split. The application half is cicd.md. Layout is terraform-project-layout.md.

Reference implementation: internal-portal. Copy the contract, not portal-only mechanics (in-workspace hcptf-* roles, bootstrap-window credential swaps, workflow-cut releases).

Ownership

Terraform creates the containers: buckets, distributions, Lambda skeletons, API Gateway, IAM, secrets shells, and the deploy SSM contract. It never writes application content.

GitHub Actions builds and publishes that content. It never creates HCP runs.

The seam is a committed bootstrap stub plus lifecycle.ignore_changes on the attributes the deploy workflow writes. A post-deploy terraform plan must be empty. That is the acceptance test.

Accounts and workspaces

Default accounts are seahaven-dev and seahaven-prod. Every app repo gets two workspaces:

Workspace Account VCS trigger Auto-apply
<repo>-dev seahaven-dev prefix terraform/** on main on
<repo>-prod seahaven-prod tag regex ^v[0-9]+\.[0-9]+\.[0-9]+$ on, after the first apply

Both workspaces are VCS-connected to main, working directory terraform, remote execution, speculative plans on for PRs. Speculative plans cover both workspaces, so the prod plan is reviewed before merge.

Staging is not in the default. A repo that already has a real staging slot may add <repo>-staging and tag vX.Y.Z-staging as a variant. Do not add staging because it feels safer.

First prod apply

The first-ever prod apply runs with auto-apply off and a human confirm. Turn auto-apply on after that apply lands. Later tag applies are automatic.

VCS file triggers

Keep file triggers on. Docs-only commits must not apply prod.

  • Dev: trigger prefix terraform/** only. App-path merges correctly skip the dev apply. GitHub Actions ships the app.
  • Prod: tag trigger ^v[0-9]+\.[0-9]+\.[0-9]+$. Tag triggering has no path filter. An app-only release still queues a prod plan; when terraform/** did not change it is a no-op apply.

Do not add Lambda source, src/, functions/, bff/, or similar to HCP trigger prefixes. Terraform does not package application content. A source-only merge is supposed to skip HCP.

The older rule (PLAT-183) that HCP must watch Lambda source paths applied to stacks where Terraform still bundled code. That is the remaining SAM/CDK-era Terraform path, not the default. New workspaces follow this page.

Record the live prefixes or patterns in the app README next to the workspace name.

Exec roles

Per-workspace variables, never a project-level variable set:

  • TFC_AWS_APPLY_ROLE_ARN → hcptf-<repo>
  • TFC_AWS_PLAN_ROLE_ARN → hcptf-<repo>-plan

Never TFC_AWS_RUN_ROLE_ARN.

Those roles are created and mutated in seahaven-org-baseline CloudFormation. Workload Terraform must not declare aws_iam_role named hcptf-* and must not ship hcp_iam.tf.

Do not swap workspace credentials through hcptf-bootstrap to change those policies. Policy changes are org-baseline PRs.

The plan role needs a specific read for every resource type the apply role creates (example: Cognito GetUserPoolMfaConfig). Narrow to named actions. Get* wildcards trip CKV_AWS_107 and block the workstation pre-push hook.

githubdeploy-<repo> lives in the workload module. The apply role is supposed to mutate that role. See cicd.md.

Deploy contract (SSM)

Terraform writes non-sensitive names that workflows read. Do not hardcode bucket names, distribution IDs, or function names in YAML.

Default (two accounts), same names in each account:

/<repo>/deploy/bucket
/<repo>/deploy/distribution-id
/<repo>/deploy/artifacts-bucket
/<repo>/deploy/<function>-function-name

Write only the parameters that repo's deployables need. A SPA without CloudFront omits distribution-id. A Lambda-only API omits bucket.

One-account leftovers namespace as /<repo>/<env>/deploy/*. New repos do not share an account across environments.

The seam

Lambda (and any other content-bearing resource Terraform must create once):

  1. Point filename / s3_key at a committed stub under terraform/bootstrap/.
  2. Ignore the attributes the deploy workflow writes:
lifecycle {
  ignore_changes = [
    filename,
    s3_bucket,
    s3_key,
    s3_object_version,
    source_code_hash,
  ]
}

Without the stub Terraform has nothing to create. Without ignore_changes every apply reverts the app.

Ship the stub and ignore_changes in the same apply so the live function is not downgraded to the stub.

Build identity (GIT_SHA, Sentry release, reported version) is inlined at build time in GitHub Actions. Do not put it in Terraform environment variables. An apply must not be able to regress the reported version.

CloudFront needs no seam when the SPA lives at the bucket root and nothing on the distribution changes per deploy.

Elastic Beanstalk is the same contract with ignore_changes = [version_label], not a third template.

Mixed PRs

Infra and app changes may share a PR. Dev may deploy the app before the dev apply finishes. Deploys are idempotent. Re-run the job.

What not to copy from portal

Portal's hcptf-internal-portal roles live in the same workspace, so changing their inline policies required a bootstrap-window credential swap. That is a portal exception. Do not add it to a new repo. Do not document it as the operator ritual.