engineering-handbook/cicd.md

377 lines
20 KiB
Markdown
Raw Normal View History

# CI/CD Pipelines
## Requirement
Every deployable repo must have a CI/CD pipeline. No manual deploys to production. If it deploys to AWS, it needs a pipeline.
## Two lanes
| Lane | Who applies infra | Who ships app content | When |
|---|---|---|---|
| **HCP Terraform** | HCP VCS, auto-apply | Per-repo `deploy-*.yaml` | Default for every new deployable |
| **SAM / CDK** | GitHub `cd-sam` / `cd-cdk` reusables | The same workflow | Remaining stacks until migrated |
GitHub Actions is the CI/CD platform in both lanes. Infra details for the default lane are in [hcp-terraform.md](hcp-terraform.md).
HCP app repos call org reusables `cd-hcp-fargate.yaml` and `cd-hcp-spa.yaml` with one caller job per GitHub Environment. Remaining SAM/CDK stacks keep `cd-sam` / `cd-cdk`. Sequential `ci-typescript-frontend.yaml` stays for remaining-lane templates until they migrate.
## HCP Terraform lane (default)
Terraform owns infrastructure and never touches application content. GitHub Actions owns application content and never creates HCP runs.
### Workflows
One workflow per deployable, kebab-case, in `.github/workflows/`. Examples: `deploy-web.yaml`, `deploy-api.yaml`. A single-deployable repo may use `deploy.yaml`.
Each Environment is its own caller job. `environment` is a `with:` input. The reusable job owns `environment:`, concurrency, OIDC, and `vars.DEPLOY_ROLE_ARN`. GitHub rejects `environment:` beside `uses:`.
| Trigger | Caller job | Notes |
|---|---|---|
| `push` to `main` | `deploy-dev` | `paths-ignore` for `terraform/**`, docs, and the other deployable's paths. Mixed app+terraform merges still deploy (GitHub skips only when **every** changed file matches the ignore list). |
| `release: published` | `deploy-prod` (or `deploy-staging`) | Human-cut GitHub Release. `ship-gate: true`. See [Releases](#releases) and [Hotfix ship path](#hotfix-ship-path). |
| `workflow_dispatch` with `environment` and `ref` | matching job | Redeploy or rollback. Empty `ref` means `github.sha`. |
Do **not** put `paths` / `paths-ignore` on tag events. A tag create often has an empty file diff, so the job never starts.
CI is a separate workflow. Converted repos trigger on `pull_request` to `main`, `hotfix/**`, and `release/**`; `merge_group`; and `push` to `hotfix/**` and `release/**`. No `push` CI on `main`. The required check is `ci-complete`. Unconverted remaining-lane repos still emit `ci / ci`.
### Caller shape
```yaml
jobs:
deploy-dev:
name: Deploy API to dev
if: github.event_name == 'push' || (github.event_name == 'workflow_dispatch' && inputs.environment == 'dev')
uses: Sea-Haven-Industries/.github/.github/workflows/cd-hcp-fargate.yaml@<sha> # vX.Y.Z
permissions: { contents: read, id-token: write }
secrets: inherit
with:
environment: dev
ref: ${{ inputs.ref }}
ssm-prefix: /<repo>/deploy
docker-platform: linux/amd64
deploy-prod:
name: Deploy API to prod
if: github.event_name == 'release' || (github.event_name == 'workflow_dispatch' && inputs.environment == 'prod')
uses: Sea-Haven-Industries/.github/.github/workflows/cd-hcp-fargate.yaml@<sha> # vX.Y.Z
permissions: { contents: read, id-token: write }
secrets: inherit
with:
environment: prod
ref: ${{ github.event.release.tag_name || inputs.ref }}
ssm-prefix: /<repo>/deploy
docker-platform: linux/amd64
ship-gate: true
```
SPA callers use `cd-hcp-spa.yaml`. Add `deploy-staging` only where that Environment exists.
### Ship-gate
`ship-gate: true` on prod and staging replaces the copied `target` job. It is legal when:
1. `compare/main...<tag>` is `behind` or `identical`, or
2. The tag is a fast-forward of the previous matching Release tag, SemVer matches that Environment (`^v[0-9]+\.[0-9]+\.[0-9]+$` or `-staging`), and it was created from `hotfix/*` or `release/*`.
Anything else fails closed, so nothing un-reviewed ships. Environment-derived values (Sentry environment, stage) come from `inputs.environment`. Do not store them as per-environment variables that can silently be unset.
### Deploy job (inside the reusable)
- `environment: ${{ inputs.environment }}`
- `concurrency: deploy-${{ inputs.ssm-prefix }}-${{ inputs.environment }}` with `cancel-in-progress: false`
- Permissions: `id-token: write`, `contents: read`
- Assumes the GitHub Environment variable `DEPLOY_ROLE_ARN`
### Worked examples
**SPA / CloudFront.** Build, `s3 sync` hashed assets with immutable cache headers, `cp` `index.html` last with `no-store`, `s3 sync --delete` to prune, `CreateInvalidation /*`. Then verify live: distribution `Deployed`, served `index.html` hash equals the built hash, cache headers, hashed assets, forbidden URLs, and `/api/health` polled for up to five minutes when the SPA depends on an API.
**Fargate.** Docker build+push tagged `$sha` and `$environment`, patch `GIT_SHA` (and optional `extra-task-env`) on the task definition, `RegisterTaskDefinition` + `UpdateService` + `services-stable`, then poll SSM `api-url` until `/api/health` reports that SHA.
**Lambda zip.** esbuild (or equivalent) with `GIT_SHA` inlined at build time. The build constant wins over any runtime env var. Upload `functions/<name>/<sha>.zip` to the artifacts bucket, `update-function-code` on each function, then poll `/api/health` until it reports that exact SHA.
Verify against live state, not action success. Every check that depends on another deployable must poll, not probe once. Parallel deployables race each other on a fresh environment.
Reference: org reusables `cd-hcp-spa.yaml` and `cd-hcp-fargate.yaml`. Inline copies remain on unconverted callers until their cutover.
### Releases
Cut by a person:
```bash
gh release create vX.Y.Z --target main --generate-notes
```
Not a workflow. Releases created with `GITHUB_TOKEN` do not fire `release: published`. A workflow-cut release would apply prod infra (HCP sees the tag) without queuing the prod app deploys.
One tag drives both prod infra and prod app. Approving the app deploys after the HCP apply lands is the operator's sequencing responsibility. SPA origin-path guards and health polls catch the common misorderings.
### Hotfix ship path
HCP, Environment `v*`, and OIDC tag refs already allow a hotfix tag that is not yet on `main`. No IAM change for this path. `ship-gate: true` accepts it when the tag is a fast-forward of the previous matching Release, SemVer matches the Environment, and it was created from `hotfix/*` or `release/*`.
```bash
git fetch --tags
git checkout -b hotfix/describe-the-break v1.2.3
# commit, push (CI on hotfix/**) or PR targeting the hotfix branch
gh release create v1.2.4 --target hotfix/describe-the-break --generate-notes
# approve prod Environment after the HCP prod apply
# merge hotfix/describe-the-break into main
```
Do not add GitFlow release trains, cherry-pick bots, or auto merge-back. Merge the hotfix branch into `main` after prod.
### GitHub Environments
| Environment | Reviewers | Deployment branch policy | Variables |
|---|---|---|---|
| `dev` | none | `main` | `DEPLOY_ROLE_ARN` |
| `prod` | required | `main` (branch) and `v*` (tag) | `DEPLOY_ROLE_ARN` |
Prod must allow both the tag pattern and `main`. Omit `v*` and release-triggered runs never reach the gate. Omit `main` and dispatch-from-main rollbacks never reach the gate.
Staging is not a default Environment. See [hcp-terraform.md](hcp-terraform.md#accounts-and-workspaces).
### GitHub deploy role
IAM role `githubdeploy-<repo>`, owned by workload Terraform.
**Trust.** OIDC `sub` stays `repo:Sea-Haven-Industries/<repo>:environment:<env>`. SHA-pinned org reusables change `job_workflow_ref` to the reusable file:
- `job_workflow_ref`: `Sea-Haven-Industries/.github/.github/workflows/cd-hcp-fargate.yaml@*` (and `cd-hcp-spa.yaml`)
- `workflow_ref`: the thin caller still `Sea-Haven-Industries/<repo>/.github/workflows/deploy-api.yaml@refs/heads/main` and `@refs/tags/v*`
Adding an app still adds `workflow_ref` for that caller. Do not pin trust to only `ref:refs/heads/main`.
**Permissions.** Only what the workflows write:
- bucket-root `s3:PutObject` / `DeleteObject` / `ListBucket` on the web bucket
- `s3:PutObject` on the artifacts prefix
- `cloudfront:CreateInvalidation` and `GetDistribution` on the one distribution
- `lambda:UpdateFunctionCode` and `GetFunction` on the named functions
- `ssm:GetParameter` on `/<repo>/deploy/*`
- Fargate callers also need ECR push, `ecs:RegisterTaskDefinition` / `UpdateService` / `Describe*`, and `iam:PassRole` on the task roles
Nothing else.
`DEPLOY_ROLE_ARN` is a GitHub Environment **variable**, not a repo secret.
### Converted CI
Converted HCP callers use parallel portions plus a `ci-complete` aggregator. Autofix is a pull-request convenience that commits with a GitHub App token. `format:check` / lint in the portions stay the fail-closed gate.
- Autofix runs on `pull_request` only, skips forks, and skips when the actor is the App. If the tree is dirty it commits `style: apply formatter` and sets `committed=true` so this SHA skips build/test. The `synchronize` run must be green. Do not `--no-verify`. Do not push to `main`.
- `merge_group` and `push` skip autofix (`result == skipped`) so those paths still run portions.
- `always()` on later jobs is required so an autofix failure does not skip `static`.
- Org secrets: `AUTOFMT_APP_ID`, `AUTOFMT_APP_PRIVATE_KEY`. The App has `contents: write` and `metadata: read`. It is not on the main-branch ruleset bypass list.
The org ruleset **CI complete** requires the check-run name `ci-complete`. Do not put portion names (`frontend / static`, `unit (1)`, …) in a ruleset. Unconverted repos stay on **main branch protection** requiring `ci / ci`. A repo is on exactly one of those rulesets. Flip membership in the same window as the workflow merge. Do not remove or retarget native GitHub merge-queue rulesets.
Python callers pass `format-command: ruff format .` and `lint-fix-command: ruff check --fix .`. Frontend passes npm scripts. Do not run `eslint --fix` unless that repo's `lint` script is already fix-safe.
## SAM / CDK lane (remaining)
Existing SAM and CDK stacks keep thin callers of org reusables until they migrate. Do not start a new deployable on this lane. Sequential `ci-typescript-frontend.yaml` remains for remaining-lane SPA templates until those repos migrate onto `ci-frontend.yaml` plus `ci-complete`.
Every remaining repo still has:
| File | Trigger | Purpose |
|---|---|---|
| `ci.yaml` | `pull_request` on `main` | Lint, typecheck, test, synth/validate |
| `deploy.yaml` | `push` on `main` | Deploy to AWS |
### CDK stacks (TypeScript)
```yaml
# .github/workflows/ci.yaml
name: CI
on:
pull_request:
branches: [main]
jobs:
ci:
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@<full-commit-sha> # vX.Y.Z
with:
node-version: "24"
# .github/workflows/deploy.yaml
name: Deploy
on:
push:
branches: [main]
jobs:
deploy:
uses: Sea-Haven-Industries/.github/.github/workflows/cd-cdk.yaml@<full-commit-sha> # vX.Y.Z
with:
node-version: "24"
secrets:
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}
```
### SAM stacks (Python)
```yaml
# .github/workflows/ci.yaml
name: CI
on:
pull_request:
branches: [main]
jobs:
ci:
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-commit-sha> # vX.Y.Z
# .github/workflows/deploy.yaml
name: Deploy
on:
push:
branches: [main]
jobs:
deploy:
uses: Sea-Haven-Industries/.github/.github/workflows/cd-sam.yaml@<full-commit-sha> # vX.Y.Z
with:
stack-name: "your-stack-name"
cfn-role-arn: "arn:aws:iam::<account-id>:role/github-cfn-execution-role"
secrets:
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}
```
`cd-sam.yaml` takes two required inputs and one required secret. The two role ARNs are not interchangeable; they are different roles with different jobs:
| Name | Kind | Purpose |
|---|---|---|
| `stack-name` | input (`with:`) | The CloudFormation stack name |
| `cfn-role-arn` | input (`with:`) | The **CloudFormation execution role** the stack is deployed *as*. This is `github-cfn-execution-role` in the repo's target account. Substitute that account's ID for `<account-id>`; the account must have the deploy substrate provisioned before the first deploy. Do not point new repos at the management account. |
| `deploy-role-arn` | secret (`secrets:`) | The **OIDC role the workflow assumes**, from the repo's `AWS_DEPLOY_ROLE_ARN` secret (see [Remaining-lane authentication](#remaining-lane-authentication)) |
Passing `cfn-role-arn` under `secrets:` fails: it is an input, so the run errors on an unexpected secret *and* a missing required input.
The remaining inputs have defaults and are only needed when a repo differs from them: `region` (`us-east-1`), `sam-template` (`template.yaml`), and `python-version` (`3.12`). The optional `parameter-overrides` secret passes `Key=Value` pairs through to `sam deploy`.
### Remaining-lane authentication
Deploy workflows authenticate to AWS via OIDC (no long-lived credentials). Each remaining SAM/CDK repo needs:
1. An IAM role named `githubdeploy-<repo-name>` with:
- OIDC trust policy for `token.actions.githubusercontent.com`
- Subject condition: `repo:Sea-Haven-Industries/<repo>:ref:refs/heads/main`
- Inline policy allowing `sts:AssumeRole` on CDK/SAM bootstrap roles
2. A repo secret `AWS_DEPLOY_ROLE_ARN` containing the role ARN
New HCP repos do not use this trust or this secret. See [GitHub deploy role](#github-deploy-role).
## Node.js Version
Always pass `node-version: "24"` to reusable workflows. Local dev uses Node 24 / npm 11 which generates lockfileVersion 3. The workflow defaults match this, but be explicit to avoid drift.
## Concurrency
### `cancel-in-progress`: false for deploys, true for CI
Cancelling a deploy midway does not roll it back. It abandons the run wherever it happens to be, which can leave a CloudFormation stack mid-update, an Elastic Beanstalk environment mid-update, a half-uploaded SPA, or a Lambda on a stub. A superseded deploy therefore queues behind the running one instead of killing it.
CI runs produce no external side effects, so when a newer commit supersedes an older one the older run should be abandoned (`cancel-in-progress: true`).
### HCP lane
Declare concurrency inside the reusable deploy job:
```yaml
concurrency:
group: deploy-${{ inputs.ssm-prefix }}-${{ inputs.environment }}
cancel-in-progress: false
```
Independent deployables must not share an `ssm-prefix`. Unconverted inline callers still key `deploy-<name>-<env>` until they move.
### Remaining SAM / CDK reusables
Reusable workflows in the central `.github` repo declare their own `concurrency` group. Consumers do not have to add one.
The `concurrency` block sits on the job, not at workflow top level. A single run can contain several jobs that must not share a group: a reusable with multiple jobs needs each one keyed separately, and a caller repo can invoke the same reusable from several jobs in one run. Where a reusable has more than one job, `${{ github.job }}` is part of the key so those jobs do not serialise against each other.
GitHub evaluates a concurrency group within the repository that owns the run, and for a reusable workflow that is the caller's repository. Two different repos calling the same reusable never contend. A group only has to be unique *inside* one repo, which is what the literal workflow-name prefix (`cd-sam-`, `cd-cdk-`, and so on) provides: it stops two different reusables in the same repo from colliding.
Every deploy group is keyed on the inputs that name what is being deployed, not just on the workflow. Keying on the workflow alone would serialise deploys that are genuinely independent.
The groups as implemented:
```yaml
# cd-sam.yaml
group: cd-sam-${{ inputs.region }}-${{ inputs.stack-name }}
# cd-cdk.yaml
group: cd-cdk-${{ inputs.region }}-${{ inputs.stacks }}-${{ inputs.stack-name }}
# cd-dotnet-eb.yaml
group: cd-dotnet-eb-${{ inputs.eb-application }}-${{ inputs.eb-environment }}
# cd-mobile-ios.yaml
group: cd-mobile-ios-${{ inputs.working-directory }}-${{ inputs.fastlane-lane }}
```
`cd-cdk` keys on `stacks`, the stack *selector*, rather than on `stack-name` alone. `stack-name` is optional there, so a multi-account caller that passes only a selector would collapse every one of its jobs into a single group.
Every component of a key is an input that is either required or always defaults, so the group can never evaluate to a bare prefix: `region` defaults, `stacks` defaults to `--all`, and `working-directory` and `fastlane-lane` both default. A key built from an input that can be empty silently merges unrelated deploys into one group.
CI groups follow the same shape and add `${{ github.ref }}` so branches do not cancel each other:
```yaml
# ci-typescript-frontend.yaml
group: ci-typescript-frontend-${{ github.workflow }}-${{ github.ref }}-${{ inputs.working-directory }}
cancel-in-progress: true
```
## Naming
- All workflow files: kebab-case
- Reusable workflow references: pinned to a full 40-character commit SHA. Never a branch or tag ref.
## Workflow Ref Pinning
Reusable workflow references are pinned to a full commit SHA of the central `.github` repo. The trailing comment names the release tag when the central repo is release-backed, or `main` when no release exists:
```yaml
# Release-backed (standard — central .github repo cuts releases):
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-commit-sha> # vX.Y.Z
# No release yet (use main only when the central repo has no release tags):
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-commit-sha> # main
```
Branch refs are mutable: a compromised or bad commit on the central repo would flow instantly into every consumer's CI and deploy path. A SHA pin turns that same change into a reviewable Renovate PR instead. The `# vX.Y.Z` comment lets readers identify which release the pin corresponds to.
Per the pinning principle, pins are for reproducibility, not for freezing time. Keep them moving with a root `renovate.json` that extends `local>Sea-Haven-Industries/.github`. Do not add `dependabot.yml`.
When adding a caller workflow by hand, pin to the commit SHA corresponding to the latest release of the central repo (`gh api /repos/Sea-Haven-Industries/.github/commits/vX.Y.Z --jq .sha`) and annotate with `# vX.Y.Z`. The commits endpoint resolves both lightweight and annotated tags to the underlying commit. Let Renovate advance the pin from there. Use `gh api /repos/Sea-Haven-Industries/.github/commits/main --jq .sha` with a `# main` comment only when the central repo has no release tags.
## When to Add a Pipeline
- When creating a new deployable project — the pipeline is part of the initial setup, not a follow-up
- When working on an existing project that lacks one — flag it and add it as part of the current work
A project is not production-ready without CI/CD. New deployables use the HCP lane.
## PR Auto-Labeling
Pull requests are auto-labeled org-wide by a reusable workflow in `.github`. The label rules live once, centrally, inside the reusable workflow itself (written to the runner at execution time), so each repo needs only a short caller and **no per-repo `labeler.yml`**:
```yaml
# .github/workflows/labeler.yml, the per-repo caller
name: Labeler
on:
pull_request:
branches: [main]
permissions:
contents: read
pull-requests: write
issues: write
jobs:
label:
uses: Sea-Haven-Industries/.github/.github/workflows/callable-labeler.yaml@<full-commit-sha> # vX.Y.Z
```
- The trigger is plain `pull_request`, not `pull_request_target`: private repos take no fork PRs, so the lower-privilege event is sufficient and avoids the pwn-request surface. Because `pull_request` runs the workflow from the merge commit, the Labeler check appears on the PR that first adds the caller — an absent or failed check means a missing permission, not expected behaviour.
- The caller MUST grant all three permissions. Reusable-workflow permissions can only be downgraded from the caller, so omitting `issues: write` (needed to create labels that don't exist yet) or any other grant causes a silent `startup_failure`.
- Adding the caller is part of new-repo provisioning.