mirror of
https://github.com/Sea-Haven-Industries/engineering-handbook.git
synced 2026-09-30 10:23:13 +00:00
Some checks are pending
ci / ci / ci (push) Waiting to run
docs: align engineering conventions for Cursor migration (PLAT-62)
205 lines
10 KiB
Markdown
205 lines
10 KiB
Markdown
# CI/CD Pipelines
|
|
|
|
## Requirement
|
|
|
|
Every deployable repo must have a CI/CD pipeline. No manual deploys to production. If it deploys to AWS, it needs a pipeline.
|
|
|
|
## Platform
|
|
|
|
GitHub Actions is the standard CI/CD platform. All pipelines use reusable workflows from the `Sea-Haven-Industries/.github` org repo (`.github/workflows/`).
|
|
|
|
## Workflow Structure
|
|
|
|
Every repo gets two thin workflow files in `.github/workflows/`:
|
|
|
|
| File | Trigger | Purpose |
|
|
|---|---|---|
|
|
| `ci.yaml` | `pull_request` on `main` | Lint, typecheck, test, synth/validate |
|
|
| `deploy.yaml` | `push` on `main` | Deploy to AWS |
|
|
|
|
### CDK Stacks (TypeScript)
|
|
|
|
```yaml
|
|
# .github/workflows/ci.yaml
|
|
name: CI
|
|
on:
|
|
pull_request:
|
|
branches: [main]
|
|
jobs:
|
|
ci:
|
|
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@<full-commit-sha> # vX.Y.Z
|
|
with:
|
|
node-version: "24"
|
|
|
|
# .github/workflows/deploy.yaml
|
|
name: Deploy
|
|
on:
|
|
push:
|
|
branches: [main]
|
|
jobs:
|
|
deploy:
|
|
uses: Sea-Haven-Industries/.github/.github/workflows/cd-cdk.yaml@<full-commit-sha> # vX.Y.Z
|
|
with:
|
|
node-version: "24"
|
|
secrets:
|
|
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}
|
|
```
|
|
|
|
### SAM Stacks (Python)
|
|
|
|
```yaml
|
|
# .github/workflows/ci.yaml
|
|
name: CI
|
|
on:
|
|
pull_request:
|
|
branches: [main]
|
|
jobs:
|
|
ci:
|
|
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-commit-sha> # vX.Y.Z
|
|
|
|
# .github/workflows/deploy.yaml
|
|
name: Deploy
|
|
on:
|
|
push:
|
|
branches: [main]
|
|
jobs:
|
|
deploy:
|
|
uses: Sea-Haven-Industries/.github/.github/workflows/cd-sam.yaml@<full-commit-sha> # vX.Y.Z
|
|
with:
|
|
stack-name: "your-stack-name"
|
|
cfn-role-arn: "arn:aws:iam::<account-id>:role/github-cfn-execution-role"
|
|
secrets:
|
|
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}
|
|
```
|
|
|
|
`cd-sam.yaml` takes two required inputs and one required secret. The two role ARNs are not interchangeable; they are different roles with different jobs:
|
|
|
|
| Name | Kind | Purpose |
|
|
|---|---|---|
|
|
| `stack-name` | input (`with:`) | The CloudFormation stack name |
|
|
| `cfn-role-arn` | input (`with:`) | The **CloudFormation execution role** the stack is deployed *as*. This is `github-cfn-execution-role` in the repo's target account. Substitute that account's ID for `<account-id>`; the account must have the deploy substrate provisioned before the first deploy. Do not point new repos at the management account. |
|
|
| `deploy-role-arn` | secret (`secrets:`) | The **OIDC role the workflow assumes**, from the repo's `AWS_DEPLOY_ROLE_ARN` secret (see [Authentication](#authentication)) |
|
|
|
|
Passing `cfn-role-arn` under `secrets:` fails: it is an input, so the run errors on an unexpected secret *and* a missing required input.
|
|
|
|
The remaining inputs have defaults and are only needed when a repo differs from them: `region` (`us-east-1`), `sam-template` (`template.yaml`), and `python-version` (`3.12`). The optional `parameter-overrides` secret passes `Key=Value` pairs through to `sam deploy`.
|
|
|
|
## Authentication
|
|
|
|
Deploy workflows authenticate to AWS via OIDC (no long-lived credentials). Each repo needs:
|
|
|
|
1. An IAM role named `githubdeploy-<repo-name>` with:
|
|
- OIDC trust policy for `token.actions.githubusercontent.com`
|
|
- Subject condition: `repo:Sea-Haven-Industries/<repo>:ref:refs/heads/main`
|
|
- Inline policy allowing `sts:AssumeRole` on CDK/SAM bootstrap roles
|
|
2. A repo secret `AWS_DEPLOY_ROLE_ARN` containing the role ARN
|
|
|
|
## Node.js Version
|
|
|
|
Always pass `node-version: "24"` to reusable workflows. Local dev uses Node 24 / npm 11 which generates lockfileVersion 3. The workflow defaults match this, but be explicit to avoid drift.
|
|
|
|
## Concurrency
|
|
|
|
Reusable workflows in the central `.github` repo declare their own `concurrency` group. Consumers do not have to add one.
|
|
|
|
### `cancel-in-progress`: false for deploys, true for CI
|
|
|
|
Deploy reusables set `cancel-in-progress: false`. Cancelling a deploy midway does not roll it back. It abandons the run wherever it happens to be, which can leave a CloudFormation stack mid-update, an Elastic Beanstalk environment mid-update, or a half-uploaded store build. A superseded deploy therefore queues behind the running one instead of killing it.
|
|
|
|
CI reusables set `cancel-in-progress: true`. A CI run produces no external side effects, so when a newer commit supersedes an older one there is nothing to protect and the older run should be abandoned.
|
|
|
|
### Groups are declared at job level
|
|
|
|
The `concurrency` block sits on the job, not at workflow top level. A single run can contain several jobs that must not share a group: a reusable with multiple jobs needs each one keyed separately, and a caller repo can invoke the same reusable from several jobs in one run. Where a reusable has more than one job, `${{ github.job }}` is part of the key so those jobs do not serialise against each other.
|
|
|
|
### Groups are scoped per repository
|
|
|
|
GitHub evaluates a concurrency group within the repository that owns the run, and for a reusable workflow that is the caller's repository. Two different repos calling the same reusable never contend. A group only has to be unique *inside* one repo, which is what the literal workflow-name prefix (`cd-sam-`, `cd-cdk-`, and so on) provides: it stops two different reusables in the same repo from colliding.
|
|
|
|
### The key must identify the deploy target
|
|
|
|
Every deploy group is keyed on the inputs that name what is being deployed, not just on the workflow. Keying on the workflow alone would serialise deploys that are genuinely independent — a repo that calls one reusable from several jobs, one per AWS account, would deploy those accounts one at a time for no reason.
|
|
|
|
The groups as implemented:
|
|
|
|
```yaml
|
|
# cd-sam.yaml
|
|
group: cd-sam-${{ inputs.region }}-${{ inputs.stack-name }}
|
|
|
|
# cd-cdk.yaml
|
|
group: cd-cdk-${{ inputs.region }}-${{ inputs.stacks }}-${{ inputs.stack-name }}
|
|
|
|
# cd-dotnet-eb.yaml
|
|
group: cd-dotnet-eb-${{ inputs.eb-application }}-${{ inputs.eb-environment }}
|
|
|
|
# cd-mobile-ios.yaml
|
|
group: cd-mobile-ios-${{ inputs.working-directory }}-${{ inputs.fastlane-lane }}
|
|
```
|
|
|
|
`cd-cdk` keys on `stacks`, the stack *selector*, rather than on `stack-name` alone. `stack-name` is optional there, so a multi-account caller that passes only a selector would collapse every one of its jobs into a single group.
|
|
|
|
Every component of a key is an input that is either required or always defaults, so the group can never evaluate to a bare prefix: `region` defaults, `stacks` defaults to `--all`, and `working-directory` and `fastlane-lane` both default. A key built from an input that can be empty silently merges unrelated deploys into one group.
|
|
|
|
CI groups follow the same shape and add `${{ github.ref }}` so branches do not cancel each other:
|
|
|
|
```yaml
|
|
# ci-typescript-frontend.yaml
|
|
group: ci-typescript-frontend-${{ github.workflow }}-${{ github.ref }}-${{ inputs.working-directory }}
|
|
cancel-in-progress: true
|
|
```
|
|
|
|
## Naming
|
|
|
|
- All workflow files: kebab-case
|
|
- Reusable workflow references: pinned to a full 40-character commit SHA — never a branch or tag ref
|
|
|
|
## Workflow Ref Pinning
|
|
|
|
Reusable workflow references are pinned to a full commit SHA of the central `.github` repo. The trailing comment names the release tag when the central repo is release-backed, or `main` when no release exists:
|
|
|
|
```yaml
|
|
# Release-backed (standard — central .github repo cuts releases):
|
|
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-commit-sha> # vX.Y.Z
|
|
|
|
# No release yet (use main only when the central repo has no release tags):
|
|
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-commit-sha> # main
|
|
```
|
|
|
|
Branch refs are mutable: a compromised or bad commit on the central repo would flow instantly into every consumer's CI and deploy path. A SHA pin turns that same change into a reviewable Dependabot PR instead. The `# vX.Y.Z` comment lets Dependabot embed changelog information and lets readers identify which release the pin corresponds to.
|
|
|
|
Per the pinning principle, pins are for reproducibility, not for freezing time. Two prerequisites keep them moving:
|
|
|
|
1. Every repo's `dependabot.yml` must include the `github-actions` ecosystem (weekly), so pin-advance PRs are opened automatically.
|
|
2. Dependabot must be granted access to the internal `.github` repo at the org level (Org Settings → Advanced Security → Global settings → "Grant Dependabot access to repositories"). Without the grant, update jobs fail with `git_dependencies_not_reachable` and pins freeze silently — consumers stop receiving central workflow fixes with no visible signal beyond the failed Dependabot run.
|
|
|
|
When adding a caller workflow by hand, pin to the commit SHA corresponding to the latest release of the central repo (`gh api /repos/Sea-Haven-Industries/.github/commits/vX.Y.Z --jq .sha`) and annotate with `# vX.Y.Z`. The commits endpoint resolves both lightweight and annotated tags to the underlying commit. Let Dependabot advance the pin from there. Use `gh api /repos/Sea-Haven-Industries/.github/commits/main --jq .sha` with a `# main` comment only when the central repo has no release tags.
|
|
|
|
## When to Add a Pipeline
|
|
|
|
- When creating a new deployable project — the pipeline is part of the initial setup, not a follow-up
|
|
- When working on an existing project that lacks one — flag it and add it as part of the current work
|
|
|
|
A project is not production-ready without CI/CD.
|
|
|
|
## PR Auto-Labeling
|
|
|
|
Pull requests are auto-labeled org-wide by a reusable workflow in `.github`. The label rules live once, centrally, inside the reusable workflow itself (written to the runner at execution time), so each repo needs only a short caller and **no per-repo `labeler.yml`**:
|
|
|
|
```yaml
|
|
# .github/workflows/labeler.yml — the per-repo caller
|
|
name: Labeler
|
|
on:
|
|
pull_request:
|
|
branches: [main]
|
|
permissions:
|
|
contents: read
|
|
pull-requests: write
|
|
issues: write
|
|
jobs:
|
|
label:
|
|
uses: Sea-Haven-Industries/.github/.github/workflows/callable-labeler.yaml@<full-commit-sha> # vX.Y.Z
|
|
```
|
|
|
|
- The trigger is plain `pull_request`, not `pull_request_target`: private repos take no fork PRs, so the lower-privilege event is sufficient and avoids the pwn-request surface. Because `pull_request` runs the workflow from the merge commit, the Labeler check appears on the PR that first adds the caller — an absent or failed check means a missing permission, not expected behaviour.
|
|
- The caller MUST grant all three permissions. Reusable-workflow permissions can only be downgraded from the caller, so omitting `issues: write` (needed to create labels that don't exist yet) or any other grant causes a silent `startup_failure`.
|
|
- Adding the caller is part of new-repo provisioning.
|