docs(cd): separate terraform infra from github app deploys

Make HCP Terraform plus GitHub Actions content CD the default for new
workloads, and keep SAM/CDK documented as the remaining path.
This commit is contained in:
Adam Moussa 2026-09-15 17:58:09 -04:00
parent aea86c0f5c
commit 85de977098
No known key found for this signature in database
15 changed files with 408 additions and 71 deletions

View file

@ -11,18 +11,20 @@ Engineering conventions and best practices for Sea Haven Industries.
- [Naming Conventions](naming-conventions.md) -- kebab-case everywhere, no exceptions
- [Development Environment](dev-environment.md) -- workstation directory layout, pyenv, Node, launchd/TCC
- [Issue Tracking](issue-tracking.md) -- Jira projects, ticket description template, ticket hygiene
- [Git Workflow](git-workflow.md) -- feature branches, incremental commits, CI-on-merge
- [Git Workflow](git-workflow.md) -- feature branches, incremental commits, merge to main then release for prod
- [Commit Messages](commit-messages.md) -- Conventional Commits, type(scope) format, explain "why"
- [Pull Requests](pull-requests.md) -- scope, title, description format, merge strategy
- [Code Review](code-review.md) -- what to look for, giving feedback, turnaround expectations
- [Code Review Rubric](code-review-rubric.md) -- BLOCK/FIX/NIT/QUESTION finding categories and output format
- [GitHub Standards](github-standards.md) -- branch defaults, repo hygiene, Dependabot
- [AWS Infrastructure](aws-infrastructure.md) -- SAM vs CDK, HCP VCS file triggers, Lambda defaults, CloudFormation
- [SAM Project Layout](sam-project-layout.md) -- standard directory structure for serverless projects
- [CDK Project Layout](cdk-project-layout.md) -- standard directory structure for CDK projects
- [Lambda Starter Template](lambda-template.md) -- minimal SAM scaffold for a new Python Lambda
- [GitHub Standards](github-standards.md) -- branch defaults, repo hygiene, Environments
- [AWS Infrastructure](aws-infrastructure.md) -- HCP Terraform default, remaining SAM/CDK, Lambda defaults
- [HCP Terraform](hcp-terraform.md) -- workspaces, seam, SSM deploy contract, exec roles
- [Terraform Project Layout](terraform-project-layout.md) -- standard directory structure for HCP app repos
- [SAM Project Layout](sam-project-layout.md) -- remaining path for existing SAM stacks
- [CDK Project Layout](cdk-project-layout.md) -- remaining path for CDK and Bedrock
- [Lambda Starter Template](lambda-template.md) -- remaining SAM scaffold for a Python Lambda
- [Secrets and Configuration](secrets-and-config.md) -- Secrets Manager vs SSM Parameter Store
- [CI/CD Pipelines](cicd.md) -- every deployable repo gets a pipeline, no manual deploys
- [CI/CD Pipelines](cicd.md) -- HCP lane default, remaining SAM/CDK reusables
- [Bedrock](bedrock.md) -- cross-region inference profiles, alias pinning, KB Docker requirement
- [Git Hooks](hooks/) -- shim for the global security pre-push hook
- [Scripts](scripts/) -- repo provisioning, automation tooling

View file

@ -2,21 +2,24 @@
## IaC Strategy
- **SAM** is the default for new serverless stacks (Lambda + API Gateway + DynamoDB)
- **CDK** only for complex infrastructure (ECS, VPCs, multi-service compositions)
- **Migrating stacks** from mgmt into seahaven-prod / seahaven-dev use **HCP Terraform** (migrate-and-convert; do not convert in place in mgmt). The authoritative per-stack steps live in the `seahaven-org-baseline` README migration checklist (plan/apply role patterns, boundary widen, artifact packaging, cutover). Reference implementation: `afi-backup-monitor` (PLAT-56).
- Every deployed resource should be managed by IaC (CloudFormation via SAM/CDK, or HCP Terraform state for migrated stacks)
- **HCP Terraform** is the default for every new deployable, including a one-Lambda repo. GitHub Actions publishes application content. See [hcp-terraform.md](hcp-terraform.md), [terraform-project-layout.md](terraform-project-layout.md), and [cicd.md](cicd.md).
- **CDK** for Bedrock Agents, Knowledge Bases, and other compositions SAM cannot express (ECS, VPCs). Do not start a new app workload on CDK in order to skip HCP.
- **SAM** remains documented for stacks that already use it. Do not start a new deployable on SAM.
- **Migrating stacks** from mgmt into `seahaven-prod` / `seahaven-dev` still use HCP Terraform (migrate-and-convert; do not convert in place in mgmt). Cut those over to the [hcp-terraform.md](hcp-terraform.md) seam (stub plus `ignore_changes`, GHA owns the zip) rather than keeping Terraform-packaged Lambda source. The org-baseline first-apply runbook still covers workspace create and `hcptf-*` roles. Reference implementation for a completed app cutover: `internal-portal`.
- Every deployed resource should be managed by IaC
- No manually-created Lambdas, roles, or other resources outside of IaC
## HCP Terraform VCS file triggers
New prod/dev HCP workspaces follow the first-apply runbook in `seahaven-org-baseline`. When the workspace is created, set VCS file triggers before the first merge to the tracked branch.
Canonical rules live in [hcp-terraform.md](hcp-terraform.md#vcs-file-triggers). New workspaces trigger on `terraform/**` only (dev) and on `vX.Y.Z` tags (prod). Do not add Lambda source trees to HCP triggers. GitHub Actions ships that content.
- Keep `file-triggers-enabled`. Do not turn file triggers off so that docs-only commits skip prod applies.
- `trigger-prefixes` are added to the working directory. `trigger-patterns` replace the working-directory filter and must include a glob for that directory.
- If the plan reads files outside the working directory (Lambda source, layers, `requirements.txt`, build scripts that copy `src/`, `functions/`, `lambda/`, or `lambdas/`), list those paths in `trigger-prefixes` or `trigger-patterns`. An empty pair with `working-directory=terraform` queues runs only for `terraform/` changes, so a source-only merge is ingested and then silently skipped.
- Record the live prefixes or patterns in the app README next to the workspace name.
- `python3 scripts/check_hcp_workspace_triggers.py` in `seahaven-org-baseline` lists file-triggered workspaces in `seahaven-mgmt`, `seahaven-prod`, and `seahaven-dev` and flags an empty prefix/pattern pair when the repo working directory references those source trees.
Keep `file-triggers-enabled`. Do not turn file triggers off.
`trigger-prefixes` are added to the working directory. `trigger-patterns` replace the working-directory filter and must include a glob for that directory.
Remaining stacks where Terraform still packages Lambda code (mgmt migrations not yet on the seam) still list those source paths in triggers, as PLAT-183 required. New repos do not. `python3 scripts/check_hcp_workspace_triggers.py` in `seahaven-org-baseline` still flags an empty prefix/pattern pair when the Terraform tree references application source. After a repo adopts the seam it should no longer reference that source, and the checker should be quiet.
Record the live prefixes or patterns in the app README next to the workspace name.
## Lambda Defaults

View file

@ -2,15 +2,14 @@
## When to Use CDK
SAM is the default for new serverless stacks (Lambda + API Gateway + DynamoDB). Use CDK when the infrastructure goes beyond what SAM handles cleanly:
HCP Terraform is the default for new app workloads. Use CDK when the infrastructure goes beyond what that layout handles cleanly:
- Bedrock Agents and Knowledge Bases (SAM lacks L2 constructs)
- ECS Fargate tasks
- VPCs, subnets, security groups
- Multi-service compositions (e.g., SES + Lambda + DynamoDB + Bedrock Agent)
- L2/L3 constructs not available in SAM (Bedrock agents, Knowledge Bases)
- Docker-based Lambda or container workloads
- Docker-based Lambda or container workloads that are not being cut over to HCP yet
If the project is a handful of Lambdas behind API Gateway, use SAM. See [sam-project-layout.md](sam-project-layout.md).
Do not start a new multi-environment app on CDK in order to skip HCP. If the project is a handful of Lambdas behind API Gateway, use HCP Terraform. Existing SAM stacks stay on SAM until migrated; see [sam-project-layout.md](sam-project-layout.md).
## Standard Directory Structure
@ -161,7 +160,7 @@ See [aws-infrastructure.md](aws-infrastructure.md#lambda-defaults), and [Node ru
## CI/CD
Same OIDC deploy role pattern as SAM. Each repo gets its own `githubdeploy-<repo-name>` IAM role -- never share deploy roles across repos.
Remaining SAM/CDK lane only. New app workloads use [cicd.md](cicd.md#hcp-terraform-lane-default). Each remaining repo still gets its own `githubdeploy-<repo-name>` IAM role. Never share deploy roles across repos.
CDK projects use the TypeScript CDK reusable workflows:
@ -191,7 +190,7 @@ jobs:
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}
```
Always pass `node-version: "24"` explicitly. Reusable workflow refs are pinned to a full 40-character commit SHA of the central `.github` repo with a trailing `# vX.Y.Z` comment (the release tag the SHA corresponds to), never a branch or tag ref; see [Workflow Ref Pinning](cicd.md#workflow-ref-pinning). Pin to the SHA for the latest release when adding a caller by hand and let Dependabot advance it. See [cicd.md](cicd.md) for the full pipeline convention.
Always pass `node-version: "24"` explicitly. Reusable workflow refs are pinned to a full 40-character commit SHA of the central `.github` repo with a trailing `# vX.Y.Z` comment (the release tag the SHA corresponds to), never a branch or tag ref; see [Workflow Ref Pinning](cicd.md#workflow-ref-pinning). Pin to the SHA for the latest release when adding a caller by hand and let Renovate advance it. See [cicd.md](cicd.md) for the full pipeline convention.
## Bedrock Agents

156
cicd.md
View file

@ -4,20 +4,113 @@
Every deployable repo must have a CI/CD pipeline. No manual deploys to production. If it deploys to AWS, it needs a pipeline.
## Platform
## Two lanes
GitHub Actions is the standard CI/CD platform. All pipelines use reusable workflows from the `Sea-Haven-Industries/.github` org repo (`.github/workflows/`).
| Lane | Who applies infra | Who ships app content | When |
|---|---|---|---|
| **HCP Terraform** | HCP VCS, auto-apply | Per-repo `deploy-*.yaml` | Default for every new deployable |
| **SAM / CDK** | GitHub `cd-sam` / `cd-cdk` reusables | The same workflow | Remaining stacks until migrated |
## Workflow Structure
GitHub Actions is the CI/CD platform in both lanes. Infra details for the default lane are in [hcp-terraform.md](hcp-terraform.md).
Every repo gets two thin workflow files in `.github/workflows/`:
Do not extract org reusable **deploy** workflows yet. Codify shared composites after a second repo copies the portal shape and the duplication is real. CI may still call org `ci-*` reusables.
## HCP Terraform lane (default)
Terraform owns infrastructure and never touches application content. GitHub Actions owns application content and never creates HCP runs.
### Workflows
One workflow per deployable, kebab-case, in `.github/workflows/`. Examples: `deploy-web.yaml`, `deploy-api.yaml`. A single-deployable repo may use `deploy.yaml`.
| Trigger | Environment | Notes |
|---|---|---|
| `push` to `main` | `dev` | `paths-ignore` for `terraform/**`, docs, and the other deployable's paths. Mixed app+terraform merges still deploy (GitHub skips only when **every** changed file matches the ignore list). |
| `release: published` | `prod` | Human-cut GitHub Release. See [Releases](#releases). |
| `workflow_dispatch` with `environment` and `ref` | chosen | Redeploy or rollback at any prior tag. |
Do **not** put `paths` / `paths-ignore` on tag events. A tag create often has an empty file diff, so the job never starts.
CI stays a separate workflow (`ci.yaml` on `pull_request` to `main`). Org `ci-*` reusables are allowed. The required check is still `ci / ci`.
### Target job
A `target` job resolves `environment` and `ref` from the event.
On `release: published` it verifies the tag is an ancestor of `main` via `compare/main...<tag>` (status `behind` or `identical`). Anything else fails closed, so nothing un-reviewed ships.
Environment-derived values (Sentry environment, stage) come from that output. Do not store them as per-environment variables that can silently be unset.
### Deploy job
- `environment: ${{ needs.target.outputs.environment }}`
- `concurrency: deploy-<name>-<env>` with `cancel-in-progress: false`
- Permissions: `id-token: write`, `contents: read`
- Assumes the GitHub Environment variable `DEPLOY_ROLE_ARN`
### Worked examples
**SPA / CloudFront.** Build, `s3 sync` hashed assets with immutable cache headers, `cp` `index.html` last with `no-store`, `s3 sync --delete` to prune, `CreateInvalidation /*`. Then verify live: distribution `Deployed`, served `index.html` hash equals the built hash, cache headers, hashed assets, forbidden URLs, and `/api/health` polled for up to five minutes when the SPA depends on an API.
**Lambda zip.** esbuild (or equivalent) with `GIT_SHA` inlined at build time. The build constant wins over any runtime env var. Upload `functions/<name>/<sha>.zip` to the artifacts bucket, `update-function-code` on each function, then poll `/api/health` until it reports that exact SHA.
Verify against live state, not action success. Every check that depends on another deployable must poll, not probe once. Parallel deployables race each other on a fresh environment.
Reference: `internal-portal` `deploy-web.yaml` and `deploy-api.yaml`.
### Releases
Cut by a person:
```bash
gh release create vX.Y.Z --target main --generate-notes
```
Not a workflow. Releases created with `GITHUB_TOKEN` do not fire `release: published`. A workflow-cut release would apply prod infra (HCP sees the tag) without queuing the prod app deploys.
One tag drives both prod infra and prod app. Approving the app deploys after the HCP apply lands is the operator's sequencing responsibility. SPA origin-path guards and health polls catch the common misorderings.
### GitHub Environments
| Environment | Reviewers | Deployment branch policy | Variables |
|---|---|---|---|
| `dev` | none | `main` | `DEPLOY_ROLE_ARN` |
| `prod` | required | `main` (branch) and `v*` (tag) | `DEPLOY_ROLE_ARN` |
Prod must allow both the tag pattern and `main`. Omit `v*` and release-triggered runs never reach the gate. Omit `main` and dispatch-from-main rollbacks never reach the gate.
Staging is not a default Environment. See [hcp-terraform.md](hcp-terraform.md#accounts-and-workspaces).
### GitHub deploy role
IAM role `githubdeploy-<repo>`, owned by workload Terraform.
**Trust.** OIDC `sub` pinned to `repo:Sea-Haven-Industries/<repo>:environment:<env>` plus `job_workflow_ref` for each deploy workflow at `refs/heads/main`. Adding a workflow means adding its ref here. That is a cross-family review change.
**Permissions.** Only what the workflows write:
- bucket-root `s3:PutObject` / `DeleteObject` / `ListBucket` on the web bucket
- `s3:PutObject` on the artifacts prefix
- `cloudfront:CreateInvalidation` and `GetDistribution` on the one distribution
- `lambda:UpdateFunctionCode` and `GetFunction` on the named functions
- `ssm:GetParameter` on `/<repo>/deploy/*`
Nothing else.
`DEPLOY_ROLE_ARN` is a GitHub Environment **variable**, not a repo secret.
## SAM / CDK lane (remaining)
Existing SAM and CDK stacks keep thin callers of org reusables until they migrate. Do not start a new deployable on this lane.
Every remaining repo still has:
| File | Trigger | Purpose |
|---|---|---|
| `ci.yaml` | `pull_request` on `main` | Lint, typecheck, test, synth/validate |
| `deploy.yaml` | `push` on `main` | Deploy to AWS |
### CDK Stacks (TypeScript)
### CDK stacks (TypeScript)
```yaml
# .github/workflows/ci.yaml
@ -45,7 +138,7 @@ jobs:
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}
```
### SAM Stacks (Python)
### SAM stacks (Python)
```yaml
# .github/workflows/ci.yaml
@ -78,15 +171,15 @@ jobs:
|---|---|---|
| `stack-name` | input (`with:`) | The CloudFormation stack name |
| `cfn-role-arn` | input (`with:`) | The **CloudFormation execution role** the stack is deployed *as*. This is `github-cfn-execution-role` in the repo's target account. Substitute that account's ID for `<account-id>`; the account must have the deploy substrate provisioned before the first deploy. Do not point new repos at the management account. |
| `deploy-role-arn` | secret (`secrets:`) | The **OIDC role the workflow assumes**, from the repo's `AWS_DEPLOY_ROLE_ARN` secret (see [Authentication](#authentication)) |
| `deploy-role-arn` | secret (`secrets:`) | The **OIDC role the workflow assumes**, from the repo's `AWS_DEPLOY_ROLE_ARN` secret (see [Remaining-lane authentication](#remaining-lane-authentication)) |
Passing `cfn-role-arn` under `secrets:` fails: it is an input, so the run errors on an unexpected secret *and* a missing required input.
The remaining inputs have defaults and are only needed when a repo differs from them: `region` (`us-east-1`), `sam-template` (`template.yaml`), and `python-version` (`3.12`). The optional `parameter-overrides` secret passes `Key=Value` pairs through to `sam deploy`.
## Authentication
### Remaining-lane authentication
Deploy workflows authenticate to AWS via OIDC (no long-lived credentials). Each repo needs:
Deploy workflows authenticate to AWS via OIDC (no long-lived credentials). Each remaining SAM/CDK repo needs:
1. An IAM role named `githubdeploy-<repo-name>` with:
- OIDC trust policy for `token.actions.githubusercontent.com`
@ -94,31 +187,41 @@ Deploy workflows authenticate to AWS via OIDC (no long-lived credentials). Each
- Inline policy allowing `sts:AssumeRole` on CDK/SAM bootstrap roles
2. A repo secret `AWS_DEPLOY_ROLE_ARN` containing the role ARN
New HCP repos do not use this trust or this secret. See [GitHub deploy role](#github-deploy-role).
## Node.js Version
Always pass `node-version: "24"` to reusable workflows. Local dev uses Node 24 / npm 11 which generates lockfileVersion 3. The workflow defaults match this, but be explicit to avoid drift.
## Concurrency
Reusable workflows in the central `.github` repo declare their own `concurrency` group. Consumers do not have to add one.
### `cancel-in-progress`: false for deploys, true for CI
Deploy reusables set `cancel-in-progress: false`. Cancelling a deploy midway does not roll it back. It abandons the run wherever it happens to be, which can leave a CloudFormation stack mid-update, an Elastic Beanstalk environment mid-update, or a half-uploaded store build. A superseded deploy therefore queues behind the running one instead of killing it.
Cancelling a deploy midway does not roll it back. It abandons the run wherever it happens to be, which can leave a CloudFormation stack mid-update, an Elastic Beanstalk environment mid-update, a half-uploaded SPA, or a Lambda on a stub. A superseded deploy therefore queues behind the running one instead of killing it.
CI reusables set `cancel-in-progress: true`. A CI run produces no external side effects, so when a newer commit supersedes an older one there is nothing to protect and the older run should be abandoned.
CI runs produce no external side effects, so when a newer commit supersedes an older one the older run should be abandoned (`cancel-in-progress: true`).
### Groups are declared at job level
### HCP lane
Declare concurrency on the deploy job:
```yaml
concurrency:
group: deploy-<name>-${{ needs.target.outputs.environment }}
cancel-in-progress: false
```
`<name>` is the deployable (`web`, `api`, …). Independent deployables must not share a group.
### Remaining SAM / CDK reusables
Reusable workflows in the central `.github` repo declare their own `concurrency` group. Consumers do not have to add one.
The `concurrency` block sits on the job, not at workflow top level. A single run can contain several jobs that must not share a group: a reusable with multiple jobs needs each one keyed separately, and a caller repo can invoke the same reusable from several jobs in one run. Where a reusable has more than one job, `${{ github.job }}` is part of the key so those jobs do not serialise against each other.
### Groups are scoped per repository
GitHub evaluates a concurrency group within the repository that owns the run, and for a reusable workflow that is the caller's repository. Two different repos calling the same reusable never contend. A group only has to be unique *inside* one repo, which is what the literal workflow-name prefix (`cd-sam-`, `cd-cdk-`, and so on) provides: it stops two different reusables in the same repo from colliding.
### The key must identify the deploy target
Every deploy group is keyed on the inputs that name what is being deployed, not just on the workflow. Keying on the workflow alone would serialise deploys that are genuinely independent — a repo that calls one reusable from several jobs, one per AWS account, would deploy those accounts one at a time for no reason.
Every deploy group is keyed on the inputs that name what is being deployed, not just on the workflow. Keying on the workflow alone would serialise deploys that are genuinely independent.
The groups as implemented:
@ -151,7 +254,7 @@ cancel-in-progress: true
## Naming
- All workflow files: kebab-case
- Reusable workflow references: pinned to a full 40-character commit SHA — never a branch or tag ref
- Reusable workflow references: pinned to a full 40-character commit SHA. Never a branch or tag ref.
## Workflow Ref Pinning
@ -165,28 +268,25 @@ uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-co
uses: Sea-Haven-Industries/.github/.github/workflows/ci-python-sam.yaml@<full-commit-sha> # main
```
Branch refs are mutable: a compromised or bad commit on the central repo would flow instantly into every consumer's CI and deploy path. A SHA pin turns that same change into a reviewable Dependabot PR instead. The `# vX.Y.Z` comment lets Dependabot embed changelog information and lets readers identify which release the pin corresponds to.
Branch refs are mutable: a compromised or bad commit on the central repo would flow instantly into every consumer's CI and deploy path. A SHA pin turns that same change into a reviewable Renovate PR instead. The `# vX.Y.Z` comment lets readers identify which release the pin corresponds to.
Per the pinning principle, pins are for reproducibility, not for freezing time. Two prerequisites keep them moving:
Per the pinning principle, pins are for reproducibility, not for freezing time. Keep them moving with a root `renovate.json` that extends `local>Sea-Haven-Industries/.github`. Do not add `dependabot.yml`.
1. Every repo's `dependabot.yml` must include the `github-actions` ecosystem (weekly), so pin-advance PRs are opened automatically.
2. Dependabot must be granted access to the internal `.github` repo at the org level (Org Settings → Advanced Security → Global settings → "Grant Dependabot access to repositories"). Without the grant, update jobs fail with `git_dependencies_not_reachable` and pins freeze silently — consumers stop receiving central workflow fixes with no visible signal beyond the failed Dependabot run.
When adding a caller workflow by hand, pin to the commit SHA corresponding to the latest release of the central repo (`gh api /repos/Sea-Haven-Industries/.github/commits/vX.Y.Z --jq .sha`) and annotate with `# vX.Y.Z`. The commits endpoint resolves both lightweight and annotated tags to the underlying commit. Let Dependabot advance the pin from there. Use `gh api /repos/Sea-Haven-Industries/.github/commits/main --jq .sha` with a `# main` comment only when the central repo has no release tags.
When adding a caller workflow by hand, pin to the commit SHA corresponding to the latest release of the central repo (`gh api /repos/Sea-Haven-Industries/.github/commits/vX.Y.Z --jq .sha`) and annotate with `# vX.Y.Z`. The commits endpoint resolves both lightweight and annotated tags to the underlying commit. Let Renovate advance the pin from there. Use `gh api /repos/Sea-Haven-Industries/.github/commits/main --jq .sha` with a `# main` comment only when the central repo has no release tags.
## When to Add a Pipeline
- When creating a new deployable project — the pipeline is part of the initial setup, not a follow-up
- When working on an existing project that lacks one — flag it and add it as part of the current work
A project is not production-ready without CI/CD.
A project is not production-ready without CI/CD. New deployables use the HCP lane.
## PR Auto-Labeling
Pull requests are auto-labeled org-wide by a reusable workflow in `.github`. The label rules live once, centrally, inside the reusable workflow itself (written to the runner at execution time), so each repo needs only a short caller and **no per-repo `labeler.yml`**:
```yaml
# .github/workflows/labeler.yml — the per-repo caller
# .github/workflows/labeler.yml, the per-repo caller
name: Labeler
on:
pull_request:

View file

@ -15,8 +15,20 @@
- Security: OWASP top 10, secrets handling, input validation at boundaries
- Sea Haven conventions: naming, secrets placement, Lambda defaults, IaC patterns
- Operational readiness: logging, error handling, monitoring, CI/CD
- HCP app CD, when the PR touches `terraform/` or `deploy-*.yaml`: stub plus `ignore_changes` present, GitHub Actions does not create HCP runs, plan role has no `Get*` wildcards, prod Environment allows `main` and `v*`
- Tests: appropriate coverage for the change
### HCP app CD BLOCK
Treat as BLOCK when the PR touches `terraform/` or `deploy-*.yaml` and any of these are true:
- Terraform would revert app content (missing bootstrap stub or missing `ignore_changes` on the attributes GitHub Actions writes)
- A GitHub Actions workflow creates or applies an HCP run
- The plan role uses `Get*` wildcards
- The prod Environment omits `v*` or `main` from its deployment branch policy
- `hcp_iam.tf` or an in-repo `hcptf-*` role is added (those belong in org-baseline)
- A workflow cuts the GitHub Release with `GITHUB_TOKEN`
## Output Format
```

View file

@ -53,19 +53,24 @@ Transition names are case-insensitive and match the issue's workflow (`#in-progr
## Deploying
CI-on-merge is the only sanctioned deploy path. A change reaches an environment by merging its PR into `main` and letting the repo's pipeline deploy, or through a `workflow_dispatch` run of that same pipeline where the repo configures one. Nobody deploys from a workstation.
Nobody deploys from a workstation. The sanctioned paths depend on the lane. See [cicd.md](cicd.md).
The standard flow for changes that deploy to AWS:
**HCP Terraform (default).** Merge to `main` deploys **dev**: HCP auto-applies if `terraform/**` changed, and GitHub Actions deploys the app unless the merge was terraform-only. **Prod** is a human GitHub Release (`gh release create vX.Y.Z --target main --generate-notes`). That tag applies prod infra and queues the prod app deploys behind Environment reviewers. Rollback is `workflow_dispatch` of the deploy workflow at a prior tag, not a Terraform revert of application content.
**SAM / CDK (remaining).** Merge to `main` (or `workflow_dispatch` of that same pipeline) is still the only path.
The standard flow for an HCP app repo:
1. Create a feature branch (`feature/add-receipt-parser`)
2. Develop and commit locally
3. Open a PR via `gh pr create`
4. Get the review, and confirm CI is green (`gh pr checks`)
4. Get the review, and confirm CI is green (`gh pr checks`). Speculative HCP plans cover both workspaces.
5. Merge the PR on GitHub
6. The pipeline deploys on the merge to `main`, or on a `workflow_dispatch` run where the repo is set up that way
7. Verify the deployed environment; if it is wrong, roll forward with a follow-up PR
6. Confirm the dev apply (if terraform changed) and the dev app deploy. Infra and app may share a PR; if the app job raced the apply, re-run it.
7. Cut prod with `gh release create`. Approve the prod Environment gate after the HCP prod apply lands. Verify live state, not action success.
8. If prod is wrong, dispatch the deploy workflow at the previous tag and approve the gate.
A local deploy puts code into an environment that no reviewed commit describes, and its result depends on whoever ran it having the right credentials and a clean working tree. The pipeline deploys a known commit with the repo's own OIDC role every time. See [cicd.md](cicd.md).
A local deploy puts code into an environment that no reviewed commit describes, and its result depends on whoever ran it having the right credentials and a clean working tree. The pipeline deploys a known commit with the repo's own OIDC role every time.
### Legacy exception: deploy-then-merge
@ -95,25 +100,32 @@ Use [Semantic Versioning](https://semver.org/) (SemVer) as the standard: `MAJOR.
### When to version
Not every project needs versioning. Apply SemVer when the project has consumers that depend on its interface:
HCP app repos always cut SemVer tags. The tag is what applies prod infra and what the prod app deploy checks out.
| Version | Don't version |
| Always version | Don't version |
|---|---|
| Published packages (npm, PyPI) | Internal SAM stacks with no external consumers |
| Public APIs with external consumers | One-off scripts and migration tools |
| Shared libraries used across repos | Internal tools used only by the team that builds them |
| HCP Terraform app repos (the prod trigger) | One-off scripts and migration tools |
| Published packages (npm, PyPI) | Remaining SAM/CDK stacks that still deploy only on merge to `main` |
| Public APIs with external consumers | |
| Shared libraries used across repos | |
| CLIs distributed to users | |
### How to tag
Create annotated tags on `main` after the PR is merged:
For HCP app repos, cut a GitHub Release from `main` after the PR is merged. Do not use a workflow. Do not `git tag` as the normal path (a tag without a Release does not fire `release: published`):
```bash
gh release create v1.2.0 --target main --generate-notes
```
For packages and libraries that are not HCP app repos, annotated tags remain fine:
```bash
git tag -a v1.2.0 -m "v1.2.0"
git push origin v1.2.0
```
Start at `v0.1.0` for new projects. Move to `v1.0.0` when the interface is stable and has external consumers.
Start at `v0.1.0` for new projects. Move to `v1.0.0` when the interface is stable.
## Git Hooks

View file

@ -122,19 +122,30 @@ updates:
Agents and automation (CI bots, Cursor agents, scripts) do not push directly to `main` unless Adam has explicitly directed it for a specific action. The default path for any automated change is a branch and a PR, same as human-authored work.
## GitHub Environments
HCP app repos use Environments as the deploy gate. See [cicd.md](cicd.md#github-environments).
| Environment | Reviewers | Deployment branch policy | Variables |
|---|---|---|---|
| `dev` | none | `main` | `DEPLOY_ROLE_ARN` |
| `prod` | required | `main` and `v*` | `DEPLOY_ROLE_ARN` |
`DEPLOY_ROLE_ARN` is an Environment variable, not a repo secret. Remaining SAM/CDK repos may still use `AWS_DEPLOY_ROLE_ARN` as a repo secret until they migrate.
## README Badges
Every repo's README carries a small badge block immediately under the H1. Use **static** badges only — dynamic badges (for example shields.io `last-commit` or `open-issues`) query the public GitHub API and render broken on private repos.
- **CI status:** the GitHub-native workflow badge (`.../actions/workflows/<ci-file>/badge.svg`), and only when a CI workflow exists. On private repos it renders only for logged-in org members — that is accepted.
- **Stack:** two to four static shields.io badges for the language and IaC/runtime (Python / TypeScript / .NET, AWS SAM / CDK), plus a Slack badge when the repo integrates Slack.
- **Stack:** two to four static shields.io badges for the language and IaC/runtime (Python / TypeScript / .NET, HCP Terraform / AWS SAM / CDK), plus a Slack badge when the repo integrates Slack.
- Keep badges static so they never go stale; version-pinned badges drift.
## Repository Topics
Every repo gets a set of lowercase, hyphenated topics so the org is filterable by stack and purpose. Draw from a consistent vocabulary:
- **Cloud / IaC:** `aws`, `sam`, `cdk`, `lambda`, `ec2`, `s3`, `cloudfront`
- **Cloud / IaC:** `aws`, `terraform`, `hcp`, `sam`, `cdk`, `lambda`, `ec2`, `s3`, `cloudfront`
- **Language:** `python`, `typescript`, `javascript`, `dotnet`, `react`, `nodejs`
- **Integration / domain:** `slack`, `bedrock`, `ai`, `security`, `internal-tool`, `documentation`

114
hcp-terraform.md Normal file
View file

@ -0,0 +1,114 @@
# HCP Terraform
Default IaC for every new deployable, including a one-Lambda repo. GitHub Actions owns application content. This page is the infrastructure half of that split. The application half is [cicd.md](cicd.md). Layout is [terraform-project-layout.md](terraform-project-layout.md).
Reference implementation: `internal-portal`. Copy the contract, not portal-only mechanics (in-workspace `hcptf-*` roles, bootstrap-window credential swaps, workflow-cut releases).
## Ownership
Terraform creates the containers: buckets, distributions, Lambda skeletons, API Gateway, IAM, secrets shells, and the deploy SSM contract. It never writes application content.
GitHub Actions builds and publishes that content. It never creates HCP runs.
The seam is a committed bootstrap stub plus `lifecycle.ignore_changes` on the attributes the deploy workflow writes. A post-deploy `terraform plan` must be empty. That is the acceptance test.
## Accounts and workspaces
Default accounts are `seahaven-dev` and `seahaven-prod`. Every app repo gets two workspaces:
| Workspace | Account | VCS trigger | Auto-apply |
|---|---|---|---|
| `<repo>-dev` | `seahaven-dev` | prefix `terraform/**` on `main` | on |
| `<repo>-prod` | `seahaven-prod` | tag regex `^v[0-9]+\.[0-9]+\.[0-9]+$` | on, after the first apply |
Both workspaces are VCS-connected to `main`, working directory `terraform`, remote execution, speculative plans on for PRs. Speculative plans cover **both** workspaces, so the prod plan is reviewed before merge.
Staging is not in the default. A repo that already has a real staging slot may add `<repo>-staging` and tag `vX.Y.Z-staging` as a variant. Do not add staging because it feels safer.
### First prod apply
The first-ever prod apply runs with auto-apply off and a human confirm. Turn auto-apply on after that apply lands. Later tag applies are automatic.
## VCS file triggers
Keep file triggers on. Docs-only commits must not apply prod.
- Dev: trigger prefix `terraform/**` only. App-path merges correctly skip the dev apply. GitHub Actions ships the app.
- Prod: tag trigger `^v[0-9]+\.[0-9]+\.[0-9]+$`. Tag triggering has no path filter. An app-only release still queues a prod plan; when `terraform/**` did not change it is a no-op apply.
Do **not** add Lambda source, `src/`, `functions/`, `bff/`, or similar to HCP trigger prefixes. Terraform does not package application content. A source-only merge is supposed to skip HCP.
The older rule (PLAT-183) that HCP must watch Lambda source paths applied to stacks where Terraform still bundled code. That is the remaining SAM/CDK-era Terraform path, not the default. New workspaces follow this page.
Record the live prefixes or patterns in the app README next to the workspace name.
## Exec roles
Per-workspace variables, never a project-level variable set:
- `TFC_AWS_APPLY_ROLE_ARN` → `hcptf-<repo>`
- `TFC_AWS_PLAN_ROLE_ARN` → `hcptf-<repo>-plan`
Never `TFC_AWS_RUN_ROLE_ARN`.
Those roles are created and mutated in `seahaven-org-baseline` CloudFormation. Workload Terraform must not declare `aws_iam_role` named `hcptf-*` and must not ship `hcp_iam.tf`.
Do not swap workspace credentials through `hcptf-bootstrap` to change those policies. Policy changes are org-baseline PRs.
The plan role needs a **specific** read for every resource type the apply role creates (example: Cognito `GetUserPoolMfaConfig`). Narrow to named actions. `Get*` wildcards trip CKV_AWS_107 and block the workstation pre-push hook.
`githubdeploy-<repo>` lives in the workload module. The apply role is supposed to mutate that role. See [cicd.md](cicd.md#github-deploy-role).
## Deploy contract (SSM)
Terraform writes non-sensitive names that workflows read. Do not hardcode bucket names, distribution IDs, or function names in YAML.
Default (two accounts), same names in each account:
```
/<repo>/deploy/bucket
/<repo>/deploy/distribution-id
/<repo>/deploy/artifacts-bucket
/<repo>/deploy/<function>-function-name
```
Write only the parameters that repo's deployables need. A SPA without CloudFront omits `distribution-id`. A Lambda-only API omits `bucket`.
One-account leftovers namespace as `/<repo>/<env>/deploy/*`. New repos do not share an account across environments.
## The seam
Lambda (and any other content-bearing resource Terraform must create once):
1. Point `filename` / `s3_key` at a committed stub under `terraform/bootstrap/`.
2. Ignore the attributes the deploy workflow writes:
```hcl
lifecycle {
ignore_changes = [
filename,
s3_bucket,
s3_key,
s3_object_version,
source_code_hash,
]
}
```
Without the stub Terraform has nothing to create. Without `ignore_changes` every apply reverts the app.
Ship the stub and `ignore_changes` in the same apply so the live function is not downgraded to the stub.
Build identity (`GIT_SHA`, Sentry release, reported version) is inlined at **build** time in GitHub Actions. Do not put it in Terraform environment variables. An apply must not be able to regress the reported version.
CloudFront needs no seam when the SPA lives at the bucket root and nothing on the distribution changes per deploy.
Elastic Beanstalk is the same contract with `ignore_changes = [version_label]`, not a third template.
## Mixed PRs
Infra and app changes may share a PR. Dev may deploy the app before the dev apply finishes. Deploys are idempotent. Re-run the job.
## What not to copy from portal
Portal's `hcptf-internal-portal` roles live in the same workspace, so changing their inline policies required a bootstrap-window credential swap. That is a portal exception. Do not add it to a new repo. Do not document it as the operator ritual.

View file

@ -1,6 +1,6 @@
# Lambda Starter Template
Minimal SAM scaffold for a new Python Lambda. Drop into `template.yaml` and adjust names. Follows the defaults in [aws-infrastructure.md](aws-infrastructure.md#lambda-defaults).
Remaining SAM scaffold for an existing Python Lambda stack. New functions belong in an HCP Terraform repo; see [terraform-project-layout.md](terraform-project-layout.md) and [aws-infrastructure.md](aws-infrastructure.md#lambda-defaults).
## Project Structure

View file

@ -17,6 +17,10 @@ No `snake_case`, `PascalCase`, or mixed styles anywhere.
| S3 bucket names | `expense-approval-bot-uploads` |
| Secrets Manager secrets | `expense-approval-bot/slack-signing` |
| Feature branches | `feature/add-receipt-parser` |
| HCP workspaces | `expense-approval-bot-dev`, `expense-approval-bot-prod` |
| HCP apply / plan roles | `hcptf-expense-approval-bot`, `hcptf-expense-approval-bot-plan` |
| GitHub deploy role | `githubdeploy-expense-approval-bot` |
| Deploy SSM prefix | `/expense-approval-bot/deploy/` |
## CDK Gotcha

View file

@ -2,7 +2,7 @@
## Scope
Each PR should represent a single logical change. If you find yourself writing "and" in the title, consider splitting it into separate PRs.
Each PR should represent a single logical change. If you find yourself writing "and" in the title, consider splitting it into separate PRs. Infra and app changes **may** share a PR on an HCP repo. Dev may deploy the app before the dev apply finishes; deploys are idempotent, re-run the job. Put sequencing (approve prod deploys after the HCP apply, rollback tag) under `## Notes`.
| Good scope | Too broad |
|---|---|
@ -52,7 +52,7 @@ State verifiable facts about the code. Do not cite the engineering handbook, and
- When the work is ready for review, not as a draft for parking incomplete work
- After verifying locally that the change works as expected
- The deploy happens after the PR merges to `main`; see [Deploying](git-workflow.md#deploying)
- Dev deploys after the PR merges to `main`. Prod is a GitHub Release, not the merge. See [Deploying](git-workflow.md#deploying)
## Merging

View file

@ -1,5 +1,7 @@
# SAM Project Layout
Remaining path for stacks that already use SAM. New deployables use [terraform-project-layout.md](terraform-project-layout.md) and [hcp-terraform.md](hcp-terraform.md).
## Standard Directory Structure
```

View file

@ -1,5 +1,9 @@
#!/usr/bin/env bash
# Provision a new Sea Haven Industries repo with all required infrastructure.
# Provision a remaining SAM or CDK repo. Do not use this for new app workloads.
# New deployables use HCP Terraform (see hcp-terraform.md). A Terraform
# provisioner (workspaces, org-baseline roles, GitHub Environments) is a
# later PLAT ticket.
#
# Usage: ./provision-repo.sh <repo-name> [sam|cdk]
#
# Creates: GitHub repo, OIDC deploy role, repo secret, security features,
@ -19,6 +23,17 @@ REGION="us-east-1"
OIDC_PROVIDER="arn:aws:iam::${ACCOUNT_ID}:oidc-provider/token.actions.githubusercontent.com"
ROLE_NAME="githubdeploy-${REPO_NAME}"
if [[ "${STACK_TYPE}" == "terraform" || "${STACK_TYPE}" == "hcp" ]]; then
echo "Error: this script does not provision HCP Terraform repos." >&2
echo "New app workloads follow hcp-terraform.md. Do not create githubdeploy-* here." >&2
exit 1
fi
if [[ "${STACK_TYPE}" != "sam" && "${STACK_TYPE}" != "cdk" ]]; then
echo "Error: stack type must be sam or cdk (remaining path only). Got: ${STACK_TYPE}" >&2
exit 1
fi
if [[ ! "$REPO_NAME" =~ ^[a-z0-9]([a-z0-9-]*[a-z0-9])?$ ]]; then
echo "Error: repo name must be kebab-case (lowercase, hyphens only, no leading/trailing hyphens)"
exit 1

View file

@ -35,6 +35,9 @@ Use Parameter Store **only** for non-sensitive configuration:
- Endpoint URLs (public, non-authenticated)
- Schedule expressions
- Channel IDs and non-sensitive identifiers
- The HCP deploy contract: names Terraform writes and GitHub Actions reads (`/<repo>/deploy/*` in each account). See [hcp-terraform.md](hcp-terraform.md#deploy-contract-ssm).
Do not put build identity (`GIT_SHA`, release labels) in Terraform-managed Lambda environment variables. Inline it at build time in GitHub Actions so an apply cannot regress the reported version.
## Lambda Pattern

View file

@ -0,0 +1,60 @@
# Terraform Project Layout
Default layout for new deployable repos. HCP Terraform applies this tree. GitHub Actions publishes application content. See [hcp-terraform.md](hcp-terraform.md) and [cicd.md](cicd.md).
## Standard Directory Structure
```
project-name/
├── terraform/
│ ├── bootstrap/ # Committed stubs Terraform can create
│ │ └── handler-stub.zip
│ ├── iam_github_deploy.tf # githubdeploy-<repo> only
│ ├── lambda.tf
│ ├── s3.tf
│ ├── ssm.tf # /<repo>/deploy/* contract
│ ├── variables.tf
│ ├── outputs.tf
│ └── versions.tf
├── src/ # Application code (GHA owns the zip)
├── scripts/ # Live-state verify scripts
├── .github/
│ └── workflows/
│ ├── ci.yaml
│ └── deploy-<name>.yaml # One workflow per deployable
└── README.md
```
Working directory in both HCP workspaces is `terraform`. Do not flatten per-env roots (`terraform/live/dev`) unless a remaining stack already has them. New repos use one tree; workspaces select the account via variables.
## File Purposes
### `terraform/bootstrap/`
A tiny committed zip (or equivalent) so Terraform can create the Lambda. GitHub Actions overwrites the live code on the first deploy. Check the stub in. Do not generate it at plan time with `data.external` or `archive_file` from application source.
### `lambda.tf` (or equivalent)
Create the skeleton. Point it at the stub. Set `lifecycle.ignore_changes` on `filename`, `s3_bucket`, `s3_key`, `s3_object_version`, and `source_code_hash`. Do not set `GIT_SHA` or a release label as a Terraform env var.
### `ssm.tf`
Write `/<repo>/deploy/*` (see [hcp-terraform.md](hcp-terraform.md#deploy-contract-ssm)). Workflows read these. YAML does not hardcode resource names.
### `iam_github_deploy.tf`
The `githubdeploy-<repo>` role only. Trust and permissions are in [cicd.md](cicd.md#github-deploy-role). Adding a deploy workflow means adding its `job_workflow_ref` here. That is a cross-family IAM change.
Do **not** add `hcp_iam.tf`. `hcptf-<repo>` and `hcptf-<repo>-plan` live in `seahaven-org-baseline`.
### Application trees (`src/`, `bff/`, `web/`, `functions/`)
Owned by GitHub Actions. HCP trigger prefixes do not include these paths.
## CI for Terraform
Repo CI runs `terraform fmt -check -recursive`, `terraform init -backend=false`, and `terraform validate`. Plans stay in HCP speculative runs. Do not have GitHub Actions call the HCP API to create or apply a run.
## Remaining path
Stacks that still package Lambda code inside Terraform (mgmt migrations that have not been cut over) keep their source paths in HCP file triggers until they adopt this layout. New repos start here.