docs(iam): codify hcp terraform migration checklist from PLAT-56 (#79)
Some checks are pending
Deploy / deploy-management (push) Waiting to run
Deploy / deploy-external-dev (push) Waiting to run
Deploy / deploy-security (push) Waiting to run
Deploy / deploy-dev (push) Waiting to run
Deploy / deploy-prod (push) Waiting to run

Expand the README playbook to steps 0–10 and document the required
plan-refresh sidecar plus prefix-scoped apply-role wildcards so the
next workload copies afi patterns instead of relearning first-apply misses.
This commit is contained in:
Adam Moussa 2026-08-05 16:14:16 -04:00 • committed by GitHub
parent 2c5c644a5e
commit 9ee4d4a3d7
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
2 changed files with 103 additions and 36 deletions

107
README.md
View file

@ -253,49 +253,102 @@ Remote. Never use HCP's "Quick setup AWS dynamic credentials" button — it
writes the single `TFC_AWS_RUN_ROLE_ARN`, which collapses the plan/apply role
split this substrate exists to enforce.
**Reference implementation:** first workload was `afi-backup-monitor` in
seahaven-prod (PLAT-56). Copy
`Sea-Haven-Industries/afi-backup-monitor` `terraform/` and the live
`hcptf-afi-backup-monitor*` / `hcptf-afi-backup-monitor-plan` statements in
this template rather than inventing new IAM shapes.
**Migration checklist (per stack, in order):**
1. Create the workspace in the target account's HCP project
(`<stack>-<env>`). Apply method **Manual**; automatic speculative plans on
if VCS-connected (CLI `terraform plan` runs are inherently speculative).
2. PR to this repo appending `hcptf-<stack>-plan` (read-only —
`arn:aws:iam::aws:policy/job-function/ViewOnlyAccess`, never
`ReadOnlyAccess`, which grants `secretsmanager:GetSecretValue`,
`s3:GetObject` and `kms:Decrypt` and would let any PR-triggered speculative
plan render secret values into HCP run output; **no** IAM writes, **no**
guardrail-policy attach)
and `hcptf-<stack>` (attaches `seahaven-hcptf-iam-management` + stack-scoped
service statements) to the substrate template. Trust: this account's
`app.terraform.io` provider; `StringEquals` on
`app.terraform.io:aud` = `aws.workload.identity` and on
`app.terraform.io:sub` =
0. **Freeze the app's SAM/CDK CD** (remove or disable the deploy workflow) so
HCP Terraform becomes the sole deploy path before the first apply. Leave
the source-account stack frozen until cutover.
1. **Secrets first.** Create exact secret shells in the target account; strip
trailing newlines/whitespace before `put-secret-value` (a trailing `\n`
breaks HTTP headers at runtime). Capture ARNs. Never put secret *values*
in Terraform state (ARN references only).
2. **HCP workspace** in the target account's project (`<stack>-<env>`). Apply
method **Manual**; automatic speculative plans on if VCS-connected;
working directory `terraform/`. (CLI `terraform plan` runs are inherently
speculative.)
3. **Substrate PR** to this repo appending `hcptf-<stack>-plan` and
`hcptf-<stack>` (see 3a/3b). Trust: this account's `app.terraform.io`
provider; `StringEquals` on `app.terraform.io:aud` =
`aws.workload.identity` and on `app.terraform.io:sub` =
`organization:seahaven:project:seahaven-<env>:workspace:<workspace>:run_phase:plan`
(or `:apply`). Exact `StringEquals` only — never `StringLike`, never a
wildcarded `run_phase` (a speculative PR plan must never hold write
credentials). **If the stack creates Lambda execution roles, this same PR
must also widen `seahaven-lambda-execution-boundary`** per the WIDENING
PATH in `lib/deploy-substrate/deploy-substrate.template.yaml`: the
PATH in `lib/deploy-substrate/deploy-substrate.template.yaml` with the
**exact** secret ARNs from step 1 (no `secret:afi-*` patterns): the
guardrail forces every Terraform-created role to carry that boundary, and
it is a fleet-wide floor with zero data-plane permissions until widened —
an unwidened migration deploys green, then every data-plane call is denied
at first invoke and async/DLQ writes are discarded silently. IAM roles and
boundary widenings = mandatory cross-family review +
`/sh-security-review` on the diff.
3. After deploy, verify: both roles exist; `hcptf-<stack>` lists
`seahaven-hcptf-iam-management` in `list-attached-role-policies`; trust
subs match the live org/project/workspace names byte-for-byte; simulate
the apply role against a `hcptf-*` ARN (expect `explicitDeny` from
`DenySelfMutation`) and against a normal stack role name (expect
`allowed`); and if step 2 widened the boundary, confirm the deployed
default version carries the stack's data-plane statements
(`aws iam get-policy-version`) — role verification alone never checks
boundary content.
4. Set **workspace-level** variables `TFC_AWS_PLAN_ROLE_ARN` +
3a. **Plan role (required for every stack):** attach
`arn:aws:iam::aws:policy/job-function/ViewOnlyAccess` (never
`ReadOnlyAccess`, which grants `secretsmanager:GetSecretValue`,
`s3:GetObject` and `kms:Decrypt` and would let any PR-triggered speculative
plan render secret values into HCP run output) **plus** a scoped
plan-refresh sidecar inline policy. ViewOnly alone is insufficient for
Terraform refresh after partial apply — it lacks `iam:GetRole`,
`events:DescribeRule`, and several Lambda/S3 reads. Sidecar minimum:
`iam:GetRole` / related reads on `role/tf-managed/<prefix>-*`;
`events:DescribeRule` (and list-targets/tags as needed) on
`rule/<prefix>-*`; `lambda:*` (or at least the Get*/List* the provider
uses) on `function:<prefix>-*` / `layer:<prefix>-*`; `s3:Get*` /
`s3:ListBucket` on the stack artifact bucket. **No** IAM writes, **no**
guardrail-policy attach on the plan role. Copy
`afi-backup-monitor-plan-refresh` on `hcptf-afi-backup-monitor-plan`.
3b. **Apply role (Lambda/EventBridge stacks):** attach
`seahaven-hcptf-iam-management` plus stack-scoped service statements.
Prefer prefix-scoped `lambda:*` on `function:<prefix>-*` /
`layer:<prefix>-*`, `events:*` on `rule/<prefix>-*`, and bucket-scoped
`s3:*` on the artifact bucket — do **not** enumerate individual provider
Get* APIs (`GetFunctionCodeSigningConfig`, `GetBucketAcl`, …); that list
lags and fails first apply. Keep list/describe-on-`*` only where the
service requires it (e.g. `lambda:ListFunctions`). Copy
`afi-backup-monitor-services` on `hcptf-afi-backup-monitor`.
4. **Deploy substrate** to `UPDATE_COMPLETE`. Verify: both roles exist;
`hcptf-<stack>` lists `seahaven-hcptf-iam-management` in
`list-attached-role-policies`; trust subs match the live
org/project/workspace names byte-for-byte; simulate the apply role against
a `hcptf-*` ARN (expect `explicitDeny` from `DenySelfMutation`) and against
a normal stack role name (expect `allowed`); and if step 3 widened the
boundary, confirm the deployed default version carries the stack's
data-plane statements (`aws iam get-policy-version`) — role verification
alone never checks boundary content. Mechanical template↔deployed policy
reconcile as for other substrate policies.
5. Set **workspace-level** variables `TFC_AWS_PLAN_ROLE_ARN` +
`TFC_AWS_APPLY_ROLE_ARN` (category env) to the verified role ARNs, plus
`TFC_AWS_PROVIDER_AUTH=true`. Never project-scoped variable sets — the
trust is pinned per workspace, so a shared set breaks every other
workspace.
5. Auto-apply stays OFF until the stack is sealed.
workspace. Auto-apply stays OFF until the stack is sealed.
6. **App Terraform PR:** every `aws_iam_role` sets `path = "/tf-managed/"` and
the boundary; package Lambda/layer zips via an account artifact S3 bucket
and `aws_s3_object` `content_base64` (HCP plan and apply run on separate
workers and do not share local `archive_file` paths — see
`afi-backup-monitor/terraform/artifacts.tf`); functions `depends_on` their
IAM policies before create; commit `.terraform.lock.hcl` with
multi-platform hashes.
7. **First Manual apply** from the HCP workspace (not local apply against
prod). Tolerate partial state on permission misses; widen the apply/plan
roles and retry. Confirm all expected resources exist in the target
account.
8. **Live-path proof:** real invoke of every critical function must hit real
external APIs / Slack (not synth or simulate alone) before cutover.
9. **Cutover + decommission:** disable source-account schedules (e.g.
EventBridge rules); observe a clean prod path; delete the source
CloudFormation/CDK stack per the decommission playbook; sweep or retain
log groups deliberately; delete source secrets last.
10. **Docs:** update Confluence AWS Architecture Map and the stack ops page;
promote durable gotchas to the convention ledger when they are general.
**HCP-side authority is AWS authority.** AWS exposes only `aud`, `sub` and
`amr` as trust-policy condition keys for a generic OIDC provider — HCP's

View file

@ -373,16 +373,30 @@ Resources:
"iam:PassedToService": "lambda.amazonaws.com"
# ---------------------------------------------------------------------------
# Per-workspace roles: afi-backup-monitor-prod (PLAT-56)
# Per-workspace role pattern (required for every future hcptf-* append)
# and first workload: afi-backup-monitor-prod (PLAT-56).
#
# First HCP Terraform workload. Trust subs are exact StringEquals on
# organization/project/workspace/run_phase — never StringLike, never a
# wildcarded run_phase. Plan role: ViewOnlyAccess (never ReadOnlyAccess,
# which grants secretsmanager:GetSecretValue) PLUS a stack-scoped refresh
# inline policy — ViewOnlyAccess omits iam:GetRole and events:DescribeRule,
# which Terraform needs to refresh state after the first apply. Apply role:
# attaches the shared guardrail plus stack-scoped Lambda / layer /
# EventBridge / Logs / artifact-bucket. Prod-only (IsProdAccount).
# Copy this shape — do not invent enumerated Get* allow-lists.
#
# Plan role (every stack):
# - Managed: ViewOnlyAccess (never ReadOnlyAccess — it grants
# secretsmanager:GetSecretValue / s3:GetObject / kms:Decrypt to
# speculative PR plans).
# - PLUS a stack-scoped plan-refresh sidecar. ViewOnly alone omits
# iam:GetRole, events:DescribeRule, and provider Lambda/S3 reads
# needed after a partial first apply.
#
# Apply role (Lambda / EventBridge stacks):
# - Attach seahaven-hcptf-iam-management.
# - Service grants: prefix-scoped lambda:* on function:<prefix>-* and
# layer:<prefix>-*, events:* on rule/<prefix>-*, and bucket-scoped
# s3:* on the stack artifact bucket. Enumerating provider Get*
# (GetFunctionCodeSigningConfig, GetBucketAcl, …) lags and fails
# first apply (PLAT-56).
#
# Trust: exact StringEquals on organization/project/workspace/run_phase —
# never StringLike, never a wildcarded run_phase. Prod-only for this
# pair (IsProdAccount). See README "Migration checklist".
# ---------------------------------------------------------------------------
HcptfAfiBackupMonitorPlanRole:
Type: AWS::IAM::Role