Callers pin `uses:` to a commit SHA of this repo. Dependabot's
github-actions updater finds a newer SHA for a pinned ref by reading the
target repo's tags and releases; this repo has 0 tags and 0 releases, so
there is nothing for it to resolve and it reports no update. The fleet's
pins have not moved as a result.
Adds a caller for the release reusable, triggered on push to main and
filtered to paths under .github/workflows/ excluding the three files no
caller consumes (ci.yaml, labeler.yaml, and this file).
Version is a patch bump from the highest existing vMAJOR.MINOR.PATCH tag,
1.0.0 when none exist. workflow_dispatch takes an explicit version for
minor and major bumps.
Runs are serialised with cancel-in-progress false: two merges landing
together would otherwise read the same highest tag, compute the same next
version, and the second would hit the reusable's existing-tag guard and
no-op, leaving that change unreleased.
release.yml / release.properties.json and ci-mobile-ios.yml /
ci-mobile-ios.properties.json are new caller templates for the two reusable
workflows added in 9389e51. Both pin the reusable to 9389e51, the commit that
introduces the workflow files.
release.yml triggers on workflow_dispatch with a required `version` input and
declares `permissions: contents: write`, which the reusable needs to push the
tag and publish the Release. ci-mobile-ios.yml triggers on pull_request and
keys its job `ci` so the check context resolves to `ci / ci`.
ci-python.yml now passes `node-version: "24"`. Its target,
ci-python-sam.yaml, declares that input at line 34 with default "24"; the
ci-static, ci-typescript-frontend, ci-node, cdk-deploy and mobile-ios-deploy
templates already pass the same value.
Both new .properties.json files carry the same five keys as the thirteen
existing ones: categories, description, filePatterns, iconName, name.
release.yaml is workflow_call-only. It normalises and validates a `version`
input against MAJOR.MINOR.PATCH, skips every mutating step when the tag or a
Release for it already exists, creates an annotated tag with `git tag -a` and
publishes a GitHub Release with `gh release create --verify-tag`. Top-level
permissions grant `contents: write` only; no id-token is requested. The
previous-tag lookup and `--generate-notes` both work in a repo with no tags.
ci-mobile-ios.yaml is workflow_call-only and pairs with cd-mobile-ios.yaml,
reusing its node-version, ruby-version, working-directory,
cache-dependency-path and fastlane-lane input names. Job `js` runs npm ci,
typecheck, optional lint and optional tests on ubuntu-latest. Job `ios-build`
runs pod install and `xcodebuild build` on macos-26 with
CODE_SIGNING_ALLOWED=NO, CODE_SIGNING_REQUIRED=NO and CODE_SIGN_IDENTITY="";
it declares no secrets and performs no upload. Job `ci` aggregates both via
`needs` so a caller job keyed `ci` reports `ci / ci`. All three jobs carry
job-level concurrency with cancel-in-progress: true. Top-level permissions are
`contents: read`.
The Fastlane branch carries a per-line `# shellcheck disable=SC2086` because
fastlane requires the platform and lane as two argv entries.
The self-CI gate ran `./actionlint -shellcheck=`, and the empty value
silently disabled the shell-linting half of the check — so every `run:`
body in the reusable workflows this repo publishes was unlinted, on the
exact path that deploys to AWS.
Measured against the pinned actionlint 1.7.12 and the shellcheck the
ubuntu-latest runner ships (0.9.0-1), the real backlog was 5 findings,
not the 4 the old comment claimed. Three were genuine and are fixed in
the shell:
- cd-cdk.yaml "Publish .NET project" (SC2046): the project path was
interpolated inline and `$(dirname ...)` was unquoted, so a path
containing whitespace split into several arguments. Now passed via
env indirection and quoted, which also removes the last inline
expression interpolation from that step.
- cd-cdk.yaml / ci-python-sam.yaml "Install Python dependencies"
(SC2044 x2): `for req in $(find ...)` word-split and globbed every
path found. Replaced with a NUL-delimited `while read` loop.
Two are deliberate and are suppressed per-line, with the reasoning in a
comment directly above:
- cd-sam.yaml `sam deploy ... $PARAMS` and cd-cdk.yaml
`cdk deploy $STACKS` (SC2086 x2) rely on word-splitting so multiple
parameter overrides / stack selectors reach the CLI as separate argv
entries. Quoting them would collapse each into a single argument and
break every parameterised or multi-stack deploy, so they keep the
unquoted expansion and carry a scoped `# shellcheck disable=SC2086`.
The gate now runs plain `./actionlint` (shellcheck defaults to the
binary on PATH) and prints `shellcheck --version` first, so the check
fails loudly if a future runner image drops it instead of quietly
linting less.
cd-dotnet-eb already serialises deploys per environment; the other three
deploy reusables had no concurrency group, so back-to-back merges could
start overlapping runs against the same target. CloudFormation rejects a
concurrent update on the same stack and cd-sam/cd-cdk pre-flight already
hard-fails on an in-progress stack, so the symptom is a failed run that
needs a manual re-run rather than a corrupted deploy. Grouping makes
those deploys queue instead.
Each group key names the thing being deployed, so independent targets in
one caller repo still deploy in parallel:
cd-sam region + stack-name (both always non-empty)
cd-cdk region + stacks + stack-name
cd-mobile-ios working-directory + fastlane-lane (both default)
cd-cdk keys on the stack selector rather than stack-name because
seahaven-org-baseline calls it from five jobs in a single run, one per
AWS account, and two of those pass no stack-name. Keying on stack-name
alone would collapse them into one group and serialise five independent
per-account deploys.
cancel-in-progress is false on all three, matching cd-dotnet-eb: unlike
CI, cancelling a deploy midway can leave infrastructure mid-update.
GitHub Actions expressions are substituted into a run body as text
before bash parses it, so a value carrying a quote, a command
substitution, or a newline becomes shell syntax rather than data.
cd-sam.yaml interpolated the parameter-overrides secret straight into
a shell test and an assignment, putting secret material into the script
body. cd-cdk.yaml interpolated the caller-supplied post-deploy-script
input into a bash invocation, which is caller-controlled command
injection rather than secret exposure.
Both now use env-var indirection, matching the STACKS precedent in the
CDK deploy step. PARAM_OVERRIDES is deliberately left unquoted at the
point of use: parameter-overrides carries multiple Key=Value pairs that
must reach sam deploy as separate argv entries, so quoting it would
collapse every override into one argument and break deploys that use
it. POST_DEPLOY_SCRIPT is a single path and is quoted.
Behaviour is otherwise unchanged. An empty parameter-overrides still
produces no --parameter-overrides flag at all, and an empty
post-deploy-script is still skipped by the step-level if condition,
which is a workflow expression and not shell.
The workflow-templates catalog offered starter workflows for only 7 of
the 12 reusable workflows in .github/workflows, so ci-python-app,
ci-typescript-frontend, ci-static, ci-dotnet and cd-mobile-ios were
invisible in the org's Actions > New workflow UI and had to be wired by
hand. Add a template + properties.json pair for each.
Each caller CI job is keyed `ci` so the check context resolves to the
`ci / ci` required by the org ruleset, and every reusable ref is pinned
to the same 40-char SHA the existing templates use. node-version: "24"
is passed on the three reusables that declare the input
(ci-typescript-frontend, ci-static, cd-mobile-ios); ci-python-app and
ci-dotnet do not declare it, so it is omitted there.
Three pieces of tooling for retired flows were still carried in this
repo, and the README documented them as if they were current.
- `.github/workflows/compliance-audit.yaml`: deprecated 2026-06-10.
Its schedule was already stripped, and its live state in the Actions
API is `disabled_manually`, so it was workflow_dispatch-only and
inert. It was also the last consumer of the `CLAUDE_CI_APP_ID` and
`CLAUDE_CI_APP_PRIVATE_KEY` org secrets.
- `scripts/rollout-review-workflow.sh`: a one-shot script that pushed a
per-repo wrapper calling `claude-code-review.yaml`. That reusable
workflow was deleted on 2026-05-13 and no longer exists in this repo
or any of the 31 org repos, so the script could only ever open PRs
for a workflow that resolves to nothing.
- README `PR Reviews`, `Scripts`, and the `claude-code-ci` GitHub App
setup steps, which described the same retired flows.
The `ANTHROPIC_API_KEY` org secret is deliberately kept in the setup
table: it is still read by the active `reviewer-eval.yml` workflow in
`open-swe`. The `CLAUDE_CI_APP_*` secrets are now unreferenced across
the org, but this change only removes their documentation. Neither the
secrets nor the App itself are touched.
Setup steps are renumbered 1-3 with no gap, and the `see §5` reference
in the IAM section is repointed to §3. actionlint 1.7.12 (the version
pinned in ci.yaml) passes clean over the remaining workflows.
ci-python-sam, ci-typescript-cdk and ci-dotnet declared no permissions
at any level, unlike every other workflow here. A reusable workflow that
declares nothing inherits the CALLER's token scopes, and these are
called from deploy repos, so a lint/test/synth job could run holding an
OIDC-mintable token it has no use for. None of the three references
GITHUB_TOKEN, github.token, gh, or any secret, so contents:read is all
they need to check out and build.
Also pass node-version explicitly in the cdk-deploy and ci-node
templates. cicd.md requires callers to pin it so lockfileVersion 3 from
local Node 24 / npm 11 cannot drift from the runner, but no template
did. Only these two targets accept the input; sam-deploy, dotnet-eb,
dependency-review and labeler do not, so they are left alone.
Verified against all 22 callers across the org that none grants
permissions omitting contents:read, so no repo's CI breaks on the
caller-cannot-be-exceeded rule.
Phase A moved the CFN execution role's boundary-gated IAM statements into
the attached seahaven-cfn-exec-iam-management managed policy, with the
escalation fixed and three Deny backstops, while deliberately leaving the
old inline iam-role-management-boundary-gated policy in place so that
deploy removed nothing. That is now redundant and this removes it.
Effective permissions are unchanged, proven statically before deploying:
of the 6 Allow statements being removed, 5 are byte-identical to the
managed policy's. The only difference is Sid IAMPutPermissionsBoundary,
where the inline copy also listed iam:DeleteRolePermissionsBoundary --
the action DenyBoundaryTampering explicitly denies, so that Allow was
already inert.
Frees the scarce budget: inline usage drops from 10,006 to 8,261 of IAM's
10,240-byte per-role limit, leaving 1,979 bytes of headroom on a role that
previously had 234.
The shared CloudFormation execution role could remove the permissions
boundary from the very roles that boundary was gating. Its
iam-role-management-boundary-gated policy allows
iam:DeleteRolePermissionsBoundary on role/* under a StringEquals
condition on iam:PermissionsBoundary -- and for a delete that condition
key reflects the boundary CURRENTLY attached to the target role, so it
matches exactly the roles the gate protects. Create a boundary-gated
role with an inline *:* policy, strip its boundary, PassRole it to
Lambda, and the result is unbounded admin in the management account.
Confirmed live with simulate-principal-policy, not inferred.
Phase A adds an attached managed policy, seahaven-cfn-exec-iam-management,
carrying the corrected statement set: the boundary-gated Allows without
iam:DeleteRolePermissionsBoundary, plus three Deny backstops --
DenyBoundaryTampering (boundary removal), DenyBoundaryPolicyEdit
(rewriting a seahaven-* policy document) and DenySelfMutation.
DenySelfMutation exists because the first draft of this fix was not
durable: the role holds iam:DetachRolePolicy, iam:DeleteRolePolicy and
iam:DeleteRole on Resource "*" with no condition, so it could detach the
Deny-carrying policy from itself in one call and reinstate the
escalation. It now cannot modify its own role, any githubdeploy-* role,
or any seahaven-* policy. Nothing legitimate needs that: the deploy
substrate's own principals are owned by this stack, which is deployed
manually with administrator credentials rather than through this role.
The change is additive. The old inline policy stays in place, so
CloudFormation removes nothing and there is no window in which the role
lacks its IAM permissions -- an explicit Deny beats an Allow anywhere in
the policy set, so the corrected version governs from the moment this
lands. Phase B removes the redundant inline copy. The split is also
required by size: inline sits at 10,006 of IAM's 10,240-byte per-role
limit, and the Deny statements do not fit there.
Also reconciles drift. The deployed role carries three logs:*MetricFilter
actions added out-of-band on 2026-06-29 and never back-ported.
afterhours-shift-manager creates an AWS::Logs::MetricFilter through this
role, so they are load-bearing; the template now matches the live policy
exactly, which keeps inline at 10,006 and stops a future write-back from
silently stripping them.
Removes the SeahavenSlackBotDeployRole resource block. Step 1 recorded
DeletionPolicy/UpdateReplacePolicy Retain in the deployed template, so
CloudFormation stops managing the resource without issuing DeleteRole
against a role that no longer exists -- confirmed from the change set,
which reports PolicyAction: Retain on a single Remove entry.
seahaven-slack-bot was decommissioned in favour of sh-mcp and the role
was deleted directly in IAM on 2026-07-23. The stack is now consistent
with reality again, and stack updates no longer fail on it.
seahaven-slack-bot was retired in favour of sh-mcp and its deploy role was
deleted directly in IAM on 2026-07-23, leaving the stack holding a
resource that no longer exists. The Outputs section resolved
!GetAtt SeahavenSlackBotDeployRole.Arn as a LIVE IAM read at the end of
every update, so the role's absence failed the whole thing:
Unable to retrieve Arn attribute for AWS::IAM::Role, with error message
The role with name githubdeploy-seahaven-slack-bot cannot be found. (404)
This is latent and invisible: the resource definition is unchanged, so it
produces no change-set entry, and change sets do not preview Outputs
resolution. A clean change set was not evidence the update would succeed.
It surfaced when the Phase A boundary-Deny change failed on it.
Step 1 of two. Removes the Output so updates stop resolving the ghost, and
records DeletionPolicy/UpdateReplacePolicy Retain so that step 2 can drop
the resource without CloudFormation issuing DeleteRole against a role that
is not there. Verified from the change set that this step touches only
DeletionPolicy and UpdateReplacePolicy -- metadata, requiresRecreation
Never -- so no IAM call is made against the missing role.
Nothing imported the Output: it had no ExportName, and no stack imports
any export from this stack.
Step 2 deletes the resource block itself.
Publishes a .NET project, packages the output as a bundle, uploads it,
creates an Elastic Beanstalk application version, and updates an
existing environment using OIDC credentials. It deploys to an
environment; it never creates one.
Two deliberate departures from the existing cd-* reusables:
- A concurrency group keyed on application+environment, with
cancel-in-progress false, so two pushes cannot deploy over each other
and an in-flight deploy is never aborted midway. The existing cd-*
workflows have no concurrency group at all.
- No input or secret is interpolated into a run: body; everything goes
through env-var indirection. The repo's actionlint runs with
shellcheck disabled, so this is a hand-maintained property.
The post-deploy check polls rather than using the CLI waiter
"elasticbeanstalk wait environment-updated": that waiter is hardcoded to
20 attempts x 20s and the CLI cannot extend it, so a slower rolling
deploy would fail the job while the deployment was still healthy. The
timeout is an input instead. The check asserts status, health and the
running version label -- Elastic Beanstalk reports a rolled-back deploy
as a healthy Ready environment running the previous version, so without
the version assertion the job would go green over a failed deploy.
Replaces the malformed cd-dotnet-eb.yaml.yml stub (doubled extension,
empty on:/jobs:). Adds the matching starter template and README entries.
The sam-deploy starter template pointed every new repo's cfn-role-arn at
the management account's execution role, silently landing new workloads
in an account frozen for workloads. The ARN is now a REPLACE-ME
placeholder with guidance to use the github-cfn-execution-role in the
repo's target account.
Per the updated handbook convention (engineering-handbook PR #18),
reusable-workflow references use full commit SHA pins with a '# main'
comment instead of the mutable @main branch ref. Templates now ship
pinned so new repos start convention-compliant; Dependabot advances
the pin after instantiation. Commented usage examples in ci-static and
ci-typescript-frontend use the <full-commit-sha> placeholder form.
Callers with an adjudicated accepted-risk advisory (suppressed with
justification in their repo-local .security-review/suppressions.json)
had no way to keep the dependency-review check green when a lockfile
diff touches a package still inside the vulnerable range. Passes the
input straight to actions/dependency-review-action. Default '' is
byte-identical to an unset action input, so existing callers are
unaffected.
First consumer: seahaven-site, allowing GHSA-mh99-v99m-4gvg
(brace-expansion, no in-range fix until @11ty/recursive-copy bumps
minimatch).
Unpinned pip install ruff let ruff 0.16.0 pick up expanded
default lint rules, breaking every caller repo without its
own ruff config. Pinning prevents implicit rule-set changes
on new ruff releases.
Refs: #87
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
payments-dashboard's BoaRawBucket (first bucket in the org with an
explicit BucketEncryption block) failed CREATE: the CFN execution
role lacked s3:PutEncryptionConfiguration. Adds the Get/Put pair to
the shared s3-management statement (bucket-level, existing * scope).
Escalation review: the role holds no kms:* actions anywhere, so the
PutEncryptionConfiguration + PutBucketPolicy combination cannot pivot
to a role-controlled KMS key; SCPs permit the action (the original
denial was identity-policy). GPT-4.1 cross-family review: FIX-level
only, dispositioned above. Stack deployed before merge per README.
Refs: payments-dashboard#76
Post-rename deploy verified green from seahaven-org-baseline (run
29355616637, both account jobs). The freed repo name must not stay
trusted (namespace-reuse window, security review IAC-02).
* Add stacks input to cd-cdk for multi-account apps
cdk deploy was hardcoded to --all, which breaks when one CDK app defines
stacks for two AWS accounts: whichever role the job assumed fails on the
other account's stacks. Callers can now pass per-job stack selectors;
default stays --all so existing callers are unaffected.
* Trust seahaven-org-baseline sub on account-baseline deploy role
Transition pair for the repo rename: OIDC sub claims carry the repo full
name, so the renamed repo cannot assume the role until its sub is
trusted. Old sub is removed after a post-rename deploy verifies green.
* Pass stacks selector via env var, not expression interpolation
Defense-in-depth from the security review: expression interpolation
into run: is pre-shell text substitution, so metacharacters in the
input would execute as script. Env-var expansion never re-parses shell
syntax; word-splitting for multiple selectors is preserved.
Add job-level concurrency (cancel-in-progress) to all four ci-python-app
jobs, matching the ci-python-sam idiom. Per-job group keys include
github.job so the parallel jobs in a single run do not share a group.
Replace sam-deploy starter-template stack-name: $default-branch (which
GitHub substitutes to the literal branch name main) with a
REPLACE-ME-stack-name placeholder, and point cfn-role-arn at the real
shared github-cfn-execution-role.