The self-CI gate ran `./actionlint -shellcheck=`, and the empty value
silently disabled the shell-linting half of the check — so every `run:`
body in the reusable workflows this repo publishes was unlinted, on the
exact path that deploys to AWS.
Measured against the pinned actionlint 1.7.12 and the shellcheck the
ubuntu-latest runner ships (0.9.0-1), the real backlog was 5 findings,
not the 4 the old comment claimed. Three were genuine and are fixed in
the shell:
- cd-cdk.yaml "Publish .NET project" (SC2046): the project path was
interpolated inline and `$(dirname ...)` was unquoted, so a path
containing whitespace split into several arguments. Now passed via
env indirection and quoted, which also removes the last inline
expression interpolation from that step.
- cd-cdk.yaml / ci-python-sam.yaml "Install Python dependencies"
(SC2044 x2): `for req in $(find ...)` word-split and globbed every
path found. Replaced with a NUL-delimited `while read` loop.
Two are deliberate and are suppressed per-line, with the reasoning in a
comment directly above:
- cd-sam.yaml `sam deploy ... $PARAMS` and cd-cdk.yaml
`cdk deploy $STACKS` (SC2086 x2) rely on word-splitting so multiple
parameter overrides / stack selectors reach the CLI as separate argv
entries. Quoting them would collapse each into a single argument and
break every parameterised or multi-stack deploy, so they keep the
unquoted expansion and carry a scoped `# shellcheck disable=SC2086`.
The gate now runs plain `./actionlint` (shellcheck defaults to the
binary on PATH) and prints `shellcheck --version` first, so the check
fails loudly if a future runner image drops it instead of quietly
linting less.
GitHub Actions expressions are substituted into a run body as text
before bash parses it, so a value carrying a quote, a command
substitution, or a newline becomes shell syntax rather than data.
cd-sam.yaml interpolated the parameter-overrides secret straight into
a shell test and an assignment, putting secret material into the script
body. cd-cdk.yaml interpolated the caller-supplied post-deploy-script
input into a bash invocation, which is caller-controlled command
injection rather than secret exposure.
Both now use env-var indirection, matching the STACKS precedent in the
CDK deploy step. PARAM_OVERRIDES is deliberately left unquoted at the
point of use: parameter-overrides carries multiple Key=Value pairs that
must reach sam deploy as separate argv entries, so quoting it would
collapse every override into one argument and break deploys that use
it. POST_DEPLOY_SCRIPT is a single path and is quoted.
Behaviour is otherwise unchanged. An empty parameter-overrides still
produces no --parameter-overrides flag at all, and an empty
post-deploy-script is still skipped by the step-level if condition,
which is a workflow expression and not shell.
ci-python-sam, ci-typescript-cdk and ci-dotnet declared no permissions
at any level, unlike every other workflow here. A reusable workflow that
declares nothing inherits the CALLER's token scopes, and these are
called from deploy repos, so a lint/test/synth job could run holding an
OIDC-mintable token it has no use for. None of the three references
GITHUB_TOKEN, github.token, gh, or any secret, so contents:read is all
they need to check out and build.
Also pass node-version explicitly in the cdk-deploy and ci-node
templates. cicd.md requires callers to pin it so lockfileVersion 3 from
local Node 24 / npm 11 cannot drift from the runner, but no template
did. Only these two targets accept the input; sam-deploy, dotnet-eb,
dependency-review and labeler do not, so they are left alone.
Verified against all 22 callers across the org that none grants
permissions omitting contents:read, so no repo's CI breaks on the
caller-cannot-be-exceeded rule.
Publishes a .NET project, packages the output as a bundle, uploads it,
creates an Elastic Beanstalk application version, and updates an
existing environment using OIDC credentials. It deploys to an
environment; it never creates one.
Two deliberate departures from the existing cd-* reusables:
- A concurrency group keyed on application+environment, with
cancel-in-progress false, so two pushes cannot deploy over each other
and an in-flight deploy is never aborted midway. The existing cd-*
workflows have no concurrency group at all.
- No input or secret is interpolated into a run: body; everything goes
through env-var indirection. The repo's actionlint runs with
shellcheck disabled, so this is a hand-maintained property.
The post-deploy check polls rather than using the CLI waiter
"elasticbeanstalk wait environment-updated": that waiter is hardcoded to
20 attempts x 20s and the CLI cannot extend it, so a slower rolling
deploy would fail the job while the deployment was still healthy. The
timeout is an input instead. The check asserts status, health and the
running version label -- Elastic Beanstalk reports a rolled-back deploy
as a healthy Ready environment running the previous version, so without
the version assertion the job would go green over a failed deploy.
Replaces the malformed cd-dotnet-eb.yaml.yml stub (doubled extension,
empty on:/jobs:). Adds the matching starter template and README entries.
Per the updated handbook convention (engineering-handbook PR #18),
reusable-workflow references use full commit SHA pins with a '# main'
comment instead of the mutable @main branch ref. Templates now ship
pinned so new repos start convention-compliant; Dependabot advances
the pin after instantiation. Commented usage examples in ci-static and
ci-typescript-frontend use the <full-commit-sha> placeholder form.
Callers with an adjudicated accepted-risk advisory (suppressed with
justification in their repo-local .security-review/suppressions.json)
had no way to keep the dependency-review check green when a lockfile
diff touches a package still inside the vulnerable range. Passes the
input straight to actions/dependency-review-action. Default '' is
byte-identical to an unset action input, so existing callers are
unaffected.
First consumer: seahaven-site, allowing GHSA-mh99-v99m-4gvg
(brace-expansion, no in-range fix until @11ty/recursive-copy bumps
minimatch).
Unpinned pip install ruff let ruff 0.16.0 pick up expanded
default lint rules, breaking every caller repo without its
own ruff config. Pinning prevents implicit rule-set changes
on new ruff releases.
Refs: #87
Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
* Add stacks input to cd-cdk for multi-account apps
cdk deploy was hardcoded to --all, which breaks when one CDK app defines
stacks for two AWS accounts: whichever role the job assumed fails on the
other account's stacks. Callers can now pass per-job stack selectors;
default stays --all so existing callers are unaffected.
* Trust seahaven-org-baseline sub on account-baseline deploy role
Transition pair for the repo rename: OIDC sub claims carry the repo full
name, so the renamed repo cannot assume the role until its sub is
trusted. Old sub is removed after a post-rename deploy verifies green.
* Pass stacks selector via env var, not expression interpolation
Defense-in-depth from the security review: expression interpolation
into run: is pre-shell text substitution, so metacharacters in the
input would execute as script. Env-var expansion never re-parses shell
syntax; word-splitting for multiple selectors is preserved.
Add job-level concurrency (cancel-in-progress) to all four ci-python-app
jobs, matching the ci-python-sam idiom. Per-job group keys include
github.job so the parallel jobs in a single run do not share a group.
Replace sam-deploy starter-template stack-name: $default-branch (which
GitHub substitutes to the literal branch name main) with a
REPLACE-ME-stack-name placeholder, and point cfn-role-arn at the real
shared github-cfn-execution-role.
actions/checkout v7.0.0 (2026-06-18) is internally an ESM rebuild plus
one behavioral change: it blocks checking out a fork PR head ref under
pull_request_target / workflow_run (PR #2454). No Sea Haven workflow uses
those triggers, so there is no reachable behavior change. The Node 24
runtime requirement already landed at v6, so v6 -> v7 carries no new
runner requirement. All runners here are GitHub-hosted (ubuntu, macos).
Covers all 16 checkout pins across 12 reusable/standalone workflows plus
the dependency-review workflow-template scaffold. Consumers on @main pick
this up automatically on merge.
Adds ci-typescript-frontend.yaml, a workflow_call reusable CI for bundled
TypeScript SPAs (Vite / React / Vue with vitest + Playwright). Existing
reusable CIs do not fit this shape: ci-static is for plain HTML sites and
ci-typescript-cdk targets CDK infra repos.
The workflow runs as a single `ci` job so callers emit the `ci / ci` status
context the org branch-protection rulesets require. Steps: a Sea Haven
standards gate (required npm scripts present, plus a changed-line guard for
AI-tool footers, hook bypasses, and hardcoded secrets), then format:check,
lint, build, unit tests, and an optional Playwright browser smoke. Every step
past the standards gate is individually toggleable, and string inputs are
passed through env to avoid expression injection.
Documents the workflow in the README reusable-workflows list.
The central label rules assumed a root-level project layout
(lib/**, bin/**, cdk/**, src/**), so monorepos that nest components
under top-level dirs (infra/, web/, mobile/, shared/) matched nothing
for those areas. PRs touching only infra/lib/** or web/** ran the
labeler green but received no label.
Label coverage:
- infra: + 'infra/**' (covers infra/lib, infra/bin, infra/cdk.json)
- app: + 'web/**', 'mobile/**', 'shared/**'
Additions are appended to the existing root paths, so single-project
repos are unaffected; deliberately avoided blanket '**/lib/**' globs
that would mislabel web/src/lib/** as infra.
Hardening rolled in while here:
- Pin actions/labeler to a commit SHA (was the floating @v6 tag)
- Add a per-PR concurrency group with a run_id fallback for non-PR
callers, so rapid pushes cancel superseded label runs
- Broaden 'ci' (.github/actions/**), 'dependencies'
(Directory.Packages.props, yarn.lock, pnpm-lock.yaml, Podfile/.lock)
and 'tests' (JS/TS .test/.spec, pytest test_*.py/conftest,
.NET *Tests.cs, Java *Test.java, Go, Ruby) globs
Caller repos must already have any label a rule can emit; actions/labeler
does not create missing labels. The org 'infra' label was backfilled
across repos separately.
Reusable CI for plain Python apps / locally-run tooling that don't deploy via
SAM or CDK. Beyond ruff lint/format + the conventions audit, it adds a
collect-only import check for a root suite whose live run needs secrets, and an
isolated full pytest run for a self-contained subproject dir (whose tests/
package would collide with the root tests/ under one rootdir).
Emits the org-required `ci / ci` via an aggregator job keyed `ci` that gates on
every other job. actionlint-clean.
Add a thin caller so the .github repo invokes its own reusable
callable-labeler.yaml on pull_request, like every consumer repo does. Without a
caller the workflow_call-only labeler never runs on .github's own PRs (this is
why #61 wasn't auto-labeled). Grants the three permissions the reusable requires
(contents:read, pull-requests:write, issues:write).
The org ruleset requires the 'ci / ci' status check on every repo, but this
.github repo only houses reusable (workflow_call) workflows and emitted no such
check — so every PR sat 'Expected — Waiting for status to be reported' and was
unmergeable, including the open Dependabot bumps (#58, #59, #60).
Add a workflow that runs actionlint (pinned, checksum-verified) over the
workflow files. The ruleset matches the required check against the JOB's
check-run name, so the job is named literally 'ci / ci' to emit that exact
context (a job named 'ci' emits context 'ci', which the UI only cosmetically
shows as 'ci / ci'). shellcheck integration is disabled for now; 4 pre-existing
run-step findings are left for a separate cleanup.
Pre-existing checkov IAM findings in oidc-deploy-roles.yaml are accepted-risk
and suppressed via machine-level security-review config, intentionally NOT
committed to this repo.
Add check-dir + build-command inputs. When build-command is set, run
npm ci + the build, then validate the built output in check-dir (e.g.
_site) instead of repo source. Without this, a site that templates its
HTML (Eleventy etc.) has no source HTML and the checks pass vacuously.
Backward-compatible: defaults (check-dir='.', build-command='') preserve
source-mode behavior for existing callers. build-command is passed via
env to avoid expression injection into the run script.
- ci-static.yaml: reusable CI for static HTML/CSS/JS sites (S3+CloudFront
repos with no build framework). Job 'ci' emits the 'ci / ci' status
context required by the org main-branch ruleset, which static sites
previously could not satisfy (only ci-dotnet/python-sam/typescript-cdk
existed). Checks: htmlhint, JSON-LD validity, sitemap well-formedness,
internal-link/asset resolution, README/.gitignore conventions.
- callable-labeler.yaml: add a 'content' rule (html/css/assets/sitemap/
robots) so static-site PRs get labeled instead of matching nothing.
Add callable-labeler.yaml, a reusable workflow that carries the label
rules inline as the single source of truth and writes them to the runner
at execution time, so caller repos need only a short caller workflow and
no per-repo labeler.yml. Triggered by callers on pull_request (private org
takes no fork PRs); requires contents:read + pull-requests:write +
issues:write on every caller so labeler@v5 can create missing labels.
Remove the workflow-templates/labeler.yml starter it supersedes (no
ruleset workflows-rule or compliance-audit reference depends on it).
Add the repo's own dependabot.yml (github-actions, weekly, grouped
minor+patch) to keep the action pins current per the Pinning Principle.
Retire the org-wide weekly Compliance Audit. Workflow is disabled in the
Actions tab and the schedule trigger is removed so it cannot run
automatically; manual workflow_dispatch is retained for archival only.
Repo compliance now runs via the Claude Code App on PRs + the engineering
handbook. The 19 open 'Compliance audit: violations found' issues it
generated are being closed.
The org ruleset now requires 1 approving review + code-owner review, so
auto-merge can never complete without a human approval — the workflow
only added a no-op (or erroring) check to every PR. Repo-level 'Allow
auto-merge' was also disabled on several repos, making the gh pr merge
--auto call fail benignly. The required-workflow rule referencing this
file was removed from org ruleset 15869156 first.
The group key ci-${{ github.workflow }}-${{ github.ref }} resolves
identically for every job in a caller workflow (github.workflow is the
caller's name in a reusable workflow), so repos calling two reusable CI
workflows from one ci.yaml (e.g. exec-aide python + typescript) had
their jobs cancel each other on every run.
Prefix each group with the reusable workflow's own filename and append
its distinguishing input (source-dirs / working-directory) so sibling
jobs get distinct groups while superseded runs of the same job still
cancel.
Callers grant no explicit permissions, so they pass the org default
read-only token. A reusable workflow cannot request more than its
caller grants, causing startup_failure on every dependency-review run.
dependency-review-action only needs contents: read when not posting
PR comments.
Replace mutable v1 tag references with immutable commit SHAs so a
compromised or force-moved tag cannot inject code into reusable
workflows. Each pin keeps a # v1 comment for readability.
- claude-code-action in compliance-audit.yaml
- ruby/setup-ruby in cd-mobile-ios.yaml (v1 branch)
Superseded CI runs on the same ref keep consuming runners and delay
feedback on the latest push. Add a job-level concurrency group keyed
on github.workflow and github.ref so a new push cancels the in-flight
CI run for that branch.
Concurrency is set at the job level rather than the workflow level
because these are workflow_call reusable workflows: workflow-level
concurrency would resolve github.workflow against the caller's context,
collapsing unrelated callers into one group. cd-* deploy workflows are
intentionally left untouched to avoid cancelling in-flight deploys.
Bump all aws-actions/configure-aws-credentials references to @v6 (org
target) across the reusable CD workflows. v6 is the verified org standard
alongside actions/checkout@v6.
Ref: engineering-handbook cicd.md (workflow standardization).
Bump all actions/checkout references to @v6 (org target). v4 runs on a
node runtime version that is being deprecated; v6 is the verified org
standard alongside configure-aws-credentials@v6.
Ref: engineering-handbook cicd.md (workflow standardization).
Add a reusable callable-dependency-review workflow that runs
actions/dependency-review-action with fail-on-severity: high, and
append a Sea Haven checklist to the PR template covering infra,
secrets, PITR, Slack, Confluence, memory, and cross-review.
The SAM template secret check grepped the 5 lines after `Environment:` for
API_KEY|SECRET|TOKEN|PASSWORD|WEBHOOK. That flags env var *names* like
`SLACK_BOT_TOKEN_SECRET: my-app/slack-token`, whose value is a Secrets
Manager id — i.e. the recommended pattern — so any well-architected
template failed CI.
Match on the value's shape instead: known inline secret formats (Slack
xox* tokens, AWS AKIA keys, GitHub gh*_/PAT tokens, sk- keys, PEM private
keys). Secrets Manager references and intrinsic functions no longer trip
it, while pasted real secrets still fail the build.