Switch fetch-config.sh from a name-prefix batch-get-secret-value
--filters scan to an explicit --secret-id-list (the 28 SECRET_VARS,
chunked at the 20/call cap). An id-list batch authorizes per-secret
ARN, so the instance role's BatchGetSecretValue moves from Resource:*
to the open-swe-<env>/* prefix and the account-wide ListSecrets grant
is dropped entirely. The box can no longer enumerate secret names
account-wide; cross-env value isolation is unchanged (GetSecretValue
was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01.
Also capture each chunk response into a variable and consume the
producer via command substitution so a failed AWS call aborts under
set -e instead of being swallowed by process substitution and
misreported as a missing required var.
Refs: OSWE-IAC-SECRETS-LIST-01
* Align workflows with Sea Haven CI/CD handbook
Bring the workflow suite in line with the handbook: bump
actions/checkout to v7 (Node 24 runtime, already standardized),
kebab-case the two snake_case workflow filenames, and add the
org-standard Labeler caller and Dependency Review gate so vulnerable
or disallowed-license deps and unlabeled PRs are caught automatically.
File renames only — job/check display names are unchanged, so the
promotion gate's REQUIRED_CHECKS and branch-protection required
checks are unaffected.
Refs: INFRA-115
* Drop Agent prefix from CI workflow + job names
The handbook names workflows for what they do (CI, Deploy, Labeler),
not the component they run, matching .github and afterhours-shift-manager.
Rename the suite to CI and its jobs to Lint / Format check / Unit tests,
and keep the promotion gate's REQUIRED_CHECKS in sync.
Refs: INFRA-115
---------
Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
T20 CD safety nets. Two gaps closed before the first real prod deploy:
1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded
dev→main unconditionally. It now hard-gates on check-dev-green.sh:
every check-run on the dev HEAD must be completed+passing AND the
Agent CI suite (lint/format/unit/E2E) must be present+success, or the
promotion blocks (fails safe on a missing/renamed check). The promote
run excludes its OWN check-run by run-id (unforgeable), never by the
mutable name "promote", so a colliding red check cannot hide. Fields
are read with a 0x1F separator so an empty conclusion (every
in_progress check) cannot shift columns. ci.yml now also runs on
push:dev so dev HEAD actually carries that signal (a PR check alone
can be admin-merged past).
2. Rollback + last-good. publish-and-deploy.sh advances
releases/last-good/ only after a successful roll (deploy.sh gates on
`systemctl is-active`), and makes releases/latest/ transactional —
reverting to the prior release if the roll fails so a replaced box
never self-deploys a broken release. New rollback.yml + rollback.sh
re-point latest at last-good (or an explicit sha) and re-fire the
deploy; prod is gated by the `prod` Environment approval, same as a
deploy. The shared fire/wait/aggregate-gate logic is factored into
roll-box.sh (used by both forward and backward rolls).
Least-privilege: drop the unused s3:DeleteObject from the app deploy
role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the
rollback fallback now depends on immutable release history staying
intact. Lifecycle expiry (not CI) handles old-version cleanup.
Gate logic unit-tested (7 cases + jq round-trip). IAM change +
release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks.
Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
Dev now runs against the dedicated seahaven-open-swe-dev org (repo openswe-dev-sandbox),
isolated from the real Sea Haven org. Two changes:
- config-store.ts: iacManagedSsm repo targeting is now per-env — dev =
seahaven-open-swe-dev/openswe-dev-sandbox, prod stays Sea-Haven-Industries/open-swe-pilot.
ALLOWED_GITHUB_ORGS tracks the env owner. Adds a dev-only SEED_USER_MAPPINGS param so the
triggering GitHub login (amoussa1229) resolves and @openswe comments aren't skipped.
- fetch-config.sh: replace the unconditional hard-pin to Sea-Haven-Industries with an owner
GUARD that HONORS the configured owner (OPENSWE_REPO_OWNER override, else the SSM value) but
forces a per-env safe org when the normalized owner is blank or the upstream langchain-ai.
Normalization (lowercase, strip whitespace, first path segment, drop dots) catches
langchain-ai/<repo>, langchain-ai., and case variants without over-blocking legit orgs
(e.g. langchain-ai-fork). Fallback org is per-env so dev can't fall back into the real org.
Reviews: GPT-4.1 cross-review APPROVE (round 1 found a path/dot bypass -> hardened, round 2 clean);
/sh-security-review authz one LOW (env-invariant fallback) -> fixed. Positive org allowlist still
enforced by the app via ALLOWED_GITHUB_ORGS.
The prior fix scoped secretsmanager:BatchGetSecretValue to the env-prefixed secret
ARN, but the live box still got AccessDenied: batch-get-secret-value invoked WITH a
name --filters is a COLLECTION call that AWS authorizes against * (a per-secret ARN
does not satisfy it). Split the statement:
- GetSecretValue + DescribeSecret stay PREFIX-scoped (secret:open-swe-<env>/*) — this
is what gates which secret VALUES the box can read (checked per-secret in the batch).
- BatchGetSecretValue + ListSecrets move to a * operation-level statement (the filtered
collection call + the list action; neither is resource-scopable for this usage).
VALUE isolation preserved (dev box still cannot read prod secret values); only secret
NAME/metadata enumeration is widened. GPT-4.1 IAM cross-review: BLOCK none, FIX none.
Suppression OSWE-IAC-SECRETS-LIST-01 updated; future hardening (explicit --secret-id-list
to drop both * grants) tracked there.
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization
fetch-config.sh materializes the box's .env via
`secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`,
but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue
is a distinct IAM action, so the call was AccessDenied and open-swe.service
crash-looped (no .env written -> ExecStartPre exit 1).
- Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement.
- Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch
call; the API has no resource-level scoping for the list action — fits the role's
stated exception). Secret VALUES stay prefix-scoped; only names are enumerable.
Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW
metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01.
Refs T7/T19 dev bring-up.
* ci: lift Node heap cap for Playwright E2E build (vite OOM)
The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite
bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) —
the same failure fixed for build-artifacts.yml in #19. Set
NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step.
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the
token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated:
CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked
systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed
PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token
never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx
default_server).
Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts);
the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates
unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the
synthesized user-data. Also add .github/scripts/** to build-artifacts paths so
script-only changes trigger a publish.
jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev",
SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B.
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
* ci: path-filtered infra CI/CD with dual OIDC roles + prod approval gate (T18)
Add the /infra half of the combined-repo pipeline (the Python agent keeps ci.yml):
- ci-infra.yml — PR check on infra/** : tsc + jest + cdk synth via the org
reusable ci-typescript-cdk.yaml (working-directory: infra).
- cd-infra.yml — push to dev/main on infra/** (or dispatch):
* job 'ci' (reusable) is the CI-green precondition (deploy needs: ci).
* deploy-dev (ref=dev, NO environment) → cdk deploy OpenSweDevStack,
assuming githubdeploy-open-swe-infra-dev (OIDC sub ref:refs/heads/dev). AUTO.
* deploy-prod (ref=main, environment: prod) → cdk deploy OpenSweProdStack,
assuming githubdeploy-open-swe-infra-prod (OIDC sub environment:prod). The
'prod' Environment's required reviewer is the manual-approval gate.
Deliberately self-contained (NOT the reusable cd-cdk.yaml) because that runs
'cdk deploy --all' — from a single-env push it would deploy the other env + the
shared IAM stack, breaking the per-env boundary. CD targets one stack per env;
the shared open-swe-iam stack is human-gated (T6), never deployed by CD.
Infra CI is enforced at the DEPLOY boundary (deploy jobs need ci), not as a
branch-protection required check — path-filtering a required check would deadlock
app-only PRs. Documented in infra/README.md along with the post-T6 prerequisites
(repo vars AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}; a 'prod' Environment w/ reviewer).
Not active until the IAM roles are applied (T6) — assuming a nonexistent role
just fails closed. App-side CD (S3 artifact + SSM) is T19.
* fix(infra): commit jest.config.js (was ignored by *.js → infra CI used Babel)
The infra/.gitignore *.js rule (for compiled CDK output) silently swept up the
hand-authored jest.config.js, so it was never committed. Local jest passed (file
present in the working tree) but CI's fresh checkout lacked it → jest fell back to
the default Babel transform → 'Cannot use import statement outside a module' on the
TypeScript test. Surfaced now because T18 is the first workflow to run infra jest
in CI. Negate the ignore for this one file and commit it.
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.
Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
(no RETAIN volume — replacement-tolerant; see ami-cache.ts).
userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
(the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
* webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
* site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).
Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.
Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:
- ConfigStore construct (infra/lib/constructs/config-store.ts):
- 28 value-LESS Secrets Manager shells open-swe-<env>/<VAR>
(RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
values are set out-of-band by put-config.sh, never in IaC/state).
- 8 IaC-managed SSM params /open-swe-<env>/<VAR> with real,
stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
- OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
OPENSWE_PUT_<VAR> env indirection; no real values committed).
Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.