Compare commits

...

17 commits

Author SHA1 Message Date
dependabot[bot]
fb42947a9f
Merge f66f86349a into 082768ff51 2026-06-27 00:32:00 -04:00
dependabot[bot]
f66f86349a
chore(deps): bump langgraph-checkpoint
Bumps the uv group with 1 update in the / directory: [langgraph-checkpoint](https://github.com/langchain-ai/langgraph).


Updates `langgraph-checkpoint` from 4.1.0 to 4.1.1
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1)

---
updated-dependencies:
- dependency-name: langgraph-checkpoint
  dependency-version: 4.1.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-06-27 04:31:56 +00:00
Adam Moussa
082768ff51
ci: print stack outputs from CDK --outputs-file (drop describe-stacks) (#26)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
cd-infra reported failure on every successful deploy: the 'Stack outputs' step ran
'aws cloudformation describe-stacks' with the githubdeploy-open-swe-infra-<env> role,
which intentionally lacks cloudformation:DescribeStacks. The cdk deploy itself succeeds
(it reads outputs via the bootstrap cfn-exec role it assumes). Switch to
'cdk deploy --outputs-file cdk-outputs.json' + cat — no extra IAM grant, and the job
goes green on actual deploy success instead of masking real failures behind a red run.
2026-06-26 20:17:17 -04:00
Adam Moussa
f2633f9bd0
fix: don't crash-loop the box when no user mapping is configured (#25)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
seed_store.sh (ExecStartPost) exited 1 when it couldn't resolve a user_mappings
entry, which — with Type=simple — fails the whole unit and crash-loops the box. This
contradicts the script's own OSWE-SEED-03 precedent (the server-not-ready path exits
0 specifically to avoid a restart loop). A missing user mapping is a seeding gap, not
an unhealthy server: the @openswe trigger just won't resolve a commenter, which a
deployment-validation env (dev) does not need.

Make it non-fatal — warn and skip the user_mappings seed; team_settings is still
seeded and the service starts. The empty MAPPINGS array makes the seed loop a no-op.
Set SEED_USER_MAPPINGS / CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL to seed it.

Verified on the dev box: service active, langgraph bound to 127.0.0.1:2024 (loopback),
/ok 200, /healthz 200.
2026-06-26 20:06:53 -04:00
Adam Moussa
a2877d46ff
fix: paginate batch-get-secret-value in fetch-config (secrets dropped past page 1) (#24)
fetch-config materialized only the FIRST page of Secrets Manager results: the AWS CLI
does NOT auto-paginate batch-get-secret-value (the script's comment claiming it does
was wrong; --no-cli-pager only disables the output pager, not API pagination). With 28
secret shells under open-swe-<env>/ the first page returned ~10 items, stranding the
rest on later pages. Required secrets that landed past page 1 (DASHBOARD_JWT_SECRET,
TOKEN_ENCRYPTION_KEY, LANGSMITH_API_KEY_PROD) were silently dropped, tripping the
FAIL-FAST 'missing required var' guard and crash-looping open-swe.service.

Follow NextToken across pages (new batch_get_secrets_tsv helper). Verified against the
live open-swe-dev secrets: now loads all 5 populated secrets (was 2). Same per-record
base64 / exact-prefix / accept_var / emit_var hardening — only the page loop is new.
2026-06-26 19:55:01 -04:00
Adam Moussa
cdcb6625b9
fix: BatchGetSecretValue must be on * for the filtered batch call (#23)
The prior fix scoped secretsmanager:BatchGetSecretValue to the env-prefixed secret
ARN, but the live box still got AccessDenied: batch-get-secret-value invoked WITH a
name --filters is a COLLECTION call that AWS authorizes against * (a per-secret ARN
does not satisfy it). Split the statement:

- GetSecretValue + DescribeSecret stay PREFIX-scoped (secret:open-swe-<env>/*) — this
  is what gates which secret VALUES the box can read (checked per-secret in the batch).
- BatchGetSecretValue + ListSecrets move to a * operation-level statement (the filtered
  collection call + the list action; neither is resource-scopable for this usage).

VALUE isolation preserved (dev box still cannot read prod secret values); only secret
NAME/metadata enumeration is widened. GPT-4.1 IAM cross-review: BLOCK none, FIX none.
Suppression OSWE-IAC-SECRETS-LIST-01 updated; future hardening (explicit --secret-id-list
to drop both * grants) tracked there.
2026-06-26 19:44:12 -04:00
Adam Moussa
cfbdcda9b7
fix: instance role BatchGetSecretValue + ListSecrets for .env materialization (#22)
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization

fetch-config.sh materializes the box's .env via
`secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`,
but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue
is a distinct IAM action, so the call was AccessDenied and open-swe.service
crash-looped (no .env written -> ExecStartPre exit 1).

- Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement.
- Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch
  call; the API has no resource-level scoping for the list action — fits the role's
  stated exception). Secret VALUES stay prefix-scoped; only names are enumerable.

Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW
metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01.

Refs T7/T19 dev bring-up.

* ci: lift Node heap cap for Playwright E2E build (vite OOM)

The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite
bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) —
the same failure fixed for build-artifacts.yml in #19. Set
NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step.
2026-06-26 19:29:50 -04:00
Adam Moussa
5fa132205b
fix(deploy): use %%...%% for CDK user-data tokens (don't collide with @@ sed) (#21)
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the
token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated:
CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked
systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed
PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token
never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx
default_server).

Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts);
the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates
unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the
synthesized user-data. Also add .github/scripts/** to build-artifacts paths so
script-only changes trigger a publish.

jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev",
SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B.
2026-06-26 19:06:37 -04:00
Adam Moussa
9e2b215f08
fix(ci): avoid tar|grep -q SIGPIPE false-failure in package-artifacts (#20)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
`tar -tzf app.tar.gz | grep -qx` under set -o pipefail fails the pipeline when
grep -q matches and exits early (SIGPIPEs tar -> 'write error' -> non-zero), a
false 'missing agent/server.py'. List the archive once into a var, then grep the
var. Same fix for the secret-guard pipe (which was also silently broken).
2026-06-26 18:57:59 -04:00
Adam Moussa
2765b65acd
fix(ci): raise Node heap for the SPA build (vite OOM'd at ~2 GB) (#19)
The build-artifacts SPA build hit Node's default ~2 GB heap cap and aborted
(JS heap OOM, exit 134) — the same memory-hungry vite build that needed an 8 GB
swapfile on-box. The runner has ~16 GB, so set NODE_OPTIONS=--max-old-space-size=8192
on both build steps.
2026-06-26 18:54:10 -04:00
Adam Moussa
404b3f6f75
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18)
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)

Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.

deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
  - the file provisioner failed uploading the templates dir ('scp: …: Is a
    directory') — a trailing-slash contents-upload needs the dest dir to exist;
    added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
    dest trailing slash.
  - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
    environment_vars never reached provision.sh (which runs under set -u and
    aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.

infra:
  - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
    from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
    exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
    now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
    discipline docs.
  - app-service.ts: machineImage → bakedOpenSweArm64().
  - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
  - cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
  - README: Baked AMI + EBS-replacement-discipline section.

tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.

NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).

* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021

Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.

* fix(deploy): GitHub App + Slack required for prod only, not dev

Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.

Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).

* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)

Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.

Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
  SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
  abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
  /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
  baked open-swe-base-arm64 AMI (folds in the held #16).

IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
  open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
  AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
  is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.

Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
  app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
  at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
  target is healthy even before the first release); deploy.sh is base64-rendered
  by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
  first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).

CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
  ui/.output/public -> spa.tar.gz), package the Python source via git archive
  (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
  the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
  push dev -> dev (auto); push main -> prod (env "prod" approval gate).

Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.

* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard

Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:

- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
  never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
  treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
  failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
  not CommandInvocations[0], so a partial failure across the brief 2-instance
  replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
  role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
  (defense in depth over .gitignore; scoped to data extensions so *_credentials.py
  source is not a false positive — verified against the real tree).

Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).

shellcheck/tsc/jest(16) clean; both stacks synth offline.

* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard

The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)

- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
  ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
  GroupDescription + rule descriptions are pure ASCII, so this fails the build
  instead of a deploy next time.

jest 18/18; tsc clean.

* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)

The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.

jest 18/18; tsc clean.

* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit

The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.

jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
Adam Moussa
70319cac0d
ci: path-filtered infra CI/CD + dual OIDC roles + prod approval gate (T18) (#15)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
* ci: path-filtered infra CI/CD with dual OIDC roles + prod approval gate (T18)

Add the /infra half of the combined-repo pipeline (the Python agent keeps ci.yml):

- ci-infra.yml — PR check on infra/** : tsc + jest + cdk synth via the org
  reusable ci-typescript-cdk.yaml (working-directory: infra).
- cd-infra.yml — push to dev/main on infra/** (or dispatch):
    * job 'ci' (reusable) is the CI-green precondition (deploy needs: ci).
    * deploy-dev  (ref=dev, NO environment)  → cdk deploy OpenSweDevStack,
      assuming githubdeploy-open-swe-infra-dev  (OIDC sub ref:refs/heads/dev). AUTO.
    * deploy-prod (ref=main, environment: prod) → cdk deploy OpenSweProdStack,
      assuming githubdeploy-open-swe-infra-prod (OIDC sub environment:prod). The
      'prod' Environment's required reviewer is the manual-approval gate.

Deliberately self-contained (NOT the reusable cd-cdk.yaml) because that runs
'cdk deploy --all' — from a single-env push it would deploy the other env + the
shared IAM stack, breaking the per-env boundary. CD targets one stack per env;
the shared open-swe-iam stack is human-gated (T6), never deployed by CD.

Infra CI is enforced at the DEPLOY boundary (deploy jobs need ci), not as a
branch-protection required check — path-filtering a required check would deadlock
app-only PRs. Documented in infra/README.md along with the post-T6 prerequisites
(repo vars AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}; a 'prod' Environment w/ reviewer).

Not active until the IAM roles are applied (T6) — assuming a nonexistent role
just fails closed. App-side CD (S3 artifact + SSM) is T19.

* fix(infra): commit jest.config.js (was ignored by *.js → infra CI used Babel)

The infra/.gitignore *.js rule (for compiled CDK output) silently swept up the
hand-authored jest.config.js, so it was never committed. Local jest passed (file
present in the working tree) but CI's fresh checkout lacked it → jest fell back to
the default Babel transform → 'Cannot use import statement outside a module' on the
TypeScript test. Surfaced now because T18 is the first workflow to run infra jest
in CI. Negate the ignore for this one file and commit it.
2026-06-26 16:11:22 -04:00
Adam Moussa
fcbdfb67aa
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.

Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
  in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
  (no RETAIN volume — replacement-tolerant; see ami-cache.ts).
  userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
  via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
  (the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
  loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
    * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
    * site     (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
  Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
  or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).

Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.

Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
Adam Moussa
aef6b26912
feat: Secrets Manager + SSM config store for open-swe (T11) (#10)
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:

- ConfigStore construct (infra/lib/constructs/config-store.ts):
  - 28 value-LESS Secrets Manager shells  open-swe-<env>/<VAR>
    (RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
    values are set out-of-band by put-config.sh, never in IaC/state).
  - 8 IaC-managed SSM params  /open-swe-<env>/<VAR>  with real,
    stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
    ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
  - OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
    own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
  literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
  OPENSWE_PUT_<VAR> env indirection; no real values committed).

Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.
2026-06-26 16:07:23 -04:00
Adam Moussa
434d0ad80c
feat(deploy): AWS-sourced fetch-config + seed_store + rotation docs (PR#7) (#8)
PR#7 of the AWS migration. deploy/seahaven/: fetch-config.sh materializes a
service-user-owned 0600 tmpfs .env from Secrets Manager + SSM (fail-fast);
seed_store.sh reseeds the in-memory store; ROTATION.md.

Incorporates T5 /sh-security-review fixes:
- seed_store no longer bash-sources the .env (closes the SH-INJ-001 RCE); uses a
  non-eval reader, jq --arg JSON bodies, and a loopback-pinned BASE.
- fetch-config: .env owned by the openswe service user (app no longer runs as
  root); DEFAULT_REPO_OWNER hard-pinned; key-identifier validation + flat-namespace
  collision detection; dropped SSM --recursive.
shellcheck + bash -n clean.
2026-06-26 15:06:41 -04:00
Adam Moussa
ce9e5fb51b
feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64,
uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data
(userDataCausesReplacement rationale), systemd unit + nginx + CW templates.

Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0);
nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre
runs fetch-config as root (+) and passes the env arg; the app runs as the
unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
2026-06-26 15:06:36 -04:00
Adam Moussa
92fd886076
feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6)
PR#1 of the AWS migration. CDK TypeScript app under /infra: stacks open-swe-iam
(four per-env GitHub-OIDC deploy roles) + open-swe-dev/-prod (per-env EC2 instance
role + AMI cache). aws-cdk-lib pinned exact 2.260.0; kebab naming Aspect + 15 tests.

IAM is synth-only (NOT deployed). Cleared the Phase-1 security gates:
- T4 GPT-4.1 IAM cross-review (StringEquals trust; cdk-hnb659fds-* wildcard kept as
  org convention; AWS-RunShellScript timeboxed to T19).
- T5 /sh-security-review: deploy roles split PER-ENV with env-scoped OIDC trust
  (dev=branch ref+tag dev, prod=environment:prod+tag prod) so a dev token cannot
  reach prod; re-verified block:false.
2026-06-26 15:06:30 -04:00
43 changed files with 8376 additions and 56 deletions

52
.github/scripts/package-artifacts.sh vendored Executable file
View file

@ -0,0 +1,52 @@
#!/usr/bin/env bash
# Package the two release artifacts (run from the repo root by build-artifacts.yml):
#
# spa.tar.gz = CONTENTS of the built SPA dir (ui/.output/public/*), so it extracts
# straight into the nginx web root with _shell.html at the root.
# app.tar.gz = the Python source the box runs `uv sync` against. `git archive`
# gives a clean tree (no node_modules, no .venv, no local cruft);
# ui/ is intentionally excluded (it ships as spa.tar.gz).
set -euo pipefail
SPA_DIR="ui/.output/public"
[ -f "${SPA_DIR}/_shell.html" ] || {
echo "ERROR: SPA build output missing ${SPA_DIR}/_shell.html (did 'bun run build' run?)" >&2
exit 1
}
echo "==> spa.tar.gz from ${SPA_DIR}"
tar -C "${SPA_DIR}" -czf spa.tar.gz .
echo "==> app.tar.gz from source (git archive HEAD)"
git archive --format=tar.gz -o app.tar.gz HEAD \
agent deploy langgraph.json pyproject.toml uv.lock README.md
# List the archive ONCE into a variable. (`tar -tzf ... | grep -q ...` is unsafe
# under `set -o pipefail`: grep -q exits on first match, SIGPIPEs tar -> "write
# error" -> the pipeline reports non-zero even though grep succeeded, a false
# failure. Listing once avoids the pipe entirely.)
APP_LIST="$(tar -tzf app.tar.gz)"
# Sanity: the box's `uv sync --frozen` needs pyproject.toml + uv.lock at the root,
# and the package itself (agent/). Fail loudly here rather than on the box.
for required in pyproject.toml uv.lock agent/server.py langgraph.json; do
printf '%s\n' "${APP_LIST}" | grep -qx "${required}" || {
echo "ERROR: app.tar.gz is missing ${required}" >&2
exit 1
}
done
# Fail-closed secret guard: the source is git-archived wholesale, so reject the
# release if a secret-shaped FILE slipped into the tracked tree (defense in depth
# on top of .gitignore — the artifact lands on the box + in S3). Scoped to data
# extensions so credential-handling *source* (e.g. team_credentials.py) is not a
# false positive.
SECRET_RE='(^|/)(\.env(\..+)?|id_rsa|.*\.(pem|key|p12|pfx)|.*(secret|credential|password|token)s?\.(json|ya?ml|txt|env|ini|cfg))$'
if printf '%s\n' "${APP_LIST}" | grep -qiE "${SECRET_RE}"; then
echo "ERROR: app.tar.gz contains a secret-shaped file — refusing to publish:" >&2
printf '%s\n' "${APP_LIST}" | grep -iE "${SECRET_RE}" >&2
exit 1
fi
echo "==> artifacts:"
ls -la spa.tar.gz app.tar.gz

75
.github/scripts/publish-and-deploy.sh vendored Executable file
View file

@ -0,0 +1,75 @@
#!/usr/bin/env bash
# Publish the packaged artifacts to the env's S3 bucket and roll the box to them.
# Run by build-artifacts.yml AFTER aws creds are configured (env: ENV, BUCKET,
# DEPLOY_DOC). Each release is stored immutably under releases/<sha>/ and mirrored
# to releases/latest/ (what the box's deploy.sh pulls).
#
# The deploy is fired by TAG (project=open-swe,env=<env>), which is exactly what
# the app deploy role's tag-scoped ssm:SendCommand allows — so this needs no
# ec2:DescribeInstances and no instance id up front.
set -euo pipefail
: "${ENV:?}" "${BUCKET:?}" "${DEPLOY_DOC:?}"
SHA="${GITHUB_SHA:?}"
[ -f app.tar.gz ] && [ -f spa.tar.gz ] || { echo "ERROR: artifacts not built" >&2; exit 1; }
echo "==> upload release ${SHA} to s3://${BUCKET}/releases/${SHA}/"
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "${f}" "s3://${BUCKET}/releases/${SHA}/${f}"
done
echo "==> mirror to releases/latest/ (server-side copy)"
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "s3://${BUCKET}/releases/${SHA}/${f}" "s3://${BUCKET}/releases/latest/${f}"
done
echo "==> fire ${DEPLOY_DOC} via SSM (tag-targeted: project=open-swe, env=${ENV})"
CMD_ID="$(aws ssm send-command \
--document-name "${DEPLOY_DOC}" \
--targets "Key=tag:project,Values=open-swe" "Key=tag:env,Values=${ENV}" \
--comment "release ${SHA}" \
--query 'Command.CommandId' --output text)"
echo "command: ${CMD_ID}"
echo "==> wait for the deploy to finish"
STATUS="Pending"
IID=""
for _ in $(seq 1 60); do
sleep 10
# CommandInvocations is empty until SSM registers the target invocation.
IID="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].InstanceId' --output text 2>/dev/null || echo None)"
[ -z "${IID}" ] || [ "${IID}" = "None" ] && continue
STATUS="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].Status' --output text 2>/dev/null || echo Pending)"
case "${STATUS}" in
Success | Failed | Cancelled | TimedOut) break ;;
esac
done
if [ -z "${IID}" ] || [ "${IID}" = "None" ]; then
echo "ERROR: no box picked up the deploy command (is a running open-swe ${ENV} box registered with SSM?)" >&2
exit 1
fi
echo "==> deploy.sh output from ${IID}:"
echo "----- stdout -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardOutputContent' --output text || true
echo "----- stderr -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardErrorContent' --output text || true
# Gate on the AGGREGATE command status (Success only if EVERY targeted invocation
# succeeded), not CommandInvocations[0] — during a userDataCausesReplacement window
# two instances can briefly share the project/env tags, and a partial failure on
# the other instance must not be reported as success.
TARGETS="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].TargetCount' --output text 2>/dev/null || echo 1)"
[ "${TARGETS}" = "1" ] || echo "WARNING: deploy fanned out to ${TARGETS} instances (expected 1)"
AGG="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].Status' --output text 2>/dev/null || echo Failed)"
echo "==> aggregate deploy status: ${AGG} (across ${TARGETS} target(s))"
[ "${AGG}" = "Success" ] || { echo "ERROR: deploy did not succeed (${AGG})" >&2; exit 1; }
echo "==> ${ENV} rolled to release ${SHA}"

123
.github/workflows/build-artifacts.yml vendored Normal file
View file

@ -0,0 +1,123 @@
name: Build & publish app artifacts
# T7 + T19 — build the release (SPA + Python source) and publish it to the per-env
# S3 artifact bucket, then roll the box to it.
#
# push to dev → publish to open-swe-dev-assets → deploy dev box (AUTO)
# push to main → publish to open-swe-prod-assets → deploy prod box (manual approval: env "prod")
#
# Two artifacts (the box's deploy.sh pulls both from releases/latest/):
# spa.tar.gz = the built dashboard SPA (vite -> ui/.output/public). Built HERE
# (not on the box) — the build is memory-heavy and the box is small.
# app.tar.gz = the Python source tree (NO ui/, NO .venv). The box runs
# `uv sync` to build a native-ARM64 venv at the real runtime path.
#
# Each release is uploaded under releases/<sha>/ (immutable, auditable) AND mirrored
# to releases/latest/ (what the box pulls). Then the open-swe-<env>-deploy SSM
# document is fired (tag-scoped to project=open-swe,env=<env>) to roll the box.
#
# OIDC subject alignment (matches the per-env app-role trust in infra/lib/config.ts):
# - publish-dev declares NO `environment:` → sub = repo:…:ref:refs/heads/dev
# - publish-prod declares `environment: prod` → sub = repo:…:environment:prod
# (also triggers the prod Environment's required-reviewer approval gate).
#
# Prerequisites:
# - repo variables AWS_DEPLOY_ROLE_APP_DEV / AWS_DEPLOY_ROLE_APP_PROD = the
# githubdeploy-open-swe-app-<env> role ARNs (open-swe-iam CfnOutputs).
# - the open-swe-<env> stack deployed (creates the bucket + the SSM deploy doc).
permissions:
contents: read
on:
push:
branches: [dev, main]
paths:
- "agent/**"
- "ui/**"
- "deploy/**"
- "langgraph.json"
- "pyproject.toml"
- "uv.lock"
- ".github/workflows/build-artifacts.yml"
- ".github/scripts/**"
workflow_dispatch:
concurrency:
# one publish+deploy per branch at a time; never cancel an in-flight release.
group: build-artifacts-${{ github.ref }}
cancel-in-progress: false
jobs:
publish-dev:
name: Publish + deploy (dev)
if: ${{ github.ref == 'refs/heads/dev' }}
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
id-token: write
contents: read
env:
ENV: dev
BUCKET: open-swe-dev-assets
DEPLOY_DOC: open-swe-dev-deploy
steps:
- uses: actions/checkout@v6
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- name: Build SPA (vite -> ui/.output/public)
working-directory: ui
# vite's bundle exceeds Node's default ~2 GB heap (the build that needed an
# 8 GB swapfile on-box); the runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"
run: |
bun install --frozen-lockfile
bun run build
- name: Package artifacts
run: bash .github/scripts/package-artifacts.sh
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_DEV }}
aws-region: us-east-1
- name: Publish to S3 + roll the box
run: bash .github/scripts/publish-and-deploy.sh
publish-prod:
name: Publish + deploy (prod)
if: ${{ github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 30
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam). Also
# makes the OIDC sub …:environment:prod (matches the prod app-role trust).
environment: prod
permissions:
id-token: write
contents: read
env:
ENV: prod
BUCKET: open-swe-prod-assets
DEPLOY_DOC: open-swe-prod-deploy
steps:
- uses: actions/checkout@v6
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- name: Build SPA (vite -> ui/.output/public)
working-directory: ui
# vite's bundle exceeds Node's default ~2 GB heap (the build that needed an
# 8 GB swapfile on-box); the runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"
run: |
bun install --frozen-lockfile
bun run build
- name: Package artifacts
run: bash .github/scripts/package-artifacts.sh
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_PROD }}
aws-region: us-east-1
- name: Publish to S3 + roll the box
run: bash .github/scripts/publish-and-deploy.sh

124
.github/workflows/cd-infra.yml vendored Normal file
View file

@ -0,0 +1,124 @@
name: Infra CD
# Path-filtered CDK deploy for /infra, per env, OIDC-only (no static keys).
#
# push to dev → CI (tsc+jest+synth) → deploy OpenSweDevStack (AUTO, CI-green-gated)
# push to main → CI → deploy OpenSweProdStack (manual approval: env "prod")
#
# Why this is NOT the reusable cd-cdk.yaml: that workflow runs `cdk deploy --all`,
# which would deploy ALL THREE stacks (incl. the OTHER env + the shared IAM stack)
# from a single-env push — breaking the per-env dev/prod boundary. So we target one
# stack explicitly per env. (Infra CI still uses the reusable ci-typescript-cdk.)
#
# The shared IAM stack (open-swe-iam — owns BOTH envs' OIDC deploy roles) is
# intentionally NOT deployed here: it is a privileged, human-gated apply (T6), so a
# routine dev push can never alter prod's deploy role.
#
# OIDC subject alignment (must match the per-env trust in infra/lib/config.ts):
# - deploy-dev declares NO `environment:` → token sub = repo:…:ref:refs/heads/dev,
# which is exactly what githubdeploy-open-swe-infra-dev trusts.
# - deploy-prod declares `environment: prod` → token sub = repo:…:environment:prod,
# which githubdeploy-open-swe-infra-prod trusts AND which triggers the GitHub
# Environment's required-reviewer (manual approval) gate.
#
# Prerequisites (post-T6, when the roles exist):
# - repo variables AWS_DEPLOY_ROLE_INFRA_DEV / AWS_DEPLOY_ROLE_INFRA_PROD = the
# githubdeploy-open-swe-infra-<env> role ARNs (open-swe-iam CfnOutputs).
# - a GitHub Environment named "prod" with Adam as a required reviewer.
permissions:
contents: read
on:
push:
branches: [dev, main]
paths:
- "infra/**"
- ".github/workflows/cd-infra.yml"
workflow_dispatch:
concurrency:
# one infra deploy per branch at a time; never cancel an in-flight deploy.
group: cd-infra-${{ github.ref }}
cancel-in-progress: false
jobs:
# CI-green precondition — re-run tsc + jest + synth on the pushed commit before
# any deploy. A failure here blocks the deploy jobs (needs: ci).
ci:
name: Infra CI (pre-deploy)
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@main
with:
node-version: "24"
working-directory: infra
cache-dependency-path: infra/package-lock.json
run-typecheck: true
run-tests: true
run-cdk-synth: true
deploy-dev:
name: Deploy open-swe-dev
needs: ci
if: ${{ github.ref == 'refs/heads/dev' }}
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
id-token: write
contents: read
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v4
with:
node-version: "24"
cache: npm
cache-dependency-path: infra/package-lock.json
- name: Install deps
working-directory: infra
run: npm ci
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_INFRA_DEV }}
aws-region: us-east-1
- name: CDK deploy (dev only)
working-directory: infra
# --outputs-file lets us print the stack outputs from CDK's own result
# (the deploy role intentionally lacks cloudformation:DescribeStacks; CDK
# gets outputs via the bootstrap cfn-exec role it assumes, so no extra grant).
run: npx cdk deploy OpenSweDevStack --require-approval never --outputs-file cdk-outputs.json
- name: Stack outputs
working-directory: infra
run: cat cdk-outputs.json
deploy-prod:
name: Deploy open-swe-prod
needs: ci
if: ${{ github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 30
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam).
# Also makes the OIDC sub …:environment:prod (matches the prod role trust).
environment: prod
permissions:
id-token: write
contents: read
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v4
with:
node-version: "24"
cache: npm
cache-dependency-path: infra/package-lock.json
- name: Install deps
working-directory: infra
run: npm ci
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_INFRA_PROD }}
aws-region: us-east-1
- name: CDK deploy (prod only)
working-directory: infra
# See deploy-dev: --outputs-file avoids needing cloudformation:DescribeStacks.
run: npx cdk deploy OpenSweProdStack --require-approval never --outputs-file cdk-outputs.json
- name: Stack outputs
working-directory: infra
run: cat cdk-outputs.json

26
.github/workflows/ci-infra.yml vendored Normal file
View file

@ -0,0 +1,26 @@
name: Infra CI
# Path-filtered CI for the /infra CDK app (TypeScript). The existing "Agent CI"
# (ci.yml) covers the Python agent; this adds tsc + jest + cdk synth for /infra so
# infra changes are gated on a PR the same way. Runs only when /infra changes.
permissions:
contents: read
on:
pull_request:
paths:
- "infra/**"
- ".github/workflows/ci-infra.yml"
jobs:
infra-ci:
name: Infra CI (tsc + jest + synth)
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@main
with:
node-version: "24"
working-directory: infra
cache-dependency-path: infra/package-lock.json
run-typecheck: true
run-tests: true
run-cdk-synth: true

View file

@ -69,6 +69,11 @@ jobs:
# ui/ SPA. The fake LLM/GitHub/Slack boundaries need no secrets.
- name: Run E2E
working-directory: tests/e2e
# Playwright's globalSetup runs the real `bun run build`, whose vite bundle
# exceeds Node's default ~2 GB heap (same OOM fixed in build-artifacts.yml).
# The runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"
run: npx playwright test
- name: Upload Playwright report
if: ${{ !cancelled() }}

5
.gitignore vendored
View file

@ -68,4 +68,7 @@ __pycache__/
*.egg-info/
.eggs/
#
# Local working docs (gitignored — survives upstream merges, never pushed)
TODO.md
# infra/cdk-outputs.json

View file

@ -0,0 +1,24 @@
{
"suppressions": [
{
"id": "OSWE-IAC-AUDIT-01",
"title": "Dev-branch infra OIDC role can assume the account-wide CDK cfn-exec admin role (cdk-hnb659fds-*), a path to mutating prod",
"file": "infra/lib/constructs/github-deploy-roles.ts",
"severity": "high",
"status": "confirmed",
"suppression_justification": "PRE-EXISTING and NOT introduced or worsened by the T7+T19 change (the assets bucket / app-role PutObject / SSM deploy doc). This is the known single-account-wide CDK cfn-exec residual already documented in infra/lib/config.ts:31-34 and the github-deploy-roles.ts construct comment, accepted at the T4 GPT-4.1 IAM cross-review and the v5 plan-review. WHO can assume each env's infra role is exact-subject scoped (StringEquals on the dev ref / prod environment); the residual is the shared account-wide cfn-exec-role that every env's infra role can reach. The tracked fix is per-env CDK bootstrap qualifiers so each env's infra role assumes its own env-scoped cfn-exec-role. Suppressed for THIS change's gate because it is out-of-diff and unchanged; surfaced to Adam for scheduling the per-env-bootstrap remediation.",
"owner": "adam@seahavenind.com",
"added": "2026-06-26"
},
{
"id": "OSWE-IAC-SECRETS-LIST-01",
"title": "EC2 instance role grants BatchGetSecretValue + ListSecrets on \"*\" (operation-level; secret-NAME enumeration account-wide)",
"file": "infra/lib/constructs/instance-role.ts",
"severity": "low",
"status": "confirmed",
"suppression_justification": "ACCEPTED LOW residual, metadata-only. fetch-config.sh materializes the .env via `batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`. With a name FILTER, both BatchGetSecretValue (a collection call) and ListSecrets are authorized by AWS against `*`, NOT a per-secret ARN — a prefix-scoped ARN AccessDenies the call (confirmed empirically on i-0af4e03e8bf70e6c3). So the two `*` grants are operation-level, not value-level. Secret VALUES remain strictly gated by the PREFIX-scoped GetSecretValue/DescribeSecret on secret:open-swe-<env>/* (GetSecretValue is checked per-secret even within the batch), so cross-env VALUE isolation is preserved; only NAMES/tags/descriptions are enumerable, within Sea Haven's own single-tenant account 328440206208. Confirmed by GPT-4.1 IAM cross-review (BLOCK: none) and the iac-iam detector (one low residual, no critical/high). Future hardening to eliminate BOTH `*` grants: switch fetch-config.sh to an explicit `--secret-id-list` (no filter), which lets BatchGetSecretValue be prefix-scoped and needs no ListSecrets.",
"owner": "adam@seahavenind.com",
"added": "2026-06-26"
}
]
}

167
deploy/ami/README.md Normal file
View file

@ -0,0 +1,167 @@
# Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(`open-swe-dev` / `open-swe-prod`) in the Open SWE → AWS migration. Builds an
**ARM64 (Graviton) Ubuntu 24.04 LTS** base AMI and provisions the box on first
boot with the stock `langgraph dev` runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo `TODO.md` ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, `seahaven-vpc` **private subnet + NAT**,
inbound **only from the ALB SG**. Runtime is **stock `langgraph dev`** (in-memory
store, `--no-reload`) + nginx + systemd. The box has **no git auth** — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), **not** on the box, so the old 8 GB-swapfile OOM hack is gone.
## Files
| Path | Purpose |
|---|---|
| `open-swe-base.pkr.hcl` | Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
| `scripts/provision.sh` | Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
| `user-data.sh` | First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
| `templates/open-swe.service` | systemd unit TEMPLATE (`@@tokens@@` rendered at boot). |
| `templates/open-swe.nginx.conf` | nginx site TEMPLATE (dashboard SPA + scoped `/dashboard/api/` proxy). |
| `templates/amazon-cloudwatch-agent.json` | CW agent config TEMPLATE — **30-day log retention**. |
`deploy/seahaven/fetch-config.sh` and `deploy/seahaven/seed_store.sh` are owned by
the parallel T10 work and ship **inside the app artifact**; this AMI wires them in
but does not author them (see "Integration contract" below).
## Build the AMI
```bash
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
```
Builds in account **328440206208 / us-east-1**. Source = latest Canonical Ubuntu
24.04 (Noble) **arm64** AMI (`source_ami_filter`, owner `099720109477`). Build host
is `t4g.medium` (ARM64). The output AMI is tagged:
```
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
```
Pinned versions live in the template `variable` defaults (`uv_version`,
`python_version`, `node_major`, the CW-agent / awscli URLs) and the
`required_plugins` block (`amazon` 1.3.6) — bump deliberately.
## AMI → `cdk.context.json` pinning contract
The CDK stacks in `/infra` (owned by T3/T12) consume the AMI **by id, pinned in the
committed `infra/cdk.context.json`** — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached `MachineImage.lookup` resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
1. `packer build` prints the new AMI id (and tags it `open-swe-base-arm64`).
2. CDK looks the AMI up with **`cachedInContext: true`** (e.g.
`MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })`),
which writes the resolved id into `infra/cdk.context.json`.
3. **`infra/cdk.context.json` is committed.** From then on every synth/deploy uses
the pinned id — no surprise replacement when a newer AMI exists.
4. To adopt a new AMI: `cdk context --reset <ami-lookup-key>` (or edit the pinned
value), commit the change, and review the cdk-diff — the PR will show
"requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (`project_open_swe_migration`) per the
"memory updated for AMI id" build criterion.
## `userDataCausesReplacement` rationale
`user-data.sh` is **provisioning-only** — it runs once at first boot and never
carries durable runtime config. CDK sets **`userDataCausesReplacement: true`** so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
**fresh on every service start** by `fetch-config.sh` (ExecStartPre) — changing a
secret or SSM value needs only a `systemctl restart open-swe.service`, not a
replacement.
## EBS discipline (binding — `feedback_inline_ebs_volumes`)
**The box holds no durable state of its own:**
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs `.env` | re-fetched at boot |
| app code + SPA | S3 `open-swe-<env>-assets` | re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by `seed_store.sh` | re-seeded at boot |
| logs | CloudWatch (30-day) — **not** a CFN resource in the stack | survive replacement |
→ **No local-only durable state ⇒ no standalone RETAIN volume is needed.** The root
volume is disposable; there is intentionally no inline data `blockDevices` to lose.
**Even so, snapshot before any replacing deploy.** Per the operational guard, before
merging/deploying any change that REPLACES the instance (`userDataCausesReplacement`,
AMI bump, instance-type change):
1. Enumerate the instance's volumes and assert **"no local-only durable state"**
(the table above is the checklist).
2. Take an **EBS snapshot of the root volume and WAIT for `state=completed`** before
letting the deploy proceed. Keep it as insurance; delete after a grace period.
3. Confirm the CloudWatch log groups are **not** CFN-managed in the stack so history
survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance
criterion). This is the enforced version — not just an assertion in the runbook.
## Integration contract (T10 — `fetch-config.sh` + `seed_store.sh`)
Both ship in the app artifact under `deploy/seahaven/` and are wired into the unit:
- **`fetch-config.sh`** (ExecStartPre, runs as `openswe`): reads `/etc/open-swe/boot.env`
(`OPENSWE_ENV`, `AWS_REGION`, `SECRETS_PREFIX=open-swe-<env>`, `SSM_PREFIX=/open-swe-<env>`,
`ENV_FILE=/run/open-swe/.env`), pulls Secrets Manager `open-swe-<env>/*` + SSM
`/open-swe-<env>/*`, and writes:
- `/run/open-swe/.env` (**0600, tmpfs**, secret-bearing app env incl. the multiline
GitHub App PEM) — loaded by langgraph/dotenv via the `${APP_DIR}/.env` symlink.
- `/run/open-swe/seed.env` (**0600, tmpfs**, simple `OPENSWE_*` vars only:
`OPENSWE_DEFAULT_REPO`, `OPENSWE_OWNER_LOGIN`, `OPENSWE_OWNER_EMAIL`, model ids) —
loaded by systemd `EnvironmentFile` so `seed_store.sh` (ExecStartPost) has them.
- It must **fail-fast** (non-zero exit) if any required value is missing, so the
unit never starts half-configured.
- **`seed_store.sh`** (ExecStartPost): existing script, reseeds `team_settings/default`
+ `user_mappings/<login>` into the in-memory store after each start.
## Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
- [ ] `cloud-init status --wait` → `done`; `/var/log/open-swe-user-data.log` ends with
"user-data done" and shows the S3 pulls + service starts.
- [ ] `systemctl is-active open-swe.service` → `active`. (If it failed, check
`ExecStartPre`/`fetch-config.sh` — fail-fast means missing config = failed unit.)
- [ ] **fetch-config fail-fast works:** `/run/open-swe/.env` exists, owner `openswe`,
mode `0600`, on tmpfs (`findmnt /run/open-swe`); `seed.env` present.
- [ ] `curl -fsS http://127.0.0.1:2024/ok` → `200` (raw LangGraph health).
- [ ] `systemctl is-active nginx` → `active`; `curl -fsS http://127.0.0.1/healthz` →
`200`; `curl -s http://127.0.0.1/threads` returns the SPA shell, **not** JSON
(proves the agent API is not proxied — the security boundary holds).
- [ ] `seed_store: done` in the journal / app.log (store reseeded).
- [ ] CloudWatch: log groups `/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`
exist with **30-day** retention and are receiving events.
- [ ] **No swapfile** (`swapon --show` empty) — the on-box SPA build is gone.
- [ ] From the ALB only: dashboard host serves the SPA; `hooks` host reaches
`/webhooks/*` on :2024 and nothing else (raw API paths hit the ALB default, not
the box).
## Assumptions
- **Artifact bucket** `open-swe-<env>-assets` (T7), with objects
`${ARTIFACT_PREFIX}/app.tar.gz` (Python app incl. `deploy/seahaven/` and a prebuilt
arm64 `.venv`) and `${ARTIFACT_PREFIX}/spa.tar.gz` (built SPA → `/var/www/open-swe`).
`ARTIFACT_PREFIX` defaults to `releases/latest`; CDK renders the concrete value.
- **Instance role** (defined in `/infra`, least-privilege per T4/T12) grants:
`s3:GetObject` on `open-swe-<env>-assets/*`; `secretsmanager:GetSecretValue` on
`open-swe-<env>/*`; `ssm:GetParameter(s)`/`GetParametersByPath` on `/open-swe-<env>/*`;
`logs:*` for the CW agent log groups + `cloudwatch:PutMetricData`; SSM Session
Manager (`ssm:UpdateInstanceInformation`, `ssmmessages:*`) for shell access.
- **CDK substitutes** the `@@OPENSWE_ENV@@`, `@@ASSETS_BUCKET@@`, `@@SERVER_NAME@@`,
`@@ARTIFACT_PREFIX@@` tokens in `user-data.sh` when rendering the launch template.
- `:2024` binds `0.0.0.0` so the ALB hooks target group can reach `/webhooks/*`; it is
reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is
never internet-exposed — the ALB hooks rule is path-scoped to `/webhooks/*`.

102
deploy/ami/deploy.sh Executable file
View file

@ -0,0 +1,102 @@
#!/usr/bin/env bash
# Open SWE app deploy — RE-RUNNABLE (first boot + every subsequent release).
#
# Pulls the current release from S3 (open-swe-<env>-assets), builds the venv
# natively on the box, and restarts the service. This is the SINGLE source of the
# app-deploy procedure; it runs in two places:
#
# 1. first boot — user-data.sh decodes this script to /opt/open-swe/bin and
# calls it ONCE (non-fatal: if no release is published yet,
# nginx is already up and the box waits for the first deploy).
# 2. every release — the `open-swe-<env>-deploy` SSM document (CI fires it after
# uploading app.tar.gz / spa.tar.gz) runs this same script.
#
# It deploys CODE + STATIC ASSETS only. Secrets/config are NOT fetched here: the
# systemd unit's ExecStartPre=fetch-config.sh materializes the tmpfs .env on every
# (re)start, fail-fast — so `systemctl restart` below is what reloads config too.
#
# Contract:
# app.tar.gz = the Python source tree (pyproject.toml + uv.lock + agent/ +
# deploy/ + langgraph.json + README.md, NO ui/, NO .venv). The venv
# is built HERE with `uv sync` so it is native ARM64 and lives at
# the real runtime path (no cross-built / non-relocatable venv).
# spa.tar.gz = the built dashboard SPA (vite output: _shell.html + assets),
# extracted to the nginx web root.
set -euo pipefail
exec > >(tee -a /var/log/open-swe/deploy.log) 2>&1
echo "==> open-swe deploy start $(date -u +%FT%TZ)"
# Non-secret pointers written by user-data.sh (env, region, bucket, artifact prefix).
# shellcheck disable=SC1091
. /etc/open-swe/boot.env
export AWS_DEFAULT_REGION="${AWS_REGION:?boot.env missing AWS_REGION}"
: "${ASSETS_BUCKET:?boot.env missing ASSETS_BUCKET}"
: "${ARTIFACT_PREFIX:?boot.env missing ARTIFACT_PREFIX}"
# Fixed layout — must match provision.sh + user-data.sh + the templates.
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
WWW_ROOT="/var/www/open-swe"
ENV_FILE="/run/open-swe/.env"
UV_BIN="/usr/local/bin/uv"
UV_PYTHON_INSTALL_DIR="/opt/uv/python" # where provision.sh pre-installed py3.12
SERVICE_HOME="/opt/open-swe"
# Benign-vs-failure distinction: on a brand-new env no release is published yet.
# Treat "app.tar.gz absent in S3" as a benign no-op (exit 0) so first boot is not a
# scary failure; ONCE a release exists, any later step failing is loud (set -e).
if ! aws s3 ls "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" >/dev/null 2>&1; then
echo "==> no release published at s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/ yet — nothing to deploy"
exit 0
fi
echo "==> pull release from s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/"
tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" "${tmp}/app.tar.gz"
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/spa.tar.gz" "${tmp}/spa.tar.gz"
# Replace app source + SPA atomically-ish: clear the dirs (drops files removed in
# this release) then extract. The venv is rebuilt below, so wiping .venv too is
# fine — uv's cache (in the service home) makes the rebuild fast.
#
# Hardening: deploy.sh runs as root, so extract with --no-same-owner
# --no-same-permissions — files take root:root + umask perms (NOT the archive's
# uid/mode), so a tarball cannot land a setuid/setgid binary or a foreign-owned
# file; the chown -R below then hands the tree to the service user. (GNU tar also
# refuses `..`-escaping members by default.) Defense-in-depth: the only writer of
# this bucket is the CI OIDC app role, but the box never trusts the archive's
# ownership/mode regardless.
echo "==> install app source -> ${APP_DIR}"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$APP_DIR" "$WWW_ROOT"
find "$APP_DIR" -mindepth 1 -delete
find "$WWW_ROOT" -mindepth 1 -delete
tar --no-same-owner --no-same-permissions -xzf "${tmp}/app.tar.gz" -C "$APP_DIR"
tar --no-same-owner --no-same-permissions -xzf "${tmp}/spa.tar.gz" -C "$WWW_ROOT"
# langgraph reads ./.env from WorkingDirectory; point it at the tmpfs file the
# systemd ExecStartPre materializes.
ln -sfn "$ENV_FILE" "${APP_DIR}/.env"
chown -R "$SERVICE_USER":"$SERVICE_USER" "$APP_DIR" "$WWW_ROOT"
echo "==> build venv natively (uv sync --frozen --no-dev)"
# Run as the service user so the venv + uv cache are owned by it. Pin the
# pre-baked interpreter dir so uv never reaches out to download Python at deploy.
cd "$APP_DIR"
sudo -u "$SERVICE_USER" env \
HOME="$SERVICE_HOME" \
UV_PYTHON_INSTALL_DIR="$UV_PYTHON_INSTALL_DIR" \
UV_CACHE_DIR="${SERVICE_HOME}/.cache/uv" \
"$UV_BIN" sync --frozen --no-dev
echo "==> restart open-swe.service + reload nginx"
# ExecStartPre=fetch-config.sh fails-fast if secrets/config are missing, so a
# restart here surfaces a bad config as a failed unit (non-zero exit below).
systemctl restart open-swe.service
nginx -t && systemctl reload nginx
if systemctl is-active --quiet open-swe.service; then
echo "==> open-swe deploy OK $(date -u +%FT%TZ)"
else
echo "!! open-swe.service is not active after deploy (check fetch-config/secrets)"
exit 1
fi

View file

@ -0,0 +1,143 @@
# Open SWE base AMI — ARM64 (Graviton) Ubuntu 24.04 LTS.
#
# Builds the immutable base image for the single EC2 instance per env
# (open-swe-dev / open-swe-prod) in seahaven-vpc. The image bakes the runtime
# (uv + Python 3.12, nginx, awscli v2, CloudWatch agent) and the service-user /
# systemd / nginx TEMPLATES. It bakes NO secrets and NO env-specific values —
# those are materialized at first boot by user-data + deploy/seahaven/fetch-config.sh
# (Secrets Manager + SSM -> root-only tmpfs .env, fail-fast).
#
# Build: packer init . && packer build open-swe-base.pkr.hcl
# The resulting AMI id is pinned in infra/cdk.context.json (CDK cachedInContext:true);
# see README.md "AMI -> cdk.context.json pinning contract".
packer {
required_version = ">= 1.11.0, < 2.0.0"
required_plugins {
amazon = {
source = "github.com/hashicorp/amazon"
version = "1.3.6"
}
}
}
variable "aws_region" {
type = string
default = "us-east-1"
}
variable "instance_type" {
type = string
default = "t4g.medium" # ARM64 (Graviton) build host; runtime instances are ~t4g.large
}
variable "ami_name_prefix" {
type = string
default = "open-swe-base-arm64"
}
# Versions baked into the image. Pin and bump deliberately.
variable "python_version" {
type = string
default = "3.12"
}
variable "node_major" {
type = string
default = "24"
}
variable "uv_version" {
type = string
default = "0.11.24"
}
variable "cloudwatch_agent_deb_url" {
type = string
default = "https://amazoncloudwatch-agent.s3.amazonaws.com/ubuntu/arm64/latest/amazon-cloudwatch-agent.deb"
}
variable "awscli_zip_url" {
type = string
default = "https://awscli.amazonaws.com/awscli-exe-linux-aarch64.zip"
}
locals {
timestamp = formatdate("YYYYMMDD-hhmmss", timestamp())
}
# Latest Canonical Ubuntu 24.04 (Noble) arm64 server image.
source "amazon-ebs" "open-swe" {
region = var.aws_region
instance_type = var.instance_type
ssh_username = "ubuntu"
ami_name = "${var.ami_name_prefix}-${local.timestamp}"
# ASCII only — AWS rejects non-ASCII in the AMI Description attribute.
ami_description = "Open SWE base - Ubuntu 24.04 arm64 + uv/py3.12 + nginx + CW agent (templates only, no secrets)"
source_ami_filter {
filters = {
name = "ubuntu/images/hvm-ssd*/ubuntu-noble-24.04-arm64-server-*"
architecture = "arm64"
root-device-type = "ebs"
virtualization-type = "hvm"
}
owners = ["099720109477"] # Canonical
most_recent = true
}
# IMDSv2 required on the build host.
metadata_options {
http_endpoint = "enabled"
http_tokens = "required"
http_put_response_hop_limit = 1
}
# gp3 root, encrypted. Runtime root size is set by CDK; this is just the build host.
launch_block_device_mappings {
device_name = "/dev/sda1"
volume_size = 20
volume_type = "gp3"
encrypted = true
delete_on_termination = true
}
tags = {
Name = "open-swe-base-arm64"
Purpose = "open-swe-runtime-base"
ManagedBy = "packer"
}
}
build {
name = "open-swe-base"
sources = ["source.amazon-ebs.open-swe"]
# Stage the boot-time templates into the image. The destination dir must exist
# BEFORE a trailing-slash (contents-only) file upload — packer's file provisioner
# does not create it, and uploading the directory itself trips scp ("Is a
# directory"). So mkdir first, then upload the contents into it.
provisioner "shell" {
inline = ["mkdir -p /tmp/open-swe-templates"]
}
provisioner "file" {
source = "${path.root}/templates/"
destination = "/tmp/open-swe-templates"
}
provisioner "shell" {
environment_vars = [
"PYTHON_VERSION=${var.python_version}",
"NODE_MAJOR=${var.node_major}",
"UV_VERSION=${var.uv_version}",
"CLOUDWATCH_AGENT_DEB_URL=${var.cloudwatch_agent_deb_url}",
"AWSCLI_ZIP_URL=${var.awscli_zip_url}",
]
# {{ .Vars }} MUST be included or the environment_vars above never reach the
# script (provision.sh runs under `set -u` and fails on the first reference).
execute_command = "chmod +x {{ .Path }}; {{ .Vars }} sudo -E bash '{{ .Path }}'"
script = "${path.root}/scripts/provision.sh"
}
}

105
deploy/ami/scripts/provision.sh Executable file
View file

@ -0,0 +1,105 @@
#!/usr/bin/env bash
# Packer provisioner for the Open SWE base AMI (ARM64 Ubuntu 24.04).
#
# Bakes the runtime + boot-time templates ONLY. No secrets, no env-specific
# values. Everything env-specific is materialized at first boot by user-data.sh
# + deploy/seahaven/fetch-config.sh.
set -euo pipefail
PYTHON_VERSION="${PYTHON_VERSION:-3.12}"
NODE_MAJOR="${NODE_MAJOR:-24}"
UV_VERSION="${UV_VERSION:-0.11.24}"
CLOUDWATCH_AGENT_DEB_URL="${CLOUDWATCH_AGENT_DEB_URL:?}"
AWSCLI_ZIP_URL="${AWSCLI_ZIP_URL:?}"
# Layout (must match user-data.sh and the templates).
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
SERVICE_HOME="/opt/open-swe"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
LOG_DIR="/var/log/open-swe"
UV_BIN="/usr/local/bin/uv"
export DEBIAN_FRONTEND=noninteractive
echo "==> apt base packages"
apt-get update -y
apt-get upgrade -y
apt-get install -y --no-install-recommends \
nginx jq curl unzip ca-certificates gnupg lsb-release \
build-essential pkg-config git acl
echo "==> awscli v2 (aarch64)"
tmp="$(mktemp -d)"
curl -fsSL "$AWSCLI_ZIP_URL" -o "$tmp/awscliv2.zip"
unzip -q "$tmp/awscliv2.zip" -d "$tmp"
"$tmp/aws/install" --update
rm -rf "$tmp"
aws --version
echo "==> CloudWatch agent (arm64)"
tmp="$(mktemp -d)"
curl -fsSL "$CLOUDWATCH_AGENT_DEB_URL" -o "$tmp/amazon-cloudwatch-agent.deb"
dpkg -i -E "$tmp/amazon-cloudwatch-agent.deb"
rm -rf "$tmp"
# Do NOT enable/start the agent during the build; user-data fetches its config
# (with env-specific log-group names + 30-day retention) and starts it at boot.
systemctl disable amazon-cloudwatch-agent.service || true
echo "==> uv ${UV_VERSION} + Python ${PYTHON_VERSION} (system-wide)"
export UV_INSTALL_DIR=/usr/local/bin
curl -fsSL "https://astral.sh/uv/${UV_VERSION}/install.sh" | env UV_NO_MODIFY_PATH=1 sh
"$UV_BIN" --version
# Pre-install the interpreter so the box never reaches out at boot to build a venv.
UV_PYTHON_INSTALL_DIR=/opt/uv/python "$UV_BIN" python install "$PYTHON_VERSION"
echo "==> node ${NODE_MAJOR} + bun (build-time UI tooling only; the SPA is built in CI)"
curl -fsSL "https://deb.nodesource.com/setup_${NODE_MAJOR}.x" | bash -
apt-get install -y --no-install-recommends nodejs
node --version
# bun installed system-wide; used only if any UI tooling must run on-box. The
# production SPA build runs in GitHub Actions -> S3 (no on-box build, no swapfile).
export BUN_INSTALL=/usr/local
curl -fsSL https://bun.sh/install | bash
/usr/local/bin/bun --version || true
echo "==> non-login service user '${SERVICE_USER}'"
if ! id "$SERVICE_USER" >/dev/null 2>&1; then
useradd --system --create-home --home-dir "$SERVICE_HOME" \
--shell /usr/sbin/nologin "$SERVICE_USER"
fi
echo "==> directories"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$SERVICE_HOME" "$APP_DIR"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$WWW_ROOT"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0750 "$LOG_DIR"
install -d -o root -g root -m 0755 "$TEMPLATE_DIR"
echo "==> stage boot-time templates into the image"
cp /tmp/open-swe-templates/* "$TEMPLATE_DIR/"
chown root:root "$TEMPLATE_DIR"/*
chmod 0644 "$TEMPLATE_DIR"/*
rm -rf /tmp/open-swe-templates
echo "==> tmpfs for the runtime .env (root/owner-only, noexec/nosuid/nodev)"
# /run is already tmpfs on Ubuntu; this is an explicit, deliberately-small mount
# scoped to the service user so the materialized .env never touches disk.
if ! grep -q '/run/open-swe' /etc/fstab; then
cat >>/etc/fstab <<EOF
tmpfs /run/open-swe tmpfs rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m 0 0
EOF
fi
echo "==> disable nginx default site (open-swe site is installed at boot)"
rm -f /etc/nginx/sites-enabled/default
systemctl enable nginx
echo "==> harden: no password auth, IMDSv2 already enforced by launch template"
# (sshd is not exposed publicly — instance is in a private subnet, SG inbound = ALB only.)
echo "==> clean apt caches"
apt-get clean
rm -rf /var/lib/apt/lists/*
echo "==> provision complete"

View file

@ -0,0 +1,55 @@
{
"agent": {
"metrics_collection_interval": 60,
"run_as_user": "root"
},
"metrics": {
"namespace": "open-swe/@@OPENSWE_ENV@@",
"append_dimensions": {
"InstanceId": "${aws:InstanceId}"
},
"metrics_collected": {
"mem": { "measurement": ["mem_used_percent"] },
"disk": {
"measurement": ["used_percent"],
"resources": ["/"]
}
}
},
"logs": {
"logs_collected": {
"files": {
"collect_list": [
{
"file_path": "/var/log/open-swe/app.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/app",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/open-swe-user-data.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/user-data",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/access.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-access",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/error.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-error",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
}
]
}
}
}
}

View file

@ -0,0 +1,62 @@
# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy.
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
#
# nginx is the SOLE ingress and security boundary (T5 OSWE-IAC-03): the backend
# binds 127.0.0.1:2024 and is NOT network-reachable. nginx proxies exactly two
# prefixes to it — /dashboard/api/* and /webhooks/* — and nothing else. The
# unauthenticated LangGraph agent API (/threads, /runs, /assistants, /store) is
# NEVER proxied; those paths return the SPA shell.
#
# Both ALB target groups (dashboard host + hooks host) point at this nginx :80,
# not at :2024 directly, so there is no path to the raw control plane even from
# inside the SG. Webhook signature verification still happens in the app (the raw
# body + GitHub/Slack/Linear signature headers are passed through unmodified).
server {
listen 80 default_server;
listen [::]:80 default_server;
server_name @@SERVER_NAME@@;
root @@WWW_ROOT@@;
index _shell.html;
# ALB target-group health check (dashboard TG).
location = /healthz { default_type text/plain; return 200 "ok\n"; }
# Dashboard API + OAuth callback -> backend webapp.
location /dashboard/api/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
client_max_body_size 10m;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 300s;
}
# Inbound webhooks (GitHub/Slack/Linear) -> backend webapp. Routed through
# nginx so :2024 stays loopback-only (T5 OSWE-IAC-03). The raw request body +
# signature headers pass through unmodified for in-app signature verification.
location /webhooks/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
# GitHub permits webhook payloads up to 25 MB; nginx's 1 MB default would
# 413 large push/PR events at the edge BEFORE in-app signature verification
# runs, silently dropping them (OSWE-T12-01). proxy_request_buffering off
# does not relax the size cap — set it explicitly.
client_max_body_size 25m;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_request_buffering off;
proxy_read_timeout 300s;
}
# Static assets + SPA shell fallback (client-side routing).
location / {
try_files $uri $uri/ /_shell.html;
}
}

View file

@ -0,0 +1,54 @@
# Open SWE — stock LangGraph dev server (3+ graphs + FastAPI webapp, :2024).
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
# In-memory runtime (--no-reload) + ExecStartPost reseed; no Aegra/Postgres.
#
# Boot contract:
# ExecStartPre = fetch-config.sh <env> -> runs as root (`+`) ONLY to materialize
# the SERVICE-USER-owned tmpfs .env (@@ENV_FILE@@) from Secrets
# Manager + SSM and chown it to @@SERVICE_USER@@, fail-fast (the
# unit does NOT start if config can't be fetched).
# ExecStart = langgraph dev (as @@SERVICE_USER@@, bound to 127.0.0.1 — nginx
# is the sole ingress; never binds 0.0.0.0).
# ExecStartPost = seed_store.sh <env> -> reseeds team_settings + user_mappings
# that the in-memory store loses on every restart.
[Unit]
Description=Open SWE stock LangGraph dev server (graphs + webapp, :@@PORT@@)
After=network-online.target
Wants=network-online.target
RequiresMountsFor=/run/open-swe
[Service]
Type=simple
User=@@SERVICE_USER@@
Group=@@SERVICE_USER@@
WorkingDirectory=@@APP_DIR@@
# No EnvironmentFile: the secret-bearing app .env (@@ENV_FILE@@) is loaded by
# langgraph/dotenv (so the multiline GitHub App PEM never hits systemd's env
# parser), and seed_store.sh reads the same .env directly (without sourcing it).
#
# ExecStartPre runs as root (`+`) so it can chown the tmpfs .env to the service
# user; the env arg (@@OPENSWE_ENV@@) selects the SSM/Secrets prefix (T5 BOOT-01).
ExecStartPre=+@@FETCH_CONFIG@@ @@OPENSWE_ENV@@
# Bind 127.0.0.1 only — nginx proxies dashboard + webhooks; :@@PORT@@ is never
# directly network-reachable (T5 OSWE-IAC-03).
ExecStart=@@VENV@@/bin/langgraph dev --host 127.0.0.1 --port @@PORT@@ --no-browser --no-reload
ExecStartPost=@@SEED_STORE@@ @@OPENSWE_ENV@@
# App logs to a file CloudWatch collects (30-day retention set in the CW config).
StandardOutput=append:/var/log/open-swe/app.log
StandardError=append:/var/log/open-swe/app.log
Restart=on-failure
RestartSec=5
TimeoutStartSec=180
# Hardening — the box holds no durable state of its own.
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
PrivateTmp=true
ReadWritePaths=/var/log/open-swe /var/www/open-swe /run/open-swe @@APP_DIR@@
[Install]
WantedBy=multi-user.target

145
deploy/ami/user-data.sh Executable file
View file

@ -0,0 +1,145 @@
#!/usr/bin/env bash
# Open SWE EC2 user-data — PROVISIONING-ONLY (runs once, at first boot).
#
# This is the rationale for `userDataCausesReplacement: true` in CDK: user-data
# does FIRST-BOOT provisioning, never durable runtime config. Editing it is a
# deliberate instance replacement. Durable runtime config is fetched fresh on
# every service start by deploy/seahaven/fetch-config.sh (ExecStartPre).
#
# The box holds NO durable state of its own:
# - secrets/config -> Secrets Manager + SSM, materialized to a tmpfs .env at boot
# - app artifact -> pulled from S3 (open-swe-<env>-assets) via the instance role
# - store state -> reseeded by seed_store.sh (ExecStartPost) on every start
# => there is no RETAIN volume to protect; replacement is tolerated. The EBS
# discipline (snapshot root + wait state=completed BEFORE any replacing deploy)
# is the safety net, not durable on-box state. See README "EBS discipline".
#
# Tokens (@@...@@) are substituted by CDK when it renders this script into the
# launch template. region is read from IMDSv2 as a fallback.
set -euo pipefail
exec > >(tee -a /var/log/open-swe-user-data.log) 2>&1
echo "==> open-swe user-data start $(date -u +%FT%TZ)"
# --- CDK-rendered values -----------------------------------------------------
# NOTE: CDK-substituted tokens use %%...%% (rendered by app-service.ts), DISTINCT
# from the @@...@@ tokens this script seds into the baked systemd/nginx templates.
# The two MUST NOT share a delimiter: a shared @@OPENSWE_ENV@@ / @@SERVER_NAME@@
# let CDK clobber the sed PATTERN, leaving the unit's token unsubstituted.
OPENSWE_ENV="%%OPENSWE_ENV%%" # dev | prod
ASSETS_BUCKET="%%ASSETS_BUCKET%%" # open-swe-<env>-assets
SERVER_NAME="%%SERVER_NAME%%" # openswe[-dev].seahaven.com
ARTIFACT_PREFIX="%%ARTIFACT_PREFIX%%" # e.g. releases/latest
# --- fixed layout (must match provision.sh + templates) ----------------------
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
VENV="${APP_DIR}/.venv"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
ENV_FILE="/run/open-swe/.env"
PORT="2024"
FETCH_CONFIG="${APP_DIR}/deploy/seahaven/fetch-config.sh"
SEED_STORE="${APP_DIR}/deploy/seahaven/seed_store.sh"
# region from IMDSv2
TOKEN="$(curl -fsS -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 300" || true)"
AWS_REGION="$(curl -fsS -H "X-aws-ec2-metadata-token: ${TOKEN}" \
http://169.254.169.254/latest/meta-data/placement/region || echo us-east-1)"
export AWS_DEFAULT_REGION="$AWS_REGION"
echo "env=${OPENSWE_ENV} region=${AWS_REGION} bucket=${ASSETS_BUCKET} host=${SERVER_NAME}"
# --- boot.env: non-secret pointers fetch-config.sh reads ---------------------
install -d -o root -g root -m 0755 /etc/open-swe
cat >/etc/open-swe/boot.env <<EOF
OPENSWE_ENV=${OPENSWE_ENV}
AWS_REGION=${AWS_REGION}
ASSETS_BUCKET=${ASSETS_BUCKET}
ARTIFACT_PREFIX=${ARTIFACT_PREFIX}
ENV_FILE=${ENV_FILE}
SECRETS_PREFIX=open-swe-${OPENSWE_ENV}
SSM_PREFIX=/open-swe-${OPENSWE_ENV}
EOF
chmod 0644 /etc/open-swe/boot.env
# --- ensure the tmpfs for the materialized .env is mounted -------------------
# (baked into /etc/fstab by the AMI; mount it now in case it isn't yet.)
install -d -o root -g root -m 0755 /run/open-swe || true
mountpoint -q /run/open-swe || mount /run/open-swe || mount -t tmpfs \
-o rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m \
tmpfs /run/open-swe
# --- install the deploy script (single source of the app-deploy procedure) ---
# deploy.sh (deploy/ami/deploy.sh) pulls the release from S3, builds the venv with
# `uv sync`, and restarts the service. CDK base64-renders the file into the
# %%DEPLOY_SH_B64%% token below so it is a normal reviewable repo file, not an
# inline heredoc. The `open-swe-<env>-deploy` SSM document runs this same script
# for every subsequent release.
echo "==> install /opt/open-swe/bin/deploy.sh"
install -d -o root -g root -m 0755 /opt/open-swe/bin
base64 -d >/opt/open-swe/bin/deploy.sh <<'DEPLOY_SH_B64'
%%DEPLOY_SH_B64%%
DEPLOY_SH_B64
chmod 0755 /opt/open-swe/bin/deploy.sh
# --- render + install the systemd unit ---------------------------------------
echo "==> install systemd unit"
sed \
-e "s|@@SERVICE_USER@@|${SERVICE_USER}|g" \
-e "s|@@APP_DIR@@|${APP_DIR}|g" \
-e "s|@@VENV@@|${VENV}|g" \
-e "s|@@PORT@@|${PORT}|g" \
-e "s|@@ENV_FILE@@|${ENV_FILE}|g" \
-e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
-e "s|@@FETCH_CONFIG@@|${FETCH_CONFIG}|g" \
-e "s|@@SEED_STORE@@|${SEED_STORE}|g" \
"${TEMPLATE_DIR}/open-swe.service" >/etc/systemd/system/open-swe.service
systemctl daemon-reload
# --- render + install the nginx site -----------------------------------------
echo "==> install nginx site"
sed \
-e "s|@@SERVER_NAME@@|${SERVER_NAME}|g" \
-e "s|@@WWW_ROOT@@|${WWW_ROOT}|g" \
-e "s|@@BACKEND_ADDR@@|127.0.0.1:${PORT}|g" \
"${TEMPLATE_DIR}/open-swe.nginx.conf" >/etc/nginx/sites-available/open-swe
ln -sfn /etc/nginx/sites-available/open-swe /etc/nginx/sites-enabled/open-swe
rm -f /etc/nginx/sites-enabled/default
nginx -t
# --- CloudWatch agent: 30-day log retention ----------------------------------
echo "==> configure CloudWatch agent (30-day retention)"
sed -e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
"${TEMPLATE_DIR}/amazon-cloudwatch-agent.json" \
>/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
/opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
-a fetch-config -m ec2 -s \
-c file:/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
# --- start nginx FIRST (the security boundary + health surface) --------------
# NOTE: intentionally NO swapfile here. The 8 GB-swapfile OOM hack existed only
# for the on-box Nitro SPA build, which now runs in GitHub Actions -> S3.
# nginx is brought up BEFORE the app is deployed so the ALB target-group health
# check (static `/healthz` -> 200) passes and the box is a healthy target even on
# the very first boot, before any release is published. open-swe.service is
# enabled (boot persistence) but STARTED by deploy.sh once the app is on disk.
echo "==> start nginx"
systemctl enable --now nginx
systemctl reload nginx
systemctl enable open-swe.service
# --- deploy the app (NON-FATAL on first boot) --------------------------------
# deploy.sh pulls the release, builds the venv, and starts open-swe.service. On a
# brand-new env no release exists yet, so this is allowed to fail WITHOUT aborting
# user-data: nginx is already up (healthy target), and the first `build-artifacts`
# run + `open-swe-<env>-deploy` SSM command will bring the app up. A failure here
# is logged, not fatal.
echo "==> initial app deploy (non-fatal if no release is published yet)"
if /opt/open-swe/bin/deploy.sh; then
echo "==> initial app deploy succeeded"
else
echo "==> no release yet (or deploy failed): open-swe.service deferred to the next SSM deploy"
fi
echo "==> open-swe user-data done $(date -u +%FT%TZ)"

View file

@ -89,6 +89,25 @@ Start a task by mentioning **`@openswe`** in a GitHub issue comment (the documen
intake — the `open-swe` *label* path needs the `Issues` event subscription). The
commenter must have a `user_mappings` entry or the run is skipped.
## AWS variant — env-sourced config (.env from Secrets Manager + SSM)
On the AWS lift-and-shift, the `.env` is **not** staged by hand. A boot hook pulls
config from AWS via the EC2 instance role and materializes a root-only `.env` on a
tmpfs before the service starts. Same stock runtime; only how `.env` is produced
changes.
| File | Role |
|---|---|
| `fetch-config.sh <dev\|prod>` | `ExecStartPre=+` hook. Reads all SSM params `/open-swe-<env>/*` + all secrets `open-swe-<env>/*`, FAIL-FAST on any missing required var, writes `/run/open-swe/.env` (tmpfs, root:root 0600), symlinks `<APP_DIR>/.env` → it. Forces `DEFAULT_REPO_OWNER`=`Sea-Haven-Industries` (never upstream `langchain-ai`). |
| `seed_store.sh <dev\|prod>` | `ExecStartPost` hook (unchanged role). Now sources the materialized `.env`, so `default_repo`/models/user-mappings come from AWS config instead of hardcoded `OPENSWE_*`. Still re-seeds the in-memory store on **every** restart. |
| `ROTATION.md` | How to rotate any secret/param (update AWS → `systemctl restart`) + per-secret notes (`TOKEN_ENCRYPTION_KEY` overlap list, GitHub App PEM, webhook signing secrets). |
Because the `.env` is root-only, the AWS `open-swe.service` runs **as root** (the
on-prem unit's `User=adam` cannot read it). See the header of `fetch-config.sh` for
the exact unit snippet and the tmpfs / `RUN_DEDICATED_TMPFS` options. Naming/secret
placement follows the T9 inventory: 29 secrets → Secrets Manager `open-swe-<env>/*`,
53 config → SSM `/open-swe-<env>/*`, each keyed by the literal env-var name.
## Aegra (deferred)
`aegra/aegra.json` + `aegra/aegra_entry.py` are the self-hosted-runtime alternative

View file

@ -0,0 +1,92 @@
# Sea Haven — Open SWE secret & config rotation
How secrets and config reach the running app, and how to rotate either one.
## How values flow at boot
```
AWS Secrets Manager open-swe-<env>/* ─┐
AWS SSM Param Store /open-swe-<env>/* ─┤── fetch-config.sh ──▶ tmpfs /run/open-swe/.env (root:root 0600)
│ (systemd ExecStartPre=+, EC2 role) │
▼ ▼
FAIL-FAST if a app symlink <APP_DIR>/.env
required var is empty python-dotenv reads at import
```
The app reads `.env` **once, at import**. There is no hot-reload of secrets.
Therefore the rotation contract is always the same two steps:
> **Rotation = (1) update the value in Secrets Manager / SSM, then (2) restart the
> service** so `fetch-config.sh` re-materializes the `.env`.
```bash
# after updating a secret/param in AWS:
sudo systemctl restart open-swe.service
# ExecStartPre=+ -> fetch-config.sh re-pulls + rewrites the tmpfs .env (fail-fast)
# ExecStartPost -> seed_store.sh re-seeds the in-memory store (team_settings + user_mappings)
```
There is **no zero-downtime path for most secrets** on the stock in-memory
runtime — a restart is required and it also wipes the in-memory store (re-seeded
by `seed_store.sh` automatically). The one secret built for zero-downtime overlap
is `TOKEN_ENCRYPTION_KEY` (see below), but even it needs the restart to load the
new key list.
## Rotating a secret (Secrets Manager)
```bash
ENV=prod # or dev
NAME=DASHBOARD_JWT_SECRET
aws secretsmanager put-secret-value \
--secret-id "open-swe-${ENV}/${NAME}" \
--secret-string 'NEW_VALUE' \
--region us-east-1
sudo systemctl restart open-swe.service # on the box
```
(`update-secret`/`put-secret-value` both create a new version; the boot hook
always reads `AWSCURRENT`.)
## Rotating a config param (SSM)
```bash
aws ssm put-parameter --overwrite \
--name "/open-swe-${ENV}/DASHBOARD_BASE_URL" \
--type String --value 'https://openswe.seahaven.com' \
--region us-east-1
sudo systemctl restart open-swe.service
```
## Per-secret rotation notes
| Secret | Rotation notes |
|---|---|
| **TOKEN_ENCRYPTION_KEY** | Fernet key(s). Supports a **comma/newline-separated list** (`agent/encryption.py`) for zero-downtime key rotation: prepend the NEW key, keep the OLD key(s) in the list. New data is encrypted with the first key; old data still decrypts with the trailing keys. After all encrypted-at-rest tokens (per-user GitHub OAuth tokens in thread metadata) have been re-encrypted/expired, drop the old key. Store the list as one secret value; `fetch-config.sh` writes it verbatim. **Never** rotate to a single new key in one step or every existing encrypted token becomes undecryptable. |
| **GITHUB_APP_PRIVATE_KEY** | Multiline PEM. Generate a new private key in the GitHub App settings (you may have **two active keys** during overlap), put the new PEM into the secret, restart, verify install-token minting + a webhook delivery, then delete the old key in GitHub. `fetch-config.sh` writes the PEM as a double-quoted multiline value (python-dotenv-safe); paste the full `-----BEGIN…-----END-----` block including newlines. |
| **GITHUB_WEBHOOK_SECRET** | Webhook HMAC. GitHub allows only **one** webhook secret per App, so this is a brief-break rotation: update the secret in AWS **and** the GitHub App webhook config, restart. Deliveries signed with the old secret during the gap will 401 (GitHub auto-redelivers). Required in **prod** (fail-fast). |
| **SLACK_SIGNING_SECRET** | Slack request-signature secret. Rotate in the Slack app config and AWS together, restart. Required in **prod** (fail-fast). A stale value silently 401s `url_verification`/events until restart (known gotcha). |
| **LINEAR_WEBHOOK_SECRET** | Linear webhook signature. Required in prod only when the Linear integration is wired (`LINEAR_API_KEY` present). Rotate in Linear + AWS together, restart. |
| **SLACK_CLIENT_SECRET / GITHUB_APP_CLIENT_SECRET** | OAuth client secrets (dashboard login / Slack OAuth). Rotate in the provider console + AWS, restart. Existing dashboard sessions are JWT-signed by `DASHBOARD_JWT_SECRET`, not these, so they survive. |
| **DASHBOARD_JWT_SECRET** | Signs dashboard session cookies. Rotating **invalidates all active sessions** (users re-login). Hard-required (RuntimeError if empty). No overlap list — single value. |
| **Model provider keys** (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GOOGLE_API_KEY`, `GROQ_API_KEY`, `FIREWORKS_API_KEY`) | Standard API-key rotation: issue new key, update AWS, restart, revoke old. The **active** provider keys (whatever `team_settings/default` seeds — currently `ANTHROPIC_API_KEY` + `OPENAI_API_KEY`) are fail-fast-required; keep `REQUIRED_PROVIDER_KEYS` in sync if you change the seeded models. |
| **LANGSMITH_API_KEY_PROD** | LangSmith key powering the `langsmith` sandbox (the only provider with working in-sandbox git/gh auth). Required when `SANDBOX_TYPE=langsmith` (fail-fast). Rotate in LangSmith + AWS, restart; existing sandboxes keep their already-injected proxy token until recycled. |
| **SLACK_BOT_TOKEN / LINEAR_API_KEY / GITHUB_PAT / EXA_API_KEY / DAYTONA_API_KEY / RUNLOOP_API_KEY / CORRIDOR_* / USER_ID_API_KEY_MAP / X_SERVICE_AUTH_JWT_SECRET** | Plain API-token rotation: update AWS, restart, revoke old at the provider. None support an overlap list. |
## Fail-fast safety
`fetch-config.sh` refuses to write the `.env` (exit 1) if any required var is
empty after a rotation — so a botched rotation (e.g. an empty `put-secret-value`)
stops the service at `ExecStartPre` instead of starting it with a partial `.env`.
The missing variable **names** are printed to the journal (values never are):
```bash
sudo journalctl -u open-swe.service -b | grep fetch-config
```
Required set enforced: `DASHBOARD_JWT_SECRET`, `TOKEN_ENCRYPTION_KEY`,
`GITHUB_APP_ID`, `GITHUB_APP_PRIVATE_KEY`, `GITHUB_APP_INSTALLATION_ID`,
`GITHUB_APP_CLIENT_ID`, `GITHUB_APP_CLIENT_SECRET`, the active provider keys
(`REQUIRED_PROVIDER_KEYS`, default `ANTHROPIC_API_KEY,OPENAI_API_KEY`), the
sandbox key for `SANDBOX_TYPE` (`langsmith` ⇒ `LANGSMITH_API_KEY_PROD` +
`DEFAULT_SANDBOX_SNAPSHOT_ID`), and — in **prod** — `GITHUB_WEBHOOK_SECRET`,
`SLACK_SIGNING_SECRET` (plus `LINEAR_WEBHOOK_SECRET` when Linear is wired).

308
deploy/seahaven/fetch-config.sh Executable file
View file

@ -0,0 +1,308 @@
#!/usr/bin/env bash
# fetch-config.sh — AWS-sourced boot hook that materializes the app's .env.
#
# The stock `langgraph dev` runtime + the Open SWE app read a plain `.env` from
# the app working directory (python-dotenv). On the AWS lift-and-shift we do NOT
# commit a .env; instead every non-sensitive value lives in SSM Parameter Store
# (`/open-swe-<env>/*`) and every secret lives in AWS Secrets Manager
# (`open-swe-<env>/*`). This hook is run by systemd BEFORE the service starts; it
# pulls both sources via the EC2 instance role (no static keys), assembles a
# single .env on a tmpfs, and writes it owned by the unprivileged service user
# `chmod 600` (T5 SC-01: the privileged pre-hook materializes the secret; the app
# itself then runs as that NON-root service user, not root).
#
# It is intentionally FAIL-FAST: if any required secret/param is missing or empty
# it prints the offending variable NAMES (never values) and exits 1, so the
# service never starts with a partial .env.
#
# ---------------------------------------------------------------------------
# Naming contract (source of truth: T9 env/secret/config inventory)
# SSM /open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# Secrets open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# i.e. the last path segment IS the literal environment-variable name. This is a
# deliberate (documented) deviation from the handbook's kebab-case value-name
# example (`my-stack/slack-signing`): a .env materializer needs a lossless,
# unambiguous round-trip from store key -> env var, and the env var name is the
# only key that guarantees that. The `open-swe-<env>` stack prefix still follows
# kebab-case per naming-conventions.md.
# ---------------------------------------------------------------------------
#
# Wiring into systemd (AWS EC2 variant):
# The unit runs as the unprivileged service user (User=openswe). ONLY the
# ExecStartPre pre-hook runs as root (the `+` prefix) so it can pull from AWS,
# write the tmpfs .env, and chown it to the service user. The app (ExecStart)
# and the seeder (ExecStartPost) then run as openswe and read the openswe-owned
# 0600 .env — the agent never runs as root (T5 SC-01). Pass the env as the
# positional arg (T5 BOOT-01):
#
# [Service]
# User=openswe
# Group=openswe
# Environment=ENV_DIR=/run/open-swe SERVICE_USER=openswe
# # ExecStartPre runs as root (+) so it can chown the .env to the service user.
# ExecStartPre=+/opt/open-swe/deploy/seahaven/fetch-config.sh prod
# ExecStart=/opt/open-swe/.venv/bin/langgraph dev --host 127.0.0.1 --port 2024 \
# --no-browser --no-reload
# ExecStartPost=/opt/open-swe/deploy/seahaven/seed_store.sh prod
#
# tmpfs: /run is already a tmpfs on systemd hosts, so ENV_DIR=/run/open-swe is
# tmpfs-backed by default (the .env never touches disk). Set RUN_DEDICATED_TMPFS=1
# to mount a private tmpfs at ENV_DIR instead. The app's CWD `.env` is a symlink
# into ENV_DIR (created idempotently below), so python-dotenv finds it unchanged.
#
# Idempotent, re-runnable on every (re)start. No secret is ever echoed.
set -euo pipefail
umask 077
# --- Inputs ------------------------------------------------------------------
ENV="${1:-${OPENSWE_ENV:-}}"
case "$ENV" in
dev | prod) ;;
*)
echo "fetch-config: ENV must be 'dev' or 'prod' (got '${ENV:-<empty>}')" >&2
echo "usage: fetch-config.sh <dev|prod> (or set OPENSWE_ENV)" >&2
exit 2
;;
esac
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
SSM_PREFIX="/open-swe-${ENV}/"
SECRET_PREFIX="open-swe-${ENV}/"
ENV_DIR="${ENV_DIR:-/run/open-swe}" # tmpfs-backed (/run) by default
ENV_FILE="${ENV_DIR}/.env"
APP_DIR="${APP_DIR:-/opt/open-swe}" # where the app + its CWD .env live
APP_ENV_LINK="${APP_DIR}/.env" # symlink -> ENV_FILE
# The unprivileged service user that runs the app and OWNS the .env (T5 SC-01).
# fetch-config runs as root (ExecStartPre=+) only to chown the secret to it.
SERVICE_USER="${SERVICE_USER:-openswe}"
SERVICE_GROUP="${SERVICE_GROUP:-${SERVICE_USER}}"
# Sea Haven override: upstream defaults DEFAULT_REPO_OWNER to "langchain-ai".
# For the Sea Haven deployment it MUST be the org. fetch-config forces this so a
# stale/blank SSM value can never point the agent at the upstream org.
SH_REPO_OWNER="${OPENSWE_REPO_OWNER:-Sea-Haven-Industries}"
for bin in aws jq; do
command -v "$bin" >/dev/null 2>&1 || { echo "fetch-config: '$bin' not found on PATH" >&2; exit 3; }
done
log() { echo "fetch-config[$ENV]: $*"; } # NAMES/counts only — never values
b64d() { base64 --decode; } # GNU coreutils on the EC2 host
# Accept a store key into VARS iff it is a valid env-var identifier and not a
# duplicate. Rejects non-identifier names (T5 SH-INJ-002 / set -e DoS hardening)
# and flat-namespace collisions (T5 SSM-05). $3 = source label for logs.
accept_var() {
local key="$1" value="$2" src="$3"
if ! [[ "$key" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
log "WARNING: skipping ${src} key with non-identifier name (rejected)"
return 0
fi
if [ -n "${VARS[$key]+set}" ]; then
echo "fetch-config[$ENV]: FAIL-FAST — duplicate key '${key}' from ${src} (flat-namespace collision)" >&2
exit 1
fi
VARS["$key"]="$value"
}
# --- tmpfs ------------------------------------------------------------------
mkdir -p "$ENV_DIR"
# Owned by the service user so the unprivileged app can traverse it (T5 SC-01).
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_DIR" 2>/dev/null || true
chmod 700 "$ENV_DIR"
if [ "${RUN_DEDICATED_TMPFS:-0}" = "1" ] && ! mountpoint -q "$ENV_DIR"; then
mount -t tmpfs -o nosuid,nodev,noexec,mode=0700,size=4m tmpfs "$ENV_DIR"
log "mounted dedicated tmpfs at $ENV_DIR"
fi
# --- Collect values into an associative array --------------------------------
declare -A VARS=()
# 1) SSM Parameter Store (non-sensitive config). NOT --recursive: the contract is
# a FLAT namespace /open-swe-<env>/<VAR>, so a non-recursive list returns exactly
# those keys and cannot collapse two nested paths onto one name (T5 SSM-05). aws
# CLI v2 auto-paginates NextToken.
log "reading SSM params under ${SSM_PREFIX} ..."
ssm_json="$(
aws ssm get-parameters-by-path \
--path "$SSM_PREFIX" \
--with-decryption \
--region "$REGION" \
--no-cli-pager \
--output json
)"
# Records are base64-encoded (name<TAB>value) so values with spaces/newlines/tabs
# survive the line-based read intact.
ssm_count=0
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}" # strip /open-swe-<env>/ prefix
[ -n "$key" ] || continue
accept_var "$key" "$value" "SSM"
ssm_count=$((ssm_count + 1))
done < <(jq -r '.Parameters[] | (.Name|@base64) + "\t" + (.Value|@base64)' <<<"$ssm_json")
log "loaded ${ssm_count} config param(s) from SSM"
# 2) Secrets Manager (sensitive values). batch-get-secret-value filters by name
# prefix; one secret per env var, SecretString = the value. The AWS CLI does NOT
# auto-paginate this operation (unlike list/get-parameters-by-path), and a single
# response caps well under the full set (~10 items/page) — so we MUST follow
# NextToken ourselves or secrets on later pages are silently dropped (they then
# surface as FAIL-FAST "missing required var"). `--no-cli-pager` only disables the
# OUTPUT pager, not API pagination. Each page's records are base64(name<TAB>value).
batch_get_secrets_tsv() {
local token="" page
while :; do
if [ -n "$token" ]; then
page="$(aws secretsmanager batch-get-secret-value \
--filters "Key=name,Values=${SECRET_PREFIX}" \
--region "$REGION" --no-cli-pager --output json --next-token "$token")"
else
page="$(aws secretsmanager batch-get-secret-value \
--filters "Key=name,Values=${SECRET_PREFIX}" \
--region "$REGION" --no-cli-pager --output json)"
fi
printf '%s' "$page" \
| jq -r '.SecretValues[] | select(.SecretString != null) | (.Name|@base64) + "\t" + (.SecretString|@base64)'
token="$(printf '%s' "$page" | jq -r '.NextToken // empty')"
[ -n "$token" ] || break
done
}
log "reading secrets under ${SECRET_PREFIX} ..."
secret_count=0
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
case "$name" in
"${SECRET_PREFIX}"*) ;; # defensive: exact-prefix only
*) continue ;;
esac
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}"
[ -n "$key" ] || continue
accept_var "$key" "$value" "Secrets"
secret_count=$((secret_count + 1))
done < <(batch_get_secrets_tsv)
log "loaded ${secret_count} secret(s) from Secrets Manager"
# --- Sea Haven DEFAULT_REPO_OWNER hard pin -----------------------------------
# T5 OSWE-OWNER-04: a HARD pin, not a deny-list. Whatever the store holds (blank,
# upstream 'langchain-ai', a case/space variant, or any other org), the owner is
# unconditionally forced to the Sea Haven org so the agent can never target the
# wrong owner.
cur_owner="${VARS[DEFAULT_REPO_OWNER]:-}"
if [ -n "$cur_owner" ] && [ "$cur_owner" != "$SH_REPO_OWNER" ]; then
log "WARNING: SSM DEFAULT_REPO_OWNER differs from the pinned org -> hard-pinning to '${SH_REPO_OWNER}'"
fi
VARS[DEFAULT_REPO_OWNER]="$SH_REPO_OWNER"
# --- FAIL-FAST: required vars -------------------------------------------------
# Hard-required regardless of mode:
required=(
DASHBOARD_JWT_SECRET # RuntimeError on startup if missing (oauth.py)
TOKEN_ENCRYPTION_KEY # Fernet key(s); decrypts per-user GitHub tokens
)
# NOTE: the GitHub App is NOT created/duplicated for dev — only prod owns the
# (single, shared) GitHub App + Slack app. So the GitHub App quintet + Slack +
# webhook-signing secrets are required for PROD only (see the prod block below).
# Dev boots without them: it has no GitHub-App/Slack/webhook integration — it is a
# deployment-validation env (boot/health/boundary), not a live-triggered agent.
# Active model-provider key(s): model selection is store-driven (team_settings),
# so fetch-config cannot infer it from .env. Default to the seeded cross-family
# pair (anthropic builder + openai reviewer). Override with a comma list.
IFS=',' read -r -a provider_keys <<<"${REQUIRED_PROVIDER_KEYS:-ANTHROPIC_API_KEY,OPENAI_API_KEY}"
for k in "${provider_keys[@]}"; do
k="${k//[[:space:]]/}"
[ -n "$k" ] && required+=("$k")
done
# Sandbox provider key(s) — depends on SANDBOX_TYPE (default langsmith).
sandbox_type="${VARS[SANDBOX_TYPE]:-langsmith}"
case "$sandbox_type" in
langsmith) required+=(LANGSMITH_API_KEY_PROD DEFAULT_SANDBOX_SNAPSHOT_ID) ;;
daytona) required+=(DAYTONA_API_KEY) ;;
runloop) required+=(RUNLOOP_API_KEY) ;;
modal | local) ;; # no key required
*) log "WARNING: unknown SANDBOX_TYPE='${sandbox_type}' — not enforcing a sandbox key" ;;
esac
# Prod-only: the GitHub App (installation-token minting + dashboard OAuth) and the
# webhook-signing secrets. Dev has no GitHub/Slack app, so none of these are
# required there; prod owns the single shared app and must have all of them.
if [ "$ENV" = "prod" ]; then
required+=(
GITHUB_APP_ID # GitHub App trio (installation-token minting) ...
GITHUB_APP_PRIVATE_KEY # ... multiline PEM ...
GITHUB_APP_INSTALLATION_ID # ... used by utils/github_app.py
GITHUB_APP_CLIENT_ID # dashboard OAuth login
GITHUB_APP_CLIENT_SECRET # dashboard OAuth login
GITHUB_WEBHOOK_SECRET # webhook signature verification
SLACK_SIGNING_SECRET # Slack webhook signature verification
)
if [ -n "${VARS[LINEAR_API_KEY]:-}" ] && [ "${OPENSWE_REQUIRE_LINEAR:-1}" = "1" ]; then
required+=(LINEAR_WEBHOOK_SECRET)
fi
fi
missing=()
for k in "${required[@]}"; do
[ -n "${VARS[$k]:-}" ] || missing+=("$k")
done
# de-dup the names for a clean report
if [ "${#missing[@]}" -gt 0 ]; then
mapfile -t missing < <(printf '%s\n' "${missing[@]}" | sort -u)
echo "fetch-config[$ENV]: FAIL-FAST — ${#missing[@]} required var(s) missing/empty:" >&2
printf ' - %s\n' "${missing[@]}" >&2
echo "fetch-config[$ENV]: refusing to write a partial .env; service will not start." >&2
exit 1
fi
# --- Write the .env atomically (root-only on tmpfs) --------------------------
# python-dotenv reads double-quoted values (incl. multiline PEMs). Its decoder
# unescapes ONLY backslash and double-quote (\\ -> \, \" -> "); it does NOT honor
# \$ or \` escapes, so escaping those would leave a spurious backslash. Escape
# exactly backslash then double-quote — real newlines stay literal (multiline OK).
# (Caveat: python-dotenv interpolates a literal `${VAR}` substring; the secret
# domain here — base64/hex/PEM keys — never contains one, so no extra guard.)
emit_var() {
local name="$1" value="$2" esc
esc="${value//\\/\\\\}"
esc="${esc//\"/\\\"}"
printf '%s="%s"\n' "$name" "$esc"
}
tmp="$(mktemp "${ENV_DIR}/.env.XXXXXX")"
chmod 600 "$tmp"
{
printf '# Generated by fetch-config.sh for env=%s at %s — DO NOT EDIT.\n' \
"$ENV" "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
printf '# Source: SSM /open-swe-%s/* + Secrets Manager open-swe-%s/*\n\n' "$ENV" "$ENV"
for k in $(printf '%s\n' "${!VARS[@]}" | sort); do
emit_var "$k" "${VARS[$k]}"
done
} >"$tmp"
mv -f "$tmp" "$ENV_FILE"
# Owned by the unprivileged service user (T5 SC-01) so the app reads it without
# running as root. fetch-config itself runs as root (ExecStartPre=+) to chown.
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_FILE"
chmod 600 "$ENV_FILE"
# Point the app's CWD .env at the tmpfs file (idempotent).
if [ "$APP_ENV_LINK" != "$ENV_FILE" ]; then
if [ -L "$APP_ENV_LINK" ] || [ ! -e "$APP_ENV_LINK" ]; then
ln -sfn "$ENV_FILE" "$APP_ENV_LINK"
elif [ "$(readlink -f "$APP_ENV_LINK" 2>/dev/null || true)" != "$(readlink -f "$ENV_FILE")" ]; then
log "WARNING: ${APP_ENV_LINK} exists and is not a symlink to ${ENV_FILE} — leaving it untouched"
fi
fi
total=$((ssm_count + secret_count))
log "wrote ${ENV_FILE} (${total} vars, sandbox=${sandbox_type}) — ${SERVICE_USER}:${SERVICE_GROUP} 0600"

155
deploy/seahaven/put-config.sh Executable file
View file

@ -0,0 +1,155 @@
#!/usr/bin/env bash
# put-config.sh — out-of-band populator for the open-swe config store.
#
# T11 (CDK) creates the RESOURCE SHELLS:
# - 28 Secrets Manager secrets open-swe-<env>/<VAR> (value-LESS shells)
# - the IaC-managed SSM params /open-swe-<env>/<VAR> (real values, owned in CDK)
# This script sets the values that CANNOT live in IaC — every secret value, plus
# the out-of-band SSM params (operationally-variable / env-specific-unknown). Run
# it AFTER `cdk deploy open-swe-<env>` and BEFORE the EC2/T12 box first boots, so
# fetch-config.sh finds every REQUIRED var populated and never writes a partial .env.
#
# The secret list below mirrors config-store.ts SECRET_VARS (28 names — the
# inventory's "29" double-counted JUDGE_ANTHROPIC_BASE_URL, which is config
# (SSM), not a secret). Keep the two lists in lockstep.
#
# SAFETY:
# - NO real secret values live in this file — every secret is a <FILL> placeholder.
# Replace <FILL...> inline at run time, pipe from a vault, or export the
# matching OPENSWE_PUT_<VAR> env var; NEVER commit real values.
# - It does NOT touch the IaC-managed SSM params (SANDBOX_TYPE, DEFAULT_REPO_OWNER,
# ALLOWED_GITHUB_ORGS, DEFAULT_REPO_NAME, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID)
# — CDK owns those; setting them here would cause drift.
# - Secrets go to Secrets Manager; the box's instance role grants read on
# open-swe-<env>/* and /open-swe-<env>/* (no kms:Decrypt — AWS-managed keys).
#
# Idempotent: put-secret-value adds a new AWSCURRENT version; put-parameter
# --overwrite updates in place.
set -euo pipefail
ENV="${1:-}"
case "$ENV" in
dev | prod) ;;
*)
echo "usage: put-config.sh <dev|prod>" >&2
exit 2
;;
esac
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
SECRET_PREFIX="open-swe-${ENV}/"
SSM_PREFIX="/open-swe-${ENV}/"
command -v aws >/dev/null 2>&1 || { echo "put-config: 'aws' not found on PATH" >&2; exit 3; }
# --- helpers -----------------------------------------------------------------
# put_secret VAR : set the value of an EXISTING secret shell open-swe-<env>/VAR.
# Resolution order for the value: $OPENSWE_PUT_<VAR> env var, else the literal
# <FILL> placeholder (which aborts so an unset secret is never silently shipped).
put_secret() {
local var="$1"
local override_name="OPENSWE_PUT_${var}"
local value="${!override_name:-<FILL>}"
if [ "$value" = "<FILL>" ]; then
echo "put-config[$ENV]: SKIP secret ${var} (no value; set ${override_name} or edit inline)" >&2
return 0
fi
aws secretsmanager put-secret-value \
--secret-id "${SECRET_PREFIX}${var}" \
--secret-string "$value" \
--region "$REGION" \
--no-cli-pager >/dev/null
echo "put-config[$ENV]: set secret ${var}"
}
# put_param VAR [TYPE] : create/update an out-of-band SSM param /open-swe-<env>/VAR.
# TYPE defaults to String; pass SecureString for anything sensitive-but-not-a-secret.
put_param() {
local var="$1" type="${2:-String}"
local override_name="OPENSWE_PUT_${var}"
local value="${!override_name:-<FILL>}"
if [ "$value" = "<FILL>" ]; then
echo "put-config[$ENV]: SKIP param ${var} (no value; set ${override_name} or edit inline)" >&2
return 0
fi
aws ssm put-parameter \
--name "${SSM_PREFIX}${var}" \
--value "$value" \
--type "$type" \
--overwrite \
--region "$REGION" \
--no-cli-pager >/dev/null
echo "put-config[$ENV]: set param ${var} (${type})"
}
# --- 1) Secrets (open-swe-<env>/<VAR>) — 28 shells from config-store.ts -------
# REQUIRED at boot (fetch-config fail-fast): DASHBOARD_JWT_SECRET,
# TOKEN_ENCRYPTION_KEY, GITHUB_APP_PRIVATE_KEY/CLIENT_SECRET, the active provider
# key(s) (ANTHROPIC_API_KEY + OPENAI_API_KEY by default), LANGSMITH_API_KEY_PROD,
# and (prod) GITHUB_WEBHOOK_SECRET + SLACK_SIGNING_SECRET.
put_secret ANTHROPIC_API_KEY # REQUIRED — primary builder provider key
put_secret OPENAI_API_KEY # REQUIRED — default reviewer provider key
put_secret DASHBOARD_JWT_SECRET # REQUIRED — dashboard session JWT signing
put_secret TOKEN_ENCRYPTION_KEY # REQUIRED — Fernet key(s) for GH-token crypto
put_secret GITHUB_APP_PRIVATE_KEY # REQUIRED — GitHub App PEM (multiline; quote it)
put_secret GITHUB_APP_CLIENT_SECRET # REQUIRED — dashboard OAuth login
put_secret LANGSMITH_API_KEY_PROD # REQUIRED (langsmith sandbox) — prod key
put_secret GITHUB_WEBHOOK_SECRET # prod-REQUIRED — GitHub webhook signature
put_secret SLACK_SIGNING_SECRET # prod-REQUIRED — Slack webhook signature
put_secret LINEAR_WEBHOOK_SECRET # required only when Linear is wired
# Optional / conditional secrets — set the ones this deployment actually uses.
put_secret CORRIDOR_API_TOKEN # optional — Corridor MCP
put_secret CORRIDOR_MCP_TOKEN # optional — Corridor MCP (alt name)
put_secret CORRIDOR_TOKEN # optional — Corridor MCP (alt name)
put_secret DAYTONA_API_KEY # only if SANDBOX_TYPE=daytona
put_secret EXA_API_KEY # optional — Exa web search
put_secret FIREWORKS_API_KEY # only if a fireworks: model is used
put_secret GITHUB_PAT # optional — PAT fallback
put_secret GOOGLE_API_KEY # only if a google_genai: model is used
put_secret GROQ_API_KEY # only if a groq: model is used
put_secret JUDGE_ANTHROPIC_API_KEY # optional — eval judge (falls back to ANTHROPIC)
put_secret LANGSMITH_API_KEY # optional — LangSmith (dev)
put_secret LANGCHAIN_API_KEY # optional — LangSmith alt name
put_secret LINEAR_API_KEY # optional — Linear API
put_secret RUNLOOP_API_KEY # only if SANDBOX_TYPE=runloop
put_secret SLACK_BOT_TOKEN # optional — Slack bot token
put_secret SLACK_CLIENT_SECRET # optional — Slack OAuth
put_secret USER_ID_API_KEY_MAP # optional — JSON map user id -> API key (per-user auth)
put_secret X_SERVICE_AUTH_JWT_SECRET # optional — service-auth JWT
# --- 2) Out-of-band SSM params (/open-swe-<env>/<VAR>) ------------------------
# These are NOT created by CDK (operationally-variable / env-specific-unknown).
# put_param creates them on first run.
put_param DEFAULT_SANDBOX_SNAPSHOT_ID # REQUIRED (langsmith) — changes every rebuild
put_param GITHUB_APP_ID # REQUIRED — GitHub App numeric id
put_param GITHUB_APP_INSTALLATION_ID # REQUIRED — GitHub App installation id
put_param GITHUB_APP_CLIENT_ID # REQUIRED — dashboard OAuth client id
put_param GITHUB_OAUTH_PROVIDER_ID # GitHub OAuth provider id
put_param LANGSMITH_TENANT_ID_PROD # LangSmith prod tenant id
put_param LANGSMITH_URL_PROD # LangSmith prod URL
put_param LANGSMITH_ENDPOINT # LangSmith API endpoint
put_param LANGSMITH_ENDPOINT_PROD # LangSmith prod API endpoint
put_param LANGSMITH_HOST_API_URL # LangSmith host API URL
put_param LANGGRAPH_URL # LangGraph server URL
put_param LANGGRAPH_URL_PROD # LangGraph server URL (prod)
put_param LANGCHAIN_REVISION_ID # LangChain revision id
put_param SLACK_CLIENT_ID # Slack OAuth client id
put_param SLACK_TEAM_ID # Slack workspace/team id
put_param SLACK_BOT_USER_ID # Slack bot user id
put_param SLACK_BOT_USERNAME # Slack bot username
put_param SLACK_REPO_OWNER # Slack default repo owner
put_param SLACK_REPO_NAME # Slack default repo name
put_param CONFIGURED_ADMINS # dashboard admin GitHub logins (comma list)
put_param OBSERVABILITY_AUTHORIZED_EMAILS # observability allowlist (comma list)
put_param PUBLIC_REPO_ORG_GATE # public-repo trigger gate (org name; empty=off)
put_param ALLOWED_GITHUB_REPOS # extra repo allowlist (comma list)
put_param LLM_FALLBACK_MODEL_ID # optional — model fallback id
put_param DATADOG_MCP_TOOLSETS # optional — Datadog MCP toolsets
put_param NOTION_MCP_CLIENT_NAME # optional — Notion MCP client name
put_param API_STANDARDS_SKILL_HANDLE # optional — API standards skill handle
put_param REPO_SNAPSHOT_BASE_IMAGE # optional — repo snapshot base image
put_param REPO_SNAPSHOT_BUILD_TIMEOUT_SECONDS # optional — snapshot build timeout
put_param REPO_SNAPSHOT_STALE_BUILD_SECONDS # optional — snapshot stale threshold
echo "put-config[$ENV]: done. Verify with fetch-config.sh ${ENV} before first boot."

View file

@ -5,70 +5,180 @@
# to it (team model settings, user mappings) is lost on every restart. This
# script idempotently re-PUTs that state and is wired as a systemd
# ExecStartPost on the open-swe.service unit so it runs after each start.
# IT MUST RE-RUN ON EVERY RESTART — the in-memory store starts empty each boot.
#
# Replace this with Postgres-backed durability (Aegra / `langgraph up`) to make
# the store survive restarts and drop this script.
#
# Configuration comes from the environment (set these in the service env or a
# sourced file alongside the app .env) — no real values are committed here:
# OPENSWE_BASE_URL default http://127.0.0.1:2024
# OPENSWE_AGENT_MODEL default anthropic:claude-opus-4-8
# OPENSWE_AGENT_EFFORT default high
# OPENSWE_REVIEWER_MODEL default openai:gpt-5.5
# OPENSWE_REVIEWER_EFFORT default high
# OPENSWE_DEFAULT_REPO e.g. your-org/your-pilot-repo (required)
# OPENSWE_OWNER_LOGIN GitHub login of the triggering owner (required)
# OPENSWE_OWNER_EMAIL work email mapped to that login (required)
# AWS-env-aware: pass the env as $1 (dev|prod). Seed values (default repo, model
# ids, user mappings) come from the fetch-config-materialized .env.
#
# SECURITY (T5 /sh-security-review):
# - SH-INJ-001: this script NEVER `source`s the .env. python-dotenv and bash
# have incompatible escaping, and a config value like `$(cmd)` would execute
# when sourced. We extract the few single-line seed keys with a non-eval
# reader (read_env) instead.
# - SH-INJ-003: the store PUT bodies are built with `jq --arg`, so values are
# always JSON-encoded (no string interpolation into a JSON heredoc).
# - SH-INJ-004: BASE is pinned to loopback — never derived from store/SSM
# config (LANGGRAPH_URL) — so a tampered value can't redirect the PUTs.
# - SC-02: not sourcing the .env means secrets are never exported into this
# script's (or curl's) environment.
#
# Seed values read from the materialized .env (set in SSM /open-swe-<env>/*):
# DEFAULT_REPO_OWNER / DEFAULT_REPO_NAME -> team_settings default_repo
# LLM_MODEL_ID -> default builder model (fallback)
# SEED_AGENT_MODEL / SEED_AGENT_EFFORT -> builder model + effort (optional)
# SEED_REVIEWER_MODEL / SEED_REVIEWER_EFFORT -> reviewer model + effort (optional)
# SEED_USER_MAPPINGS -> "login:email,login:email" (optional)
# CONFIGURED_ADMINS -> "login,email" fallback for the mapping
# Legacy OPENSWE_* process-env overrides are still honored (highest precedence).
set -euo pipefail
BASE="${OPENSWE_BASE_URL:-http://127.0.0.1:2024}"
AGENT_MODEL="${OPENSWE_AGENT_MODEL:-anthropic:claude-opus-4-8}"
AGENT_EFFORT="${OPENSWE_AGENT_EFFORT:-high}"
REVIEWER_MODEL="${OPENSWE_REVIEWER_MODEL:-openai:gpt-5.5}"
REVIEWER_EFFORT="${OPENSWE_REVIEWER_EFFORT:-high}"
DEFAULT_REPO="${OPENSWE_DEFAULT_REPO:?set OPENSWE_DEFAULT_REPO=owner/repo}"
OWNER_LOGIN="${OPENSWE_OWNER_LOGIN:?set OPENSWE_OWNER_LOGIN=github-login}"
OWNER_EMAIL="${OPENSWE_OWNER_EMAIL:?set OPENSWE_OWNER_EMAIL=work-email}"
ENV="${1:-${OPENSWE_ENV:-}}"
case "$ENV" in
dev | prod | "") ;; # empty allowed: pure-env / on-prem backward-compat mode
*)
echo "seed_store: ENV must be 'dev' or 'prod' (got '$ENV')" >&2
exit 2
;;
esac
ENV_DIR="${ENV_DIR:-/run/open-swe}"
ENV_FILE="${ENV_FILE:-${ENV_DIR}/.env}"
# read_env KEY -> prints the value of a SINGLE-LINE `KEY="..."` entry from the
# materialized .env WITHOUT shell evaluation (SH-INJ-001 fix). Seed keys are
# simple single-line values; multiline secrets (e.g. the PEM) are never read
# here. Returns empty if the key is absent/unreadable.
read_env() {
local key="$1" line
[ -r "$ENV_FILE" ] || return 0
line="$(grep -m1 -- "^${key}=" "$ENV_FILE" 2>/dev/null || true)"
[ -n "$line" ] || return 0
line="${line#*=}"
# strip one layer of surrounding double quotes (python-dotenv double-quoted form)
if [ "${line#\"}" != "$line" ]; then line="${line%\"}"; line="${line#\"}"; fi
# reverse python-dotenv double-quote escaping (only \" and \\ are escaped)
line="${line//\\\"/\"}"; line="${line//\\\\/\\}"
printf '%s' "$line"
}
# Process-env override (legacy/on-prem) -> .env value -> default.
pick() { # pick DEFAULT OVERRIDE_VALUE FILE_KEY...
local def="$1" override="$2"; shift 2
if [ -n "$override" ]; then printf '%s' "$override"; return; fi
local k v
for k in "$@"; do v="$(read_env "$k")"; [ -n "$v" ] && { printf '%s' "$v"; return; }; done
printf '%s' "$def"
}
# BASE is loopback-pinned (SH-INJ-004): this on-box seeder only talks to the
# local server; OPENSWE_PORT may override the port but never the host.
BASE="http://127.0.0.1:${OPENSWE_PORT:-2024}"
AGENT_MODEL="$(pick 'anthropic:claude-opus-4-8' "${OPENSWE_AGENT_MODEL:-}" SEED_AGENT_MODEL LLM_MODEL_ID)"
AGENT_EFFORT="$(pick 'high' "${OPENSWE_AGENT_EFFORT:-}" SEED_AGENT_EFFORT)"
REVIEWER_MODEL="$(pick 'openai:gpt-5.5' "${OPENSWE_REVIEWER_MODEL:-}" SEED_REVIEWER_MODEL)"
REVIEWER_EFFORT="$(pick 'high' "${OPENSWE_REVIEWER_EFFORT:-}" SEED_REVIEWER_EFFORT)"
# default_repo = owner/name from AWS config (DEFAULT_REPO_OWNER is hard-pinned
# away from upstream by fetch-config.sh).
REPO_OWNER="$(pick '' '' DEFAULT_REPO_OWNER)"
REPO_NAME="$(pick '' '' DEFAULT_REPO_NAME)"
if [ -n "${OPENSWE_DEFAULT_REPO:-}" ]; then
DEFAULT_REPO="$OPENSWE_DEFAULT_REPO"
elif [ -n "$REPO_OWNER" ] && [ -n "$REPO_NAME" ]; then
DEFAULT_REPO="${REPO_OWNER}/${REPO_NAME}"
else
echo "seed_store: set OPENSWE_DEFAULT_REPO=owner/repo (or DEFAULT_REPO_OWNER + DEFAULT_REPO_NAME in .env)" >&2
exit 1
fi
# user_mappings: explicit SEED_USER_MAPPINGS ("login:email,..."), then legacy
# OPENSWE_OWNER_LOGIN/EMAIL, then parse CONFIGURED_ADMINS ("login,email").
SEED_MAP="$(pick '' "${SEED_USER_MAPPINGS:-}" SEED_USER_MAPPINGS)"
ADMINS="$(pick '' "${CONFIGURED_ADMINS:-}" CONFIGURED_ADMINS)"
declare -a MAPPINGS=()
if [ -n "$SEED_MAP" ]; then
IFS=',' read -r -a _pairs <<<"$SEED_MAP"
for p in "${_pairs[@]}"; do
p="${p//[[:space:]]/}"
[ -n "$p" ] && MAPPINGS+=("$p")
done
elif [ -n "${OPENSWE_OWNER_LOGIN:-}" ] && [ -n "${OPENSWE_OWNER_EMAIL:-}" ]; then
MAPPINGS+=("${OPENSWE_OWNER_LOGIN}:${OPENSWE_OWNER_EMAIL}")
elif [ -n "$ADMINS" ]; then
_login="" _email=""
IFS=',' read -r -a _toks <<<"$ADMINS"
for t in "${_toks[@]}"; do
t="${t//[[:space:]]/}"
[ -z "$t" ] && continue
case "$t" in
*@*) [ -z "$_email" ] && _email="$t" ;;
*) [ -z "$_login" ] && _login="$t" ;;
esac
done
[ -n "$_login" ] && [ -n "$_email" ] && MAPPINGS+=("${_login}:${_email}")
fi
# No user mapping is NON-FATAL (OSWE-SEED-03 precedent: never fail the unit into a
# restart loop over a seeding gap — same as the server-not-ready path below). The
# server itself is healthy; an unseeded user_mappings table only means the @openswe
# trigger won't resolve a commenter, which a deployment-validation env (e.g. dev)
# does not need. team_settings is still seeded. Set SEED_USER_MAPPINGS (or
# CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL) to seed the mapping when wanted.
if [ "${#MAPPINGS[@]}" -eq 0 ]; then
echo "seed_store: no user mapping resolved — skipping user_mappings seed (set SEED_USER_MAPPINGS or OPENSWE_OWNER_LOGIN/EMAIL to enable)" >&2
fi
NOW="$(date -u +%Y-%m-%dT%H:%M:%S+00:00)"
# Wait for the server to accept requests (up to ~60s).
# Wait for the server to accept requests (up to ~60s). Authoritative (OSWE-SEED-03):
# if it never comes up, log and exit 0 — do NOT fail the unit into a restart loop.
READY=0
for _ in $(seq 1 30); do
[ "$(curl -s -o /dev/null -w '%{http_code}' "$BASE/ok" || true)" = "200" ] && break
if [ "$(curl -s -o /dev/null -w '%{http_code}' "$BASE/ok" || true)" = "200" ]; then READY=1; break; fi
sleep 2
done
if [ "$READY" -ne 1 ]; then
echo "seed_store: server not ready at $BASE after ~60s; skipping seed (will reseed on next restart)" >&2
exit 0
fi
# 1) team_settings/default — builder + reviewer models (NOT read from env by the
# app; the store value wins over LLM_MODEL_ID). gpt-4.1 is NOT in this fork's
# SUPPORTED_MODELS, so the reviewer uses gpt-5.5 (cross-family vs the builder).
curl -s -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d @- <<JSON
{"namespace":["team_settings"],"key":"default","value":{
"review_draft_prs": false,
"pr_summaries": true,
"review_trace_links": true,
"org_guidelines": null,
"default_agent_model": "$AGENT_MODEL",
"default_agent_reasoning_effort": "$AGENT_EFFORT",
"default_agent_subagent_model": "$AGENT_MODEL",
"default_agent_subagent_reasoning_effort": "$AGENT_EFFORT",
"default_repo": "$DEFAULT_REPO",
"default_reviewer_model": "$REVIEWER_MODEL",
"default_reviewer_reasoning_effort": "$REVIEWER_EFFORT",
"default_reviewer_subagent_model": "$REVIEWER_MODEL",
"default_reviewer_subagent_reasoning_effort": "$REVIEWER_EFFORT",
"default_grouping_model": null,
"default_grouping_reasoning_effort": null,
"default_chat_model": null,
"default_chat_reasoning_effort": null,
"updated_at": "$NOW"
}}
JSON
# 1) team_settings/default — JSON built with jq --arg (SH-INJ-003 fix).
team_body="$(jq -n \
--arg am "$AGENT_MODEL" --arg ae "$AGENT_EFFORT" \
--arg rm "$REVIEWER_MODEL" --arg re "$REVIEWER_EFFORT" \
--arg repo "$DEFAULT_REPO" --arg now "$NOW" \
'{namespace:["team_settings"],key:"default",value:{
review_draft_prs:false, pr_summaries:true, review_trace_links:true,
org_guidelines:null,
default_agent_model:$am, default_agent_reasoning_effort:$ae,
default_agent_subagent_model:$am, default_agent_subagent_reasoning_effort:$ae,
default_repo:$repo,
default_reviewer_model:$rm, default_reviewer_reasoning_effort:$re,
default_reviewer_subagent_model:$rm, default_reviewer_subagent_reasoning_effort:$re,
default_grouping_model:null, default_grouping_reasoning_effort:null,
default_chat_model:null, default_chat_reasoning_effort:null,
updated_at:$now}}')"
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$team_body" >/dev/null \
|| echo "seed_store: WARN team_settings PUT failed (will reseed next restart)" >&2
# 2) user_mappings/<login> — required, or the @openswe trigger ignores the commenter.
curl -s -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d @- <<JSON
{"namespace":["user_mappings"],"key":"$OWNER_LOGIN","value":{
"github_login":"$OWNER_LOGIN","work_email":"$OWNER_EMAIL","slack_user_id":null,
"source":"slack_oauth","status":"active","created_at":"$NOW","updated_at":"$NOW"
}}
JSON
for pair in "${MAPPINGS[@]}"; do
login="${pair%%:*}"
email="${pair#*:}"
if [ -z "$login" ] || [ -z "$email" ] || [ "$login" = "$pair" ]; then
echo "seed_store: skipping malformed mapping '$pair' (want login:email)" >&2
continue
fi
map_body="$(jq -n --arg login "$login" --arg email "$email" --arg now "$NOW" \
'{namespace:["user_mappings"],key:$login,value:{
github_login:$login, work_email:$email, slack_user_id:null,
source:"slack_oauth", status:"active", created_at:$now, updated_at:$now}}')"
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$map_body" >/dev/null \
|| echo "seed_store: WARN user_mapping PUT failed for $login" >&2
done
echo "seed_store: done at $NOW"
echo "seed_store: done at $NOW (env=${ENV:-none}, repo=$DEFAULT_REPO, mappings=${#MAPPINGS[@]})"

21
infra/.gitignore vendored Normal file
View file

@ -0,0 +1,21 @@
# CDK / build output
cdk.out/
*.js
*.d.ts
*.js.map
# ...but jest.config.js is hand-authored config, not build output — keep it.
!jest.config.js
# deps
node_modules/
# env
.env
# coverage
coverage/
# NOTE: cdk.context.json IS committed on purpose (pins the AMI / lookups so
# deploys are reproducible and don't implicitly pick up a newer AMI — see
# lib/constructs/ami-cache.ts and feedback_inline_ebs_volumes).
!cdk.context.json

301
infra/README.md Normal file
View file

@ -0,0 +1,301 @@
# open-swe infra (CDK TypeScript)
AWS infrastructure for the Open SWE → AWS migration. **Synth-only at this stage —
nothing here is deployed yet.** All IAM is applied only after the Phase-1 security
gate (T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review`) clears (T6).
## Layout
```
infra/
├── bin/
│ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
│ ├── aspects/
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
│ └── constructs/
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│ └── ami-cache.ts # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore
```
## Stacks
| Stack name (kebab) | Construct | Contents |
|---|---|---|
| `open-swe-iam` | `OpenSweIamStack` | Account-level shared GitHub OIDC deploy roles (singletons). |
| `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role`, config store (T11), and the EC2 box + ALB ingress (T12, `AppService`). |
| `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role`, config store, and the EC2 box + ALB ingress. |
Account `328440206208`, region `us-east-1`. Stack names are set explicitly so CDK
never defaults to PascalCase; resource names follow `open-swe-<env>-*`.
> The two env stacks (`open-swe-dev` / `open-swe-prod`) are the required pair. The
> shared OIDC deploy roles are account-wide singletons (one `RoleName` each), so
> they live in their own dedicated `open-swe-iam` stack rather than being
> duplicated across the env stacks — and that stack deploys first (see ordering).
## IAM roles defined (unapplied)
- **`githubdeploy-open-swe-infra`** — GitHub OIDC role for CDK/CFN infra deploys.
Trust scoped to `repo:Sea-Haven-Industries/open-swe` on the `main`/`dev`
branches only. Permission is the org-standard CDK pattern: `sts:AssumeRole` on
the CDK bootstrap roles (`cdk-hnb659fds-*`) — the real CFN/IAM/resource scope
lives in the bootstrap `cfn-exec-role`, not in this role.
- **`githubdeploy-open-swe-app`** — GitHub OIDC role for app deploys. Tag-scoped
`ssm:SendCommand` (instances tagged `project=open-swe` + `env in {dev,prod}`) +
read-only access to the `open-swe-<env>-assets` S3 artifact buckets.
- **`open-swe-<env>-instance-role`** — EC2 instance role, least-privilege: read
`open-swe-<env>-assets` (S3), read `/open-swe-<env>/*` (SSM), read
`open-swe-<env>/*` (Secrets Manager), put `/open-swe/<env>/*` CloudWatch Logs,
plus `AmazonSSMManagedInstanceCore` for SSM agent registration. No admin.
The GitHub OIDC provider already exists account-wide (created for seahaven-site);
it is referenced by ARN, never re-created.
## Kebab-case naming Aspect
`KebabNamingAspect` (applied app-wide in `bin/app.ts`) fails synth via
`Annotations.addError` when a stack name or an explicit physical resource name
(`RoleName`, `BucketName`, …) is not kebab-case. Path-style names (Secrets
Manager `a/b`, SSM `/a/b`, log groups `/aws/.../x`) are validated per `/`-segment.
CDK logical construct ids are intentionally NOT validated (they are conventionally
PascalCase). Covered by `test/kebab-naming-aspect.test.ts`.
## Config store (Secrets Manager + SSM shells — T11)
`ConfigStore` (`lib/constructs/config-store.ts`, one per env from `OpenSweStack`)
renders the resource shells the boot hook `deploy/seahaven/fetch-config.sh` reads.
The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var
names as the last path segment — `open-swe-<env>/<VAR>` for secrets,
`/open-swe-<env>/<VAR>` (FLAT) for config — because fetch-config strips the prefix
and exports that segment verbatim.
Three buckets:
1. **Secret shells (Secrets Manager) — 27 secrets.** Created value-LESS (an L1
`CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which
CloudFormation creates as an empty secret). The real value is set **out-of-band**
(`put-config.sh`) — CDK never owns it, so a later `cdk deploy` can never clobber
it. `UpdateReplacePolicy/DeletionPolicy: Retain` so a teardown can't destroy
operator-set secret material. AWS-managed key (no CMK — matches the instance
role, which omits `kms:Decrypt`).
> The T9 header says "29 secrets" but its table enumerates **27** distinct VAR
> names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29.
> **Confirm** the 27-vs-29 count (code-only candidates not in the table:
> `USER_ID_API_KEY_MAP`, `JUDGE_ANTHROPIC_BASE_URL`).
2. **IaC-managed SSM config — 8 params, real values owned in code:**
| Param | dev | prod |
|---|---|---|
| `SANDBOX_TYPE` | `langsmith` | `langsmith` |
| `DEFAULT_REPO_OWNER` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `DEFAULT_REPO_NAME` | `open-swe-pilot` *(confirm)* | `open-swe-pilot` *(confirm)* |
| `DASHBOARD_BASE_URL` | `https://openswe-dev.seahaven.com` *(confirm host)* | `https://openswe.seahaven.com` *(confirm host)* |
| `DASHBOARD_API_BASE_URL` | same as base | same as base |
| `DASHBOARD_ALLOWED_ORIGINS` | same as base | same as base |
| `LLM_MODEL_ID` | `anthropic:claude-opus-4-8` *(confirm)* | `anthropic:claude-opus-4-8` *(confirm)* |
3. **Out-of-band SSM config — NOT created by CDK.** Operationally-variable or
env-specific-unknown values listed in `OUT_OF_BAND_SSM` and populated by
`put-config.sh`. The keystone is `DEFAULT_SANDBOX_SNAPSHOT_ID` (changes on every
snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the
GitHub App ids, Slack ids, and LangSmith tenant/urls.
### Kebab-Aspect deviation
`KebabNamingAspect` exempts `AWS::SecretsManager::Secret` and `AWS::SSM::Parameter`
from the kebab check (see the `KEBAB_EXEMPT_RESOURCE_TYPES` set) — the UPPER_SNAKE
env-var segment is a required, documented deviation for a lossless store→env
round-trip. Every other explicitly-named resource is still validated. Covered by a
dedicated case in `test/kebab-naming-aspect.test.ts`.
### Deploy ordering (values BEFORE the box boots)
The shells are synth-able now (T11). Population is out-of-band and happens **after**
`cdk deploy open-swe-<env>` but **before** the EC2/T12 box first boots:
```bash
cdk deploy open-swe-<env> # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod> # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
```
`put-config.sh` ships `<FILL>` placeholders only (no real secret values committed);
provide each value inline, via `OPENSWE_PUT_<VAR>` env vars, or from a vault. It does
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
## Compute + ingress (`AppService` — T12)
`AppService` (`lib/constructs/app-service.ts`, one per env from `OpenSweStack`)
builds the box and its path to the internet. A **single** internet-facing ALB
(`app/seahaven-com`) and a **single** VPC are shared with the on-prem
`seahaven-site` stack, so open-swe **imports** the VPC, the ALB security group
(`sg-0b0301deed193258a`), the `:443` listener, and the public `seahaven.com`
zone — and never owns/mutates them. It **adds**:
- **One ARM64 EC2 box** (`open-swe-<env>-box`, `t4g.medium` dev / `t4g.large`
prod) in **private1 (us-east-1a)** — same AZ as the single NAT for in-AZ egress.
`requireImdsv2`, gp3 **encrypted** root, `deleteOnTermination` (no RETAIN
volume — see below). `userDataCausesReplacement: true`; user-data is rendered
from `deploy/ami/user-data.sh`.
- **A standalone instance SG** reachable **only** from the shared ALB SG on `:80`
(nginx). Egress open (NAT). The ALB SG is opened to the box via a **standalone
`CfnSecurityGroupEgress`** so the imported (on-prem-owned) SG is never mutated.
- **A target group → instance `:80`** (nginx is the sole ingress; the LangGraph
control plane stays on loopback `:2024`). Health check `GET /healthz`.
- **Two listener rules** on the imported `:443` listener, both → the same TG:
- **Webhooks** (priority **2** dev / **3** prod): `host ∈ {openswe-<env>, hooks-<env>}.seahaven.com` **AND** path `/webhooks/*`.
- **Site** (priority **10** dev / **11** prod): `host = openswe-<env>.seahaven.com` (dashboard SPA + `/dashboard/api/`).
- **Route53 alias records** `openswe[-dev]` + `hooks[-dev]` → the shared ALB.
- **Four CloudWatch log groups** (`/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`) at **30-day** retention (IaC-owned; mirrors the CW-agent config).
### Listener-rule ordering (load-bearing)
The shared listener already has a **host-agnostic** `/webhooks/*` PATH rule at
**priority 5** (on-prem). ALB rules are first-match by ascending priority, so the
open-swe webhook rule **must** sit below 5 or every `…/webhooks/*` request (any
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
a host condition, so it does **not** steal the on-prem hosts' webhooks. The
dashboard "site" rule carries no path that collides with rule 5, so it sits at
10/11.
**Cross-stack coordination (T13 review).** The `seahaven-site` (on-prem) and
`open-swe` stacks both add resources to the *same imported* listener and ALB SG.
This is safe: each stack owns only the resources it declares (its own logical
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
and the standalone `CfnSecurityGroupEgress` never mutates the shared SG's own
definition (the pattern on-prem itself uses). The one shared namespace that needs
care is **listener-rule priority** (globally unique per listener; a collision is
a fail-*safe* deploy error, not silent drift). Ownership — keep disjoint:
`seahaven-site` = **4-7 + default**; `open-swe` = **2, 3, 10, 11**. open-swe's
webhook rules are host-scoped to its own `*.seahaven.com` hosts, so they never
match an on-prem `seahavenind.com` host.
### Security review (T5/T12 `/sh-security-review`)
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
One **confirmed medium** (OSWE-T12-01: nginx's 1 MB default `client_max_body_size`
would 413 large GitHub webhooks before in-app signature verification) is fixed in
`open-swe.nginx.conf` (`25m` on `/webhooks/`, `10m` on `/dashboard/api/`). The
hooks hostname is scoped to `/webhooks/*` only (OSWE-T12-02 hygiene). An
X-Forwarded-For spoof candidate was **killed** — no code trusts the leftmost XFF.
No confirmed critical/high; no block.
## Baked AMI + EBS-replacement discipline
`bakedOpenSweArm64()` (in `lib/constructs/ami-cache.ts`) pins the custom
**open-swe-base-arm64** image by EXACT id (`BAKED_OPEN_SWE_AMI_ID`) via
`MachineImage.genericLinux({ "us-east-1": "<ami-id>" })` — no SSM lookup, so synth
and deploy are fully offline/deterministic. The image is built by
`deploy/ami/open-swe-base.pkr.hcl` (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW
agent + boot templates, **no secrets**); the box's `user-data.sh` assumes that
baked layout (`/opt/open-swe`, `openswe` user, nginx, CW agent).
Pinning by exact id (vs a `most_recent` name filter) is what prevents a routine
deploy from silently swapping the AMI → **EC2 instance replacement** (the
file-share data-loss root cause — memory `feedback_inline_ebs_volumes`).
- `userDataCausesReplacement: true` is **deliberate** — user-data is
provisioning-only and carries no durable state.
- **No durable state on the box → no RETAIN volume.** The in-memory langgraph
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The goal is
replacement-*tolerance*, not avoidance.
- **Snapshot-before-replace** still applies operationally: before any replacing
deploy snapshot the root volume and wait `state=completed`, and re-verify "no
local-only durable state" first.
Refresh the AMI deliberately:
```bash
cd deploy/ami && packer build open-swe-base.pkr.hcl # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack # WILL show "requires replacement"
```
> `cdk.context.json` is `{}` — nothing is resolved via context anymore (the AMI is
> a static id pin), so synth makes no live AWS call.
## Commands
```bash
npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test # jest — naming Aspect
```
## CI/CD (T18 — `.github/workflows/ci-infra.yml` + `cd-infra.yml`)
Path-filtered, OIDC-only (no static keys). The Python agent keeps its own
`ci.yml` ("Agent CI"); these two add the `/infra` half.
| Workflow | Trigger | Does |
|---|---|---|
| `ci-infra.yml` | PR touching `infra/**` | `tsc` + `jest` + `cdk synth` (reusable `ci-typescript-cdk.yaml`). |
| `cd-infra.yml` | push to `dev`/`main` touching `infra/**`, or dispatch | CI (pre-deploy) → per-env `cdk deploy`. |
`cd-infra.yml` flow:
- **push to `dev`** → CI green → **auto** `cdk deploy OpenSweDevStack` (assumes
`githubdeploy-open-swe-infra-dev`; the job declares **no** `environment:`, so the
OIDC subject is `…:ref:refs/heads/dev` — matching that role's trust).
- **push to `main`** → CI green → `cdk deploy OpenSweProdStack` behind the
**`prod` GitHub Environment** (required reviewer = Adam). The `environment: prod`
declaration both fires the manual-approval gate and makes the OIDC subject
`…:environment:prod` — matching `githubdeploy-open-swe-infra-prod`'s trust.
**Why not the reusable `cd-cdk.yaml`:** it runs `cdk deploy --all`, which from a
single-env push would deploy the *other* env + the shared IAM stack — breaking the
per-env boundary. So CD targets one stack explicitly per env. The shared
`open-swe-iam` stack is **not** deployed by CD (privileged, human-gated — T6).
**Gating note:** infra CI is enforced at the *deploy* boundary (`cd-infra`'s
`deploy-*` jobs `needs: ci`), not as a branch-protection required check —
path-filtering a *required* check would deadlock app-only PRs (a skipped required
check never satisfies). Making `Infra CI` a required check later needs a
skip-aware shim or dropping its path filter.
**Prerequisites (set post-T6, when the roles exist):**
- repo **variables** `AWS_DEPLOY_ROLE_INFRA_DEV` / `AWS_DEPLOY_ROLE_INFRA_PROD`
= the `githubdeploy-open-swe-infra-<env>` role ARNs (`open-swe-iam` outputs).
- a GitHub **Environment** named `prod` with Adam as a required reviewer.
> App-side CD (CI → S3 artifact → SSM deploy via `githubdeploy-open-swe-app-<env>`)
> is **T19**, not here.
## Deploy ordering (when the gate clears — NOT yet)
1. **`open-swe-iam` first** — apply the IAM stack (T6, human-gated), then set the
repo `AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}` variables from its role-ARN outputs and
configure the `prod` Environment reviewer (BLOCK#3).
2. **Security gate** — T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review` on
the synth; resolve every confirmed critical/high.
3. **IAM applied** (T6) — only after the gate.
4. Env stacks: first `open-swe-dev` (T14, manual validate), then CD auto-deploys
dev on push; `open-swe-prod` (T21) behind the `prod` Environment approval.
## Version policy
`aws-cdk-lib` is pinned EXACT (`2.260.0`) — no `^`/`~`. Dependabot keeps it
current; CI (`npm ci` + `cdk synth`) + dependency review gate each bump. See
`aws-infrastructure.md` "CDK Version Policy" and memory
`feedback_cdk_lib_bundled_deps`.

35
infra/bin/app.ts Normal file
View file

@ -0,0 +1,35 @@
#!/usr/bin/env node
import "source-map-support/register";
import * as cdk from "aws-cdk-lib";
import { ACCOUNT, REGION } from "../lib/config";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
import { KebabNamingAspect } from "../lib/aspects/kebab-naming-aspect";
const app = new cdk.App();
const env = { account: ACCOUNT, region: REGION };
// Account-level shared OIDC deploy roles (singletons). Deployed FIRST.
new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env,
});
// The two env stacks — explicit kebab-case stackName (never let CDK default to
// PascalCase), env-parameterised so resources are `open-swe-<env>-*`.
new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env,
envName: "dev",
});
new OpenSweStack(app, "OpenSweProdStack", {
stackName: "open-swe-prod",
env,
envName: "prod",
});
// Fail synth on any non-kebab-case explicit resource/stack name.
cdk.Aspects.of(app).add(new KebabNamingAspect());
app.synth();

1
infra/cdk.context.json Normal file
View file

@ -0,0 +1 @@
{}

21
infra/cdk.json Normal file
View file

@ -0,0 +1,21 @@
{
"app": "npx ts-node --prefer-ts-exts bin/app.ts",
"watch": {
"include": ["**"],
"exclude": [
"README.md",
"cdk*.json",
"**/*.d.ts",
"**/*.js",
"tsconfig.json",
"package*.json",
"node_modules",
"cdk.out"
]
},
"context": {
"@aws-cdk/aws-lambda:recognizeLayerVersion": true,
"@aws-cdk/core:checkSecretUsage": true,
"@aws-cdk/core:target-partitions": ["aws"]
}
}

9
infra/jest.config.js Normal file
View file

@ -0,0 +1,9 @@
module.exports = {
testEnvironment: "node",
roots: ["<rootDir>/test"],
testMatch: ["**/*.test.ts"],
preset: "ts-jest",
transform: {
"^.+\\.tsx?$": ["ts-jest", { tsconfig: "tsconfig.json" }],
},
};

View file

@ -0,0 +1,119 @@
import { Annotations, CfnResource, IAspect, Stack, Token } from "aws-cdk-lib";
import { IConstruct } from "constructs";
/**
* One "/"-delimited segment must be lower kebab-case: `a-b-c`, digits allowed.
*/
const KEBAB_SEGMENT = /^[a-z0-9]+(-[a-z0-9]+)*$/;
/**
* CloudFormation property keys that carry an *explicit physical name*. The
* codegen'd L1 stores these either camelCased (`roleName`) or CFN-cased
* (`RoleName`) depending on the construct, so the aspect matches keys
* case-insensitively.
*
* We deliberately validate physical NAMES + the stack name only — not CDK
* logical construct ids (those are conventionally PascalCase, e.g.
* `InfraDeployRole`, and validating them would be wrong).
*/
const NAME_PROPERTY_KEYS = [
"RoleName",
"BucketName",
"FunctionName",
"TableName",
"LogGroupName",
"QueueName",
"TopicName",
"SecretName",
"StreamName",
"RepositoryName",
"DBInstanceIdentifier",
"DBClusterIdentifier",
"StateMachineName",
"RuleName",
"UserPoolName",
];
// NOTE: `PolicyName` is intentionally NOT checked — CDK auto-generates inline
// `DefaultPolicy` names (e.g. "InstanceRoleDefaultPolicyF15F...") from the
// logical id; those are not explicit, user-controlled physical names and are
// outside the naming convention's scope.
const NAME_KEYS_LC = new Set(NAME_PROPERTY_KEYS.map((k) => k.toLowerCase()));
/**
* Resource types whose physical NAME is a REQUIRED deviation from kebab-case:
* the open-swe config store names secrets `open-swe-<env>/<ENV_VAR_NAME>` and SSM
* params `/open-swe-<env>/<ENV_VAR_NAME>`, where the last segment is the LITERAL
* UPPER_SNAKE environment-variable name. The boot hook
* (deploy/seahaven/fetch-config.sh) strips the prefix and exports that segment
* verbatim, so a lossless store→env round-trip needs the exact env-var name —
* it cannot be kebab-cased. These two resource types are therefore exempt; every
* OTHER explicitly-named resource is still validated. (The `open-swe-<env>`
* prefix is code-generated from `prefix(env)` and is always kebab-case.)
*/
const KEBAB_EXEMPT_RESOURCE_TYPES = new Set([
"AWS::SecretsManager::Secret",
"AWS::SSM::Parameter",
]);
/**
* `true` when every non-empty "/"-delimited segment is kebab-case.
*
* Path-style names are tolerated so the same check works for Secrets Manager
* (`open-swe-dev/foo`), SSM params (`/open-swe-dev/foo`) and log groups
* (`/open-swe/dev/agent`): each segment is validated independently, and a
* leading slash (empty first segment) is ignored.
*/
export function isKebabCase(value: string): boolean {
return value
.split("/")
.filter((seg) => seg.length > 0)
.every((seg) => KEBAB_SEGMENT.test(seg));
}
/**
* Aspect that FAILS synth (`Annotations.addError`) when an explicitly-named
* resource — or a stack name — is not kebab-case. Enforces the org naming
* convention (naming-conventions.md) deterministically at synth time so a
* non-conforming name can never reach a deploy. Wired in bin/app.ts via
* `Aspects.of(app).add(new KebabNamingAspect())`.
*/
export class KebabNamingAspect implements IAspect {
public visit(node: IConstruct): void {
if (node instanceof Stack) {
const name = node.stackName;
if (!Token.isUnresolved(name) && !isKebabCase(name)) {
Annotations.of(node).addError(
`Stack name "${name}" is not kebab-case (open-swe naming convention).`,
);
}
return;
}
if (node instanceof CfnResource) {
// The config store's Secret/Parameter names carry the literal UPPER_SNAKE
// env-var name per the fetch-config naming contract — a required deviation.
if (KEBAB_EXEMPT_RESOURCE_TYPES.has(node.cfnResourceType)) {
return;
}
// `_cfnProperties` is the props as set on the L1; resolve to collapse any
// intrinsic tokens (refs/getatt) so only literal strings are checked.
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const raw = (node as any)._cfnProperties ?? {};
const resolved = Stack.of(node).resolve(raw) ?? {};
for (const [key, value] of Object.entries(resolved)) {
if (
NAME_KEYS_LC.has(key.toLowerCase()) &&
typeof value === "string" &&
!Token.isUnresolved(value) &&
!isKebabCase(value)
) {
Annotations.of(node).addError(
`Resource "${node.node.path}" property ${key}="${value}" is not kebab-case ` +
`(open-swe naming convention).`,
);
}
}
}
}
}

47
infra/lib/config.ts Normal file
View file

@ -0,0 +1,47 @@
/**
* Shared, non-sensitive constants for the open-swe infra app.
* Account / region are locked per the migration spec (TODO.md "Architecture (locked)").
*/
export const ACCOUNT = "328440206208";
export const REGION = "us-east-1";
export const GITHUB_ORG = "Sea-Haven-Industries";
export const GITHUB_REPO = "open-swe";
export type EnvName = "dev" | "prod";
/** `open-swe-dev` / `open-swe-prod` — kebab-case stack + resource prefix. */
export const prefix = (env: EnvName): string => `open-swe-${env}`;
/**
* Per-ENV GitHub OIDC trust subject for the deploy roles (T5 OSWE-IAC-01/02 fix:
* the dev/prod boundary is enforced in the IAM trust, not by convention).
*
* - `dev` → the `dev` integration branch ref (auto-deploy on push to dev).
* - `prod` → the **GitHub `prod` Environment** subject. A workflow can only mint
* a token with sub `…:environment:prod` by declaring `environment: prod`,
* which triggers the Environment's manual-approval gate (Adam, T18). So the
* prod approval is now expressed at the IAM layer: a dev-branch token can
* never assume a prod deploy role.
*
* Each env gets its OWN infra + app role (githubdeploy-open-swe-{infra,app}-<env>)
* so a dev token cannot reach prod. Exact subject → StringEquals (no `*`).
*
* Residual (documented): CDK's single account-wide `cfn-exec-role` means the dev
* INFRA role can still technically `cdk deploy open-swe-prod`; the workflow only
* ever targets its own env stack, and prod's environment-gated role is the
* approved path. Per-env bootstrap qualifiers would close this fully (future).
*/
export const oidcSubject = (env: EnvName): string =>
env === "prod"
? `repo:${GITHUB_ORG}/${GITHUB_REPO}:environment:prod`
: `repo:${GITHUB_ORG}/${GITHUB_REPO}:ref:refs/heads/dev`;
/**
* The GitHub Actions OIDC provider already exists account-wide (created for
* seahaven-site; see .github/oidc-deploy-roles.yaml `CreateOIDCProvider=false`).
* Reference it by ARN — never create a duplicate `AWS::IAM::OIDCProvider`
* (CloudFormation rejects a second provider for the same URL).
*/
export const GITHUB_OIDC_PROVIDER_ARN = `arn:aws:iam::${ACCOUNT}:oidc-provider/token.actions.githubusercontent.com`;

View file

@ -0,0 +1,35 @@
import * as ec2 from "aws-cdk-lib/aws-ec2";
import { REGION } from "../config";
/**
* The baked open-swe base AMI (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW agent
* + boot templates — NO secrets), produced by `deploy/ami/open-swe-base.pkr.hcl`.
* Pinned by EXACT id (not a name filter) so synth/deploy is fully offline and
* deterministic.
*
* Built 2026-06-26 from open-swe-base-arm64-20260626-203433.
*
* ── EBS / AMI replacement discipline (memory feedback_inline_ebs_volumes) ──
*
* Refresh DELIBERATELY: `cd deploy/ami && packer build open-swe-base.pkr.hcl`,
* then update this id. A new id → EC2 instance REPLACEMENT. Pinning by exact id
* (vs a `most_recent` name filter) is what prevents a routine deploy from silently
* swapping the AMI — the root cause of the file-share data-loss incidents
* (5/15, 5/27, 6/5).
*
* `userDataCausesReplacement: true` (AppService) is likewise DELIBERATE: user-data
* is provisioning-only and the box holds NO durable state (the langgraph store is
* in-memory, rebuilt every boot from S3 + Secrets Manager / SSM), so there is
* intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The design
* goal is replacement-TOLERANCE, not avoidance.
*
* Operational guard before ANY replacing deploy (AMI / userData / instance-type):
* snapshot the root volume AND wait `state=completed`, re-verify "no local-only
* durable state", and review the `cdk diff` replacement at PR time.
*/
export const BAKED_OPEN_SWE_AMI_ID = "ami-00080084502093021";
/** The baked open-swe base image, pinned by id (offline, deterministic). */
export function bakedOpenSweArm64(): ec2.IMachineImage {
return ec2.MachineImage.genericLinux({ [REGION]: BAKED_OPEN_SWE_AMI_ID });
}

View file

@ -0,0 +1,347 @@
import * as fs from "fs";
import * as path from "path";
import * as cdk from "aws-cdk-lib";
import * as ec2 from "aws-cdk-lib/aws-ec2";
import * as elbv2 from "aws-cdk-lib/aws-elasticloadbalancingv2";
import * as elbTargets from "aws-cdk-lib/aws-elasticloadbalancingv2-targets";
import * as logs from "aws-cdk-lib/aws-logs";
import * as route53 from "aws-cdk-lib/aws-route53";
import * as iam from "aws-cdk-lib/aws-iam";
import * as ssm from "aws-cdk-lib/aws-ssm";
import { Construct } from "constructs";
import { EnvName, prefix } from "../config";
import { bakedOpenSweArm64 } from "./ami-cache";
/**
* Shared seahaven-vpc + internet-facing ALB facts (read-only recon 2026-06-26;
* scratchpad/T12-infra-facts.md). A SINGLE VPC and a SINGLE shared ALB front
* both the on-prem `seahaven-site` stack and open-swe. We IMPORT every one of
* these and NEVER own them - open-swe only ADDS its own instance SG, a standalone
* ALB-egress rule, listener rules, a target group, and DNS records.
*
* ── Cross-stack coordination on the SHARED listener + ALB SG (T13 review) ──
* Two CDK stacks (seahaven-site, open-swe) add resources to the same imported
* `:443` listener and ALB SG. This is safe because each stack owns ONLY the
* resources it declares (its own logical ids): an on-prem `cdk deploy` computes a
* changeset over its own template and cannot delete rules/egress it never
* declared. The standalone-egress pattern is what on-prem itself uses
* (sgr-0c57812752a3bca13), so it does not mutate the shared SG's own definition.
*
* The ONE shared namespace that REQUIRES coordination is listener-rule PRIORITY
* (globally unique per listener; a collision is a fail-SAFE deploy error, not
* silent drift). Ownership map - keep these disjoint when editing either stack:
* - seahaven-site (on-prem): priorities 4-7 + default.
* - open-swe: priorities 2, 3 (webhooks) and 10, 11 (site).
* open-swe's webhook rules are HOST-scoped to its own *.seahaven.com hosts, so
* they never match (let alone "steal") any seahavenind.com / on-prem host.
*/
const SHARED = {
vpcId: "vpc-0d3d4b67bd0cf8a68",
availabilityZones: ["us-east-1a", "us-east-1b"],
// Private subnets host the EC2 box. The single NAT gateway lives in 1a, so the
// box is pinned to private1 (1a) for in-AZ NAT egress (no cross-AZ data $).
privateSubnetIds: ["subnet-04e38c507e96f1926", "subnet-0a0b4fc6f296dfba5"],
instanceSubnetId: "subnet-04e38c507e96f1926",
instanceAz: "us-east-1a",
albDnsName: "seahaven-com-1856441924.us-east-1.elb.amazonaws.com",
albCanonicalHostedZoneId: "Z35SXDOTRQ7X7K",
albSecurityGroupId: "sg-0b0301deed193258a",
httpsListenerArn:
"arn:aws:elasticloadbalancing:us-east-1:328440206208:listener/app/seahaven-com/222c3257354ab559/bab8bcf0da0e2927",
publicZoneId: "Z06652411XKH89KTZD3XA",
publicZoneName: "seahaven.com",
} as const;
/**
* Per-env public hostnames, listener-rule priorities, and instance size.
*
* ── Listener-rule ordering hazard (load-bearing) ──
* The shared listener already has a HOST-AGNOSTIC `/webhooks/*` PATH rule at
* priority 5 (the on-prem seahaven-site stack owns it). ALB rules are first-match
* by ASCENDING priority, so a `…/webhooks/*` request to our host would match
* rule 5 (priority 5) and be forwarded to the on-prem target BEFORE any host rule
* at 10+. Therefore our webhook rule MUST sit below priority 5. The catch-all
* "site" rule (dashboard SPA + /dashboard/api/) carries no path that collides
* with rule 5, so it can sit at any free higher number (10/11). Free priorities
* confirmed by recon: 1-3 and 8+ (4=forgejo, 5=/webhooks/*, 6/7=seahavenind).
*/
const ENV_NET: Record<
EnvName,
{
dashboardHost: string;
hooksHost: string;
webhookPriority: number;
sitePriority: number;
instanceType: string;
}
> = {
dev: {
dashboardHost: "openswe-dev.seahaven.com",
hooksHost: "hooks-dev.seahaven.com",
webhookPriority: 2,
sitePriority: 10,
instanceType: "t4g.medium",
},
prod: {
dashboardHost: "openswe.seahaven.com",
hooksHost: "hooks.seahaven.com",
webhookPriority: 3,
sitePriority: 11,
instanceType: "t4g.large",
},
};
export interface AppServiceProps {
readonly envName: EnvName;
/** Least-privilege EC2 instance role (per-env; from InstanceRole). */
readonly instanceRole: iam.IRole;
/** S3 artifact key prefix the box pulls app.tar.gz / spa.tar.gz from. */
readonly artifactPrefix?: string;
}
/**
* The open-swe compute + ingress wiring for one env (T12):
* - one ARM64 EC2 box in private1 (1a), replacement-tolerant (no RETAIN volume),
* - a standalone instance SG reachable ONLY from the shared ALB SG on :80,
* - a target group -> instance:80 (nginx is the sole ingress; :2024 stays loopback),
* - two listener rules on the imported :443 listener (webhooks below the on-prem
* path rule; site catch-all above it), both -> the same TG,
* - Route53 alias records for both hostnames -> the shared ALB,
* - IaC-owned CloudWatch log groups at 30-day retention.
*
* Everything ALB/VPC/zone-side is IMPORTED. Synth is offline: the AMI is the
* cdk.context.json-pinned AL2023 ARM64 placeholder until the baked
* open-swe-base-arm64 id is pinned before the first real deploy.
*/
export class AppService extends Construct {
public readonly instance: ec2.Instance;
public readonly targetGroup: elbv2.ApplicationTargetGroup;
/** Name of the SSM document CI fires to roll the box to the latest release. */
public readonly deployDocumentName: string;
constructor(scope: Construct, id: string, props: AppServiceProps) {
super(scope, id);
const env = props.envName;
const p = prefix(env);
const net = ENV_NET[env];
const artifactPrefix = props.artifactPrefix ?? "releases/latest";
// Import the shared VPC with explicit attributes (no fromLookup -> offline synth).
const vpc = ec2.Vpc.fromVpcAttributes(this, "Vpc", {
vpcId: SHARED.vpcId,
availabilityZones: [...SHARED.availabilityZones],
privateSubnetIds: [...SHARED.privateSubnetIds],
});
// Standalone instance SG. Egress open (NAT path); ingress only from the ALB SG.
const instanceSg = new ec2.SecurityGroup(this, "InstanceSg", {
vpc,
securityGroupName: `${p}-instance-sg`,
description: `${p} instance SG - ingress only from the shared ALB SG on :80; egress via NAT.`,
allowAllOutbound: true,
});
instanceSg.addIngressRule(
ec2.Peer.securityGroupId(SHARED.albSecurityGroupId),
ec2.Port.tcp(80),
`${p}: shared ALB SG to nginx :80`,
);
// Open the IMPORTED ALB SG to our instance via a STANDALONE egress rule, so we
// never mutate the ALB SG's own (on-prem-owned) definition.
new ec2.CfnSecurityGroupEgress(this, "AlbToInstanceEgress", {
groupId: SHARED.albSecurityGroupId,
ipProtocol: "tcp",
fromPort: 80,
toPort: 80,
destinationSecurityGroupId: instanceSg.securityGroupId,
description: `${p}: ALB to instance nginx :80`,
});
// The app-deploy procedure (deploy/ami/deploy.sh) is a normal reviewable repo
// file; CDK base64-encodes it (single line - no `$`/regex-special chars in the
// base64 alphabet) and renders it into user-data's @@DEPLOY_SH_B64@@ token, so
// user-data writes it verbatim to /opt/open-swe/bin/deploy.sh at first boot.
// The same file is run by the open-swe-<env>-deploy SSM document on every
// release - a single source of truth for "pull release, build venv, restart".
const deployShPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "deploy.sh");
// Minify before embedding: strip full-line comments + blank lines (keep the
// shebang) so the base64 fits EC2's 25.6 KB user-data limit. The repo file
// keeps its comments; only the on-box copy is minified. deploy.sh becomes
// opaque base64 here, so this never affects user-data's heredoc parsing.
const deployShMin = fs
.readFileSync(deployShPath, "utf8")
.split("\n")
.filter((line, i) => i === 0 || (!/^\s*#/.test(line) && line.trim() !== ""))
.join("\n");
const deployShB64 = Buffer.from(deployShMin, "utf8").toString("base64");
// Render the provisioning script's @@tokens@@ into the instance user-data.
// userDataCausesReplacement makes a bootstrap change roll a fresh box (the box
// holds no durable state - see ami-cache.ts / user-data.sh). Editing deploy.sh
// therefore also rolls the box (its base64 is embedded here) - acceptable: the
// box is replacement-tolerant, and ongoing releases never touch user-data.
const userDataPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "user-data.sh");
const userData = ec2.UserData.custom(
fs
.readFileSync(userDataPath, "utf8")
// %%...%% tokens are CDK-substituted here; they are DELIBERATELY a
// different delimiter from the @@...@@ tokens user-data.sh seds into the
// baked systemd/nginx templates, so CDK can never clobber a sed pattern
// (a shared @@OPENSWE_ENV@@/@@SERVER_NAME@@ left the unit unsubstituted).
.replace(/%%OPENSWE_ENV%%/g, env)
.replace(/%%ASSETS_BUCKET%%/g, `${p}-assets`)
.replace(/%%SERVER_NAME%%/g, net.dashboardHost)
.replace(/%%ARTIFACT_PREFIX%%/g, artifactPrefix)
.replace(/%%DEPLOY_SH_B64%%/g, deployShB64),
);
this.instance = new ec2.Instance(this, "Instance", {
vpc,
vpcSubnets: {
subnets: [
ec2.Subnet.fromSubnetAttributes(this, "InstanceSubnet", {
subnetId: SHARED.instanceSubnetId,
availabilityZone: SHARED.instanceAz,
}),
],
},
instanceType: new ec2.InstanceType(net.instanceType),
// The baked open-swe base AMI (deploy/ami packer build) - ARM64 Ubuntu 24.04
// with the /opt/open-swe layout, openswe user, nginx, and CW agent that
// user-data.sh assumes. Pinned by exact id (see ami-cache.ts); refresh by
// rebuilding and updating BAKED_OPEN_SWE_AMI_ID.
machineImage: bakedOpenSweArm64(),
role: props.instanceRole,
securityGroup: instanceSg,
userData,
userDataCausesReplacement: true,
requireImdsv2: true,
instanceName: `${p}-box`,
blockDevices: [
{
deviceName: "/dev/xvda",
// gp3 encrypted root; deleteOnTermination (no durable on-box state ->
// intentionally NO standalone RETAIN volume; see ami-cache.ts).
volume: ec2.BlockDeviceVolume.ebs(30, {
volumeType: ec2.EbsDeviceVolumeType.GP3,
encrypted: true,
deleteOnTermination: true,
}),
},
],
});
// SSM deploy document (open-swe-<env>-deploy): runs the baked
// /opt/open-swe/bin/deploy.sh to pull the latest release + restart. CI fires it
// (tag-scoped to project=open-swe,env=<env>) after uploading a release, so the
// app deploy role needs SendCommand ONLY on this document - NOT on the generic
// AWS-RunShellScript (closes the T4 BLOCK#3 arbitrary-shell timebox).
this.deployDocumentName = `${p}-deploy`;
new ssm.CfnDocument(this, "DeployDoc", {
name: this.deployDocumentName,
documentType: "Command",
documentFormat: "YAML",
updateMethod: "NewVersion",
content: {
schemaVersion: "2.2",
description: `Roll the ${p} box to the latest published release (runs /opt/open-swe/bin/deploy.sh).`,
mainSteps: [
{
action: "aws:runShellScript",
name: "deploy",
inputs: {
// Fixed command - no parameters, so nothing untrusted is interpolated
// into the shell. The script itself reads /etc/open-swe/boot.env.
runCommand: ["bash /opt/open-swe/bin/deploy.sh"],
},
},
],
},
});
// Target group -> instance:80 (nginx). Health check hits nginx's /healthz
// (returns 200; the dashboard TG health path defined in open-swe.nginx.conf).
this.targetGroup = new elbv2.ApplicationTargetGroup(this, "Tg", {
vpc,
targetGroupName: `${p}-tg`,
port: 80,
protocol: elbv2.ApplicationProtocol.HTTP,
targetType: elbv2.TargetType.INSTANCE,
targets: [new elbTargets.InstanceTarget(this.instance)],
deregistrationDelay: cdk.Duration.seconds(15),
healthCheck: {
path: "/healthz",
healthyHttpCodes: "200",
interval: cdk.Duration.seconds(30),
timeout: cdk.Duration.seconds(5),
healthyThresholdCount: 2,
unhealthyThresholdCount: 3,
},
});
// Import the shared :443 listener (with its ALB SG) and ADD our two rules.
const albSg = ec2.SecurityGroup.fromSecurityGroupId(this, "AlbSg", SHARED.albSecurityGroupId, {
mutable: false,
});
const listener = elbv2.ApplicationListener.fromApplicationListenerAttributes(this, "HttpsListener", {
listenerArn: SHARED.httpsListenerArn,
securityGroup: albSg,
});
// (1) Webhooks - accepted on EITHER host (integrations may target either), and
// MUST be below the on-prem path-only rule 5 (see ENV_NET note).
new elbv2.ApplicationListenerRule(this, "WebhooksRule", {
listener,
priority: net.webhookPriority,
conditions: [
elbv2.ListenerCondition.hostHeaders([net.dashboardHost, net.hooksHost]),
elbv2.ListenerCondition.pathPatterns(["/webhooks/*"]),
],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
// (2) Dashboard SPA + /dashboard/api/ (OAuth) - DASHBOARD host ONLY. The hooks
// host intentionally serves nothing but /webhooks/* (rule 1), so the OAuth /
// dashboard surface stays single-origin (OSWE-T12-02). Non-webhook paths on the
// hooks host fall through to the on-prem default.
new elbv2.ApplicationListenerRule(this, "SiteRule", {
listener,
priority: net.sitePriority,
conditions: [elbv2.ListenerCondition.hostHeaders([net.dashboardHost])],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
// Route53 ALIAS records -> the shared ALB, for both hostnames.
const zone = route53.HostedZone.fromHostedZoneAttributes(this, "PublicZone", {
hostedZoneId: SHARED.publicZoneId,
zoneName: SHARED.publicZoneName,
});
const albAlias: route53.IAliasRecordTarget = {
bind: () => ({
dnsName: SHARED.albDnsName,
hostedZoneId: SHARED.albCanonicalHostedZoneId,
}),
};
for (const [label, host] of [
["Dashboard", net.dashboardHost],
["Hooks", net.hooksHost],
] as const) {
new route53.ARecord(this, `${label}Alias`, {
zone,
recordName: host,
target: route53.RecordTarget.fromAlias(albAlias),
comment: `${p} ${label.toLowerCase()} -> shared seahaven-com ALB`,
});
}
// IaC-owned CloudWatch log groups at 30-day retention. Names mirror the
// CloudWatch-agent config (deploy/ami/templates/amazon-cloudwatch-agent.json);
// owning them here makes retention declarative rather than agent-set. Logs are
// not durable state -> DESTROY on stack delete.
for (const suffix of ["app", "user-data", "nginx-access", "nginx-error"]) {
new logs.LogGroup(this, `Log-${suffix}`, {
logGroupName: `/open-swe/${env}/${suffix}`,
retention: logs.RetentionDays.ONE_MONTH,
removalPolicy: cdk.RemovalPolicy.DESTROY,
});
}
}
}

View file

@ -0,0 +1,58 @@
import * as cdk from "aws-cdk-lib";
import * as s3 from "aws-cdk-lib/aws-s3";
import { Construct } from "constructs";
import { EnvName, prefix } from "../config";
/**
* The per-env S3 artifact bucket (`open-swe-<env>-assets`) the box pulls its
* release from (T7). CI builds the SPA + packages the app source and uploads
* `app.tar.gz` / `spa.tar.gz` under `releases/<sha>/` + `releases/latest/`
* (`build-artifacts.yml`, via the `githubdeploy-open-swe-app-<env>` OIDC role);
* the box pulls `releases/latest/*` at boot / on deploy via its instance role.
*
* The bucket holds ONLY build artifacts — no secrets (those live in Secrets
* Manager + SSM), no durable runtime state (the langgraph store is in-memory and
* rebuilt every boot). It is therefore safe to treat as reproducible-from-CI, but
* we RETAIN it on stack delete so an accidental `cdk destroy` cannot strand the
* box with no artifact to pull on its next replacement.
*
* Security posture (locked, reviewed in T7):
* - `BLOCK_ALL` public access (this is an internal artifact store; ALB/nginx is
* the only public surface — never S3 directly).
* - SSE-S3 encryption at rest + `enforceSSL` (deny any non-TLS request).
* - versioned, so a bad release can be rolled back to the previous object
* version (the last-good-artifact story in T19); a lifecycle rule expires
* NONcurrent versions after 30 days so history does not grow unbounded.
* - aborts incomplete multipart uploads after 7 days (cost hygiene).
*
* The name is the load-bearing contract: `instance-role.ts` (read), the app
* deploy role in `github-deploy-roles.ts` (write), and `user-data.sh` /
* `deploy.sh` (`@@ASSETS_BUCKET@@`) all reference `open-swe-<env>-assets` by
* literal name, so it is set explicitly here rather than auto-generated.
*/
export class AssetsBucket extends Construct {
public readonly bucket: s3.Bucket;
constructor(scope: Construct, id: string, envName: EnvName) {
super(scope, id);
const p = prefix(envName);
this.bucket = new s3.Bucket(this, "Bucket", {
bucketName: `${p}-assets`,
blockPublicAccess: s3.BlockPublicAccess.BLOCK_ALL,
encryption: s3.BucketEncryption.S3_MANAGED,
enforceSSL: true,
versioned: true,
// Artifacts are reproducible from CI, but RETAIN protects against an
// accidental stack delete leaving the box with nothing to pull (see above).
removalPolicy: cdk.RemovalPolicy.RETAIN,
lifecycleRules: [
{
id: "expire-noncurrent-artifact-versions",
noncurrentVersionExpiration: cdk.Duration.days(30),
abortIncompleteMultipartUploadAfter: cdk.Duration.days(7),
},
],
});
}
}

View file

@ -0,0 +1,235 @@
import * as cdk from "aws-cdk-lib";
import * as secretsmanager from "aws-cdk-lib/aws-secretsmanager";
import * as ssm from "aws-cdk-lib/aws-ssm";
import { Construct } from "constructs";
import { EnvName } from "../config";
/**
* Config / secret "shells" for the boot hook (`deploy/seahaven/fetch-config.sh`).
*
* The naming contract (source of truth: the T9 env/secret/config inventory + the
* fetch-config header) is LITERAL env-var names as the last path segment:
*
* Secrets open-swe-<env>/<ENV_VAR_NAME> (AWS Secrets Manager)
* Config /open-swe-<env>/<ENV_VAR_NAME> (AWS SSM Parameter Store, FLAT)
*
* fetch-config reads secrets with `batch-get-secret-value --filters
* Key=name,Values=open-swe-<env>/` and config with `get-parameters-by-path
* --path /open-swe-<env>/` (NON-recursive), then strips the prefix so the last
* segment IS the exported variable name. So these resources MUST carry the
* UPPER_SNAKE env-var name verbatim — which is why both resource types are
* exempted from the kebab-naming Aspect (see aspects/kebab-naming-aspect.ts).
*
* Three buckets:
*
* 1. SECRETS_SHELLS — the 29 secrets. Created as value-LESS shells (an L1
* `CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which
* CloudFormation creates as an empty secret with no version). The real value
* is set out-of-band via `deploy/seahaven/put-config.sh` (put-secret-value)
* BEFORE the box boots. Because CDK never owns the value, a later
* `cdk deploy` can never clobber the operator-set value. AWS-managed key
* (alias/aws/secretsmanager) — no CMK, matching the instance role which
* deliberately omits kms:Decrypt.
*
* 2. IAC_MANAGED_SSM — stable / derivable config. Real values are owned here in
* IaC (one StringParameter each) so they are reproducible and reviewed.
*
* 3. Out-of-band SSM (NOT created here) — operationally-variable or
* env-specific-unknown config (e.g. DEFAULT_SANDBOX_SNAPSHOT_ID, which
* changes on every snapshot rebuild and would be clobbered by a deploy if it
* were CDK-managed; GitHub App ids; Slack ids; LangSmith tenant/urls). These
* are listed in OUT_OF_BAND_SSM purely for documentation and are set by
* `put-config.sh`, never by CDK.
*/
/** The 29 Secrets Manager secret VAR names (T9 inventory SECRETS table). */
export const SECRET_VARS: readonly string[] = [
"ANTHROPIC_API_KEY",
"CORRIDOR_API_TOKEN",
"CORRIDOR_MCP_TOKEN",
"CORRIDOR_TOKEN",
"DASHBOARD_JWT_SECRET",
"DAYTONA_API_KEY",
"EXA_API_KEY",
"FIREWORKS_API_KEY",
"GITHUB_APP_CLIENT_SECRET",
"GITHUB_APP_PRIVATE_KEY",
"GITHUB_PAT",
"GITHUB_WEBHOOK_SECRET",
"GOOGLE_API_KEY",
"GROQ_API_KEY",
"JUDGE_ANTHROPIC_API_KEY",
"LANGSMITH_API_KEY",
"LANGSMITH_API_KEY_PROD",
"LANGCHAIN_API_KEY",
"LINEAR_API_KEY",
"LINEAR_WEBHOOK_SECRET",
"OPENAI_API_KEY",
"RUNLOOP_API_KEY",
"SLACK_BOT_TOKEN",
"SLACK_CLIENT_SECRET",
"SLACK_SIGNING_SECRET",
"TOKEN_ENCRYPTION_KEY",
"USER_ID_API_KEY_MAP",
"X_SERVICE_AUTH_JWT_SECRET",
] as const;
// 28 secret shells. The T9 inventory header said "29" vs 27 enumerated; reconciled
// (Adam confirm 2026-06-26): USER_ID_API_KEY_MAP (maps user ids -> API keys; flagged
// sensitive by the T5 security review) is a SECRET and is included here.
// JUDGE_ANTHROPIC_BASE_URL is a URL (non-sensitive config, eval-only) -> SSM/default,
// NOT a secret. So the inventory's "29" was effectively a miscount.
/** Short, value-free descriptions for the secret shells (no secret material). */
const SECRET_DESCRIPTIONS: Record<string, string> = {
ANTHROPIC_API_KEY: "Claude LLM API key (primary builder provider).",
CORRIDOR_API_TOKEN: "Corridor MCP token (optional).",
CORRIDOR_MCP_TOKEN: "Corridor MCP token alt name (optional).",
CORRIDOR_TOKEN: "Corridor MCP token alt name (optional).",
DASHBOARD_JWT_SECRET: "JWT signing secret for dashboard session cookies (REQUIRED).",
DAYTONA_API_KEY: "Daytona sandbox key (only if SANDBOX_TYPE=daytona).",
EXA_API_KEY: "Exa web-search key (optional).",
FIREWORKS_API_KEY: "Fireworks LLM key (only if a fireworks: model is used).",
GITHUB_APP_CLIENT_SECRET: "GitHub App OAuth client secret (dashboard login).",
GITHUB_APP_PRIVATE_KEY: "GitHub App private key PEM (installation-token minting).",
GITHUB_PAT: "GitHub PAT fallback (optional).",
GITHUB_WEBHOOK_SECRET: "GitHub webhook signature secret (prod-required).",
GOOGLE_API_KEY: "Google GenAI key (only if a google_genai: model is used).",
GROQ_API_KEY: "Groq LLM key (only if a groq: model is used).",
JUDGE_ANTHROPIC_API_KEY: "Eval judge key (optional; falls back to ANTHROPIC_API_KEY).",
LANGSMITH_API_KEY: "LangSmith key (dev).",
LANGSMITH_API_KEY_PROD: "LangSmith key (prod / deployed sandbox).",
LANGCHAIN_API_KEY: "LangSmith key alt name (fallback).",
LINEAR_API_KEY: "Linear API key (optional).",
LINEAR_WEBHOOK_SECRET: "Linear webhook signature secret (required when Linear is wired).",
OPENAI_API_KEY: "OpenAI key (primary reviewer provider).",
RUNLOOP_API_KEY: "Runloop sandbox key (only if SANDBOX_TYPE=runloop).",
SLACK_BOT_TOKEN: "Slack bot token (optional).",
SLACK_CLIENT_SECRET: "Slack OAuth client secret.",
SLACK_SIGNING_SECRET: "Slack webhook signing secret (prod-required).",
TOKEN_ENCRYPTION_KEY: "Fernet key(s) for per-user GitHub-token encryption (REQUIRED).",
USER_ID_API_KEY_MAP: "JSON map of user id -> API key for per-user auth (optional, sensitive).",
X_SERVICE_AUTH_JWT_SECRET: "Service-auth JWT secret (optional).",
};
/**
* IaC-managed SSM config: stable / derivable values owned in code, per env.
* Values are functions of envName so dev/prod render correct hosts.
*
* Anything operationally-variable or env-specific-unknown is deliberately NOT
* here — see OUT_OF_BAND_SSM.
*/
export function iacManagedSsm(env: EnvName): Record<string, string> {
// Public host = seahaven.com (the migration's new AWS public face; confirmed by
// recon: seahaven.com Route53 zone + *.seahaven.com ACM cert are live on the ALB
// — distinct from the on-prem seahavenind.com). dev = openswe-dev, prod = openswe.
const host = `https://openswe${env === "dev" ? "-dev" : ""}.seahaven.com`;
return {
// Sandbox provider — plan keeps stock langsmith (T9). Stable.
SANDBOX_TYPE: "langsmith",
// Sea Haven org pin. fetch-config ALSO hard-pins this at boot, but owning the
// real value here keeps SSM self-consistent rather than blank.
DEFAULT_REPO_OWNER: "Sea-Haven-Industries",
// Repo allowlist (comma list) — the Sea Haven org. Stable/derivable.
ALLOWED_GITHUB_ORGS: "Sea-Haven-Industries",
// Pilot repo (project memory: Sea-Haven-Industries/open-swe-pilot).
DEFAULT_REPO_NAME: "open-swe-pilot",
// Dashboard URLs — derived from the public host.
DASHBOARD_BASE_URL: host,
DASHBOARD_API_BASE_URL: host,
DASHBOARD_ALLOWED_ORIGINS: host,
// Primary builder model (project memory team_settings: anthropic:claude-opus-4-8).
LLM_MODEL_ID: "anthropic:claude-opus-4-8",
};
}
/**
* Out-of-band SSM config: NOT created by CDK. Listed for documentation and for
* `put-config.sh` to populate before the box boots. Each MUST stay out of IaC
* because its value is operationally-variable or env-specific and unknown at
* synth time — making it CDK-managed would either clobber the operator value on
* the next deploy (e.g. DEFAULT_SANDBOX_SNAPSHOT_ID) or hardcode a secret-ish id.
*/
export const OUT_OF_BAND_SSM: readonly string[] = [
// Sandbox snapshot id — changes on EVERY snapshot rebuild. MUST NOT be
// CDK-managed or a deploy clobbers it. Required for SANDBOX_TYPE=langsmith.
"DEFAULT_SANDBOX_SNAPSHOT_ID",
// GitHub App identifiers — set when the per-env GitHub App is created.
"GITHUB_APP_ID",
"GITHUB_APP_CLIENT_ID",
"GITHUB_APP_INSTALLATION_ID",
"GITHUB_OAUTH_PROVIDER_ID",
// LangSmith deployment coordinates (prod tenant/urls/endpoints).
"LANGSMITH_TENANT_ID_PROD",
"LANGSMITH_URL_PROD",
"LANGSMITH_ENDPOINT",
"LANGSMITH_ENDPOINT_PROD",
"LANGSMITH_HOST_API_URL",
"LANGGRAPH_URL",
"LANGGRAPH_URL_PROD",
"LANGCHAIN_REVISION_ID",
// Slack workspace ids — set after the Slack app is installed.
"SLACK_CLIENT_ID",
"SLACK_TEAM_ID",
"SLACK_BOT_USER_ID",
"SLACK_BOT_USERNAME",
"SLACK_REPO_OWNER",
"SLACK_REPO_NAME",
// Access / observability allowlists — operator-curated.
"CONFIGURED_ADMINS",
"OBSERVABILITY_AUTHORIZED_EMAILS",
"PUBLIC_REPO_ORG_GATE",
"ALLOWED_GITHUB_REPOS",
// Optional integrations + tuning knobs (left to code defaults unless set).
"LLM_FALLBACK_MODEL_ID",
"DATADOG_MCP_TOOLSETS",
"NOTION_MCP_CLIENT_NAME",
"API_STANDARDS_SKILL_HANDLE",
"REPO_SNAPSHOT_BASE_IMAGE",
"REPO_SNAPSHOT_BUILD_TIMEOUT_SECONDS",
"REPO_SNAPSHOT_STALE_BUILD_SECONDS",
] as const;
export interface ConfigStoreProps {
readonly envName: EnvName;
}
/**
* Per-env Secrets Manager + SSM Parameter Store shells the boot hook reads.
* Instantiated from OpenSweStack. Synth-able now (T11); values populated
* out-of-band BEFORE the EC2/T12 deploy. See infra/README.md "Config store".
*/
export class ConfigStore extends Construct {
public readonly secrets: secretsmanager.CfnSecret[] = [];
public readonly params: ssm.StringParameter[] = [];
constructor(scope: Construct, id: string, props: ConfigStoreProps) {
super(scope, id);
const env = props.envName;
// --- 1) Secret shells (value-LESS; populated out-of-band) ----------------
for (const varName of SECRET_VARS) {
const secret = new secretsmanager.CfnSecret(this, `Secret-${varName}`, {
name: `open-swe-${env}/${varName}`,
description: SECRET_DESCRIPTIONS[varName] ?? `open-swe ${varName}`,
// Deliberately NO secretString / generateSecretString: CloudFormation
// creates an empty secret, so the out-of-band value is never clobbered.
});
// RETAIN: a stack teardown must not destroy operator-set secret material.
secret.applyRemovalPolicy(cdk.RemovalPolicy.RETAIN);
this.secrets.push(secret);
}
// --- 2) IaC-managed SSM config (real, derivable values) ------------------
const managed = iacManagedSsm(env);
for (const [varName, value] of Object.entries(managed)) {
this.params.push(
new ssm.StringParameter(this, `Param-${varName}`, {
parameterName: `/open-swe-${env}/${varName}`,
stringValue: value,
description: `IaC-managed open-swe ${varName} (${env}).`,
tier: ssm.ParameterTier.STANDARD,
}),
);
}
}
}

View file

@ -0,0 +1,157 @@
import * as iam from "aws-cdk-lib/aws-iam";
import { Construct } from "constructs";
import {
ACCOUNT,
EnvName,
GITHUB_OIDC_PROVIDER_ARN,
REGION,
oidcSubject,
} from "../config";
/**
* Per-ENV GitHub Actions OIDC deploy roles. Created ONCE per env in the
* dedicated `open-swe-iam` stack. Two roles per env, per the locked architecture's
* "dual OIDC roles":
*
* - githubdeploy-open-swe-infra-<env> → CFN/IAM (CDK) deploys of that env's stack
* - githubdeploy-open-swe-app-<env> → app deploys (env-tag-scoped SSM + S3 read)
*
* T5 OSWE-IAC-01/02 fix: roles are split per env and the trust subject is
* env-scoped (dev = dev branch ref; prod = the GitHub `prod` Environment subject,
* so the manual-approval gate is IAM-enforced). A dev-branch token therefore
* cannot SendCommand to the prod box nor assume a prod deploy role.
*
* Reviewed at T4 (GPT-4.1 IAM cross-review) + T5 (/sh-security-review) and
* deployed FIRST (BLOCK#3 "OIDC-role-first" ordering) before any other infra or
* secrets CI step.
*/
export class GithubDeployRoles extends Construct {
public readonly infraRole: iam.Role;
public readonly appRole: iam.Role;
constructor(scope: Construct, id: string, envName: EnvName) {
super(scope, id);
// The provider already exists account-wide — reference, never re-create.
const provider = iam.OpenIdConnectProvider.fromOpenIdConnectProviderArn(
this,
"GithubOidcProvider",
GITHUB_OIDC_PROVIDER_ARN,
);
// T4 BLOCK#2 + T5 IAC-01/02: exact env-scoped subject via StringEquals (no
// StringLike, no `*`). prod = environment:prod (manual-approval gate),
// dev = the dev branch ref.
const trust = new iam.WebIdentityPrincipal(provider.openIdConnectProviderArn, {
StringEquals: {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": oidcSubject(envName),
},
});
// ---- githubdeploy-open-swe-infra-<env> --------------------------------
this.infraRole = new iam.Role(this, "InfraDeployRole", {
roleName: `githubdeploy-open-swe-infra-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for CDK deploys of the open-swe-${envName} infra stack (assumes CDK bootstrap roles).`,
});
// Org-standard CDK deploy pattern (mirrors githubdeploy-seahaven-account-
// baseline / -forgejo / -apm-wo-analysis): the deploy role only needs to
// assume the CDK bootstrap roles. The actual CloudFormation + IAM + resource
// permissions are exercised by the bootstrap `cfn-exec-role`, whose scope is
// owned by the CDKToolkit stack — NOT granted directly here.
//
// T4 BLOCK#1: GPT-4.1 flagged the `cdk-hnb659fds-*` wildcard and recommended
// enumerating the four exact ARNs. ACCEPTED EXCEPTION (Adam, 2026-06-26): kept
// as the verified org-wide convention (githubdeploy-seahaven-account-baseline
// uses the identical wildcard). Only `cdk bootstrap` creates roles with this
// prefix, so practical escalation risk is low.
// T5 residual (OSWE-IAC-02): the single account-wide cfn-exec-role means the
// dev infra role can technically deploy any stack; per-env trust gates WHO can
// assume, and the prod role requires the environment:prod approval. Per-env
// bootstrap qualifiers would close the residual fully (future hardening).
this.infraRole.addToPolicy(
new iam.PolicyStatement({
sid: "AssumeCdkBootstrapRoles",
actions: ["sts:AssumeRole"],
resources: [`arn:aws:iam::${ACCOUNT}:role/cdk-hnb659fds-*`],
}),
);
// ---- githubdeploy-open-swe-app-<env> ----------------------------------
this.appRole = new iam.Role(this, "AppDeployRole", {
roleName: `githubdeploy-open-swe-app-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for open-swe-${envName} app deploys: env-tag-scoped ssm:SendCommand + read of the ${envName} S3 artifact bucket.`,
});
// T5 OSWE-IAC-01 fix: SendCommand only to instances tagged project=open-swe
// AND env=<this env> (a SINGLE value, not {dev,prod}). The dev app role can
// never command the prod box and vice versa — env isolation in IAM.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandTagScoped",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ec2:${REGION}:${ACCOUNT}:instance/*`],
conditions: {
StringEquals: {
"ssm:resourceTag/project": "open-swe",
"ssm:resourceTag/env": envName,
},
},
}),
);
// SendCommand also has to reference the command document. Scope to this env's
// open-swe deploy document ONLY.
// T4 BLOCK#3 (CLOSED at T19): GPT-4.1 flagged AWS-RunShellScript as an
// arbitrary-shell escalation path. The dedicated `open-swe-${envName}-deploy`
// SSM document (app-service.ts) now runs the fixed, parameter-less command
// `bash /opt/open-swe/bin/deploy.sh`, so AWS-RunShellScript is dropped here:
// this role can run ONLY that one document, and only on its own env's box
// (tag-scoped by the SsmSendCommandTagScoped statement above).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandDocuments",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:document/open-swe-${envName}-deploy`],
}),
);
// Poll command results. These read actions do not support resource-level
// scoping, so `*` is required by the API (T4 FIX: API limitation, documented).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmReadCommandStatus",
actions: [
"ssm:GetCommandInvocation",
"ssm:ListCommands",
"ssm:ListCommandInvocations",
],
resources: ["*"],
}),
);
// Read+WRITE access to THIS env's artifact bucket only (T19): the
// build-artifacts workflow uploads app.tar.gz / spa.tar.gz under releases/*,
// then fires the deploy document so the box pulls them via its instance role.
// Object actions are scoped to releases/* (the only prefix CI writes), and to
// THIS env's bucket — a dev token can never write the prod bucket. No
// bucket-level mutation (no PutBucket*/Delete bucket) — that stays with CDK.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ReadWriteArtifactObjects",
actions: ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets/releases/*`],
}),
);
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ListArtifactBucket",
actions: ["s3:ListBucket", "s3:GetBucketLocation"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
}),
);
}
}

View file

@ -0,0 +1,119 @@
import * as iam from "aws-cdk-lib/aws-iam";
import { Construct } from "constructs";
import { ACCOUNT, EnvName, REGION, prefix } from "../config";
/**
* Least-privilege EC2 instance role for the open-swe box (one per env).
*
* Grants exactly what the boot/runtime flow needs and NOTHING ELSE — no admin,
* no `*` resources except where the AWS action genuinely has no resource-level
* scoping. Per-env so the dev box can never read prod secrets/config and vice
* versa. Reviewed at T4 (GPT-4.1 IAM cross-review) / T5 (/sh-security-review)
* before it is ever deployed (T6).
*/
export class InstanceRole extends Construct {
public readonly role: iam.Role;
constructor(scope: Construct, id: string, env: EnvName) {
super(scope, id);
const p = prefix(env);
this.role = new iam.Role(this, "Role", {
roleName: `${p}-instance-role`,
assumedBy: new iam.ServicePrincipal("ec2.amazonaws.com"),
description: `EC2 instance role for the ${p} open-swe box (least-privilege).`,
});
// AWS-managed: lets the SSM agent register the instance and RECEIVE the
// app-deploy `ssm:SendCommand` from githubdeploy-open-swe-app. This is the
// standard Session-Manager / RunCommand grant and is the only managed
// policy on the role. DELIBERATE — flag for T4 confirmation.
this.role.addManagedPolicy(
iam.ManagedPolicy.fromAwsManagedPolicyName("AmazonSSMManagedInstanceCore"),
);
// Read the build artifact from the env's S3 asset bucket (deploy = pull).
// Scoped to releases/* — the only prefix CI writes and the box pulls — so a
// compromised box (or stolen IMDS creds) cannot read anything else that might
// ever land in the bucket (least-privilege; mirrors the app role's write scope).
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadArtifactObjects",
actions: ["s3:GetObject"],
resources: [`arn:aws:s3:::${p}-assets/releases/*`],
}),
);
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ListArtifactBucket",
actions: ["s3:ListBucket", "s3:GetBucketLocation"],
resources: [`arn:aws:s3:::${p}-assets`],
}),
);
// Read non-sensitive config from SSM Parameter Store under /open-swe-<env>/*.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadSsmConfig",
actions: ["ssm:GetParameter", "ssm:GetParameters", "ssm:GetParametersByPath"],
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:parameter/${p}/*`],
}),
);
// VALUE access — Secrets Manager under open-swe-<env>/*. Secret ARNs carry a
// random 6-char suffix, hence the trailing `*`. This is the statement that
// actually gates which secret VALUES the box can read: prefix-scoped, so the
// dev box can never read prod secret values (and vice versa). GetSecretValue is
// checked per-secret even when the value is returned via the batch call below.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadSecretValues",
actions: ["secretsmanager:GetSecretValue", "secretsmanager:DescribeSecret"],
resources: [`arn:aws:secretsmanager:${REGION}:${ACCOUNT}:secret:${p}/*`],
}),
);
// OPERATION-level grants for fetch-config.sh's prefix-FILTERED batch read
// (`batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`). Both of
// these are collection operations that AWS authorizes against `*`, NOT a
// per-secret ARN: BatchGetSecretValue with a filter is a collection call (a
// prefix-scoped ARN does NOT satisfy it — it AccessDenies), and ListSecrets is a
// list action with no resource-level scoping at all. Neither returns or widens
// VALUE access: a secret's value is still only returned when the prefix-scoped
// GetSecretValue above allows it, so cross-env VALUE isolation is preserved. The
// residual is metadata-only (the box can ENUMERATE secret names account-wide).
// Future hardening to drop both `*` grants: switch fetch-config to an explicit
// `--secret-id-list` (no filter), which lets BatchGetSecretValue be prefix-scoped
// and needs no ListSecrets. Tracked as OSWE-IAC-SECRETS-LIST-01.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "SecretsBatchListOps",
actions: ["secretsmanager:BatchGetSecretValue", "secretsmanager:ListSecrets"],
resources: ["*"],
}),
);
// NOTE (T11): SSM SecureString + Secrets Manager here are assumed to use the
// AWS-managed keys (alias/aws/ssm, alias/aws/secretsmanager) for which the
// service grants Decrypt implicitly — so NO kms:Decrypt is granted. If T11
// moves these to a customer CMK, add a scoped `kms:Decrypt` on that key ARN
// ONLY (not `*`).
// Ship application logs to CloudWatch Logs under /open-swe/<env>/*.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "PutAppLogs",
actions: [
"logs:CreateLogGroup",
"logs:CreateLogStream",
"logs:PutLogEvents",
"logs:DescribeLogStreams",
],
resources: [
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*`,
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*:*`,
],
}),
);
}
}

View file

@ -0,0 +1,38 @@
import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { GithubDeployRoles } from "./constructs/github-deploy-roles";
/**
* Account-level IAM stack: the per-ENV GitHub OIDC deploy roles
* (githubdeploy-open-swe-{infra,app}-{dev,prod} — four roles).
*
* T5 OSWE-IAC-01/02 fix: roles are split per env with env-scoped OIDC trust, so
* a dev-branch token cannot reach prod (prod roles require the GitHub
* `prod` Environment manual-approval gate). They live in this dedicated stack
* rather than the env stacks because IAM roles are global and this stack ships
* FIRST (TODO.md BLOCK#3): the infra OIDC roles + the repo deploy-role-ARN
* secrets must exist before any infra/secrets CI step. Synth-only until the
* Phase-1 security gate (T4 + T5) clears (T6).
*/
export class OpenSweIamStack extends cdk.Stack {
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
super(scope, id, props);
const dev = new GithubDeployRoles(this, "DeployRolesDev", "dev");
const prod = new GithubDeployRoles(this, "DeployRolesProd", "prod");
cdk.Tags.of(this).add("project", "open-swe");
cdk.Tags.of(this).add("ManagedBy", "cdk");
const out = (id: string, role: { roleName?: string }, env: string, kind: string) =>
new cdk.CfnOutput(this, id, {
value: `arn:aws:iam::${this.account}:role/${role.roleName}`,
description: `OIDC role ARN for ${env} ${kind} deploys — set as the ${env} deploy-role secret.`,
});
out("InfraDeployRoleDevArn", dev.infraRole, "dev", "infra (CDK)");
out("AppDeployRoleDevArn", dev.appRole, "dev", "app (tag-scoped SSM + S3)");
out("InfraDeployRoleProdArn", prod.infraRole, "prod", "infra (CDK)");
out("AppDeployRoleProdArn", prod.appRole, "prod", "app (tag-scoped SSM + S3)");
}
}

View file

@ -0,0 +1,90 @@
import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { EnvName, prefix } from "./config";
import { AppService } from "./constructs/app-service";
import { AssetsBucket } from "./constructs/assets-bucket";
import { ConfigStore } from "./constructs/config-store";
import { InstanceRole } from "./constructs/instance-role";
import { BAKED_OPEN_SWE_AMI_ID } from "./constructs/ami-cache";
export interface OpenSweStackProps extends cdk.StackProps {
/** open-swe environment — drives the `open-swe-<env>-*` resource naming. */
readonly envName: EnvName;
}
/**
* Per-env open-swe stack (`open-swe-dev` / `open-swe-prod`). Resource names are
* prefixed `open-swe-<env>-*`.
*
* Composes: the per-env least-privilege instance role (T6), the Secrets/SSM
* config store (T11), and the compute + ingress wiring (T12, AppService — EC2
* box, instance SG, target group, imported-listener rules, Route53 aliases,
* 30-day log groups). The shared VPC and ALB are imported, never owned. Synth is
* offline (AMI is the cdk.context.json-pinned placeholder until T12-deploy).
*/
export class OpenSweStack extends cdk.Stack {
public readonly instanceRole: InstanceRole;
public readonly configStore: ConfigStore;
public readonly assetsBucket: AssetsBucket;
public readonly appService: AppService;
constructor(scope: Construct, id: string, props: OpenSweStackProps) {
super(scope, id, props);
const envName = props.envName;
const p = prefix(envName);
cdk.Tags.of(this).add("project", "open-swe");
cdk.Tags.of(this).add("env", envName);
cdk.Tags.of(this).add("ManagedBy", "cdk");
// Per-env least-privilege EC2 instance role (open-swe-<env>-instance-role).
this.instanceRole = new InstanceRole(this, "Instance", envName);
// Secrets Manager + SSM Parameter Store shells the boot hook reads
// (deploy/seahaven/fetch-config.sh). Secret shells are value-less and
// populated out-of-band; IaC-managed SSM params carry real derivable values.
// The instance role already grants read on open-swe-<env>/* + /open-swe-<env>/*.
this.configStore = new ConfigStore(this, "Config", { envName });
// T7: the S3 artifact bucket (open-swe-<env>-assets) CI uploads releases to
// and the box pulls app.tar.gz / spa.tar.gz from. The instance role already
// grants read on it by name; the app deploy role grants write.
this.assetsBucket = new AssetsBucket(this, "Assets", envName);
// Surface the baked open-swe base AMI id the box runs on (pinned by id in
// ami-cache.ts; refreshed by a deliberate packer rebuild → replacement).
new cdk.CfnOutput(this, "BakedAmiId", {
value: BAKED_OPEN_SWE_AMI_ID,
description: "Baked open-swe-base-arm64 AMI id consumed by the EC2 instance.",
});
// T12: compute + ingress. Imports the shared seahaven-vpc + ALB and adds the
// env's EC2 box, instance SG, target group, listener rules, DNS, log groups.
this.appService = new AppService(this, "App", {
envName,
instanceRole: this.instanceRole.role,
});
new cdk.CfnOutput(this, "InstanceRoleArn", {
value: this.instanceRole.role.roleArn,
description: `${p} EC2 instance role ARN.`,
});
new cdk.CfnOutput(this, "InstanceId", {
value: this.appService.instance.instanceId,
description: `${p} EC2 instance id.`,
});
new cdk.CfnOutput(this, "TargetGroupArn", {
value: this.appService.targetGroup.targetGroupArn,
description: `${p} ALB target group ARN (→ instance:80 nginx).`,
});
new cdk.CfnOutput(this, "AssetsBucketName", {
value: this.assetsBucket.bucket.bucketName,
description: `${p} S3 artifact bucket (CI uploads releases; box pulls).`,
});
new cdk.CfnOutput(this, "DeployDocumentName", {
value: this.appService.deployDocumentName,
description: `${p} SSM document that rolls the box to the latest release.`,
});
}
}

4487
infra/package-lock.json generated Normal file

File diff suppressed because it is too large Load diff

31
infra/package.json Normal file
View file

@ -0,0 +1,31 @@
{
"name": "open-swe-infra",
"version": "1.0.0",
"description": "Open SWE AWS infrastructure (CDK TypeScript) — open-swe-dev / open-swe-prod stacks + shared OIDC deploy roles.",
"private": true,
"bin": {
"open-swe-infra": "bin/app.js"
},
"scripts": {
"build": "tsc",
"cdk": "cdk",
"synth": "cdk synth",
"diff": "cdk diff",
"test": "jest"
},
"devDependencies": {
"@types/jest": "^29.5.14",
"@types/node": "^24.0.0",
"@types/source-map-support": "^0.5.10",
"aws-cdk": "^2.1029.0",
"jest": "^29.7.0",
"source-map-support": "^0.5.21",
"ts-jest": "^29.2.5",
"ts-node": "^10.9.2",
"typescript": "~5.6.3"
},
"dependencies": {
"aws-cdk-lib": "2.260.0",
"constructs": "^10.0.0"
}
}

View file

@ -0,0 +1,87 @@
import * as cdk from "aws-cdk-lib";
import { Template } from "aws-cdk-lib/assertions";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
/**
* Several AWS APIs reject non-ASCII in fields that `cdk synth` happily emits and
* `tsc` happily compiles — so a stray em-dash/arrow only blows up at DEPLOY time
* (e.g. EC2 SecurityGroup GroupDescription: "Character sets beyond ASCII are not
* supported"). This has bitten us twice (the AMI Description, then the instance-SG
* description). This test fails the build at synth time instead.
*
* Scope: the EC2 fields with a documented ASCII/restricted-charset constraint —
* SecurityGroup GroupDescription and ingress/egress rule descriptions. (CloudFormation
* Output descriptions + Route53 comments accept UTF-8, so they are not asserted.)
*
* EC2 rule descriptions are stricter than ASCII: the allowed set is
* `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*` — note it EXCLUDES `<` and `>`, which is why a
* naive em-dash -> "->" replacement still fails at deploy. We assert that exact set.
*/
// Characters NOT in the EC2 description allowed set.
const DISALLOWED = /[^a-zA-Z0-9. _:/()#,@[\]+=&;{}!$*-]/;
function synthDev(): Record<string, { Type: string; Properties?: Record<string, unknown> }> {
const app = new cdk.App();
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
return Template.fromStack(dev).toJSON().Resources;
}
describe("ASCII-only EC2 description fields", () => {
const resources = synthDev();
it("SecurityGroup GroupDescription is ASCII", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::SecurityGroup") continue;
const desc = (r.Properties?.GroupDescription as string) ?? "";
expect(DISALLOWED.test(desc) ? `${id}: ${desc}` : "ascii").toBe("ascii");
}
});
it("SecurityGroup ingress/egress rule descriptions are ASCII", () => {
for (const [id, r] of Object.entries(resources)) {
const props = r.Properties ?? {};
const groups: Array<{ Description?: string }> = [];
if (Array.isArray(props.SecurityGroupIngress)) groups.push(...props.SecurityGroupIngress);
if (Array.isArray(props.SecurityGroupEgress)) groups.push(...props.SecurityGroupEgress);
// Standalone AWS::EC2::SecurityGroupEgress / ...Ingress resources.
if (r.Type === "AWS::EC2::SecurityGroupEgress" || r.Type === "AWS::EC2::SecurityGroupIngress") {
groups.push(props as { Description?: string });
}
for (const rule of groups) {
const desc = rule.Description ?? "";
expect(DISALLOWED.test(desc) ? `${id}: ${desc}` : "ascii").toBe("ascii");
}
}
});
// EC2 caps base64-encoded user-data at 25600 bytes; CDK + tsc don't check it, so
// an oversized boot script (e.g. an embedded deploy.sh) only fails at deploy.
it("EC2 user-data fits the 25600-byte encoded limit", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::Instance") continue;
const ud = (r.Properties?.UserData as { "Fn::Base64"?: string }) ?? {};
const script = typeof ud["Fn::Base64"] === "string" ? ud["Fn::Base64"] : "";
const encoded = Buffer.from(script, "utf8").toString("base64").length;
expect(`${id}: ${encoded} bytes`).toBe(encoded < 25600 ? `${id}: ${encoded} bytes` : "OVER 25600");
}
});
// CDK substitutes %%...%% tokens in user-data at synth. Any %%TOKEN%% left in the
// rendered script means a token wasn't wired in app-service.ts (the @@...@@ tokens
// are intentional — user-data seds those into the baked templates at boot).
it("user-data has no unresolved %%CDK%% tokens", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::Instance") continue;
const ud = (r.Properties?.UserData as { "Fn::Base64"?: string }) ?? {};
const script = typeof ud["Fn::Base64"] === "string" ? ud["Fn::Base64"] : "";
const leftover = script.match(/%%[A-Z0-9_]+%%/g) ?? [];
expect(`${id}: ${leftover.join(",")}`).toBe(`${id}: `);
}
});
});

View file

@ -0,0 +1,109 @@
import * as cdk from "aws-cdk-lib";
import { Annotations, Match, Template } from "aws-cdk-lib/assertions";
import * as iam from "aws-cdk-lib/aws-iam";
import { KebabNamingAspect, isKebabCase } from "../lib/aspects/kebab-naming-aspect";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
describe("isKebabCase", () => {
it.each([
"open-swe-dev",
"open-swe-prod-instance-role",
"githubdeploy-open-swe-infra",
"open-swe-dev/slack-signing", // Secrets Manager path
"/open-swe-dev/feature-flag", // SSM param path
"/open-swe/dev/agent", // log group path
"abc123",
])("accepts conforming name %s", (name) => {
expect(isKebabCase(name)).toBe(true);
});
it.each([
"OpenSweDev",
"open_swe_dev",
"openSweDev",
"Open-Swe-Dev",
"open-swe-dev/SlackSigning",
])("rejects non-conforming name %s", (name) => {
expect(isKebabCase(name)).toBe(false);
});
});
describe("KebabNamingAspect", () => {
it("passes the real app stacks (no errors)", () => {
const app = new cdk.App();
cdk.Aspects.of(app).add(new KebabNamingAspect());
new OpenSweIamStack(app, "OpenSweIamStack", { stackName: "open-swe-iam", env: ENV });
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
const prod = new OpenSweStack(app, "OpenSweProdStack", {
stackName: "open-swe-prod",
env: ENV,
envName: "prod",
});
for (const s of [dev, prod]) {
Annotations.fromStack(s).hasNoError("*", Match.anyValue());
}
});
it("exempts Secrets Manager + SSM names that carry the literal env-var segment", () => {
// The config store names a secret open-swe-dev/ANTHROPIC_API_KEY and a param
// /open-swe-dev/SANDBOX_TYPE — the UPPER_SNAKE last segment is a REQUIRED
// deviation from kebab (fetch-config naming contract). Must NOT be flagged.
const app = new cdk.App();
cdk.Aspects.of(app).add(new KebabNamingAspect());
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
const tpl = Template.fromStack(dev);
// Sanity: the shells actually render with the literal env-var names.
tpl.hasResourceProperties("AWS::SecretsManager::Secret", {
Name: "open-swe-dev/ANTHROPIC_API_KEY",
});
tpl.hasResourceProperties("AWS::SSM::Parameter", {
Name: "/open-swe-dev/SANDBOX_TYPE",
Value: "langsmith",
});
Annotations.fromStack(dev).hasNoError("*", Match.anyValue());
});
it("flags a deliberately non-kebab-case resource name", () => {
const app = new cdk.App();
const stack = new cdk.Stack(app, "ConformingStackId", { stackName: "open-swe-test", env: ENV });
cdk.Aspects.of(stack).add(new KebabNamingAspect());
// Deliberately bad physical name — must be flagged.
new iam.Role(stack, "BadlyNamedRole", {
roleName: "OpenSweBadRole",
assumedBy: new iam.ServicePrincipal("ec2.amazonaws.com"),
});
Annotations.fromStack(stack).hasError(
"*",
Match.stringLikeRegexp("not kebab-case"),
);
});
it("flags a deliberately non-kebab-case stack name", () => {
const app = new cdk.App();
// PascalCase stackName — the convention CDK defaults to and that we forbid.
const stack = new cdk.Stack(app, "BadStack", { stackName: "OpenSweBadStack", env: ENV });
cdk.Aspects.of(stack).add(new KebabNamingAspect());
Annotations.fromStack(stack).hasError(
"*",
Match.stringLikeRegexp("Stack name .* is not kebab-case"),
);
});
});

24
infra/tsconfig.json Normal file
View file

@ -0,0 +1,24 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "commonjs",
"lib": ["ES2022"],
"types": ["node", "jest"],
"declaration": true,
"strict": true,
"noImplicitAny": true,
"strictNullChecks": true,
"noImplicitReturns": true,
"noFallthroughCasesInSwitch": true,
"inlineSourceMap": true,
"inlineSources": true,
"strictPropertyInitialization": false,
"outDir": "./cdk.out",
"rootDir": ".",
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"resolveJsonModule": true,
"esModuleInterop": true
},
"exclude": ["node_modules", "cdk.out"]
}

6
uv.lock generated
View file

@ -1614,15 +1614,15 @@ wheels = [
[[package]]
name = "langgraph-checkpoint"
version = "4.1.0"
version = "4.1.1"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "langchain-core" },
{ name = "ormsgpack" },
]
sdist = { url = "https://files.pythonhosted.org/packages/02/b4/6005c5dd88ad484fe6235d4c43a0d2cee7e91b08ad85a180985c2662df87/langgraph_checkpoint-4.1.0.tar.gz", hash = "sha256:e5bb304e30fc1363ac8fcb5f7dee5ca2185d77fe475b0d01de2c5f91324c2c21", size = 181942, upload-time = "2026-05-12T03:33:49.888Z" }
sdist = { url = "https://files.pythonhosted.org/packages/83/47/886af6f886f0bff2273164a45f008694e48a96ff3cd25ff0228f2aa9480e/langgraph_checkpoint-4.1.1.tar.gz", hash = "sha256:6c2bdb530c91f91d7d9c1bd100925d0fc4f498d418c17f3587d1526279482a25", size = 184020, upload-time = "2026-05-22T16:57:38.503Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/93/74/d3be2b41955e20ccd624dba5f6fe9d38dcee385ba470a6e13ed86732fc86/langgraph_checkpoint-4.1.0-py3-none-any.whl", hash = "sha256:8bc2a0466a20c38b865ce6671b42093fd5c041133f32351cae4222e0eeaf7fb5", size = 56047, upload-time = "2026-05-12T03:33:48.548Z" },
{ url = "https://files.pythonhosted.org/packages/bd/b4/71425e3e38be92611300b9cc5e46a5bf98ab23f5ea8a75b73d02a2f1413c/langgraph_checkpoint-4.1.1-py3-none-any.whl", hash = "sha256:25d29144b082827218e7bc3f1e9b0566a4bb007895cd6cc26f66a8428739f56e", size = 56212, upload-time = "2026-05-22T16:57:37.203Z" },
]
[[package]]