Commit graph

13 commits

Author SHA1 Message Date
Adam Moussa
f379fbdaa9
docs: document RETAIN secret-shell orphan gotcha (#52)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
RETAIN + a fixed secret name means a failed FIRST create leaves empty
secret shells behind when the stack rolls back. The shells keep the
global `open-swe-<env>/<VAR>` names, so every later create fails with
`AlreadyExists`, and a plain delete-secret keeps the name reserved for
the recovery window rather than freeing it.

Record the trap and the force-delete recovery (only for empty shells)
in the config-store construct and the infra README so the next
teardown/rebuild, secret logical-id change, or new-env stand-up does
not rediscover it the hard way.

Prod's first deploy hit this on 2026-06-29: 28 orphaned shells from an
earlier failed create reserved the names and had to be force-deleted
before the stack would create.
2026-06-29 10:51:49 -04:00
Adam Moussa
86b4859589
fix: env-scope the EC2 launch template name (unblocks prod deploy) (#51)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
requireImdsv2:true makes CDK auto-create a launch template named from the
construct id ('Instance' -> 'InstanceLaunchTemplate') with no env qualifier,
so OpenSweDevStack and OpenSweProdStack both render
LaunchTemplateName: InstanceLaunchTemplate. dev created it first (the live
dev box runs on it); the prod first-deploy then failed with
InvalidLaunchTemplateName.AlreadyExistsException and the whole stack rolled
back.

Force a per-env LT name (open-swe-<env>-lt) via an aspect (the LT is created
at synth time by the requireImdsv2 handling, not in the constructor), and
rename the instance's launch-template REFERENCE in lockstep so CFN still
resolves it. synth-verified: dev=open-swe-dev-lt, prod=open-swe-prod-lt on
both the LT resource and the instance reference; version GetAtt preserved.

NOTE: deploying this renames dev's LT -> one-time dev box replacement
(stateless; boots from the baked AMI + pulls releases/latest). prod then
creates open-swe-prod-lt cleanly.
2026-06-29 00:51:16 -04:00
Adam Moussa
3c69dd9de6
fix: BatchGetSecretValue must be granted on * (corrects #48, fixes dev crash-loop) (#50)
Some checks are pending
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
#48 (OSWE-IAC-SECRETS-LIST-01) scoped secretsmanager:BatchGetSecretValue
to the open-swe-<env>/* ARN on the theory that an explicit --secret-id-list
batch authorizes per-secret. That is FALSE: BatchGetSecretValue is a
collection action AWS authorizes against the account (*), regardless of
--filters vs --secret-id-list. A prefix-scoped grant AccessDenies the whole
call. The dev box passed right after #48 only because the prior broad grant
had not finished propagating; once it lapsed, fetch-config got AccessDenied
-> loaded 0 secrets -> FAIL-FAST -> open-swe.service crash-loop. Verified on
the live dev box (i-0af4e03e8bf70e6c3): the exact call returned
'not authorized to perform: secretsmanager:BatchGetSecretValue'; restoring
the * grant recovered it.

Move BatchGetSecretValue back to Resource:* (its own statement); keep
GetSecretValue + DescribeSecret prefix-scoped (those gate VALUE access, so
cross-env isolation holds). The surviving win from #48: --secret-id-list
needs no name filter, so ListSecrets stays dropped -> no account-wide name
enumeration. fetch-config.sh is unchanged (--secret-id-list is correct).

The /sh-security-review finding OSWE-IAC-IAM-01 called this out and was
wrongly refuted; the reference_secretsmanager_batch_get memory was wrong.
2026-06-28 22:08:02 -04:00
Adam Moussa
faae9a685b
Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Switch fetch-config.sh from a name-prefix batch-get-secret-value
--filters scan to an explicit --secret-id-list (the 28 SECRET_VARS,
chunked at the 20/call cap). An id-list batch authorizes per-secret
ARN, so the instance role's BatchGetSecretValue moves from Resource:*
to the open-swe-<env>/* prefix and the account-wide ListSecrets grant
is dropped entirely. The box can no longer enumerate secret names
account-wide; cross-env value isolation is unchanged (GetSecretValue
was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01.

Also capture each chunk response into a variable and consume the
producer via command substitution so a failed AWS call aborts under
set -e instead of being swallowed by process substitution and
misreported as a missing required var.

Refs: OSWE-IAC-SECRETS-LIST-01
2026-06-28 20:45:17 -04:00
Adam Moussa
3224d7abf4
ci: gate dev→prod promotion on green checks + add rollback safety net (#28)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
Agent CI / Agent lint (push) Waiting to run
Agent CI / Agent format check (push) Waiting to run
Agent CI / Agent unit tests (push) Waiting to run
Agent CI / Playwright E2E (push) Waiting to run
T20 CD safety nets. Two gaps closed before the first real prod deploy:

1. Promotion gate. promote_dev_to_prod.yml previously fast-forwarded
   dev→main unconditionally. It now hard-gates on check-dev-green.sh:
   every check-run on the dev HEAD must be completed+passing AND the
   Agent CI suite (lint/format/unit/E2E) must be present+success, or the
   promotion blocks (fails safe on a missing/renamed check). The promote
   run excludes its OWN check-run by run-id (unforgeable), never by the
   mutable name "promote", so a colliding red check cannot hide. Fields
   are read with a 0x1F separator so an empty conclusion (every
   in_progress check) cannot shift columns. ci.yml now also runs on
   push:dev so dev HEAD actually carries that signal (a PR check alone
   can be admin-merged past).

2. Rollback + last-good. publish-and-deploy.sh advances
   releases/last-good/ only after a successful roll (deploy.sh gates on
   `systemctl is-active`), and makes releases/latest/ transactional —
   reverting to the prior release if the roll fails so a replaced box
   never self-deploys a broken release. New rollback.yml + rollback.sh
   re-point latest at last-good (or an explicit sha) and re-fire the
   deploy; prod is gated by the `prod` Environment approval, same as a
   deploy. The shared fire/wait/aggregate-gate logic is factored into
   roll-box.sh (used by both forward and backward rolls).

Least-privilege: drop the unused s3:DeleteObject from the app deploy
role — publish/rollback/deploy only Get+Put (S3-to-S3 copy), and the
rollback fallback now depends on immutable release history staying
intact. Lifecycle expiry (not CI) handles old-version cleanup.

Gate logic unit-tested (7 cases + jq round-trip). IAM change +
release-safety control cross-reviewed by GPT-4.1: APPROVE, no blocks.

Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
2026-06-27 20:21:59 -04:00
Adam Moussa
5e30bc6be2
Point dev agent at seahaven-open-swe-dev org + un-pin the owner guard (#27)
Dev now runs against the dedicated seahaven-open-swe-dev org (repo openswe-dev-sandbox),
isolated from the real Sea Haven org. Two changes:

- config-store.ts: iacManagedSsm repo targeting is now per-env — dev =
  seahaven-open-swe-dev/openswe-dev-sandbox, prod stays Sea-Haven-Industries/open-swe-pilot.
  ALLOWED_GITHUB_ORGS tracks the env owner. Adds a dev-only SEED_USER_MAPPINGS param so the
  triggering GitHub login (amoussa1229) resolves and @openswe comments aren't skipped.

- fetch-config.sh: replace the unconditional hard-pin to Sea-Haven-Industries with an owner
  GUARD that HONORS the configured owner (OPENSWE_REPO_OWNER override, else the SSM value) but
  forces a per-env safe org when the normalized owner is blank or the upstream langchain-ai.
  Normalization (lowercase, strip whitespace, first path segment, drop dots) catches
  langchain-ai/<repo>, langchain-ai., and case variants without over-blocking legit orgs
  (e.g. langchain-ai-fork). Fallback org is per-env so dev can't fall back into the real org.

Reviews: GPT-4.1 cross-review APPROVE (round 1 found a path/dot bypass -> hardened, round 2 clean);
/sh-security-review authz one LOW (env-invariant fallback) -> fixed. Positive org allowlist still
enforced by the app via ALLOWED_GITHUB_ORGS.
2026-06-27 19:29:03 -04:00
Adam Moussa
cdcb6625b9
fix: BatchGetSecretValue must be on * for the filtered batch call (#23)
The prior fix scoped secretsmanager:BatchGetSecretValue to the env-prefixed secret
ARN, but the live box still got AccessDenied: batch-get-secret-value invoked WITH a
name --filters is a COLLECTION call that AWS authorizes against * (a per-secret ARN
does not satisfy it). Split the statement:

- GetSecretValue + DescribeSecret stay PREFIX-scoped (secret:open-swe-<env>/*) — this
  is what gates which secret VALUES the box can read (checked per-secret in the batch).
- BatchGetSecretValue + ListSecrets move to a * operation-level statement (the filtered
  collection call + the list action; neither is resource-scopable for this usage).

VALUE isolation preserved (dev box still cannot read prod secret values); only secret
NAME/metadata enumeration is widened. GPT-4.1 IAM cross-review: BLOCK none, FIX none.
Suppression OSWE-IAC-SECRETS-LIST-01 updated; future hardening (explicit --secret-id-list
to drop both * grants) tracked there.
2026-06-26 19:44:12 -04:00
Adam Moussa
cfbdcda9b7
fix: instance role BatchGetSecretValue + ListSecrets for .env materialization (#22)
* fix(infra): grant instance role BatchGetSecretValue + ListSecrets for .env materialization

fetch-config.sh materializes the box's .env via
`secretsmanager batch-get-secret-value --filters Key=name,Values=open-swe-<env>/`,
but the instance role only granted GetSecretValue/DescribeSecret. BatchGetSecretValue
is a distinct IAM action, so the call was AccessDenied and open-swe.service
crash-looped (no .env written -> ExecStartPre exit 1).

- Add secretsmanager:BatchGetSecretValue to the prefix-scoped ReadSecrets statement.
- Add secretsmanager:ListSecrets on * (required by the name-prefix filtered batch
  call; the API has no resource-level scoping for the list action — fits the role's
  stated exception). Secret VALUES stay prefix-scoped; only names are enumerable.

Reviews: GPT-4.1 IAM cross-review BLOCK=none; /sh-security-review iac-iam one LOW
metadata residual (no critical/high), recorded as OSWE-IAC-SECRETS-LIST-01.

Refs T7/T19 dev bring-up.

* ci: lift Node heap cap for Playwright E2E build (vite OOM)

The E2E job's Playwright globalSetup runs the real `bun run build`, whose vite
bundle exceeds Node's default ~2 GB heap and OOMs (JavaScript heap out of memory) —
the same failure fixed for build-artifacts.yml in #19. Set
NODE_OPTIONS=--max-old-space-size=8192 on the Run E2E step.
2026-06-26 19:29:50 -04:00
Adam Moussa
5fa132205b
fix(deploy): use %%...%% for CDK user-data tokens (don't collide with @@ sed) (#21)
The systemd unit booted with a literal `@@OPENSWE_ENV@@` (fetch-config.sh got the
token, not "dev" -> exit 2 -> crash-loop) because user-data.sh is double-templated:
CDK substitutes @@tokens@@ AND user-data seds @@tokens@@ into the baked
systemd/nginx files. CDK's `.replace(/@@OPENSWE_ENV@@/g, "dev")` clobbered the sed
PATTERN (`s|@@OPENSWE_ENV@@|...|` -> `s|dev|...|`, a no-op), so the unit's token
never got replaced. Same collision hit @@SERVER_NAME@@ (masked by nginx
default_server).

Fix: CDK tokens move to a DISTINCT delimiter %%...%% (rendered in app-service.ts);
the @@...@@ tokens stay for the baked-template seds. No AMI rebuild (templates
unchanged). Add a guard test asserting no unresolved %%CDK%% token survives in the
synthesized user-data. Also add .github/scripts/** to build-artifacts paths so
script-only changes trigger a publish.

jest 20/20; tsc + shellcheck clean; rendered user-data: OPENSWE_ENV="dev",
SERVER_NAME="openswe-dev.seahaven.com", @@ sed patterns preserved, 16872 B.
2026-06-26 19:06:37 -04:00
Adam Moussa
404b3f6f75
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18)
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)

Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.

deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
  - the file provisioner failed uploading the templates dir ('scp: …: Is a
    directory') — a trailing-slash contents-upload needs the dest dir to exist;
    added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
    dest trailing slash.
  - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
    environment_vars never reached provision.sh (which runs under set -u and
    aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.

infra:
  - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
    from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
    exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
    now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
    discipline docs.
  - app-service.ts: machineImage → bakedOpenSweArm64().
  - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
  - cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
  - README: Baked AMI + EBS-replacement-discipline section.

tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.

NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).

* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021

Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.

* fix(deploy): GitHub App + Slack required for prod only, not dev

Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.

Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).

* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)

Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.

Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
  SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
  abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
  /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
  baked open-swe-base-arm64 AMI (folds in the held #16).

IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
  open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
  AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
  is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.

Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
  app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
  at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
  target is healthy even before the first release); deploy.sh is base64-rendered
  by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
  first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).

CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
  ui/.output/public -> spa.tar.gz), package the Python source via git archive
  (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
  the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
  push dev -> dev (auto); push main -> prod (env "prod" approval gate).

Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.

* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard

Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:

- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
  never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
  treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
  failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
  not CommandInvocations[0], so a partial failure across the brief 2-instance
  replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
  role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
  (defense in depth over .gitignore; scoped to data extensions so *_credentials.py
  source is not a false positive — verified against the real tree).

Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).

shellcheck/tsc/jest(16) clean; both stacks synth offline.

* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard

The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)

- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
  ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
  GroupDescription + rule descriptions are pure ASCII, so this fails the build
  instead of a deploy next time.

jest 18/18; tsc clean.

* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)

The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.

jest 18/18; tsc clean.

* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit

The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.

jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
Adam Moussa
fcbdfb67aa
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.

Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
  in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
  (no RETAIN volume — replacement-tolerant; see ami-cache.ts).
  userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
  via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
  (the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
  loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
    * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
    * site     (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
  Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
  or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).

Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.

Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
Adam Moussa
aef6b26912
feat: Secrets Manager + SSM config store for open-swe (T11) (#10)
Create the per-env config surface the EC2 box reads at boot via
fetch-config.sh / seed_store.sh:

- ConfigStore construct (infra/lib/constructs/config-store.ts):
  - 28 value-LESS Secrets Manager shells  open-swe-<env>/<VAR>
    (RemovalPolicy.RETAIN, no SecretString/generateSecretString — real
    values are set out-of-band by put-config.sh, never in IaC/state).
  - 8 IaC-managed SSM params  /open-swe-<env>/<VAR>  with real,
    stable/derivable values (SANDBOX_TYPE, DEFAULT_REPO_OWNER/NAME,
    ALLOWED_GITHUB_ORGS, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID).
  - OUT_OF_BAND_SSM documents the ~30 params CDK intentionally does NOT
    own (operationally-variable / env-specific-unknown).
- Wire ConfigStore into OpenSwe<Env>Stack.
- KebabNamingAspect: exempt Secrets Manager + SSM names, which carry the
  literal UPPER_SNAKE env-var segment (open-swe-dev/DASHBOARD_JWT_SECRET).
- deploy/seahaven/put-config.sh: out-of-band populator (placeholders only,
  OPENSWE_PUT_<VAR> env indirection; no real values committed).

Synth-only; not deployed. Instance-role read grants on open-swe-<env>/*
already exist from T6 — no IAM/trust changes here.
2026-06-26 16:07:23 -04:00
Adam Moussa
92fd886076
feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6)
PR#1 of the AWS migration. CDK TypeScript app under /infra: stacks open-swe-iam
(four per-env GitHub-OIDC deploy roles) + open-swe-dev/-prod (per-env EC2 instance
role + AMI cache). aws-cdk-lib pinned exact 2.260.0; kebab naming Aspect + 15 tests.

IAM is synth-only (NOT deployed). Cleared the Phase-1 security gates:
- T4 GPT-4.1 IAM cross-review (StringEquals trust; cdk-hnb659fds-* wildcard kept as
  org convention; AWS-RunShellScript timeboxed to T19).
- T5 /sh-security-review: deploy roles split PER-ENV with env-scoped OIDC trust
  (dev=branch ref+tag dev, prod=environment:prod+tag prod) so a dev token cannot
  reach prod; re-verified block:false.
2026-06-26 15:06:30 -04:00