# open-swe infra (CDK TypeScript) AWS infrastructure for the Open SWE → AWS migration. **Synth-only at this stage — nothing here is deployed yet.** All IAM is applied only after the Phase-1 security gate (T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review`) clears (T6). ## Layout ``` infra/ ├── bin/ │ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect ├── lib/ │ ├── config.ts # account/region/org constants, env type, OIDC trust subjects │ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles │ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService) │ ├── aspects/ │ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name │ └── constructs/ │ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app │ ├── instance-role.ts # open-swe--instance-role (least-privilege) │ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11) │ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12) │ └── ami-cache.ts # baked open-swe AMI pin (by id) + EBS/replacement docs ├── test/ │ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones ├── cdk.json ├── cdk.context.json # COMMITTED — {} (AMI is a static id pin; no lookups) ├── package.json # aws-cdk-lib pinned EXACT (2.260.0) ├── tsconfig.json ├── jest.config.js └── .gitignore ``` ## Stacks | Stack name (kebab) | Construct | Contents | |---|---|---| | `open-swe-iam` | `OpenSweIamStack` | Account-level shared GitHub OIDC deploy roles (singletons). | | `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role`, config store (T11), and the EC2 box + ALB ingress (T12, `AppService`). | | `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role`, config store, and the EC2 box + ALB ingress. | Account `328440206208`, region `us-east-1`. Stack names are set explicitly so CDK never defaults to PascalCase; resource names follow `open-swe--*`. > The two env stacks (`open-swe-dev` / `open-swe-prod`) are the required pair. The > shared OIDC deploy roles are account-wide singletons (one `RoleName` each), so > they live in their own dedicated `open-swe-iam` stack rather than being > duplicated across the env stacks — and that stack deploys first (see ordering). ## IAM roles defined (unapplied) - **`githubdeploy-open-swe-infra`** — GitHub OIDC role for CDK/CFN infra deploys. Trust scoped to `repo:Sea-Haven-Industries/open-swe` on the `main`/`dev` branches only. Permission is the org-standard CDK pattern: `sts:AssumeRole` on the CDK bootstrap roles (`cdk-hnb659fds-*`) — the real CFN/IAM/resource scope lives in the bootstrap `cfn-exec-role`, not in this role. - **`githubdeploy-open-swe-app`** — GitHub OIDC role for app deploys. Tag-scoped `ssm:SendCommand` (instances tagged `project=open-swe` + `env in {dev,prod}`) + read-only access to the `open-swe--assets` S3 artifact buckets. - **`open-swe--instance-role`** — EC2 instance role, least-privilege: read `open-swe--assets` (S3), read `/open-swe-/*` (SSM), read `open-swe-/*` (Secrets Manager), put `/open-swe//*` CloudWatch Logs, plus `AmazonSSMManagedInstanceCore` for SSM agent registration. No admin. The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created. ## Kebab-case naming Aspect `KebabNamingAspect` (applied app-wide in `bin/app.ts`) fails synth via `Annotations.addError` when a stack name or an explicit physical resource name (`RoleName`, `BucketName`, …) is not kebab-case. Path-style names (Secrets Manager `a/b`, SSM `/a/b`, log groups `/aws/.../x`) are validated per `/`-segment. CDK logical construct ids are intentionally NOT validated (they are conventionally PascalCase). Covered by `test/kebab-naming-aspect.test.ts`. ## Config store (Secrets Manager + SSM shells — T11) `ConfigStore` (`lib/constructs/config-store.ts`, one per env from `OpenSweStack`) renders the resource shells the boot hook `deploy/seahaven/fetch-config.sh` reads. The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var names as the last path segment — `open-swe-/` for secrets, `/open-swe-/` (FLAT) for config — because fetch-config strips the prefix and exports that segment verbatim. Three buckets: 1. **Secret shells (Secrets Manager) — 27 secrets.** Created value-LESS (an L1 `CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which CloudFormation creates as an empty secret). The real value is set **out-of-band** (`put-config.sh`) — CDK never owns it, so a later `cdk deploy` can never clobber it. `UpdateReplacePolicy/DeletionPolicy: Retain` so a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omits `kms:Decrypt`). > The T9 header says "29 secrets" but its table enumerates **27** distinct VAR > names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. > **Confirm** the 27-vs-29 count (code-only candidates not in the table: > `USER_ID_API_KEY_MAP`, `JUDGE_ANTHROPIC_BASE_URL`). 2. **IaC-managed SSM config — 8 params, real values owned in code:** | Param | dev | prod | |---|---|---| | `SANDBOX_TYPE` | `langsmith` | `langsmith` | | `DEFAULT_REPO_OWNER` | `Sea-Haven-Industries` | `Sea-Haven-Industries` | | `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` | `Sea-Haven-Industries` | | `DEFAULT_REPO_NAME` | `open-swe-pilot` *(confirm)* | `open-swe-pilot` *(confirm)* | | `DASHBOARD_BASE_URL` | `https://openswe-dev.seahaven.com` *(confirm host)* | `https://openswe.seahaven.com` *(confirm host)* | | `DASHBOARD_API_BASE_URL` | same as base | same as base | | `DASHBOARD_ALLOWED_ORIGINS` | same as base | same as base | | `LLM_MODEL_ID` | `anthropic:claude-opus-4-8` *(confirm)* | `anthropic:claude-opus-4-8` *(confirm)* | 3. **Out-of-band SSM config — NOT created by CDK.** Operationally-variable or env-specific-unknown values listed in `OUT_OF_BAND_SSM` and populated by `put-config.sh`. The keystone is `DEFAULT_SANDBOX_SNAPSHOT_ID` (changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls. ### Kebab-Aspect deviation `KebabNamingAspect` exempts `AWS::SecretsManager::Secret` and `AWS::SSM::Parameter` from the kebab check (see the `KEBAB_EXEMPT_RESOURCE_TYPES` set) — the UPPER_SNAKE env-var segment is a required, documented deviation for a lossless store→env round-trip. Every other explicitly-named resource is still validated. Covered by a dedicated case in `test/kebab-naming-aspect.test.ts`. ### Deploy ordering (values BEFORE the box boots) The shells are synth-able now (T11). Population is out-of-band and happens **after** `cdk deploy open-swe-` but **before** the EC2/T12 box first boots: ```bash cdk deploy open-swe- # creates the 27 secret shells + 8 IaC params deploy/seahaven/put-config.sh # sets the 27 secret values + out-of-band SSM deploy/seahaven/fetch-config.sh # (on the box) fail-fast verify before first start ``` `put-config.sh` ships `` placeholders only (no real secret values committed); provide each value inline, via `OPENSWE_PUT_` env vars, or from a vault. It does NOT touch the IaC-managed params (CDK owns those — editing them here would drift). ## Compute + ingress (`AppService` — T12) `AppService` (`lib/constructs/app-service.ts`, one per env from `OpenSweStack`) builds the box and its path to the internet. A **single** internet-facing ALB (`app/seahaven-com`) and a **single** VPC are shared with the on-prem `seahaven-site` stack, so open-swe **imports** the VPC, the ALB security group (`sg-0b0301deed193258a`), the `:443` listener, and the public `seahaven.com` zone — and never owns/mutates them. It **adds**: - **One ARM64 EC2 box** (`open-swe--box`, `t4g.medium` dev / `t4g.large` prod) in **private1 (us-east-1a)** — same AZ as the single NAT for in-AZ egress. `requireImdsv2`, gp3 **encrypted** root, `deleteOnTermination` (no RETAIN volume — see below). `userDataCausesReplacement: true`; user-data is rendered from `deploy/ami/user-data.sh`. - **A standalone instance SG** reachable **only** from the shared ALB SG on `:80` (nginx). Egress open (NAT). The ALB SG is opened to the box via a **standalone `CfnSecurityGroupEgress`** so the imported (on-prem-owned) SG is never mutated. - **A target group → instance `:80`** (nginx is the sole ingress; the LangGraph control plane stays on loopback `:2024`). Health check `GET /healthz`. - **Two listener rules** on the imported `:443` listener, both → the same TG: - **Webhooks** (priority **2** dev / **3** prod): `host ∈ {openswe-, hooks-}.seahaven.com` **AND** path `/webhooks/*`. - **Site** (priority **10** dev / **11** prod): `host = openswe-.seahaven.com` (dashboard SPA + `/dashboard/api/`). - **Route53 alias records** `openswe[-dev]` + `hooks[-dev]` → the shared ALB. - **Four CloudWatch log groups** (`/open-swe//{app,user-data,nginx-access,nginx-error}`) at **30-day** retention (IaC-owned; mirrors the CW-agent config). ### Listener-rule ordering (load-bearing) The shared listener already has a **host-agnostic** `/webhooks/*` PATH rule at **priority 5** (on-prem). ALB rules are first-match by ascending priority, so the open-swe webhook rule **must** sit below 5 or every `…/webhooks/*` request (any host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs a host condition, so it does **not** steal the on-prem hosts' webhooks. The dashboard "site" rule carries no path that collides with rule 5, so it sits at 10/11. **Cross-stack coordination (T13 review).** The `seahaven-site` (on-prem) and `open-swe` stacks both add resources to the *same imported* listener and ALB SG. This is safe: each stack owns only the resources it declares (its own logical ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa, and the standalone `CfnSecurityGroupEgress` never mutates the shared SG's own definition (the pattern on-prem itself uses). The one shared namespace that needs care is **listener-rule priority** (globally unique per listener; a collision is a fail-*safe* deploy error, not silent drift). Ownership — keep disjoint: `seahaven-site` = **4-7 + default**; `open-swe` = **2, 3, 10, 11**. open-swe's webhook rules are host-scoped to its own `*.seahaven.com` hosts, so they never match an on-prem `seahavenind.com` host. ### Security review (T5/T12 `/sh-security-review`) The T12 surface was run through the detector-fan-out + proof-or-kill verifier. One **confirmed medium** (OSWE-T12-01: nginx's 1 MB default `client_max_body_size` would 413 large GitHub webhooks before in-app signature verification) is fixed in `open-swe.nginx.conf` (`25m` on `/webhooks/`, `10m` on `/dashboard/api/`). The hooks hostname is scoped to `/webhooks/*` only (OSWE-T12-02 hygiene). An X-Forwarded-For spoof candidate was **killed** — no code trusts the leftmost XFF. No confirmed critical/high; no block. ## Baked AMI + EBS-replacement discipline `bakedOpenSweArm64()` (in `lib/constructs/ami-cache.ts`) pins the custom **open-swe-base-arm64** image by EXACT id (`BAKED_OPEN_SWE_AMI_ID`) via `MachineImage.genericLinux({ "us-east-1": "" })` — no SSM lookup, so synth and deploy are fully offline/deterministic. The image is built by `deploy/ami/open-swe-base.pkr.hcl` (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW agent + boot templates, **no secrets**); the box's `user-data.sh` assumes that baked layout (`/opt/open-swe`, `openswe` user, nginx, CW agent). Pinning by exact id (vs a `most_recent` name filter) is what prevents a routine deploy from silently swapping the AMI → **EC2 instance replacement** (the file-share data-loss root cause — memory `feedback_inline_ebs_volumes`). - `userDataCausesReplacement: true` is **deliberate** — user-data is provisioning-only and carries no durable state. - **No durable state on the box → no RETAIN volume.** The in-memory langgraph store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The goal is replacement-*tolerance*, not avoidance. - **Snapshot-before-replace** still applies operationally: before any replacing deploy snapshot the root volume and wait `state=completed`, and re-verify "no local-only durable state" first. Refresh the AMI deliberately: ```bash cd deploy/ami && packer build open-swe-base.pkr.hcl # prints the new ami-… id # update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts cd infra && npx cdk diff OpenSweDevStack # WILL show "requires replacement" ``` > `cdk.context.json` is `{}` — nothing is resolved via context anymore (the AMI is > a static id pin), so synth makes no live AWS call. ## Commands ```bash npm install npx cdk synth open-swe-iam npx cdk synth open-swe-dev npx cdk synth open-swe-prod npm test # jest — naming Aspect ``` ## CI/CD (T18 — `.github/workflows/ci-infra.yml` + `cd-infra.yml`) Path-filtered, OIDC-only (no static keys). The Python agent keeps its own `ci.yml` ("Agent CI"); these two add the `/infra` half. | Workflow | Trigger | Does | |---|---|---| | `ci-infra.yml` | PR touching `infra/**` | `tsc` + `jest` + `cdk synth` (reusable `ci-typescript-cdk.yaml`). | | `cd-infra.yml` | push to `dev`/`main` touching `infra/**`, or dispatch | CI (pre-deploy) → per-env `cdk deploy`. | `cd-infra.yml` flow: - **push to `dev`** → CI green → **auto** `cdk deploy OpenSweDevStack` (assumes `githubdeploy-open-swe-infra-dev`; the job declares **no** `environment:`, so the OIDC subject is `…:ref:refs/heads/dev` — matching that role's trust). - **push to `main`** → CI green → `cdk deploy OpenSweProdStack` behind the **`prod` GitHub Environment** (required reviewer = Adam). The `environment: prod` declaration both fires the manual-approval gate and makes the OIDC subject `…:environment:prod` — matching `githubdeploy-open-swe-infra-prod`'s trust. **Why not the reusable `cd-cdk.yaml`:** it runs `cdk deploy --all`, which from a single-env push would deploy the *other* env + the shared IAM stack — breaking the per-env boundary. So CD targets one stack explicitly per env. The shared `open-swe-iam` stack is **not** deployed by CD (privileged, human-gated — T6). **Gating note:** infra CI is enforced at the *deploy* boundary (`cd-infra`'s `deploy-*` jobs `needs: ci`), not as a branch-protection required check — path-filtering a *required* check would deadlock app-only PRs (a skipped required check never satisfies). Making `Infra CI` a required check later needs a skip-aware shim or dropping its path filter. **Prerequisites (set post-T6, when the roles exist):** - repo **variables** `AWS_DEPLOY_ROLE_INFRA_DEV` / `AWS_DEPLOY_ROLE_INFRA_PROD` = the `githubdeploy-open-swe-infra-` role ARNs (`open-swe-iam` outputs). - a GitHub **Environment** named `prod` with Adam as a required reviewer. > App-side CD (CI → S3 artifact → SSM deploy via `githubdeploy-open-swe-app-`) > is **T19**, not here. ## Deploy ordering (when the gate clears — NOT yet) 1. **`open-swe-iam` first** — apply the IAM stack (T6, human-gated), then set the repo `AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}` variables from its role-ARN outputs and configure the `prod` Environment reviewer (BLOCK#3). 2. **Security gate** — T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review` on the synth; resolve every confirmed critical/high. 3. **IAM applied** (T6) — only after the gate. 4. Env stacks: first `open-swe-dev` (T14, manual validate), then CD auto-deploys dev on push; `open-swe-prod` (T21) behind the `prod` Environment approval. ## Version policy `aws-cdk-lib` is pinned EXACT (`2.260.0`) — no `^`/`~`. Dependabot keeps it current; CI (`npm ci` + `cdk synth`) + dependency review gate each bump. See `aws-infrastructure.md` "CDK Version Policy" and memory `feedback_cdk_lib_bundled_deps`.