# open-swe infra (CDK TypeScript) AWS infrastructure for the Open SWE → AWS migration. **Synth-only at this stage — nothing here is deployed yet.** All IAM is applied only after the Phase-1 security gate (T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review`) clears (T6). ## Layout ``` infra/ ├── bin/ │ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect ├── lib/ │ ├── config.ts # account/region/org constants, env type, OIDC trust subjects │ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles │ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService) │ ├── aspects/ │ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name │ └── constructs/ │ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app │ ├── instance-role.ts # open-swe--instance-role (least-privilege) │ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11) │ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12) │ └── ami-cache.ts # cached ARM64 AL2023 helper + EBS/AMI discipline docs ├── test/ │ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones ├── cdk.json ├── cdk.context.json # COMMITTED — pins the AMI (see AMI cache discipline) ├── package.json # aws-cdk-lib pinned EXACT (2.260.0) ├── tsconfig.json ├── jest.config.js └── .gitignore ``` ## Stacks | Stack name (kebab) | Construct | Contents | |---|---|---| | `open-swe-iam` | `OpenSweIamStack` | Account-level shared GitHub OIDC deploy roles (singletons). | | `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role`, config store (T11), and the EC2 box + ALB ingress (T12, `AppService`). | | `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role`, config store, and the EC2 box + ALB ingress. | Account `328440206208`, region `us-east-1`. Stack names are set explicitly so CDK never defaults to PascalCase; resource names follow `open-swe--*`. > The two env stacks (`open-swe-dev` / `open-swe-prod`) are the required pair. The > shared OIDC deploy roles are account-wide singletons (one `RoleName` each), so > they live in their own dedicated `open-swe-iam` stack rather than being > duplicated across the env stacks — and that stack deploys first (see ordering). ## IAM roles defined (unapplied) - **`githubdeploy-open-swe-infra`** — GitHub OIDC role for CDK/CFN infra deploys. Trust scoped to `repo:Sea-Haven-Industries/open-swe` on the `main`/`dev` branches only. Permission is the org-standard CDK pattern: `sts:AssumeRole` on the CDK bootstrap roles (`cdk-hnb659fds-*`) — the real CFN/IAM/resource scope lives in the bootstrap `cfn-exec-role`, not in this role. - **`githubdeploy-open-swe-app`** — GitHub OIDC role for app deploys. Tag-scoped `ssm:SendCommand` (instances tagged `project=open-swe` + `env in {dev,prod}`) + read-only access to the `open-swe--assets` S3 artifact buckets. - **`open-swe--instance-role`** — EC2 instance role, least-privilege: read `open-swe--assets` (S3), read `/open-swe-/*` (SSM), read `open-swe-/*` (Secrets Manager), put `/open-swe//*` CloudWatch Logs, plus `AmazonSSMManagedInstanceCore` for SSM agent registration. No admin. The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created. ## Kebab-case naming Aspect `KebabNamingAspect` (applied app-wide in `bin/app.ts`) fails synth via `Annotations.addError` when a stack name or an explicit physical resource name (`RoleName`, `BucketName`, …) is not kebab-case. Path-style names (Secrets Manager `a/b`, SSM `/a/b`, log groups `/aws/.../x`) are validated per `/`-segment. CDK logical construct ids are intentionally NOT validated (they are conventionally PascalCase). Covered by `test/kebab-naming-aspect.test.ts`. ## Config store (Secrets Manager + SSM shells — T11) `ConfigStore` (`lib/constructs/config-store.ts`, one per env from `OpenSweStack`) renders the resource shells the boot hook `deploy/seahaven/fetch-config.sh` reads. The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var names as the last path segment — `open-swe-/` for secrets, `/open-swe-/` (FLAT) for config — because fetch-config strips the prefix and exports that segment verbatim. Three buckets: 1. **Secret shells (Secrets Manager) — 27 secrets.** Created value-LESS (an L1 `CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which CloudFormation creates as an empty secret). The real value is set **out-of-band** (`put-config.sh`) — CDK never owns it, so a later `cdk deploy` can never clobber it. `UpdateReplacePolicy/DeletionPolicy: Retain` so a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omits `kms:Decrypt`). > The T9 header says "29 secrets" but its table enumerates **27** distinct VAR > names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. > **Confirm** the 27-vs-29 count (code-only candidates not in the table: > `USER_ID_API_KEY_MAP`, `JUDGE_ANTHROPIC_BASE_URL`). 2. **IaC-managed SSM config — 8 params, real values owned in code:** | Param | dev | prod | |---|---|---| | `SANDBOX_TYPE` | `langsmith` | `langsmith` | | `DEFAULT_REPO_OWNER` | `Sea-Haven-Industries` | `Sea-Haven-Industries` | | `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` | `Sea-Haven-Industries` | | `DEFAULT_REPO_NAME` | `open-swe-pilot` *(confirm)* | `open-swe-pilot` *(confirm)* | | `DASHBOARD_BASE_URL` | `https://openswe-dev.seahaven.com` *(confirm host)* | `https://openswe.seahaven.com` *(confirm host)* | | `DASHBOARD_API_BASE_URL` | same as base | same as base | | `DASHBOARD_ALLOWED_ORIGINS` | same as base | same as base | | `LLM_MODEL_ID` | `anthropic:claude-opus-4-8` *(confirm)* | `anthropic:claude-opus-4-8` *(confirm)* | 3. **Out-of-band SSM config — NOT created by CDK.** Operationally-variable or env-specific-unknown values listed in `OUT_OF_BAND_SSM` and populated by `put-config.sh`. The keystone is `DEFAULT_SANDBOX_SNAPSHOT_ID` (changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls. ### Kebab-Aspect deviation `KebabNamingAspect` exempts `AWS::SecretsManager::Secret` and `AWS::SSM::Parameter` from the kebab check (see the `KEBAB_EXEMPT_RESOURCE_TYPES` set) — the UPPER_SNAKE env-var segment is a required, documented deviation for a lossless store→env round-trip. Every other explicitly-named resource is still validated. Covered by a dedicated case in `test/kebab-naming-aspect.test.ts`. ### Deploy ordering (values BEFORE the box boots) The shells are synth-able now (T11). Population is out-of-band and happens **after** `cdk deploy open-swe-` but **before** the EC2/T12 box first boots: ```bash cdk deploy open-swe- # creates the 27 secret shells + 8 IaC params deploy/seahaven/put-config.sh # sets the 27 secret values + out-of-band SSM deploy/seahaven/fetch-config.sh # (on the box) fail-fast verify before first start ``` `put-config.sh` ships `` placeholders only (no real secret values committed); provide each value inline, via `OPENSWE_PUT_` env vars, or from a vault. It does NOT touch the IaC-managed params (CDK owns those — editing them here would drift). ## Compute + ingress (`AppService` — T12) `AppService` (`lib/constructs/app-service.ts`, one per env from `OpenSweStack`) builds the box and its path to the internet. A **single** internet-facing ALB (`app/seahaven-com`) and a **single** VPC are shared with the on-prem `seahaven-site` stack, so open-swe **imports** the VPC, the ALB security group (`sg-0b0301deed193258a`), the `:443` listener, and the public `seahaven.com` zone — and never owns/mutates them. It **adds**: - **One ARM64 EC2 box** (`open-swe--box`, `t4g.medium` dev / `t4g.large` prod) in **private1 (us-east-1a)** — same AZ as the single NAT for in-AZ egress. `requireImdsv2`, gp3 **encrypted** root, `deleteOnTermination` (no RETAIN volume — see below). `userDataCausesReplacement: true`; user-data is rendered from `deploy/ami/user-data.sh`. - **A standalone instance SG** reachable **only** from the shared ALB SG on `:80` (nginx). Egress open (NAT). The ALB SG is opened to the box via a **standalone `CfnSecurityGroupEgress`** so the imported (on-prem-owned) SG is never mutated. - **A target group → instance `:80`** (nginx is the sole ingress; the LangGraph control plane stays on loopback `:2024`). Health check `GET /healthz`. - **Two listener rules** on the imported `:443` listener, both → the same TG: - **Webhooks** (priority **2** dev / **3** prod): `host ∈ {openswe-, hooks-}.seahaven.com` **AND** path `/webhooks/*`. - **Site** (priority **10** dev / **11** prod): `host = openswe-.seahaven.com` (dashboard SPA + `/dashboard/api/`). - **Route53 alias records** `openswe[-dev]` + `hooks[-dev]` → the shared ALB. - **Four CloudWatch log groups** (`/open-swe//{app,user-data,nginx-access,nginx-error}`) at **30-day** retention (IaC-owned; mirrors the CW-agent config). ### Listener-rule ordering (load-bearing) The shared listener already has a **host-agnostic** `/webhooks/*` PATH rule at **priority 5** (on-prem). ALB rules are first-match by ascending priority, so the open-swe webhook rule **must** sit below 5 or every `…/webhooks/*` request (any host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs a host condition, so it does **not** steal the on-prem hosts' webhooks. The dashboard "site" rule carries no path that collides with rule 5, so it sits at 10/11. **Cross-stack coordination (T13 review).** The `seahaven-site` (on-prem) and `open-swe` stacks both add resources to the *same imported* listener and ALB SG. This is safe: each stack owns only the resources it declares (its own logical ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa, and the standalone `CfnSecurityGroupEgress` never mutates the shared SG's own definition (the pattern on-prem itself uses). The one shared namespace that needs care is **listener-rule priority** (globally unique per listener; a collision is a fail-*safe* deploy error, not silent drift). Ownership — keep disjoint: `seahaven-site` = **4-7 + default**; `open-swe` = **2, 3, 10, 11**. open-swe's webhook rules are host-scoped to its own `*.seahaven.com` hosts, so they never match an on-prem `seahavenind.com` host. ### Security review (T5/T12 `/sh-security-review`) The T12 surface was run through the detector-fan-out + proof-or-kill verifier. One **confirmed medium** (OSWE-T12-01: nginx's 1 MB default `client_max_body_size` would 413 large GitHub webhooks before in-app signature verification) is fixed in `open-swe.nginx.conf` (`25m` on `/webhooks/`, `10m` on `/dashboard/api/`). The hooks hostname is scoped to `/webhooks/*` only (OSWE-T12-02 hygiene). An X-Forwarded-For spoof candidate was **killed** — no code trusts the leftmost XFF. No confirmed critical/high; no block. ## AMI cache discipline (EBS-fix plumbing — consumed by T12 `AppService`) `cachedArm64AmazonLinux2023()` (in `lib/constructs/ami-cache.ts`) returns an ARM64 Amazon Linux 2023 image with `cachedInContext: true`, so the resolved AMI id is pinned in the committed `cdk.context.json`. Without the pin, every deploy could pick up a newer AL2023 release → AMI change → **EC2 instance replacement** (the file-share data-loss root cause — memory `feedback_inline_ebs_volumes`). Design intent (now consumed by `AppService`): - `userDataCausesReplacement: true` is the **deliberate** choice — user-data is provisioning-only and carries no durable state. - **No durable state on the box → no RETAIN volume.** The in-memory langgraph store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The goal is replacement-*tolerance*, not avoidance. - **Snapshot-before-replace** still applies operationally: before any replacing deploy snapshot the root volume and wait `state=completed`, and re-verify "no local-only durable state" first. Refresh the AMI pin deliberately: ```bash cdk context --reset 'ssm:account=328440206208:parameterName=/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64:region=us-east-1' cdk synth # review the diff — it WILL show "requires replacement" ``` > The committed `cdk.context.json` ships a dummy-but-valid-shaped AMI id > (`ami-00000000000000000`) so `cdk synth` resolves the cache locally without any > live AWS call. **Before the first real deploy**, repoint `AppService` to the > baked `open-swe-base-arm64` AMI and pin its real id — the placeholder is > intentionally un-bootable on stock AL2023. ## Commands ```bash npm install npx cdk synth open-swe-iam npx cdk synth open-swe-dev npx cdk synth open-swe-prod npm test # jest — naming Aspect ``` ## Deploy ordering (when the gate clears — NOT yet) 1. **`open-swe-iam` first** — create `githubdeploy-open-swe-infra` + set the repo `AWS_DEPLOY_ROLE_ARN` secret before any infra/secrets CI step (BLOCK#3). 2. **Security gate** — T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review` on the synth; resolve every confirmed critical/high. 3. **IAM applied** (T6) — only after the gate. 4. Env stacks (`open-swe-dev`, then `open-swe-prod`) build out at T12+, prod gated by a GitHub Environment manual approval. ## Version policy `aws-cdk-lib` is pinned EXACT (`2.260.0`) — no `^`/`~`. Dependabot keeps it current; CI (`npm ci` + `cdk synth`) + dependency review gate each bump. See `aws-infrastructure.md` "CDK Version Policy" and memory `feedback_cdk_lib_bundled_deps`.