open-swe/infra/README.md
Adam Moussa fcbdfb67aa
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.

Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
  in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
  (no RETAIN volume — replacement-tolerant; see ami-cache.ts).
  userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
  via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
  (the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
  loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
    * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
    * site     (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
  Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
  or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).

Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.

Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00

258 lines
14 KiB
Markdown

# open-swe infra (CDK TypeScript)
AWS infrastructure for the Open SWE → AWS migration. **Synth-only at this stage —
nothing here is deployed yet.** All IAM is applied only after the Phase-1 security
gate (T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review`) clears (T6).
## Layout
```
infra/
├── bin/
│ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
│ ├── aspects/
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
│ └── constructs/
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│ └── ami-cache.ts # cached ARM64 AL2023 helper + EBS/AMI discipline docs
├── test/
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json # COMMITTED — pins the AMI (see AMI cache discipline)
├── package.json # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore
```
## Stacks
| Stack name (kebab) | Construct | Contents |
|---|---|---|
| `open-swe-iam` | `OpenSweIamStack` | Account-level shared GitHub OIDC deploy roles (singletons). |
| `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role`, config store (T11), and the EC2 box + ALB ingress (T12, `AppService`). |
| `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role`, config store, and the EC2 box + ALB ingress. |
Account `328440206208`, region `us-east-1`. Stack names are set explicitly so CDK
never defaults to PascalCase; resource names follow `open-swe-<env>-*`.
> The two env stacks (`open-swe-dev` / `open-swe-prod`) are the required pair. The
> shared OIDC deploy roles are account-wide singletons (one `RoleName` each), so
> they live in their own dedicated `open-swe-iam` stack rather than being
> duplicated across the env stacks — and that stack deploys first (see ordering).
## IAM roles defined (unapplied)
- **`githubdeploy-open-swe-infra`** — GitHub OIDC role for CDK/CFN infra deploys.
Trust scoped to `repo:Sea-Haven-Industries/open-swe` on the `main`/`dev`
branches only. Permission is the org-standard CDK pattern: `sts:AssumeRole` on
the CDK bootstrap roles (`cdk-hnb659fds-*`) — the real CFN/IAM/resource scope
lives in the bootstrap `cfn-exec-role`, not in this role.
- **`githubdeploy-open-swe-app`** — GitHub OIDC role for app deploys. Tag-scoped
`ssm:SendCommand` (instances tagged `project=open-swe` + `env in {dev,prod}`) +
read-only access to the `open-swe-<env>-assets` S3 artifact buckets.
- **`open-swe-<env>-instance-role`** — EC2 instance role, least-privilege: read
`open-swe-<env>-assets` (S3), read `/open-swe-<env>/*` (SSM), read
`open-swe-<env>/*` (Secrets Manager), put `/open-swe/<env>/*` CloudWatch Logs,
plus `AmazonSSMManagedInstanceCore` for SSM agent registration. No admin.
The GitHub OIDC provider already exists account-wide (created for seahaven-site);
it is referenced by ARN, never re-created.
## Kebab-case naming Aspect
`KebabNamingAspect` (applied app-wide in `bin/app.ts`) fails synth via
`Annotations.addError` when a stack name or an explicit physical resource name
(`RoleName`, `BucketName`, …) is not kebab-case. Path-style names (Secrets
Manager `a/b`, SSM `/a/b`, log groups `/aws/.../x`) are validated per `/`-segment.
CDK logical construct ids are intentionally NOT validated (they are conventionally
PascalCase). Covered by `test/kebab-naming-aspect.test.ts`.
## Config store (Secrets Manager + SSM shells — T11)
`ConfigStore` (`lib/constructs/config-store.ts`, one per env from `OpenSweStack`)
renders the resource shells the boot hook `deploy/seahaven/fetch-config.sh` reads.
The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var
names as the last path segment — `open-swe-<env>/<VAR>` for secrets,
`/open-swe-<env>/<VAR>` (FLAT) for config — because fetch-config strips the prefix
and exports that segment verbatim.
Three buckets:
1. **Secret shells (Secrets Manager) — 27 secrets.** Created value-LESS (an L1
`CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which
CloudFormation creates as an empty secret). The real value is set **out-of-band**
(`put-config.sh`) — CDK never owns it, so a later `cdk deploy` can never clobber
it. `UpdateReplacePolicy/DeletionPolicy: Retain` so a teardown can't destroy
operator-set secret material. AWS-managed key (no CMK — matches the instance
role, which omits `kms:Decrypt`).
> The T9 header says "29 secrets" but its table enumerates **27** distinct VAR
> names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29.
> **Confirm** the 27-vs-29 count (code-only candidates not in the table:
> `USER_ID_API_KEY_MAP`, `JUDGE_ANTHROPIC_BASE_URL`).
2. **IaC-managed SSM config — 8 params, real values owned in code:**
| Param | dev | prod |
|---|---|---|
| `SANDBOX_TYPE` | `langsmith` | `langsmith` |
| `DEFAULT_REPO_OWNER` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `DEFAULT_REPO_NAME` | `open-swe-pilot` *(confirm)* | `open-swe-pilot` *(confirm)* |
| `DASHBOARD_BASE_URL` | `https://openswe-dev.seahaven.com` *(confirm host)* | `https://openswe.seahaven.com` *(confirm host)* |
| `DASHBOARD_API_BASE_URL` | same as base | same as base |
| `DASHBOARD_ALLOWED_ORIGINS` | same as base | same as base |
| `LLM_MODEL_ID` | `anthropic:claude-opus-4-8` *(confirm)* | `anthropic:claude-opus-4-8` *(confirm)* |
3. **Out-of-band SSM config — NOT created by CDK.** Operationally-variable or
env-specific-unknown values listed in `OUT_OF_BAND_SSM` and populated by
`put-config.sh`. The keystone is `DEFAULT_SANDBOX_SNAPSHOT_ID` (changes on every
snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the
GitHub App ids, Slack ids, and LangSmith tenant/urls.
### Kebab-Aspect deviation
`KebabNamingAspect` exempts `AWS::SecretsManager::Secret` and `AWS::SSM::Parameter`
from the kebab check (see the `KEBAB_EXEMPT_RESOURCE_TYPES` set) — the UPPER_SNAKE
env-var segment is a required, documented deviation for a lossless store→env
round-trip. Every other explicitly-named resource is still validated. Covered by a
dedicated case in `test/kebab-naming-aspect.test.ts`.
### Deploy ordering (values BEFORE the box boots)
The shells are synth-able now (T11). Population is out-of-band and happens **after**
`cdk deploy open-swe-<env>` but **before** the EC2/T12 box first boots:
```bash
cdk deploy open-swe-<env> # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod> # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
```
`put-config.sh` ships `<FILL>` placeholders only (no real secret values committed);
provide each value inline, via `OPENSWE_PUT_<VAR>` env vars, or from a vault. It does
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
## Compute + ingress (`AppService` — T12)
`AppService` (`lib/constructs/app-service.ts`, one per env from `OpenSweStack`)
builds the box and its path to the internet. A **single** internet-facing ALB
(`app/seahaven-com`) and a **single** VPC are shared with the on-prem
`seahaven-site` stack, so open-swe **imports** the VPC, the ALB security group
(`sg-0b0301deed193258a`), the `:443` listener, and the public `seahaven.com`
zone — and never owns/mutates them. It **adds**:
- **One ARM64 EC2 box** (`open-swe-<env>-box`, `t4g.medium` dev / `t4g.large`
prod) in **private1 (us-east-1a)** — same AZ as the single NAT for in-AZ egress.
`requireImdsv2`, gp3 **encrypted** root, `deleteOnTermination` (no RETAIN
volume — see below). `userDataCausesReplacement: true`; user-data is rendered
from `deploy/ami/user-data.sh`.
- **A standalone instance SG** reachable **only** from the shared ALB SG on `:80`
(nginx). Egress open (NAT). The ALB SG is opened to the box via a **standalone
`CfnSecurityGroupEgress`** so the imported (on-prem-owned) SG is never mutated.
- **A target group → instance `:80`** (nginx is the sole ingress; the LangGraph
control plane stays on loopback `:2024`). Health check `GET /healthz`.
- **Two listener rules** on the imported `:443` listener, both → the same TG:
- **Webhooks** (priority **2** dev / **3** prod): `host ∈ {openswe-<env>, hooks-<env>}.seahaven.com` **AND** path `/webhooks/*`.
- **Site** (priority **10** dev / **11** prod): `host = openswe-<env>.seahaven.com` (dashboard SPA + `/dashboard/api/`).
- **Route53 alias records** `openswe[-dev]` + `hooks[-dev]` → the shared ALB.
- **Four CloudWatch log groups** (`/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`) at **30-day** retention (IaC-owned; mirrors the CW-agent config).
### Listener-rule ordering (load-bearing)
The shared listener already has a **host-agnostic** `/webhooks/*` PATH rule at
**priority 5** (on-prem). ALB rules are first-match by ascending priority, so the
open-swe webhook rule **must** sit below 5 or every `…/webhooks/*` request (any
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
a host condition, so it does **not** steal the on-prem hosts' webhooks. The
dashboard "site" rule carries no path that collides with rule 5, so it sits at
10/11.
**Cross-stack coordination (T13 review).** The `seahaven-site` (on-prem) and
`open-swe` stacks both add resources to the *same imported* listener and ALB SG.
This is safe: each stack owns only the resources it declares (its own logical
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
and the standalone `CfnSecurityGroupEgress` never mutates the shared SG's own
definition (the pattern on-prem itself uses). The one shared namespace that needs
care is **listener-rule priority** (globally unique per listener; a collision is
a fail-*safe* deploy error, not silent drift). Ownership — keep disjoint:
`seahaven-site` = **4-7 + default**; `open-swe` = **2, 3, 10, 11**. open-swe's
webhook rules are host-scoped to its own `*.seahaven.com` hosts, so they never
match an on-prem `seahavenind.com` host.
### Security review (T5/T12 `/sh-security-review`)
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
One **confirmed medium** (OSWE-T12-01: nginx's 1 MB default `client_max_body_size`
would 413 large GitHub webhooks before in-app signature verification) is fixed in
`open-swe.nginx.conf` (`25m` on `/webhooks/`, `10m` on `/dashboard/api/`). The
hooks hostname is scoped to `/webhooks/*` only (OSWE-T12-02 hygiene). An
X-Forwarded-For spoof candidate was **killed** — no code trusts the leftmost XFF.
No confirmed critical/high; no block.
## AMI cache discipline (EBS-fix plumbing — consumed by T12 `AppService`)
`cachedArm64AmazonLinux2023()` (in `lib/constructs/ami-cache.ts`) returns an
ARM64 Amazon Linux 2023 image with `cachedInContext: true`, so the resolved AMI
id is pinned in the committed `cdk.context.json`. Without the pin, every deploy
could pick up a newer AL2023 release → AMI change → **EC2 instance replacement**
(the file-share data-loss root cause — memory `feedback_inline_ebs_volumes`).
Design intent (now consumed by `AppService`):
- `userDataCausesReplacement: true` is the **deliberate** choice — user-data is
provisioning-only and carries no durable state.
- **No durable state on the box → no RETAIN volume.** The in-memory langgraph
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The goal is
replacement-*tolerance*, not avoidance.
- **Snapshot-before-replace** still applies operationally: before any replacing
deploy snapshot the root volume and wait `state=completed`, and re-verify "no
local-only durable state" first.
Refresh the AMI pin deliberately:
```bash
cdk context --reset 'ssm:account=328440206208:parameterName=/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64:region=us-east-1'
cdk synth # review the diff — it WILL show "requires replacement"
```
> The committed `cdk.context.json` ships a dummy-but-valid-shaped AMI id
> (`ami-00000000000000000`) so `cdk synth` resolves the cache locally without any
> live AWS call. **Before the first real deploy**, repoint `AppService` to the
> baked `open-swe-base-arm64` AMI and pin its real id — the placeholder is
> intentionally un-bootable on stock AL2023.
## Commands
```bash
npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test # jest — naming Aspect
```
## Deploy ordering (when the gate clears — NOT yet)
1. **`open-swe-iam` first** — create `githubdeploy-open-swe-infra` + set the repo
`AWS_DEPLOY_ROLE_ARN` secret before any infra/secrets CI step (BLOCK#3).
2. **Security gate** — T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review` on
the synth; resolve every confirmed critical/high.
3. **IAM applied** (T6) — only after the gate.
4. Env stacks (`open-swe-dev`, then `open-swe-prod`) build out at T12+, prod gated
by a GitHub Environment manual approval.
## Version policy
`aws-cdk-lib` is pinned EXACT (`2.260.0`) — no `^`/`~`. Dependabot keeps it
current; CI (`npm ci` + `cdk synth`) + dependency review gate each bump. See
`aws-infrastructure.md` "CDK Version Policy" and memory
`feedback_cdk_lib_bundled_deps`.