open-swe/infra
Adam Moussa fcbdfb67aa
feat: open-swe dev/prod compute + ALB ingress (T12) (#14)
AppService construct wires the per-env EC2 box and its internet path. The
seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the
on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and
never owned/mutated; open-swe only ADDS its own resources.

Per env (open-swe-stack.ts → AppService):
- ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a,
  in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination
  (no RETAIN volume — replacement-tolerant; see ami-cache.ts).
  userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh.
- Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress
  via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress
  (the imported, on-prem-owned SG is never mutated).
- Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays
  loopback). Health check GET /healthz.
- Two rules on the imported :443 listener, both → the TG:
    * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/*
    * site     (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api)
  Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5)
  or it would steal every webhook — first-match-by-ascending-priority.
- Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB.
- 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config).

Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed
medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large
GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on
/dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene);
XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed
critical/high.

Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked
open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean.
Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
..
bin feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
lib feat: open-swe dev/prod compute + ALB ingress (T12) (#14) 2026-06-26 16:10:23 -04:00
test feat: Secrets Manager + SSM config store for open-swe (T11) (#10) 2026-06-26 16:07:23 -04:00
.gitignore feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
cdk.context.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
cdk.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
package-lock.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
package.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
README.md feat: open-swe dev/prod compute + ALB ingress (T12) (#14) 2026-06-26 16:10:23 -04:00
tsconfig.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00

open-swe infra (CDK TypeScript)

AWS infrastructure for the Open SWE → AWS migration. Synth-only at this stage — nothing here is deployed yet. All IAM is applied only after the Phase-1 security gate (T4 GPT-4.1 IAM cross-review + T5 /sh-security-review) clears (T6).

Layout

infra/
├── bin/
│   └── app.ts                       # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│   ├── config.ts                    # account/region/org constants, env type, OIDC trust subjects
│   ├── open-swe-iam-stack.ts        # account-level: shared OIDC deploy roles
│   ├── open-swe-stack.ts            # per-env stack (instance role + config store + AppService)
│   ├── aspects/
│   │   └── kebab-naming-aspect.ts   # fails synth on any non-kebab-case explicit name
│   └── constructs/
│       ├── github-deploy-roles.ts   # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│       ├── instance-role.ts         # open-swe-<env>-instance-role (least-privilege)
│       ├── config-store.ts          # Secrets Manager + SSM Parameter Store shells (T11)
│       ├── app-service.ts           # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│       └── ami-cache.ts             # cached ARM64 AL2023 helper + EBS/AMI discipline docs
├── test/
│   └── kebab-naming-aspect.test.ts  # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json                 # COMMITTED — pins the AMI (see AMI cache discipline)
├── package.json                     # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore

Stacks

Stack name (kebab) Construct Contents
open-swe-iam OpenSweIamStack Account-level shared GitHub OIDC deploy roles (singletons).
open-swe-dev OpenSweStack (envName: dev) open-swe-dev-instance-role, config store (T11), and the EC2 box + ALB ingress (T12, AppService).
open-swe-prod OpenSweStack (envName: prod) open-swe-prod-instance-role, config store, and the EC2 box + ALB ingress.

Account 328440206208, region us-east-1. Stack names are set explicitly so CDK never defaults to PascalCase; resource names follow open-swe-<env>-*.

The two env stacks (open-swe-dev / open-swe-prod) are the required pair. The shared OIDC deploy roles are account-wide singletons (one RoleName each), so they live in their own dedicated open-swe-iam stack rather than being duplicated across the env stacks — and that stack deploys first (see ordering).

IAM roles defined (unapplied)

  • githubdeploy-open-swe-infra — GitHub OIDC role for CDK/CFN infra deploys. Trust scoped to repo:Sea-Haven-Industries/open-swe on the main/dev branches only. Permission is the org-standard CDK pattern: sts:AssumeRole on the CDK bootstrap roles (cdk-hnb659fds-*) — the real CFN/IAM/resource scope lives in the bootstrap cfn-exec-role, not in this role.
  • githubdeploy-open-swe-app — GitHub OIDC role for app deploys. Tag-scoped ssm:SendCommand (instances tagged project=open-swe + env in {dev,prod}) + read-only access to the open-swe-<env>-assets S3 artifact buckets.
  • open-swe-<env>-instance-role — EC2 instance role, least-privilege: read open-swe-<env>-assets (S3), read /open-swe-<env>/* (SSM), read open-swe-<env>/* (Secrets Manager), put /open-swe/<env>/* CloudWatch Logs, plus AmazonSSMManagedInstanceCore for SSM agent registration. No admin.

The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created.

Kebab-case naming Aspect

KebabNamingAspect (applied app-wide in bin/app.ts) fails synth via Annotations.addError when a stack name or an explicit physical resource name (RoleName, BucketName, …) is not kebab-case. Path-style names (Secrets Manager a/b, SSM /a/b, log groups /aws/.../x) are validated per /-segment. CDK logical construct ids are intentionally NOT validated (they are conventionally PascalCase). Covered by test/kebab-naming-aspect.test.ts.

Config store (Secrets Manager + SSM shells — T11)

ConfigStore (lib/constructs/config-store.ts, one per env from OpenSweStack) renders the resource shells the boot hook deploy/seahaven/fetch-config.sh reads. The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var names as the last path segment — open-swe-<env>/<VAR> for secrets, /open-swe-<env>/<VAR> (FLAT) for config — because fetch-config strips the prefix and exports that segment verbatim.

Three buckets:

  1. Secret shells (Secrets Manager) — 27 secrets. Created value-LESS (an L1 CfnSecret with NEITHER secretString NOR generateSecretString, which CloudFormation creates as an empty secret). The real value is set out-of-band (put-config.sh) — CDK never owns it, so a later cdk deploy can never clobber it. UpdateReplacePolicy/DeletionPolicy: Retain so a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omits kms:Decrypt).

    The T9 header says "29 secrets" but its table enumerates 27 distinct VAR names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. Confirm the 27-vs-29 count (code-only candidates not in the table: USER_ID_API_KEY_MAP, JUDGE_ANTHROPIC_BASE_URL).

  2. IaC-managed SSM config — 8 params, real values owned in code:

    Param dev prod
    SANDBOX_TYPE langsmith langsmith
    DEFAULT_REPO_OWNER Sea-Haven-Industries Sea-Haven-Industries
    ALLOWED_GITHUB_ORGS Sea-Haven-Industries Sea-Haven-Industries
    DEFAULT_REPO_NAME open-swe-pilot (confirm) open-swe-pilot (confirm)
    DASHBOARD_BASE_URL https://openswe-dev.seahaven.com (confirm host) https://openswe.seahaven.com (confirm host)
    DASHBOARD_API_BASE_URL same as base same as base
    DASHBOARD_ALLOWED_ORIGINS same as base same as base
    LLM_MODEL_ID anthropic:claude-opus-4-8 (confirm) anthropic:claude-opus-4-8 (confirm)
  3. Out-of-band SSM config — NOT created by CDK. Operationally-variable or env-specific-unknown values listed in OUT_OF_BAND_SSM and populated by put-config.sh. The keystone is DEFAULT_SANDBOX_SNAPSHOT_ID (changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls.

Kebab-Aspect deviation

KebabNamingAspect exempts AWS::SecretsManager::Secret and AWS::SSM::Parameter from the kebab check (see the KEBAB_EXEMPT_RESOURCE_TYPES set) — the UPPER_SNAKE env-var segment is a required, documented deviation for a lossless store→env round-trip. Every other explicitly-named resource is still validated. Covered by a dedicated case in test/kebab-naming-aspect.test.ts.

Deploy ordering (values BEFORE the box boots)

The shells are synth-able now (T11). Population is out-of-band and happens after cdk deploy open-swe-<env> but before the EC2/T12 box first boots:

cdk deploy open-swe-<env>                  # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod>   # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start

put-config.sh ships <FILL> placeholders only (no real secret values committed); provide each value inline, via OPENSWE_PUT_<VAR> env vars, or from a vault. It does NOT touch the IaC-managed params (CDK owns those — editing them here would drift).

Compute + ingress (AppService — T12)

AppService (lib/constructs/app-service.ts, one per env from OpenSweStack) builds the box and its path to the internet. A single internet-facing ALB (app/seahaven-com) and a single VPC are shared with the on-prem seahaven-site stack, so open-swe imports the VPC, the ALB security group (sg-0b0301deed193258a), the :443 listener, and the public seahaven.com zone — and never owns/mutates them. It adds:

  • One ARM64 EC2 box (open-swe-<env>-box, t4g.medium dev / t4g.large prod) in private1 (us-east-1a) — same AZ as the single NAT for in-AZ egress. requireImdsv2, gp3 encrypted root, deleteOnTermination (no RETAIN volume — see below). userDataCausesReplacement: true; user-data is rendered from deploy/ami/user-data.sh.
  • A standalone instance SG reachable only from the shared ALB SG on :80 (nginx). Egress open (NAT). The ALB SG is opened to the box via a standalone CfnSecurityGroupEgress so the imported (on-prem-owned) SG is never mutated.
  • A target group → instance :80 (nginx is the sole ingress; the LangGraph control plane stays on loopback :2024). Health check GET /healthz.
  • Two listener rules on the imported :443 listener, both → the same TG:
    • Webhooks (priority 2 dev / 3 prod): host ∈ {openswe-<env>, hooks-<env>}.seahaven.com AND path /webhooks/*.
    • Site (priority 10 dev / 11 prod): host = openswe-<env>.seahaven.com (dashboard SPA + /dashboard/api/).
  • Route53 alias records openswe[-dev] + hooks[-dev] → the shared ALB.
  • Four CloudWatch log groups (/open-swe/<env>/{app,user-data,nginx-access,nginx-error}) at 30-day retention (IaC-owned; mirrors the CW-agent config).

Listener-rule ordering (load-bearing)

The shared listener already has a host-agnostic /webhooks/* PATH rule at priority 5 (on-prem). ALB rules are first-match by ascending priority, so the open-swe webhook rule must sit below 5 or every …/webhooks/* request (any host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs a host condition, so it does not steal the on-prem hosts' webhooks. The dashboard "site" rule carries no path that collides with rule 5, so it sits at 10/11.

Cross-stack coordination (T13 review). The seahaven-site (on-prem) and open-swe stacks both add resources to the same imported listener and ALB SG. This is safe: each stack owns only the resources it declares (its own logical ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa, and the standalone CfnSecurityGroupEgress never mutates the shared SG's own definition (the pattern on-prem itself uses). The one shared namespace that needs care is listener-rule priority (globally unique per listener; a collision is a fail-safe deploy error, not silent drift). Ownership — keep disjoint: seahaven-site = 4-7 + default; open-swe = 2, 3, 10, 11. open-swe's webhook rules are host-scoped to its own *.seahaven.com hosts, so they never match an on-prem seahavenind.com host.

Security review (T5/T12 /sh-security-review)

The T12 surface was run through the detector-fan-out + proof-or-kill verifier. One confirmed medium (OSWE-T12-01: nginx's 1 MB default client_max_body_size would 413 large GitHub webhooks before in-app signature verification) is fixed in open-swe.nginx.conf (25m on /webhooks/, 10m on /dashboard/api/). The hooks hostname is scoped to /webhooks/* only (OSWE-T12-02 hygiene). An X-Forwarded-For spoof candidate was killed — no code trusts the leftmost XFF. No confirmed critical/high; no block.

AMI cache discipline (EBS-fix plumbing — consumed by T12 AppService)

cachedArm64AmazonLinux2023() (in lib/constructs/ami-cache.ts) returns an ARM64 Amazon Linux 2023 image with cachedInContext: true, so the resolved AMI id is pinned in the committed cdk.context.json. Without the pin, every deploy could pick up a newer AL2023 release → AMI change → EC2 instance replacement (the file-share data-loss root cause — memory feedback_inline_ebs_volumes).

Design intent (now consumed by AppService):

  • userDataCausesReplacement: true is the deliberate choice — user-data is provisioning-only and carries no durable state.
  • No durable state on the box → no RETAIN volume. The in-memory langgraph store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is intentionally no standalone ec2.Volume + removalPolicy.RETAIN. The goal is replacement-tolerance, not avoidance.
  • Snapshot-before-replace still applies operationally: before any replacing deploy snapshot the root volume and wait state=completed, and re-verify "no local-only durable state" first.

Refresh the AMI pin deliberately:

cdk context --reset 'ssm:account=328440206208:parameterName=/aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-arm64:region=us-east-1'
cdk synth   # review the diff — it WILL show "requires replacement"

The committed cdk.context.json ships a dummy-but-valid-shaped AMI id (ami-00000000000000000) so cdk synth resolves the cache locally without any live AWS call. Before the first real deploy, repoint AppService to the baked open-swe-base-arm64 AMI and pin its real id — the placeholder is intentionally un-bootable on stock AL2023.

Commands

npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test            # jest — naming Aspect

Deploy ordering (when the gate clears — NOT yet)

  1. open-swe-iam first — create githubdeploy-open-swe-infra + set the repo AWS_DEPLOY_ROLE_ARN secret before any infra/secrets CI step (BLOCK#3).
  2. Security gate — T4 GPT-4.1 IAM cross-review + T5 /sh-security-review on the synth; resolve every confirmed critical/high.
  3. IAM applied (T6) — only after the gate.
  4. Env stacks (open-swe-dev, then open-swe-prod) build out at T12+, prod gated by a GitHub Environment manual approval.

Version policy

aws-cdk-lib is pinned EXACT (2.260.0) — no ^/~. Dependabot keeps it current; CI (npm ci + cdk synth) + dependency review gate each bump. See aws-infrastructure.md "CDK Version Policy" and memory feedback_cdk_lib_bundled_deps.