open-swe/infra
Adam Moussa faae9a685b
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48)
Switch fetch-config.sh from a name-prefix batch-get-secret-value
--filters scan to an explicit --secret-id-list (the 28 SECRET_VARS,
chunked at the 20/call cap). An id-list batch authorizes per-secret
ARN, so the instance role's BatchGetSecretValue moves from Resource:*
to the open-swe-<env>/* prefix and the account-wide ListSecrets grant
is dropped entirely. The box can no longer enumerate secret names
account-wide; cross-env value isolation is unchanged (GetSecretValue
was already prefix-scoped). Resolves OSWE-IAC-SECRETS-LIST-01.

Also capture each chunk response into a variable and consume the
producer via command substitution so a failed AWS call aborts under
set -e instead of being swallowed by process substitution and
misreported as a missing required var.

Refs: OSWE-IAC-SECRETS-LIST-01
2026-06-28 20:45:17 -04:00
..
bin feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
lib Scope secrets fetch to --secret-id-list; drop ListSecrets grant (#48) 2026-06-28 20:45:17 -04:00
test fix(deploy): use %%...%% for CDK user-data tokens (don't collide with @@ sed) (#21) 2026-06-26 19:06:37 -04:00
.gitignore ci: path-filtered infra CI/CD + dual OIDC roles + prod approval gate (T18) (#15) 2026-06-26 16:11:22 -04:00
cdk.context.json feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) 2026-06-26 18:49:09 -04:00
cdk.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
jest.config.js ci: path-filtered infra CI/CD + dual OIDC roles + prod approval gate (T18) (#15) 2026-06-26 16:11:22 -04:00
package-lock.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
package.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00
README.md ci: align workflows with Sea Haven CI/CD handbook (#29) 2026-06-27 21:47:25 -04:00
tsconfig.json feat(infra): add /infra CDK scaffold + per-env OIDC/instance IAM role defs (#6) 2026-06-26 15:06:30 -04:00

open-swe infra (CDK TypeScript)

AWS infrastructure for the Open SWE → AWS migration. Synth-only at this stage — nothing here is deployed yet. All IAM is applied only after the Phase-1 security gate (T4 GPT-4.1 IAM cross-review + T5 /sh-security-review) clears (T6).

Layout

infra/
├── bin/
│   └── app.ts                       # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│   ├── config.ts                    # account/region/org constants, env type, OIDC trust subjects
│   ├── open-swe-iam-stack.ts        # account-level: shared OIDC deploy roles
│   ├── open-swe-stack.ts            # per-env stack (instance role + config store + AppService)
│   ├── aspects/
│   │   └── kebab-naming-aspect.ts   # fails synth on any non-kebab-case explicit name
│   └── constructs/
│       ├── github-deploy-roles.ts   # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│       ├── instance-role.ts         # open-swe-<env>-instance-role (least-privilege)
│       ├── config-store.ts          # Secrets Manager + SSM Parameter Store shells (T11)
│       ├── app-service.ts           # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│       └── ami-cache.ts             # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│   └── kebab-naming-aspect.test.ts  # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json                 # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json                     # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore

Stacks

Stack name (kebab) Construct Contents
open-swe-iam OpenSweIamStack Account-level shared GitHub OIDC deploy roles (singletons).
open-swe-dev OpenSweStack (envName: dev) open-swe-dev-instance-role, config store (T11), and the EC2 box + ALB ingress (T12, AppService).
open-swe-prod OpenSweStack (envName: prod) open-swe-prod-instance-role, config store, and the EC2 box + ALB ingress.

Account 328440206208, region us-east-1. Stack names are set explicitly so CDK never defaults to PascalCase; resource names follow open-swe-<env>-*.

The two env stacks (open-swe-dev / open-swe-prod) are the required pair. The shared OIDC deploy roles are account-wide singletons (one RoleName each), so they live in their own dedicated open-swe-iam stack rather than being duplicated across the env stacks — and that stack deploys first (see ordering).

IAM roles defined (unapplied)

  • githubdeploy-open-swe-infra — GitHub OIDC role for CDK/CFN infra deploys. Trust scoped to repo:Sea-Haven-Industries/open-swe on the main/dev branches only. Permission is the org-standard CDK pattern: sts:AssumeRole on the CDK bootstrap roles (cdk-hnb659fds-*) — the real CFN/IAM/resource scope lives in the bootstrap cfn-exec-role, not in this role.
  • githubdeploy-open-swe-app — GitHub OIDC role for app deploys. Tag-scoped ssm:SendCommand (instances tagged project=open-swe + env in {dev,prod}) + read-only access to the open-swe-<env>-assets S3 artifact buckets.
  • open-swe-<env>-instance-role — EC2 instance role, least-privilege: read open-swe-<env>-assets (S3), read /open-swe-<env>/* (SSM), read open-swe-<env>/* (Secrets Manager), put /open-swe/<env>/* CloudWatch Logs, plus AmazonSSMManagedInstanceCore for SSM agent registration. No admin.

The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created.

Kebab-case naming Aspect

KebabNamingAspect (applied app-wide in bin/app.ts) fails synth via Annotations.addError when a stack name or an explicit physical resource name (RoleName, BucketName, …) is not kebab-case. Path-style names (Secrets Manager a/b, SSM /a/b, log groups /aws/.../x) are validated per /-segment. CDK logical construct ids are intentionally NOT validated (they are conventionally PascalCase). Covered by test/kebab-naming-aspect.test.ts.

Config store (Secrets Manager + SSM shells — T11)

ConfigStore (lib/constructs/config-store.ts, one per env from OpenSweStack) renders the resource shells the boot hook deploy/seahaven/fetch-config.sh reads. The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var names as the last path segment — open-swe-<env>/<VAR> for secrets, /open-swe-<env>/<VAR> (FLAT) for config — because fetch-config strips the prefix and exports that segment verbatim.

Three buckets:

  1. Secret shells (Secrets Manager) — 27 secrets. Created value-LESS (an L1 CfnSecret with NEITHER secretString NOR generateSecretString, which CloudFormation creates as an empty secret). The real value is set out-of-band (put-config.sh) — CDK never owns it, so a later cdk deploy can never clobber it. UpdateReplacePolicy/DeletionPolicy: Retain so a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omits kms:Decrypt).

    The T9 header says "29 secrets" but its table enumerates 27 distinct VAR names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. Confirm the 27-vs-29 count (code-only candidates not in the table: USER_ID_API_KEY_MAP, JUDGE_ANTHROPIC_BASE_URL).

  2. IaC-managed SSM config — 8 params, real values owned in code:

    Param dev prod
    SANDBOX_TYPE langsmith langsmith
    DEFAULT_REPO_OWNER Sea-Haven-Industries Sea-Haven-Industries
    ALLOWED_GITHUB_ORGS Sea-Haven-Industries Sea-Haven-Industries
    DEFAULT_REPO_NAME open-swe-pilot (confirm) open-swe-pilot (confirm)
    DASHBOARD_BASE_URL https://openswe-dev.seahaven.com (confirm host) https://openswe.seahaven.com (confirm host)
    DASHBOARD_API_BASE_URL same as base same as base
    DASHBOARD_ALLOWED_ORIGINS same as base same as base
    LLM_MODEL_ID anthropic:claude-opus-4-8 (confirm) anthropic:claude-opus-4-8 (confirm)
  3. Out-of-band SSM config — NOT created by CDK. Operationally-variable or env-specific-unknown values listed in OUT_OF_BAND_SSM and populated by put-config.sh. The keystone is DEFAULT_SANDBOX_SNAPSHOT_ID (changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls.

Kebab-Aspect deviation

KebabNamingAspect exempts AWS::SecretsManager::Secret and AWS::SSM::Parameter from the kebab check (see the KEBAB_EXEMPT_RESOURCE_TYPES set) — the UPPER_SNAKE env-var segment is a required, documented deviation for a lossless store→env round-trip. Every other explicitly-named resource is still validated. Covered by a dedicated case in test/kebab-naming-aspect.test.ts.

Deploy ordering (values BEFORE the box boots)

The shells are synth-able now (T11). Population is out-of-band and happens after cdk deploy open-swe-<env> but before the EC2/T12 box first boots:

cdk deploy open-swe-<env>                  # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod>   # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start

put-config.sh ships <FILL> placeholders only (no real secret values committed); provide each value inline, via OPENSWE_PUT_<VAR> env vars, or from a vault. It does NOT touch the IaC-managed params (CDK owns those — editing them here would drift).

Compute + ingress (AppService — T12)

AppService (lib/constructs/app-service.ts, one per env from OpenSweStack) builds the box and its path to the internet. A single internet-facing ALB (app/seahaven-com) and a single VPC are shared with the on-prem seahaven-site stack, so open-swe imports the VPC, the ALB security group (sg-0b0301deed193258a), the :443 listener, and the public seahaven.com zone — and never owns/mutates them. It adds:

  • One ARM64 EC2 box (open-swe-<env>-box, t4g.medium dev / t4g.large prod) in private1 (us-east-1a) — same AZ as the single NAT for in-AZ egress. requireImdsv2, gp3 encrypted root, deleteOnTermination (no RETAIN volume — see below). userDataCausesReplacement: true; user-data is rendered from deploy/ami/user-data.sh.
  • A standalone instance SG reachable only from the shared ALB SG on :80 (nginx). Egress open (NAT). The ALB SG is opened to the box via a standalone CfnSecurityGroupEgress so the imported (on-prem-owned) SG is never mutated.
  • A target group → instance :80 (nginx is the sole ingress; the LangGraph control plane stays on loopback :2024). Health check GET /healthz.
  • Two listener rules on the imported :443 listener, both → the same TG:
    • Webhooks (priority 2 dev / 3 prod): host ∈ {openswe-<env>, hooks-<env>}.seahaven.com AND path /webhooks/*.
    • Site (priority 10 dev / 11 prod): host = openswe-<env>.seahaven.com (dashboard SPA + /dashboard/api/).
  • Route53 alias records openswe[-dev] + hooks[-dev] → the shared ALB.
  • Four CloudWatch log groups (/open-swe/<env>/{app,user-data,nginx-access,nginx-error}) at 30-day retention (IaC-owned; mirrors the CW-agent config).

Listener-rule ordering (load-bearing)

The shared listener already has a host-agnostic /webhooks/* PATH rule at priority 5 (on-prem). ALB rules are first-match by ascending priority, so the open-swe webhook rule must sit below 5 or every …/webhooks/* request (any host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs a host condition, so it does not steal the on-prem hosts' webhooks. The dashboard "site" rule carries no path that collides with rule 5, so it sits at 10/11.

Cross-stack coordination (T13 review). The seahaven-site (on-prem) and open-swe stacks both add resources to the same imported listener and ALB SG. This is safe: each stack owns only the resources it declares (its own logical ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa, and the standalone CfnSecurityGroupEgress never mutates the shared SG's own definition (the pattern on-prem itself uses). The one shared namespace that needs care is listener-rule priority (globally unique per listener; a collision is a fail-safe deploy error, not silent drift). Ownership — keep disjoint: seahaven-site = 4-7 + default; open-swe = 2, 3, 10, 11. open-swe's webhook rules are host-scoped to its own *.seahaven.com hosts, so they never match an on-prem seahavenind.com host.

Security review (T5/T12 /sh-security-review)

The T12 surface was run through the detector-fan-out + proof-or-kill verifier. One confirmed medium (OSWE-T12-01: nginx's 1 MB default client_max_body_size would 413 large GitHub webhooks before in-app signature verification) is fixed in open-swe.nginx.conf (25m on /webhooks/, 10m on /dashboard/api/). The hooks hostname is scoped to /webhooks/* only (OSWE-T12-02 hygiene). An X-Forwarded-For spoof candidate was killed — no code trusts the leftmost XFF. No confirmed critical/high; no block.

Baked AMI + EBS-replacement discipline

bakedOpenSweArm64() (in lib/constructs/ami-cache.ts) pins the custom open-swe-base-arm64 image by EXACT id (BAKED_OPEN_SWE_AMI_ID) via MachineImage.genericLinux({ "us-east-1": "<ami-id>" }) — no SSM lookup, so synth and deploy are fully offline/deterministic. The image is built by deploy/ami/open-swe-base.pkr.hcl (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW agent + boot templates, no secrets); the box's user-data.sh assumes that baked layout (/opt/open-swe, openswe user, nginx, CW agent).

Pinning by exact id (vs a most_recent name filter) is what prevents a routine deploy from silently swapping the AMI → EC2 instance replacement (the file-share data-loss root cause — memory feedback_inline_ebs_volumes).

  • userDataCausesReplacement: true is deliberate — user-data is provisioning-only and carries no durable state.
  • No durable state on the box → no RETAIN volume. The in-memory langgraph store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is intentionally no standalone ec2.Volume + removalPolicy.RETAIN. The goal is replacement-tolerance, not avoidance.
  • Snapshot-before-replace still applies operationally: before any replacing deploy snapshot the root volume and wait state=completed, and re-verify "no local-only durable state" first.

Refresh the AMI deliberately:

cd deploy/ami && packer build open-swe-base.pkr.hcl   # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack               # WILL show "requires replacement"

cdk.context.json is {} — nothing is resolved via context anymore (the AMI is a static id pin), so synth makes no live AWS call.

Commands

npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test            # jest — naming Aspect

CI/CD (T18 — .github/workflows/ci-infra.yml + cd-infra.yml)

Path-filtered, OIDC-only (no static keys). The Python agent keeps its own ci.yml ("CI"); these two add the /infra half.

Workflow Trigger Does
ci-infra.yml PR touching infra/** tsc + jest + cdk synth (reusable ci-typescript-cdk.yaml).
cd-infra.yml push to dev/main touching infra/**, or dispatch CI (pre-deploy) → per-env cdk deploy.

cd-infra.yml flow:

  • push to dev → CI green → auto cdk deploy OpenSweDevStack (assumes githubdeploy-open-swe-infra-dev; the job declares no environment:, so the OIDC subject is …:ref:refs/heads/dev — matching that role's trust).
  • push to main → CI green → cdk deploy OpenSweProdStack behind the prod GitHub Environment (required reviewer = Adam). The environment: prod declaration both fires the manual-approval gate and makes the OIDC subject …:environment:prod — matching githubdeploy-open-swe-infra-prod's trust.

Why not the reusable cd-cdk.yaml: it runs cdk deploy --all, which from a single-env push would deploy the other env + the shared IAM stack — breaking the per-env boundary. So CD targets one stack explicitly per env. The shared open-swe-iam stack is not deployed by CD (privileged, human-gated — T6).

Gating note: infra CI is enforced at the deploy boundary (cd-infra's deploy-* jobs needs: ci), not as a branch-protection required check — path-filtering a required check would deadlock app-only PRs (a skipped required check never satisfies). Making Infra CI a required check later needs a skip-aware shim or dropping its path filter.

Prerequisites (set post-T6, when the roles exist):

  • repo variables AWS_DEPLOY_ROLE_INFRA_DEV / AWS_DEPLOY_ROLE_INFRA_PROD = the githubdeploy-open-swe-infra-<env> role ARNs (open-swe-iam outputs).
  • a GitHub Environment named prod with Adam as a required reviewer.

App-side CD (CI → S3 artifact → SSM deploy via githubdeploy-open-swe-app-<env>) is T19, not here.

Deploy ordering (when the gate clears — NOT yet)

  1. open-swe-iam first — apply the IAM stack (T6, human-gated), then set the repo AWS_DEPLOY_ROLE_INFRA_{DEV,PROD} variables from its role-ARN outputs and configure the prod Environment reviewer (BLOCK#3).
  2. Security gate — T4 GPT-4.1 IAM cross-review + T5 /sh-security-review on the synth; resolve every confirmed critical/high.
  3. IAM applied (T6) — only after the gate.
  4. Env stacks: first open-swe-dev (T14, manual validate), then CD auto-deploys dev on push; open-swe-prod (T21) behind the prod Environment approval.

Version policy

aws-cdk-lib is pinned EXACT (2.260.0) — no ^/~. Dependabot keeps it current; CI (npm ci + cdk synth) + dependency review gate each bump. See aws-infrastructure.md "CDK Version Policy" and memory feedback_cdk_lib_bundled_deps.