open-swe/infra/README.md
Adam Moussa e9499e49b8
fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1/OSWE-IAC-01) (#55)
* fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1)

Dev synthesizes against the oswedev qualifier and the dev infra deploy role
is scoped to cdk-oswedev-* — it can no longer assume the default hnb659fds
bootstrap roles whose admin cfn-exec-role deploys prod, closing the cross-env
escalation (OSWE-IAC-01). Prod stays on the default qualifier.

* test: assert per-env bootstrap qualifier isolation + document (B-1)
2026-06-29 12:39:23 -04:00

19 KiB
Raw Blame History

open-swe infra (CDK TypeScript)

AWS infrastructure for the Open SWE → AWS migration. Synth-only at this stage — nothing here is deployed yet. All IAM is applied only after the Phase-1 security gate (T4 GPT-4.1 IAM cross-review + T5 /sh-security-review) clears (T6).

Layout

infra/
├── bin/
│   └── app.ts                       # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│   ├── config.ts                    # account/region/org constants, env type, OIDC trust subjects
│   ├── open-swe-iam-stack.ts        # account-level: shared OIDC deploy roles
│   ├── open-swe-stack.ts            # per-env stack (instance role + config store + AppService)
│   ├── aspects/
│   │   └── kebab-naming-aspect.ts   # fails synth on any non-kebab-case explicit name
│   └── constructs/
│       ├── github-deploy-roles.ts   # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│       ├── instance-role.ts         # open-swe-<env>-instance-role (least-privilege)
│       ├── config-store.ts          # Secrets Manager + SSM Parameter Store shells (T11)
│       ├── app-service.ts           # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│       └── ami-cache.ts             # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│   └── kebab-naming-aspect.test.ts  # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json                 # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json                     # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore

Stacks

Stack name (kebab) Construct Contents
open-swe-iam OpenSweIamStack Account-level shared GitHub OIDC deploy roles (singletons).
open-swe-dev OpenSweStack (envName: dev) open-swe-dev-instance-role, config store (T11), and the EC2 box + ALB ingress (T12, AppService).
open-swe-prod OpenSweStack (envName: prod) open-swe-prod-instance-role, config store, and the EC2 box + ALB ingress.

Account 328440206208, region us-east-1. Stack names are set explicitly so CDK never defaults to PascalCase; resource names follow open-swe-<env>-*.

The two env stacks (open-swe-dev / open-swe-prod) are the required pair. The shared OIDC deploy roles are account-wide singletons (one RoleName each), so they live in their own dedicated open-swe-iam stack rather than being duplicated across the env stacks — and that stack deploys first (see ordering).

IAM roles defined (unapplied)

  • githubdeploy-open-swe-infra — GitHub OIDC role for CDK/CFN infra deploys. Trust scoped to repo:Sea-Haven-Industries/open-swe on the main/dev branches only. Permission is the org-standard CDK pattern: sts:AssumeRole on the CDK bootstrap roles (cdk-hnb659fds-*) — the real CFN/IAM/resource scope lives in the bootstrap cfn-exec-role, not in this role.
  • githubdeploy-open-swe-app — GitHub OIDC role for app deploys. Tag-scoped ssm:SendCommand (instances tagged project=open-swe + env in {dev,prod}) + read-only access to the open-swe-<env>-assets S3 artifact buckets.
  • open-swe-<env>-instance-role — EC2 instance role, least-privilege: read open-swe-<env>-assets (S3), read /open-swe-<env>/* (SSM), read open-swe-<env>/* (Secrets Manager), put /open-swe/<env>/* CloudWatch Logs, plus AmazonSSMManagedInstanceCore for SSM agent registration. No admin.

The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created.

CDK bootstrap qualifiers — per-env deploy isolation (B-1 / OSWE-IAC-01)

Each env's infra deploy role may assume only its own bootstrap qualifier's roles, so a dev-branch token can never assume the bootstrap roles whose admin cfn-exec-role deploys prod (closing the cross-env escalation that bypassed prod's Environment approval gate). Mapping lives in config.ts:bootstrapQualifier:

Env Qualifier Toolkit stack Infra role assumes
dev oswedev CDKToolkit-oswedev cdk-oswedev-*
prod hnb659fds (default) CDKToolkit cdk-hnb659fds-*

The dev stack synthesizes with DefaultStackSynthesizer({ qualifier: "oswedev" }) (bin/app.ts); prod uses the default. Bootstrap a new env qualifier with:

npx cdk bootstrap --qualifier <qual> --toolkit-stack-name CDKToolkit-<qual> \
  --cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess \
  aws://328440206208/us-east-1

Deploy order matters when changing an env's qualifier: bootstrap the new qualifier and deploy the env stack onto it before re-scoping that env's infra role in open-swe-iam — otherwise a pipeline deploy with the re-scoped role would fail to assume the not-yet-targeted bootstrap roles.

Kebab-case naming Aspect

KebabNamingAspect (applied app-wide in bin/app.ts) fails synth via Annotations.addError when a stack name or an explicit physical resource name (RoleName, BucketName, …) is not kebab-case. Path-style names (Secrets Manager a/b, SSM /a/b, log groups /aws/.../x) are validated per /-segment. CDK logical construct ids are intentionally NOT validated (they are conventionally PascalCase). Covered by test/kebab-naming-aspect.test.ts.

Config store (Secrets Manager + SSM shells — T11)

ConfigStore (lib/constructs/config-store.ts, one per env from OpenSweStack) renders the resource shells the boot hook deploy/seahaven/fetch-config.sh reads. The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var names as the last path segment — open-swe-<env>/<VAR> for secrets, /open-swe-<env>/<VAR> (FLAT) for config — because fetch-config strips the prefix and exports that segment verbatim.

Three buckets:

  1. Secret shells (Secrets Manager) — 27 secrets. Created value-LESS (an L1 CfnSecret with NEITHER secretString NOR generateSecretString, which CloudFormation creates as an empty secret). The real value is set out-of-band (put-config.sh) — CDK never owns it, so a later cdk deploy can never clobber it. UpdateReplacePolicy/DeletionPolicy: Retain so a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omits kms:Decrypt).

    The T9 header says "29 secrets" but its table enumerates 27 distinct VAR names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. Confirm the 27-vs-29 count (code-only candidates not in the table: USER_ID_API_KEY_MAP, JUDGE_ANTHROPIC_BASE_URL).

    ⚠️ Retain + fixed name orphans these shells on a failed FIRST create. If the stack's initial create fails and rolls back, Retain keeps the shells instead of deleting them. The stack is then gone, but the secrets survive, still holding the global open-swe-<env>/<VAR> names — so every later create fails with AlreadyExists. A plain delete-secret does not clear it (the name stays reserved for the 7–30 day recovery window). This bites on a teardown/rebuild, a secret logical-id change/refactor, or standing up a new env — never on routine updates of an already-created stack. Recovery — before re-creating the stack, force-delete the empty orphans so the names free immediately:

    aws secretsmanager list-secrets --region us-east-1 \
      --filters Key=name,Values=open-swe-<env>/ \
      --query 'SecretList[].Name' --output text | tr '\t' '\n' | while read -r n; do
        aws secretsmanager delete-secret --secret-id "$n" \
          --region us-east-1 --force-delete-without-recovery
      done
    

    Force-delete only shells with no value version — a populated secret holds real operator material. (Hit on prod 2026-06-29; see PR #51's deploy failure.)

  2. IaC-managed SSM config — 8 params, real values owned in code:

    Param dev prod
    SANDBOX_TYPE langsmith langsmith
    DEFAULT_REPO_OWNER Sea-Haven-Industries Sea-Haven-Industries
    ALLOWED_GITHUB_ORGS Sea-Haven-Industries Sea-Haven-Industries
    DEFAULT_REPO_NAME open-swe-pilot (confirm) open-swe-pilot (confirm)
    DASHBOARD_BASE_URL https://openswe-dev.seahaven.com (confirm host) https://openswe.seahaven.com (confirm host)
    DASHBOARD_API_BASE_URL same as base same as base
    DASHBOARD_ALLOWED_ORIGINS same as base same as base
    LLM_MODEL_ID anthropic:claude-opus-4-8 (confirm) anthropic:claude-opus-4-8 (confirm)
  3. Out-of-band SSM config — NOT created by CDK. Operationally-variable or env-specific-unknown values listed in OUT_OF_BAND_SSM and populated by put-config.sh. The keystone is DEFAULT_SANDBOX_SNAPSHOT_ID (changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls.

Kebab-Aspect deviation

KebabNamingAspect exempts AWS::SecretsManager::Secret and AWS::SSM::Parameter from the kebab check (see the KEBAB_EXEMPT_RESOURCE_TYPES set) — the UPPER_SNAKE env-var segment is a required, documented deviation for a lossless store→env round-trip. Every other explicitly-named resource is still validated. Covered by a dedicated case in test/kebab-naming-aspect.test.ts.

Deploy ordering (values BEFORE the box boots)

The shells are synth-able now (T11). Population is out-of-band and happens after cdk deploy open-swe-<env> but before the EC2/T12 box first boots:

cdk deploy open-swe-<env>                  # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod>   # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start

put-config.sh ships <FILL> placeholders only (no real secret values committed); provide each value inline, via OPENSWE_PUT_<VAR> env vars, or from a vault. It does NOT touch the IaC-managed params (CDK owns those — editing them here would drift).

Compute + ingress (AppService — T12)

AppService (lib/constructs/app-service.ts, one per env from OpenSweStack) builds the box and its path to the internet. A single internet-facing ALB (app/seahaven-com) and a single VPC are shared with the on-prem seahaven-site stack, so open-swe imports the VPC, the ALB security group (sg-0b0301deed193258a), the :443 listener, and the public seahaven.com zone — and never owns/mutates them. It adds:

  • One ARM64 EC2 box (open-swe-<env>-box, t4g.medium dev / t4g.large prod) in private1 (us-east-1a) — same AZ as the single NAT for in-AZ egress. requireImdsv2, gp3 encrypted root, deleteOnTermination (no RETAIN volume — see below). userDataCausesReplacement: true; user-data is rendered from deploy/ami/user-data.sh.
  • A standalone instance SG reachable only from the shared ALB SG on :80 (nginx). Egress open (NAT). The ALB SG is opened to the box via a standalone CfnSecurityGroupEgress so the imported (on-prem-owned) SG is never mutated.
  • A target group → instance :80 (nginx is the sole ingress; the LangGraph control plane stays on loopback :2024). Health check GET /healthz.
  • Two listener rules on the imported :443 listener, both → the same TG:
    • Webhooks (priority 2 dev / 3 prod): host ∈ {openswe-<env>, hooks-<env>}.seahaven.com AND path /webhooks/*.
    • Site (priority 10 dev / 11 prod): host = openswe-<env>.seahaven.com (dashboard SPA + /dashboard/api/).
  • Route53 alias records openswe[-dev] + hooks[-dev] → the shared ALB.
  • Four CloudWatch log groups (/open-swe/<env>/{app,user-data,nginx-access,nginx-error}) at 30-day retention (IaC-owned; mirrors the CW-agent config).

Listener-rule ordering (load-bearing)

The shared listener already has a host-agnostic /webhooks/* PATH rule at priority 5 (on-prem). ALB rules are first-match by ascending priority, so the open-swe webhook rule must sit below 5 or every …/webhooks/* request (any host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs a host condition, so it does not steal the on-prem hosts' webhooks. The dashboard "site" rule carries no path that collides with rule 5, so it sits at 10/11.

Cross-stack coordination (T13 review). The seahaven-site (on-prem) and open-swe stacks both add resources to the same imported listener and ALB SG. This is safe: each stack owns only the resources it declares (its own logical ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa, and the standalone CfnSecurityGroupEgress never mutates the shared SG's own definition (the pattern on-prem itself uses). The one shared namespace that needs care is listener-rule priority (globally unique per listener; a collision is a fail-safe deploy error, not silent drift). Ownership — keep disjoint: seahaven-site = 4-7 + default; open-swe = 2, 3, 10, 11. open-swe's webhook rules are host-scoped to its own *.seahaven.com hosts, so they never match an on-prem seahavenind.com host.

Security review (T5/T12 /sh-security-review)

The T12 surface was run through the detector-fan-out + proof-or-kill verifier. One confirmed medium (OSWE-T12-01: nginx's 1 MB default client_max_body_size would 413 large GitHub webhooks before in-app signature verification) is fixed in open-swe.nginx.conf (25m on /webhooks/, 10m on /dashboard/api/). The hooks hostname is scoped to /webhooks/* only (OSWE-T12-02 hygiene). An X-Forwarded-For spoof candidate was killed — no code trusts the leftmost XFF. No confirmed critical/high; no block.

Baked AMI + EBS-replacement discipline

bakedOpenSweArm64() (in lib/constructs/ami-cache.ts) pins the custom open-swe-base-arm64 image by EXACT id (BAKED_OPEN_SWE_AMI_ID) via MachineImage.genericLinux({ "us-east-1": "<ami-id>" }) — no SSM lookup, so synth and deploy are fully offline/deterministic. The image is built by deploy/ami/open-swe-base.pkr.hcl (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW agent + boot templates, no secrets); the box's user-data.sh assumes that baked layout (/opt/open-swe, openswe user, nginx, CW agent).

Pinning by exact id (vs a most_recent name filter) is what prevents a routine deploy from silently swapping the AMI → EC2 instance replacement (the file-share data-loss root cause — memory feedback_inline_ebs_volumes).

  • userDataCausesReplacement: true is deliberate — user-data is provisioning-only and carries no durable state.
  • No durable state on the box → no RETAIN volume. The in-memory langgraph store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is intentionally no standalone ec2.Volume + removalPolicy.RETAIN. The goal is replacement-tolerance, not avoidance.
  • Snapshot-before-replace still applies operationally: before any replacing deploy snapshot the root volume and wait state=completed, and re-verify "no local-only durable state" first.

Refresh the AMI deliberately:

cd deploy/ami && packer build open-swe-base.pkr.hcl   # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack               # WILL show "requires replacement"

cdk.context.json is {} — nothing is resolved via context anymore (the AMI is a static id pin), so synth makes no live AWS call.

Commands

npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test            # jest — naming Aspect

CI/CD (T18 — .github/workflows/ci-infra.yml + cd-infra.yml)

Path-filtered, OIDC-only (no static keys). The Python agent keeps its own ci.yml ("CI"); these two add the /infra half.

Workflow Trigger Does
ci-infra.yml PR touching infra/** tsc + jest + cdk synth (reusable ci-typescript-cdk.yaml).
cd-infra.yml push to dev/main touching infra/**, or dispatch CI (pre-deploy) → per-env cdk deploy.

cd-infra.yml flow:

  • push to dev → CI green → auto cdk deploy OpenSweDevStack (assumes githubdeploy-open-swe-infra-dev; the job declares no environment:, so the OIDC subject is …:ref:refs/heads/dev — matching that role's trust).
  • push to main → CI green → cdk deploy OpenSweProdStack behind the prod GitHub Environment (required reviewer = Adam). The environment: prod declaration both fires the manual-approval gate and makes the OIDC subject …:environment:prod — matching githubdeploy-open-swe-infra-prod's trust.

Why not the reusable cd-cdk.yaml: it runs cdk deploy --all, which from a single-env push would deploy the other env + the shared IAM stack — breaking the per-env boundary. So CD targets one stack explicitly per env. The shared open-swe-iam stack is not deployed by CD (privileged, human-gated — T6).

Gating note: infra CI is enforced at the deploy boundary (cd-infra's deploy-* jobs needs: ci), not as a branch-protection required check — path-filtering a required check would deadlock app-only PRs (a skipped required check never satisfies). Making Infra CI a required check later needs a skip-aware shim or dropping its path filter.

Prerequisites (set post-T6, when the roles exist):

  • repo variables AWS_DEPLOY_ROLE_INFRA_DEV / AWS_DEPLOY_ROLE_INFRA_PROD = the githubdeploy-open-swe-infra-<env> role ARNs (open-swe-iam outputs).
  • a GitHub Environment named prod with Adam as a required reviewer.

App-side CD (CI → S3 artifact → SSM deploy via githubdeploy-open-swe-app-<env>) is T19, not here.

Deploy ordering (when the gate clears — NOT yet)

  1. open-swe-iam first — apply the IAM stack (T6, human-gated), then set the repo AWS_DEPLOY_ROLE_INFRA_{DEV,PROD} variables from its role-ARN outputs and configure the prod Environment reviewer (BLOCK#3).
  2. Security gate — T4 GPT-4.1 IAM cross-review + T5 /sh-security-review on the synth; resolve every confirmed critical/high.
  3. IAM applied (T6) — only after the gate.
  4. Env stacks: first open-swe-dev (T14, manual validate), then CD auto-deploys dev on push; open-swe-prod (T21) behind the prod Environment approval.

Version policy

aws-cdk-lib is pinned EXACT (2.260.0) — no ^/~. Dependabot keeps it current; CI (npm ci + cdk synth) + dependency review gate each bump. See aws-infrastructure.md "CDK Version Policy" and memory feedback_cdk_lib_bundled_deps.