* fix: isolate dev CDK deploys on their own bootstrap qualifier (B-1) Dev synthesizes against the oswedev qualifier and the dev infra deploy role is scoped to cdk-oswedev-* — it can no longer assume the default hnb659fds bootstrap roles whose admin cfn-exec-role deploys prod, closing the cross-env escalation (OSWE-IAC-01). Prod stays on the default qualifier. * test: assert per-env bootstrap qualifier isolation + document (B-1) |
||
|---|---|---|
| .. | ||
| bin | ||
| lib | ||
| test | ||
| .gitignore | ||
| cdk.context.json | ||
| cdk.json | ||
| jest.config.js | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| tsconfig.json | ||
open-swe infra (CDK TypeScript)
AWS infrastructure for the Open SWE → AWS migration. Synth-only at this stage —
nothing here is deployed yet. All IAM is applied only after the Phase-1 security
gate (T4 GPT-4.1 IAM cross-review + T5 /sh-security-review) clears (T6).
Layout
infra/
├── bin/
│ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
│ ├── aspects/
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
│ └── constructs/
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│ └── ami-cache.ts # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore
Stacks
| Stack name (kebab) | Construct | Contents |
|---|---|---|
open-swe-iam |
OpenSweIamStack |
Account-level shared GitHub OIDC deploy roles (singletons). |
open-swe-dev |
OpenSweStack (envName: dev) |
open-swe-dev-instance-role, config store (T11), and the EC2 box + ALB ingress (T12, AppService). |
open-swe-prod |
OpenSweStack (envName: prod) |
open-swe-prod-instance-role, config store, and the EC2 box + ALB ingress. |
Account 328440206208, region us-east-1. Stack names are set explicitly so CDK
never defaults to PascalCase; resource names follow open-swe-<env>-*.
The two env stacks (
open-swe-dev/open-swe-prod) are the required pair. The shared OIDC deploy roles are account-wide singletons (oneRoleNameeach), so they live in their own dedicatedopen-swe-iamstack rather than being duplicated across the env stacks — and that stack deploys first (see ordering).
IAM roles defined (unapplied)
githubdeploy-open-swe-infra— GitHub OIDC role for CDK/CFN infra deploys. Trust scoped torepo:Sea-Haven-Industries/open-sweon themain/devbranches only. Permission is the org-standard CDK pattern:sts:AssumeRoleon the CDK bootstrap roles (cdk-hnb659fds-*) — the real CFN/IAM/resource scope lives in the bootstrapcfn-exec-role, not in this role.githubdeploy-open-swe-app— GitHub OIDC role for app deploys. Tag-scopedssm:SendCommand(instances taggedproject=open-swe+env in {dev,prod}) + read-only access to theopen-swe-<env>-assetsS3 artifact buckets.open-swe-<env>-instance-role— EC2 instance role, least-privilege: readopen-swe-<env>-assets(S3), read/open-swe-<env>/*(SSM), readopen-swe-<env>/*(Secrets Manager), put/open-swe/<env>/*CloudWatch Logs, plusAmazonSSMManagedInstanceCorefor SSM agent registration. No admin.
The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created.
CDK bootstrap qualifiers — per-env deploy isolation (B-1 / OSWE-IAC-01)
Each env's infra deploy role may assume only its own bootstrap qualifier's
roles, so a dev-branch token can never assume the bootstrap roles whose admin
cfn-exec-role deploys prod (closing the cross-env escalation that bypassed
prod's Environment approval gate). Mapping lives in config.ts:bootstrapQualifier:
| Env | Qualifier | Toolkit stack | Infra role assumes |
|---|---|---|---|
| dev | oswedev |
CDKToolkit-oswedev |
cdk-oswedev-* |
| prod | hnb659fds (default) |
CDKToolkit |
cdk-hnb659fds-* |
The dev stack synthesizes with DefaultStackSynthesizer({ qualifier: "oswedev" })
(bin/app.ts); prod uses the default. Bootstrap a new env qualifier with:
npx cdk bootstrap --qualifier <qual> --toolkit-stack-name CDKToolkit-<qual> \
--cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess \
aws://328440206208/us-east-1
Deploy order matters when changing an env's qualifier: bootstrap the new
qualifier and deploy the env stack onto it before re-scoping that env's infra
role in open-swe-iam — otherwise a pipeline deploy with the re-scoped role would
fail to assume the not-yet-targeted bootstrap roles.
Kebab-case naming Aspect
KebabNamingAspect (applied app-wide in bin/app.ts) fails synth via
Annotations.addError when a stack name or an explicit physical resource name
(RoleName, BucketName, …) is not kebab-case. Path-style names (Secrets
Manager a/b, SSM /a/b, log groups /aws/.../x) are validated per /-segment.
CDK logical construct ids are intentionally NOT validated (they are conventionally
PascalCase). Covered by test/kebab-naming-aspect.test.ts.
Config store (Secrets Manager + SSM shells — T11)
ConfigStore (lib/constructs/config-store.ts, one per env from OpenSweStack)
renders the resource shells the boot hook deploy/seahaven/fetch-config.sh reads.
The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var
names as the last path segment — open-swe-<env>/<VAR> for secrets,
/open-swe-<env>/<VAR> (FLAT) for config — because fetch-config strips the prefix
and exports that segment verbatim.
Three buckets:
-
Secret shells (Secrets Manager) — 27 secrets. Created value-LESS (an L1
CfnSecretwith NEITHERsecretStringNORgenerateSecretString, which CloudFormation creates as an empty secret). The real value is set out-of-band (put-config.sh) — CDK never owns it, so a latercdk deploycan never clobber it.UpdateReplacePolicy/DeletionPolicy: Retainso a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omitskms:Decrypt).The T9 header says "29 secrets" but its table enumerates 27 distinct VAR names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. Confirm the 27-vs-29 count (code-only candidates not in the table:
USER_ID_API_KEY_MAP,JUDGE_ANTHROPIC_BASE_URL).⚠️
Retain+ fixed name orphans these shells on a failed FIRST create. If the stack's initial create fails and rolls back,Retainkeeps the shells instead of deleting them. The stack is then gone, but the secrets survive, still holding the globalopen-swe-<env>/<VAR>names — so every later create fails withAlreadyExists. A plaindelete-secretdoes not clear it (the name stays reserved for the 7–30 day recovery window). This bites on a teardown/rebuild, a secret logical-id change/refactor, or standing up a new env — never on routine updates of an already-created stack. Recovery — before re-creating the stack, force-delete the empty orphans so the names free immediately:aws secretsmanager list-secrets --region us-east-1 \ --filters Key=name,Values=open-swe-<env>/ \ --query 'SecretList[].Name' --output text | tr '\t' '\n' | while read -r n; do aws secretsmanager delete-secret --secret-id "$n" \ --region us-east-1 --force-delete-without-recovery doneForce-delete only shells with no value version — a populated secret holds real operator material. (Hit on prod 2026-06-29; see PR #51's deploy failure.)
-
IaC-managed SSM config — 8 params, real values owned in code:
Param dev prod SANDBOX_TYPElangsmithlangsmithDEFAULT_REPO_OWNERSea-Haven-IndustriesSea-Haven-IndustriesALLOWED_GITHUB_ORGSSea-Haven-IndustriesSea-Haven-IndustriesDEFAULT_REPO_NAMEopen-swe-pilot(confirm)open-swe-pilot(confirm)DASHBOARD_BASE_URLhttps://openswe-dev.seahaven.com(confirm host)https://openswe.seahaven.com(confirm host)DASHBOARD_API_BASE_URLsame as base same as base DASHBOARD_ALLOWED_ORIGINSsame as base same as base LLM_MODEL_IDanthropic:claude-opus-4-8(confirm)anthropic:claude-opus-4-8(confirm) -
Out-of-band SSM config — NOT created by CDK. Operationally-variable or env-specific-unknown values listed in
OUT_OF_BAND_SSMand populated byput-config.sh. The keystone isDEFAULT_SANDBOX_SNAPSHOT_ID(changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls.
Kebab-Aspect deviation
KebabNamingAspect exempts AWS::SecretsManager::Secret and AWS::SSM::Parameter
from the kebab check (see the KEBAB_EXEMPT_RESOURCE_TYPES set) — the UPPER_SNAKE
env-var segment is a required, documented deviation for a lossless store→env
round-trip. Every other explicitly-named resource is still validated. Covered by a
dedicated case in test/kebab-naming-aspect.test.ts.
Deploy ordering (values BEFORE the box boots)
The shells are synth-able now (T11). Population is out-of-band and happens after
cdk deploy open-swe-<env> but before the EC2/T12 box first boots:
cdk deploy open-swe-<env> # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod> # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
put-config.sh ships <FILL> placeholders only (no real secret values committed);
provide each value inline, via OPENSWE_PUT_<VAR> env vars, or from a vault. It does
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
Compute + ingress (AppService — T12)
AppService (lib/constructs/app-service.ts, one per env from OpenSweStack)
builds the box and its path to the internet. A single internet-facing ALB
(app/seahaven-com) and a single VPC are shared with the on-prem
seahaven-site stack, so open-swe imports the VPC, the ALB security group
(sg-0b0301deed193258a), the :443 listener, and the public seahaven.com
zone — and never owns/mutates them. It adds:
- One ARM64 EC2 box (
open-swe-<env>-box,t4g.mediumdev /t4g.largeprod) in private1 (us-east-1a) — same AZ as the single NAT for in-AZ egress.requireImdsv2, gp3 encrypted root,deleteOnTermination(no RETAIN volume — see below).userDataCausesReplacement: true; user-data is rendered fromdeploy/ami/user-data.sh. - A standalone instance SG reachable only from the shared ALB SG on
:80(nginx). Egress open (NAT). The ALB SG is opened to the box via a standaloneCfnSecurityGroupEgressso the imported (on-prem-owned) SG is never mutated. - A target group → instance
:80(nginx is the sole ingress; the LangGraph control plane stays on loopback:2024). Health checkGET /healthz. - Two listener rules on the imported
:443listener, both → the same TG:- Webhooks (priority 2 dev / 3 prod):
host ∈ {openswe-<env>, hooks-<env>}.seahaven.comAND path/webhooks/*. - Site (priority 10 dev / 11 prod):
host = openswe-<env>.seahaven.com(dashboard SPA +/dashboard/api/).
- Webhooks (priority 2 dev / 3 prod):
- Route53 alias records
openswe[-dev]+hooks[-dev]→ the shared ALB. - Four CloudWatch log groups (
/open-swe/<env>/{app,user-data,nginx-access,nginx-error}) at 30-day retention (IaC-owned; mirrors the CW-agent config).
Listener-rule ordering (load-bearing)
The shared listener already has a host-agnostic /webhooks/* PATH rule at
priority 5 (on-prem). ALB rules are first-match by ascending priority, so the
open-swe webhook rule must sit below 5 or every …/webhooks/* request (any
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
a host condition, so it does not steal the on-prem hosts' webhooks. The
dashboard "site" rule carries no path that collides with rule 5, so it sits at
10/11.
Cross-stack coordination (T13 review). The seahaven-site (on-prem) and
open-swe stacks both add resources to the same imported listener and ALB SG.
This is safe: each stack owns only the resources it declares (its own logical
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
and the standalone CfnSecurityGroupEgress never mutates the shared SG's own
definition (the pattern on-prem itself uses). The one shared namespace that needs
care is listener-rule priority (globally unique per listener; a collision is
a fail-safe deploy error, not silent drift). Ownership — keep disjoint:
seahaven-site = 4-7 + default; open-swe = 2, 3, 10, 11. open-swe's
webhook rules are host-scoped to its own *.seahaven.com hosts, so they never
match an on-prem seahavenind.com host.
Security review (T5/T12 /sh-security-review)
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
One confirmed medium (OSWE-T12-01: nginx's 1 MB default client_max_body_size
would 413 large GitHub webhooks before in-app signature verification) is fixed in
open-swe.nginx.conf (25m on /webhooks/, 10m on /dashboard/api/). The
hooks hostname is scoped to /webhooks/* only (OSWE-T12-02 hygiene). An
X-Forwarded-For spoof candidate was killed — no code trusts the leftmost XFF.
No confirmed critical/high; no block.
Baked AMI + EBS-replacement discipline
bakedOpenSweArm64() (in lib/constructs/ami-cache.ts) pins the custom
open-swe-base-arm64 image by EXACT id (BAKED_OPEN_SWE_AMI_ID) via
MachineImage.genericLinux({ "us-east-1": "<ami-id>" }) — no SSM lookup, so synth
and deploy are fully offline/deterministic. The image is built by
deploy/ami/open-swe-base.pkr.hcl (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW
agent + boot templates, no secrets); the box's user-data.sh assumes that
baked layout (/opt/open-swe, openswe user, nginx, CW agent).
Pinning by exact id (vs a most_recent name filter) is what prevents a routine
deploy from silently swapping the AMI → EC2 instance replacement (the
file-share data-loss root cause — memory feedback_inline_ebs_volumes).
userDataCausesReplacement: trueis deliberate — user-data is provisioning-only and carries no durable state.- No durable state on the box → no RETAIN volume. The in-memory langgraph
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
intentionally no standalone
ec2.Volume+removalPolicy.RETAIN. The goal is replacement-tolerance, not avoidance. - Snapshot-before-replace still applies operationally: before any replacing
deploy snapshot the root volume and wait
state=completed, and re-verify "no local-only durable state" first.
Refresh the AMI deliberately:
cd deploy/ami && packer build open-swe-base.pkr.hcl # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack # WILL show "requires replacement"
cdk.context.jsonis{}— nothing is resolved via context anymore (the AMI is a static id pin), so synth makes no live AWS call.
Commands
npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test # jest — naming Aspect
CI/CD (T18 — .github/workflows/ci-infra.yml + cd-infra.yml)
Path-filtered, OIDC-only (no static keys). The Python agent keeps its own
ci.yml ("CI"); these two add the /infra half.
| Workflow | Trigger | Does |
|---|---|---|
ci-infra.yml |
PR touching infra/** |
tsc + jest + cdk synth (reusable ci-typescript-cdk.yaml). |
cd-infra.yml |
push to dev/main touching infra/**, or dispatch |
CI (pre-deploy) → per-env cdk deploy. |
cd-infra.yml flow:
- push to
dev→ CI green → autocdk deploy OpenSweDevStack(assumesgithubdeploy-open-swe-infra-dev; the job declares noenvironment:, so the OIDC subject is…:ref:refs/heads/dev— matching that role's trust). - push to
main→ CI green →cdk deploy OpenSweProdStackbehind theprodGitHub Environment (required reviewer = Adam). Theenvironment: proddeclaration both fires the manual-approval gate and makes the OIDC subject…:environment:prod— matchinggithubdeploy-open-swe-infra-prod's trust.
Why not the reusable cd-cdk.yaml: it runs cdk deploy --all, which from a
single-env push would deploy the other env + the shared IAM stack — breaking the
per-env boundary. So CD targets one stack explicitly per env. The shared
open-swe-iam stack is not deployed by CD (privileged, human-gated — T6).
Gating note: infra CI is enforced at the deploy boundary (cd-infra's
deploy-* jobs needs: ci), not as a branch-protection required check —
path-filtering a required check would deadlock app-only PRs (a skipped required
check never satisfies). Making Infra CI a required check later needs a
skip-aware shim or dropping its path filter.
Prerequisites (set post-T6, when the roles exist):
- repo variables
AWS_DEPLOY_ROLE_INFRA_DEV/AWS_DEPLOY_ROLE_INFRA_PROD= thegithubdeploy-open-swe-infra-<env>role ARNs (open-swe-iamoutputs). - a GitHub Environment named
prodwith Adam as a required reviewer.
App-side CD (CI → S3 artifact → SSM deploy via
githubdeploy-open-swe-app-<env>) is T19, not here.
Deploy ordering (when the gate clears — NOT yet)
open-swe-iamfirst — apply the IAM stack (T6, human-gated), then set the repoAWS_DEPLOY_ROLE_INFRA_{DEV,PROD}variables from its role-ARN outputs and configure theprodEnvironment reviewer (BLOCK#3).- Security gate — T4 GPT-4.1 IAM cross-review + T5
/sh-security-reviewon the synth; resolve every confirmed critical/high. - IAM applied (T6) — only after the gate.
- Env stacks: first
open-swe-dev(T14, manual validate), then CD auto-deploys dev on push;open-swe-prod(T21) behind theprodEnvironment approval.
Version policy
aws-cdk-lib is pinned EXACT (2.260.0) — no ^/~. Dependabot keeps it
current; CI (npm ci + cdk synth) + dependency review gate each bump. See
aws-infrastructure.md "CDK Version Policy" and memory
feedback_cdk_lib_bundled_deps.