* Align workflows with Sea Haven CI/CD handbook Bring the workflow suite in line with the handbook: bump actions/checkout to v7 (Node 24 runtime, already standardized), kebab-case the two snake_case workflow filenames, and add the org-standard Labeler caller and Dependency Review gate so vulnerable or disallowed-license deps and unlabeled PRs are caught automatically. File renames only — job/check display names are unchanged, so the promotion gate's REQUIRED_CHECKS and branch-protection required checks are unaffected. Refs: INFRA-115 * Drop Agent prefix from CI workflow + job names The handbook names workflows for what they do (CI, Deploy, Labeler), not the component they run, matching .github and afterhours-shift-manager. Rename the suite to CI and its jobs to Lint / Format check / Unit tests, and keep the promotion gate's REQUIRED_CHECKS in sync. Refs: INFRA-115 --------- Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| bin | ||
| lib | ||
| test | ||
| .gitignore | ||
| cdk.context.json | ||
| cdk.json | ||
| jest.config.js | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| tsconfig.json | ||
open-swe infra (CDK TypeScript)
AWS infrastructure for the Open SWE → AWS migration. Synth-only at this stage —
nothing here is deployed yet. All IAM is applied only after the Phase-1 security
gate (T4 GPT-4.1 IAM cross-review + T5 /sh-security-review) clears (T6).
Layout
infra/
├── bin/
│ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
│ ├── aspects/
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
│ └── constructs/
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│ └── ami-cache.ts # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore
Stacks
| Stack name (kebab) | Construct | Contents |
|---|---|---|
open-swe-iam |
OpenSweIamStack |
Account-level shared GitHub OIDC deploy roles (singletons). |
open-swe-dev |
OpenSweStack (envName: dev) |
open-swe-dev-instance-role, config store (T11), and the EC2 box + ALB ingress (T12, AppService). |
open-swe-prod |
OpenSweStack (envName: prod) |
open-swe-prod-instance-role, config store, and the EC2 box + ALB ingress. |
Account 328440206208, region us-east-1. Stack names are set explicitly so CDK
never defaults to PascalCase; resource names follow open-swe-<env>-*.
The two env stacks (
open-swe-dev/open-swe-prod) are the required pair. The shared OIDC deploy roles are account-wide singletons (oneRoleNameeach), so they live in their own dedicatedopen-swe-iamstack rather than being duplicated across the env stacks — and that stack deploys first (see ordering).
IAM roles defined (unapplied)
githubdeploy-open-swe-infra— GitHub OIDC role for CDK/CFN infra deploys. Trust scoped torepo:Sea-Haven-Industries/open-sweon themain/devbranches only. Permission is the org-standard CDK pattern:sts:AssumeRoleon the CDK bootstrap roles (cdk-hnb659fds-*) — the real CFN/IAM/resource scope lives in the bootstrapcfn-exec-role, not in this role.githubdeploy-open-swe-app— GitHub OIDC role for app deploys. Tag-scopedssm:SendCommand(instances taggedproject=open-swe+env in {dev,prod}) + read-only access to theopen-swe-<env>-assetsS3 artifact buckets.open-swe-<env>-instance-role— EC2 instance role, least-privilege: readopen-swe-<env>-assets(S3), read/open-swe-<env>/*(SSM), readopen-swe-<env>/*(Secrets Manager), put/open-swe/<env>/*CloudWatch Logs, plusAmazonSSMManagedInstanceCorefor SSM agent registration. No admin.
The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created.
Kebab-case naming Aspect
KebabNamingAspect (applied app-wide in bin/app.ts) fails synth via
Annotations.addError when a stack name or an explicit physical resource name
(RoleName, BucketName, …) is not kebab-case. Path-style names (Secrets
Manager a/b, SSM /a/b, log groups /aws/.../x) are validated per /-segment.
CDK logical construct ids are intentionally NOT validated (they are conventionally
PascalCase). Covered by test/kebab-naming-aspect.test.ts.
Config store (Secrets Manager + SSM shells — T11)
ConfigStore (lib/constructs/config-store.ts, one per env from OpenSweStack)
renders the resource shells the boot hook deploy/seahaven/fetch-config.sh reads.
The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var
names as the last path segment — open-swe-<env>/<VAR> for secrets,
/open-swe-<env>/<VAR> (FLAT) for config — because fetch-config strips the prefix
and exports that segment verbatim.
Three buckets:
-
Secret shells (Secrets Manager) — 27 secrets. Created value-LESS (an L1
CfnSecretwith NEITHERsecretStringNORgenerateSecretString, which CloudFormation creates as an empty secret). The real value is set out-of-band (put-config.sh) — CDK never owns it, so a latercdk deploycan never clobber it.UpdateReplacePolicy/DeletionPolicy: Retainso a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omitskms:Decrypt).The T9 header says "29 secrets" but its table enumerates 27 distinct VAR names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. Confirm the 27-vs-29 count (code-only candidates not in the table:
USER_ID_API_KEY_MAP,JUDGE_ANTHROPIC_BASE_URL). -
IaC-managed SSM config — 8 params, real values owned in code:
Param dev prod SANDBOX_TYPElangsmithlangsmithDEFAULT_REPO_OWNERSea-Haven-IndustriesSea-Haven-IndustriesALLOWED_GITHUB_ORGSSea-Haven-IndustriesSea-Haven-IndustriesDEFAULT_REPO_NAMEopen-swe-pilot(confirm)open-swe-pilot(confirm)DASHBOARD_BASE_URLhttps://openswe-dev.seahaven.com(confirm host)https://openswe.seahaven.com(confirm host)DASHBOARD_API_BASE_URLsame as base same as base DASHBOARD_ALLOWED_ORIGINSsame as base same as base LLM_MODEL_IDanthropic:claude-opus-4-8(confirm)anthropic:claude-opus-4-8(confirm) -
Out-of-band SSM config — NOT created by CDK. Operationally-variable or env-specific-unknown values listed in
OUT_OF_BAND_SSMand populated byput-config.sh. The keystone isDEFAULT_SANDBOX_SNAPSHOT_ID(changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls.
Kebab-Aspect deviation
KebabNamingAspect exempts AWS::SecretsManager::Secret and AWS::SSM::Parameter
from the kebab check (see the KEBAB_EXEMPT_RESOURCE_TYPES set) — the UPPER_SNAKE
env-var segment is a required, documented deviation for a lossless store→env
round-trip. Every other explicitly-named resource is still validated. Covered by a
dedicated case in test/kebab-naming-aspect.test.ts.
Deploy ordering (values BEFORE the box boots)
The shells are synth-able now (T11). Population is out-of-band and happens after
cdk deploy open-swe-<env> but before the EC2/T12 box first boots:
cdk deploy open-swe-<env> # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod> # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
put-config.sh ships <FILL> placeholders only (no real secret values committed);
provide each value inline, via OPENSWE_PUT_<VAR> env vars, or from a vault. It does
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
Compute + ingress (AppService — T12)
AppService (lib/constructs/app-service.ts, one per env from OpenSweStack)
builds the box and its path to the internet. A single internet-facing ALB
(app/seahaven-com) and a single VPC are shared with the on-prem
seahaven-site stack, so open-swe imports the VPC, the ALB security group
(sg-0b0301deed193258a), the :443 listener, and the public seahaven.com
zone — and never owns/mutates them. It adds:
- One ARM64 EC2 box (
open-swe-<env>-box,t4g.mediumdev /t4g.largeprod) in private1 (us-east-1a) — same AZ as the single NAT for in-AZ egress.requireImdsv2, gp3 encrypted root,deleteOnTermination(no RETAIN volume — see below).userDataCausesReplacement: true; user-data is rendered fromdeploy/ami/user-data.sh. - A standalone instance SG reachable only from the shared ALB SG on
:80(nginx). Egress open (NAT). The ALB SG is opened to the box via a standaloneCfnSecurityGroupEgressso the imported (on-prem-owned) SG is never mutated. - A target group → instance
:80(nginx is the sole ingress; the LangGraph control plane stays on loopback:2024). Health checkGET /healthz. - Two listener rules on the imported
:443listener, both → the same TG:- Webhooks (priority 2 dev / 3 prod):
host ∈ {openswe-<env>, hooks-<env>}.seahaven.comAND path/webhooks/*. - Site (priority 10 dev / 11 prod):
host = openswe-<env>.seahaven.com(dashboard SPA +/dashboard/api/).
- Webhooks (priority 2 dev / 3 prod):
- Route53 alias records
openswe[-dev]+hooks[-dev]→ the shared ALB. - Four CloudWatch log groups (
/open-swe/<env>/{app,user-data,nginx-access,nginx-error}) at 30-day retention (IaC-owned; mirrors the CW-agent config).
Listener-rule ordering (load-bearing)
The shared listener already has a host-agnostic /webhooks/* PATH rule at
priority 5 (on-prem). ALB rules are first-match by ascending priority, so the
open-swe webhook rule must sit below 5 or every …/webhooks/* request (any
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
a host condition, so it does not steal the on-prem hosts' webhooks. The
dashboard "site" rule carries no path that collides with rule 5, so it sits at
10/11.
Cross-stack coordination (T13 review). The seahaven-site (on-prem) and
open-swe stacks both add resources to the same imported listener and ALB SG.
This is safe: each stack owns only the resources it declares (its own logical
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
and the standalone CfnSecurityGroupEgress never mutates the shared SG's own
definition (the pattern on-prem itself uses). The one shared namespace that needs
care is listener-rule priority (globally unique per listener; a collision is
a fail-safe deploy error, not silent drift). Ownership — keep disjoint:
seahaven-site = 4-7 + default; open-swe = 2, 3, 10, 11. open-swe's
webhook rules are host-scoped to its own *.seahaven.com hosts, so they never
match an on-prem seahavenind.com host.
Security review (T5/T12 /sh-security-review)
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
One confirmed medium (OSWE-T12-01: nginx's 1 MB default client_max_body_size
would 413 large GitHub webhooks before in-app signature verification) is fixed in
open-swe.nginx.conf (25m on /webhooks/, 10m on /dashboard/api/). The
hooks hostname is scoped to /webhooks/* only (OSWE-T12-02 hygiene). An
X-Forwarded-For spoof candidate was killed — no code trusts the leftmost XFF.
No confirmed critical/high; no block.
Baked AMI + EBS-replacement discipline
bakedOpenSweArm64() (in lib/constructs/ami-cache.ts) pins the custom
open-swe-base-arm64 image by EXACT id (BAKED_OPEN_SWE_AMI_ID) via
MachineImage.genericLinux({ "us-east-1": "<ami-id>" }) — no SSM lookup, so synth
and deploy are fully offline/deterministic. The image is built by
deploy/ami/open-swe-base.pkr.hcl (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW
agent + boot templates, no secrets); the box's user-data.sh assumes that
baked layout (/opt/open-swe, openswe user, nginx, CW agent).
Pinning by exact id (vs a most_recent name filter) is what prevents a routine
deploy from silently swapping the AMI → EC2 instance replacement (the
file-share data-loss root cause — memory feedback_inline_ebs_volumes).
userDataCausesReplacement: trueis deliberate — user-data is provisioning-only and carries no durable state.- No durable state on the box → no RETAIN volume. The in-memory langgraph
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
intentionally no standalone
ec2.Volume+removalPolicy.RETAIN. The goal is replacement-tolerance, not avoidance. - Snapshot-before-replace still applies operationally: before any replacing
deploy snapshot the root volume and wait
state=completed, and re-verify "no local-only durable state" first.
Refresh the AMI deliberately:
cd deploy/ami && packer build open-swe-base.pkr.hcl # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack # WILL show "requires replacement"
cdk.context.jsonis{}— nothing is resolved via context anymore (the AMI is a static id pin), so synth makes no live AWS call.
Commands
npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test # jest — naming Aspect
CI/CD (T18 — .github/workflows/ci-infra.yml + cd-infra.yml)
Path-filtered, OIDC-only (no static keys). The Python agent keeps its own
ci.yml ("CI"); these two add the /infra half.
| Workflow | Trigger | Does |
|---|---|---|
ci-infra.yml |
PR touching infra/** |
tsc + jest + cdk synth (reusable ci-typescript-cdk.yaml). |
cd-infra.yml |
push to dev/main touching infra/**, or dispatch |
CI (pre-deploy) → per-env cdk deploy. |
cd-infra.yml flow:
- push to
dev→ CI green → autocdk deploy OpenSweDevStack(assumesgithubdeploy-open-swe-infra-dev; the job declares noenvironment:, so the OIDC subject is…:ref:refs/heads/dev— matching that role's trust). - push to
main→ CI green →cdk deploy OpenSweProdStackbehind theprodGitHub Environment (required reviewer = Adam). Theenvironment: proddeclaration both fires the manual-approval gate and makes the OIDC subject…:environment:prod— matchinggithubdeploy-open-swe-infra-prod's trust.
Why not the reusable cd-cdk.yaml: it runs cdk deploy --all, which from a
single-env push would deploy the other env + the shared IAM stack — breaking the
per-env boundary. So CD targets one stack explicitly per env. The shared
open-swe-iam stack is not deployed by CD (privileged, human-gated — T6).
Gating note: infra CI is enforced at the deploy boundary (cd-infra's
deploy-* jobs needs: ci), not as a branch-protection required check —
path-filtering a required check would deadlock app-only PRs (a skipped required
check never satisfies). Making Infra CI a required check later needs a
skip-aware shim or dropping its path filter.
Prerequisites (set post-T6, when the roles exist):
- repo variables
AWS_DEPLOY_ROLE_INFRA_DEV/AWS_DEPLOY_ROLE_INFRA_PROD= thegithubdeploy-open-swe-infra-<env>role ARNs (open-swe-iamoutputs). - a GitHub Environment named
prodwith Adam as a required reviewer.
App-side CD (CI → S3 artifact → SSM deploy via
githubdeploy-open-swe-app-<env>) is T19, not here.
Deploy ordering (when the gate clears — NOT yet)
open-swe-iamfirst — apply the IAM stack (T6, human-gated), then set the repoAWS_DEPLOY_ROLE_INFRA_{DEV,PROD}variables from its role-ARN outputs and configure theprodEnvironment reviewer (BLOCK#3).- Security gate — T4 GPT-4.1 IAM cross-review + T5
/sh-security-reviewon the synth; resolve every confirmed critical/high. - IAM applied (T6) — only after the gate.
- Env stacks: first
open-swe-dev(T14, manual validate), then CD auto-deploys dev on push;open-swe-prod(T21) behind theprodEnvironment approval.
Version policy
aws-cdk-lib is pinned EXACT (2.260.0) — no ^/~. Dependabot keeps it
current; CI (npm ci + cdk synth) + dependency review gate each bump. See
aws-infrastructure.md "CDK Version Policy" and memory
feedback_cdk_lib_bundled_deps.