* fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only. |
||
|---|---|---|
| .. | ||
| bin | ||
| lib | ||
| test | ||
| .gitignore | ||
| cdk.context.json | ||
| cdk.json | ||
| jest.config.js | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| tsconfig.json | ||
open-swe infra (CDK TypeScript)
AWS infrastructure for the Open SWE → AWS migration. Synth-only at this stage —
nothing here is deployed yet. All IAM is applied only after the Phase-1 security
gate (T4 GPT-4.1 IAM cross-review + T5 /sh-security-review) clears (T6).
Layout
infra/
├── bin/
│ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
│ ├── aspects/
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
│ └── constructs/
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│ └── ami-cache.ts # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore
Stacks
| Stack name (kebab) | Construct | Contents |
|---|---|---|
open-swe-iam |
OpenSweIamStack |
Account-level shared GitHub OIDC deploy roles (singletons). |
open-swe-dev |
OpenSweStack (envName: dev) |
open-swe-dev-instance-role, config store (T11), and the EC2 box + ALB ingress (T12, AppService). |
open-swe-prod |
OpenSweStack (envName: prod) |
open-swe-prod-instance-role, config store, and the EC2 box + ALB ingress. |
Account 328440206208, region us-east-1. Stack names are set explicitly so CDK
never defaults to PascalCase; resource names follow open-swe-<env>-*.
The two env stacks (
open-swe-dev/open-swe-prod) are the required pair. The shared OIDC deploy roles are account-wide singletons (oneRoleNameeach), so they live in their own dedicatedopen-swe-iamstack rather than being duplicated across the env stacks — and that stack deploys first (see ordering).
IAM roles defined (unapplied)
githubdeploy-open-swe-infra— GitHub OIDC role for CDK/CFN infra deploys. Trust scoped torepo:Sea-Haven-Industries/open-sweon themain/devbranches only. Permission is the org-standard CDK pattern:sts:AssumeRoleon the CDK bootstrap roles (cdk-hnb659fds-*) — the real CFN/IAM/resource scope lives in the bootstrapcfn-exec-role, not in this role.githubdeploy-open-swe-app— GitHub OIDC role for app deploys. Tag-scopedssm:SendCommand(instances taggedproject=open-swe+env in {dev,prod}) + read-only access to theopen-swe-<env>-assetsS3 artifact buckets.open-swe-<env>-instance-role— EC2 instance role, least-privilege: readopen-swe-<env>-assets(S3), read/open-swe-<env>/*(SSM), readopen-swe-<env>/*(Secrets Manager), put/open-swe/<env>/*CloudWatch Logs, plusAmazonSSMManagedInstanceCorefor SSM agent registration. No admin.
The GitHub OIDC provider already exists account-wide (created for seahaven-site); it is referenced by ARN, never re-created.
Kebab-case naming Aspect
KebabNamingAspect (applied app-wide in bin/app.ts) fails synth via
Annotations.addError when a stack name or an explicit physical resource name
(RoleName, BucketName, …) is not kebab-case. Path-style names (Secrets
Manager a/b, SSM /a/b, log groups /aws/.../x) are validated per /-segment.
CDK logical construct ids are intentionally NOT validated (they are conventionally
PascalCase). Covered by test/kebab-naming-aspect.test.ts.
Config store (Secrets Manager + SSM shells — T11)
ConfigStore (lib/constructs/config-store.ts, one per env from OpenSweStack)
renders the resource shells the boot hook deploy/seahaven/fetch-config.sh reads.
The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var
names as the last path segment — open-swe-<env>/<VAR> for secrets,
/open-swe-<env>/<VAR> (FLAT) for config — because fetch-config strips the prefix
and exports that segment verbatim.
Three buckets:
-
Secret shells (Secrets Manager) — 27 secrets. Created value-LESS (an L1
CfnSecretwith NEITHERsecretStringNORgenerateSecretString, which CloudFormation creates as an empty secret). The real value is set out-of-band (put-config.sh) — CDK never owns it, so a latercdk deploycan never clobber it.UpdateReplacePolicy/DeletionPolicy: Retainso a teardown can't destroy operator-set secret material. AWS-managed key (no CMK — matches the instance role, which omitskms:Decrypt).The T9 header says "29 secrets" but its table enumerates 27 distinct VAR names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29. Confirm the 27-vs-29 count (code-only candidates not in the table:
USER_ID_API_KEY_MAP,JUDGE_ANTHROPIC_BASE_URL).⚠️
Retain+ fixed name orphans these shells on a failed FIRST create. If the stack's initial create fails and rolls back,Retainkeeps the shells instead of deleting them. The stack is then gone, but the secrets survive, still holding the globalopen-swe-<env>/<VAR>names — so every later create fails withAlreadyExists. A plaindelete-secretdoes not clear it (the name stays reserved for the 7–30 day recovery window). This bites on a teardown/rebuild, a secret logical-id change/refactor, or standing up a new env — never on routine updates of an already-created stack. Recovery — before re-creating the stack, force-delete the empty orphans so the names free immediately:aws secretsmanager list-secrets --region us-east-1 \ --filters Key=name,Values=open-swe-<env>/ \ --query 'SecretList[].Name' --output text | tr '\t' '\n' | while read -r n; do aws secretsmanager delete-secret --secret-id "$n" \ --region us-east-1 --force-delete-without-recovery doneForce-delete only shells with no value version — a populated secret holds real operator material. (Hit on prod 2026-06-29; see PR #51's deploy failure.)
-
IaC-managed SSM config — 8 params, real values owned in code:
Param dev prod SANDBOX_TYPElangsmithlangsmithDEFAULT_REPO_OWNERSea-Haven-IndustriesSea-Haven-IndustriesALLOWED_GITHUB_ORGSSea-Haven-IndustriesSea-Haven-IndustriesDEFAULT_REPO_NAMEopen-swe-pilot(confirm)open-swe-pilot(confirm)DASHBOARD_BASE_URLhttps://openswe-dev.seahaven.com(confirm host)https://openswe.seahaven.com(confirm host)DASHBOARD_API_BASE_URLsame as base same as base DASHBOARD_ALLOWED_ORIGINSsame as base same as base LLM_MODEL_IDanthropic:claude-opus-4-8(confirm)anthropic:claude-opus-4-8(confirm) -
Out-of-band SSM config — NOT created by CDK. Operationally-variable or env-specific-unknown values listed in
OUT_OF_BAND_SSMand populated byput-config.sh. The keystone isDEFAULT_SANDBOX_SNAPSHOT_ID(changes on every snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the GitHub App ids, Slack ids, and LangSmith tenant/urls.
Kebab-Aspect deviation
KebabNamingAspect exempts AWS::SecretsManager::Secret and AWS::SSM::Parameter
from the kebab check (see the KEBAB_EXEMPT_RESOURCE_TYPES set) — the UPPER_SNAKE
env-var segment is a required, documented deviation for a lossless store→env
round-trip. Every other explicitly-named resource is still validated. Covered by a
dedicated case in test/kebab-naming-aspect.test.ts.
Deploy ordering (values BEFORE the box boots)
The shells are synth-able now (T11). Population is out-of-band and happens after
cdk deploy open-swe-<env> but before the EC2/T12 box first boots:
cdk deploy open-swe-<env> # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod> # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
put-config.sh ships <FILL> placeholders only (no real secret values committed);
provide each value inline, via OPENSWE_PUT_<VAR> env vars, or from a vault. It does
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
Compute + ingress (AppService — T12)
AppService (lib/constructs/app-service.ts, one per env from OpenSweStack)
builds the box and its path to the internet. A single internet-facing ALB
(app/seahaven-com) and a single VPC are shared with the on-prem
seahaven-site stack, so open-swe imports the VPC, the ALB security group
(sg-0b0301deed193258a), the :443 listener, and the public seahaven.com
zone — and never owns/mutates them. It adds:
- One ARM64 EC2 box (
open-swe-<env>-box,t4g.mediumdev /t4g.largeprod) in private1 (us-east-1a) — same AZ as the single NAT for in-AZ egress.requireImdsv2, gp3 encrypted root,deleteOnTermination(no RETAIN volume — see below).userDataCausesReplacement: true; user-data is rendered fromdeploy/ami/user-data.sh. - A standalone instance SG reachable only from the shared ALB SG on
:80(nginx). Egress open (NAT). The ALB SG is opened to the box via a standaloneCfnSecurityGroupEgressso the imported (on-prem-owned) SG is never mutated. - A target group → instance
:80(nginx is the sole ingress; the LangGraph control plane stays on loopback:2024). Health checkGET /healthz. - Two listener rules on the imported
:443listener, both → the same TG:- Webhooks (priority 2 dev / 3 prod):
host ∈ {openswe-<env>, hooks-<env>}.seahaven.comAND path/webhooks/*. - Site (priority 10 dev / 11 prod):
host = openswe-<env>.seahaven.com(dashboard SPA +/dashboard/api/).
- Webhooks (priority 2 dev / 3 prod):
- Route53 alias records
openswe[-dev]+hooks[-dev]→ the shared ALB. - Four CloudWatch log groups (
/open-swe/<env>/{app,user-data,nginx-access,nginx-error}) at 30-day retention (IaC-owned; mirrors the CW-agent config).
Listener-rule ordering (load-bearing)
The shared listener already has a host-agnostic /webhooks/* PATH rule at
priority 5 (on-prem). ALB rules are first-match by ascending priority, so the
open-swe webhook rule must sit below 5 or every …/webhooks/* request (any
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
a host condition, so it does not steal the on-prem hosts' webhooks. The
dashboard "site" rule carries no path that collides with rule 5, so it sits at
10/11.
Cross-stack coordination (T13 review). The seahaven-site (on-prem) and
open-swe stacks both add resources to the same imported listener and ALB SG.
This is safe: each stack owns only the resources it declares (its own logical
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
and the standalone CfnSecurityGroupEgress never mutates the shared SG's own
definition (the pattern on-prem itself uses). The one shared namespace that needs
care is listener-rule priority (globally unique per listener; a collision is
a fail-safe deploy error, not silent drift). Ownership — keep disjoint:
seahaven-site = 4-7 + default; open-swe = 2, 3, 10, 11. open-swe's
webhook rules are host-scoped to its own *.seahaven.com hosts, so they never
match an on-prem seahavenind.com host.
Security review (T5/T12 /sh-security-review)
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
One confirmed medium (OSWE-T12-01: nginx's 1 MB default client_max_body_size
would 413 large GitHub webhooks before in-app signature verification) is fixed in
open-swe.nginx.conf (25m on /webhooks/, 10m on /dashboard/api/). The
hooks hostname is scoped to /webhooks/* only (OSWE-T12-02 hygiene). An
X-Forwarded-For spoof candidate was killed — no code trusts the leftmost XFF.
No confirmed critical/high; no block.
Baked AMI + EBS-replacement discipline
bakedOpenSweArm64() (in lib/constructs/ami-cache.ts) pins the custom
open-swe-base-arm64 image by EXACT id (BAKED_OPEN_SWE_AMI_ID) via
MachineImage.genericLinux({ "us-east-1": "<ami-id>" }) — no SSM lookup, so synth
and deploy are fully offline/deterministic. The image is built by
deploy/ami/open-swe-base.pkr.hcl (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW
agent + boot templates, no secrets); the box's user-data.sh assumes that
baked layout (/opt/open-swe, openswe user, nginx, CW agent).
Pinning by exact id (vs a most_recent name filter) is what prevents a routine
deploy from silently swapping the AMI → EC2 instance replacement (the
file-share data-loss root cause — memory feedback_inline_ebs_volumes).
userDataCausesReplacement: trueis deliberate — user-data is provisioning-only and carries no durable state.- No durable state on the box → no RETAIN volume. The in-memory langgraph
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
intentionally no standalone
ec2.Volume+removalPolicy.RETAIN. The goal is replacement-tolerance, not avoidance. - Snapshot-before-replace still applies operationally: before any replacing
deploy snapshot the root volume and wait
state=completed, and re-verify "no local-only durable state" first.
Refresh the AMI deliberately:
cd deploy/ami && packer build open-swe-base.pkr.hcl # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack # WILL show "requires replacement"
cdk.context.jsonis{}— nothing is resolved via context anymore (the AMI is a static id pin), so synth makes no live AWS call.
Commands
npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test # jest — naming Aspect
CI/CD (T18 — .github/workflows/ci-infra.yml + cd-infra.yml)
Path-filtered, OIDC-only (no static keys). The Python agent keeps its own
ci.yml ("CI"); these two add the /infra half.
| Workflow | Trigger | Does |
|---|---|---|
ci-infra.yml |
PR touching infra/** |
tsc + jest + cdk synth (reusable ci-typescript-cdk.yaml). |
cd-infra.yml |
push to dev/main touching infra/**, or dispatch |
CI (pre-deploy) → per-env cdk deploy. |
cd-infra.yml flow:
- push to
dev→ CI green → autocdk deploy OpenSweDevStack(assumesgithubdeploy-open-swe-infra-dev; the job declares noenvironment:, so the OIDC subject is…:ref:refs/heads/dev— matching that role's trust). - push to
main→ CI green →cdk deploy OpenSweProdStackbehind theprodGitHub Environment (required reviewer = Adam). Theenvironment: proddeclaration both fires the manual-approval gate and makes the OIDC subject…:environment:prod— matchinggithubdeploy-open-swe-infra-prod's trust.
Why not the reusable cd-cdk.yaml: it runs cdk deploy --all, which from a
single-env push would deploy the other env + the shared IAM stack — breaking the
per-env boundary. So CD targets one stack explicitly per env. The shared
open-swe-iam stack is not deployed by CD (privileged, human-gated — T6).
Gating note: infra CI is enforced at the deploy boundary (cd-infra's
deploy-* jobs needs: ci), not as a branch-protection required check —
path-filtering a required check would deadlock app-only PRs (a skipped required
check never satisfies). Making Infra CI a required check later needs a
skip-aware shim or dropping its path filter.
Prerequisites (set post-T6, when the roles exist):
- repo variables
AWS_DEPLOY_ROLE_INFRA_DEV/AWS_DEPLOY_ROLE_INFRA_PROD= thegithubdeploy-open-swe-infra-<env>role ARNs (open-swe-iamoutputs). - a GitHub Environment named
prodwith Adam as a required reviewer.
App-side CD (CI → S3 artifact → SSM deploy via
githubdeploy-open-swe-app-<env>) is T19, not here.
Deploy ordering (when the gate clears — NOT yet)
open-swe-iamfirst — apply the IAM stack (T6, human-gated), then set the repoAWS_DEPLOY_ROLE_INFRA_{DEV,PROD}variables from its role-ARN outputs and configure theprodEnvironment reviewer (BLOCK#3).- Security gate — T4 GPT-4.1 IAM cross-review + T5
/sh-security-reviewon the synth; resolve every confirmed critical/high. - IAM applied (T6) — only after the gate.
- Env stacks: first
open-swe-dev(T14, manual validate), then CD auto-deploys dev on push;open-swe-prod(T21) behind theprodEnvironment approval.
Version policy
aws-cdk-lib is pinned EXACT (2.260.0) — no ^/~. Dependabot keeps it
current; CI (npm ci + cdk synth) + dependency review gate each bump. See
aws-infrastructure.md "CDK Version Policy" and memory
feedback_cdk_lib_bundled_deps.