open-swe/infra/lib/constructs/app-service.ts

369 lines
16 KiB
TypeScript
Raw Normal View History

feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
import * as fs from "fs";
import * as path from "path";
import * as cdk from "aws-cdk-lib";
import * as ec2 from "aws-cdk-lib/aws-ec2";
import * as elbv2 from "aws-cdk-lib/aws-elasticloadbalancingv2";
import * as elbTargets from "aws-cdk-lib/aws-elasticloadbalancingv2-targets";
import * as logs from "aws-cdk-lib/aws-logs";
import * as route53 from "aws-cdk-lib/aws-route53";
import * as iam from "aws-cdk-lib/aws-iam";
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
import * as ssm from "aws-cdk-lib/aws-ssm";
import { Construct, IConstruct } from "constructs";
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
import { EnvName, prefix } from "../config";
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
import { bakedOpenSweArm64 } from "./ami-cache";
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
/**
* Shared seahaven-vpc + internet-facing ALB facts (read-only recon 2026-06-26;
* scratchpad/T12-infra-facts.md). A SINGLE VPC and a SINGLE shared ALB front
* both the on-prem `seahaven-site` stack and open-swe. We IMPORT every one of
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
* these and NEVER own them - open-swe only ADDS its own instance SG, a standalone
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
* ALB-egress rule, listener rules, a target group, and DNS records.
*
* ── Cross-stack coordination on the SHARED listener + ALB SG (T13 review) ──
* Two CDK stacks (seahaven-site, open-swe) add resources to the same imported
* `:443` listener and ALB SG. This is safe because each stack owns ONLY the
* resources it declares (its own logical ids): an on-prem `cdk deploy` computes a
* changeset over its own template and cannot delete rules/egress it never
* declared. The standalone-egress pattern is what on-prem itself uses
* (sgr-0c57812752a3bca13), so it does not mutate the shared SG's own definition.
*
* The ONE shared namespace that REQUIRES coordination is listener-rule PRIORITY
* (globally unique per listener; a collision is a fail-SAFE deploy error, not
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
* silent drift). Ownership map - keep these disjoint when editing either stack:
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
* - seahaven-site (on-prem): priorities 4-7 + default.
* - open-swe: priorities 2, 3 (webhooks) and 10, 11 (site).
* open-swe's webhook rules are HOST-scoped to its own *.seahaven.com hosts, so
* they never match (let alone "steal") any seahavenind.com / on-prem host.
*/
const SHARED = {
vpcId: "vpc-0d3d4b67bd0cf8a68",
availabilityZones: ["us-east-1a", "us-east-1b"],
// Private subnets host the EC2 box. The single NAT gateway lives in 1a, so the
// box is pinned to private1 (1a) for in-AZ NAT egress (no cross-AZ data $).
privateSubnetIds: ["subnet-04e38c507e96f1926", "subnet-0a0b4fc6f296dfba5"],
instanceSubnetId: "subnet-04e38c507e96f1926",
instanceAz: "us-east-1a",
albDnsName: "seahaven-com-1856441924.us-east-1.elb.amazonaws.com",
albCanonicalHostedZoneId: "Z35SXDOTRQ7X7K",
albSecurityGroupId: "sg-0b0301deed193258a",
httpsListenerArn:
"arn:aws:elasticloadbalancing:us-east-1:328440206208:listener/app/seahaven-com/222c3257354ab559/bab8bcf0da0e2927",
publicZoneId: "Z06652411XKH89KTZD3XA",
publicZoneName: "seahaven.com",
} as const;
/**
* Per-env public hostnames, listener-rule priorities, and instance size.
*
* ── Listener-rule ordering hazard (load-bearing) ──
* The shared listener already has a HOST-AGNOSTIC `/webhooks/*` PATH rule at
* priority 5 (the on-prem seahaven-site stack owns it). ALB rules are first-match
* by ASCENDING priority, so a `…/webhooks/*` request to our host would match
* rule 5 (priority 5) and be forwarded to the on-prem target BEFORE any host rule
* at 10+. Therefore our webhook rule MUST sit below priority 5. The catch-all
* "site" rule (dashboard SPA + /dashboard/api/) carries no path that collides
* with rule 5, so it can sit at any free higher number (10/11). Free priorities
* confirmed by recon: 1-3 and 8+ (4=forgejo, 5=/webhooks/*, 6/7=seahavenind).
*/
const ENV_NET: Record<
EnvName,
{
dashboardHost: string;
hooksHost: string;
webhookPriority: number;
sitePriority: number;
instanceType: string;
}
> = {
dev: {
dashboardHost: "openswe-dev.seahaven.com",
hooksHost: "hooks-dev.seahaven.com",
webhookPriority: 2,
sitePriority: 10,
instanceType: "t4g.medium",
},
prod: {
dashboardHost: "openswe.seahaven.com",
hooksHost: "hooks.seahaven.com",
webhookPriority: 3,
sitePriority: 11,
instanceType: "t4g.large",
},
};
export interface AppServiceProps {
readonly envName: EnvName;
/** Least-privilege EC2 instance role (per-env; from InstanceRole). */
readonly instanceRole: iam.IRole;
/** S3 artifact key prefix the box pulls app.tar.gz / spa.tar.gz from. */
readonly artifactPrefix?: string;
}
/**
* The open-swe compute + ingress wiring for one env (T12):
* - one ARM64 EC2 box in private1 (1a), replacement-tolerant (no RETAIN volume),
* - a standalone instance SG reachable ONLY from the shared ALB SG on :80,
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
* - a target group -> instance:80 (nginx is the sole ingress; :2024 stays loopback),
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
* - two listener rules on the imported :443 listener (webhooks below the on-prem
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
* path rule; site catch-all above it), both -> the same TG,
* - Route53 alias records for both hostnames -> the shared ALB,
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
* - IaC-owned CloudWatch log groups at 30-day retention.
*
* Everything ALB/VPC/zone-side is IMPORTED. Synth is offline: the AMI is the
* cdk.context.json-pinned AL2023 ARM64 placeholder until the baked
* open-swe-base-arm64 id is pinned before the first real deploy.
*/
export class AppService extends Construct {
public readonly instance: ec2.Instance;
public readonly targetGroup: elbv2.ApplicationTargetGroup;
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
/** Name of the SSM document CI fires to roll the box to the latest release. */
public readonly deployDocumentName: string;
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
constructor(scope: Construct, id: string, props: AppServiceProps) {
super(scope, id);
const env = props.envName;
const p = prefix(env);
const net = ENV_NET[env];
const artifactPrefix = props.artifactPrefix ?? "releases/latest";
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// Import the shared VPC with explicit attributes (no fromLookup -> offline synth).
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
const vpc = ec2.Vpc.fromVpcAttributes(this, "Vpc", {
vpcId: SHARED.vpcId,
availabilityZones: [...SHARED.availabilityZones],
privateSubnetIds: [...SHARED.privateSubnetIds],
});
// Standalone instance SG. Egress open (NAT path); ingress only from the ALB SG.
const instanceSg = new ec2.SecurityGroup(this, "InstanceSg", {
vpc,
securityGroupName: `${p}-instance-sg`,
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
description: `${p} instance SG - ingress only from the shared ALB SG on :80; egress via NAT.`,
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
allowAllOutbound: true,
});
instanceSg.addIngressRule(
ec2.Peer.securityGroupId(SHARED.albSecurityGroupId),
ec2.Port.tcp(80),
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
`${p}: shared ALB SG to nginx :80`,
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
);
// Open the IMPORTED ALB SG to our instance via a STANDALONE egress rule, so we
// never mutate the ALB SG's own (on-prem-owned) definition.
new ec2.CfnSecurityGroupEgress(this, "AlbToInstanceEgress", {
groupId: SHARED.albSecurityGroupId,
ipProtocol: "tcp",
fromPort: 80,
toPort: 80,
destinationSecurityGroupId: instanceSg.securityGroupId,
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
description: `${p}: ALB to instance nginx :80`,
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
});
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// The app-deploy procedure (deploy/ami/deploy.sh) is a normal reviewable repo
// file; CDK base64-encodes it (single line - no `$`/regex-special chars in the
// base64 alphabet) and renders it into user-data's @@DEPLOY_SH_B64@@ token, so
// user-data writes it verbatim to /opt/open-swe/bin/deploy.sh at first boot.
// The same file is run by the open-swe-<env>-deploy SSM document on every
// release - a single source of truth for "pull release, build venv, restart".
const deployShPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "deploy.sh");
// Minify before embedding: strip full-line comments + blank lines (keep the
// shebang) so the base64 fits EC2's 25.6 KB user-data limit. The repo file
// keeps its comments; only the on-box copy is minified. deploy.sh becomes
// opaque base64 here, so this never affects user-data's heredoc parsing.
const deployShMin = fs
.readFileSync(deployShPath, "utf8")
.split("\n")
.filter((line, i) => i === 0 || (!/^\s*#/.test(line) && line.trim() !== ""))
.join("\n");
const deployShB64 = Buffer.from(deployShMin, "utf8").toString("base64");
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
// Render the provisioning script's @@tokens@@ into the instance user-data.
// userDataCausesReplacement makes a bootstrap change roll a fresh box (the box
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// holds no durable state - see ami-cache.ts / user-data.sh). Editing deploy.sh
// therefore also rolls the box (its base64 is embedded here) - acceptable: the
// box is replacement-tolerant, and ongoing releases never touch user-data.
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
const userDataPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "user-data.sh");
const userData = ec2.UserData.custom(
fs
.readFileSync(userDataPath, "utf8")
// %%...%% tokens are CDK-substituted here; they are DELIBERATELY a
// different delimiter from the @@...@@ tokens user-data.sh seds into the
// baked systemd/nginx templates, so CDK can never clobber a sed pattern
// (a shared @@OPENSWE_ENV@@/@@SERVER_NAME@@ left the unit unsubstituted).
.replace(/%%OPENSWE_ENV%%/g, env)
.replace(/%%ASSETS_BUCKET%%/g, `${p}-assets`)
.replace(/%%SERVER_NAME%%/g, net.dashboardHost)
.replace(/%%ARTIFACT_PREFIX%%/g, artifactPrefix)
.replace(/%%DEPLOY_SH_B64%%/g, deployShB64),
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
);
this.instance = new ec2.Instance(this, "Instance", {
vpc,
vpcSubnets: {
subnets: [
ec2.Subnet.fromSubnetAttributes(this, "InstanceSubnet", {
subnetId: SHARED.instanceSubnetId,
availabilityZone: SHARED.instanceAz,
}),
],
},
instanceType: new ec2.InstanceType(net.instanceType),
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// The baked open-swe base AMI (deploy/ami packer build) - ARM64 Ubuntu 24.04
// with the /opt/open-swe layout, openswe user, nginx, and CW agent that
// user-data.sh assumes. Pinned by exact id (see ami-cache.ts); refresh by
// rebuilding and updating BAKED_OPEN_SWE_AMI_ID.
machineImage: bakedOpenSweArm64(),
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
role: props.instanceRole,
securityGroup: instanceSg,
userData,
userDataCausesReplacement: true,
requireImdsv2: true,
instanceName: `${p}-box`,
blockDevices: [
{
deviceName: "/dev/xvda",
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// gp3 encrypted root; deleteOnTermination (no durable on-box state ->
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
// intentionally NO standalone RETAIN volume; see ami-cache.ts).
volume: ec2.BlockDeviceVolume.ebs(30, {
volumeType: ec2.EbsDeviceVolumeType.GP3,
encrypted: true,
deleteOnTermination: true,
}),
},
],
});
// `requireImdsv2: true` makes CDK auto-create a launch template, and it names
// that LT from the construct id ("Instance" -> "InstanceLaunchTemplate") with NO
// env qualifier — so OpenSweDevStack and OpenSweProdStack both want the identical
// LT name and the second env to deploy fails with
// InvalidLaunchTemplateName.AlreadyExistsException (prod rollback, 2026-06-29).
// Force a per-env LT name. Done via an aspect because the LT is created at synth
// time by the requireImdsv2 handling, not in this constructor.
cdk.Aspects.of(this.instance).add({
visit(node: IConstruct) {
if (node instanceof ec2.CfnLaunchTemplate) {
node.launchTemplateName = `${p}-lt`;
}
// The instance references the LT BY NAME, so the reference must be renamed in
// lockstep (preserve the GetAtt version) or CFN can't find the template.
if (node instanceof ec2.CfnInstance && node.launchTemplate) {
const spec = node.launchTemplate as ec2.CfnInstance.LaunchTemplateSpecificationProperty;
node.launchTemplate = { ...spec, launchTemplateName: `${p}-lt` };
}
},
});
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// SSM deploy document (open-swe-<env>-deploy): runs the baked
// /opt/open-swe/bin/deploy.sh to pull the latest release + restart. CI fires it
// (tag-scoped to project=open-swe,env=<env>) after uploading a release, so the
// app deploy role needs SendCommand ONLY on this document - NOT on the generic
// AWS-RunShellScript (closes the T4 BLOCK#3 arbitrary-shell timebox).
this.deployDocumentName = `${p}-deploy`;
new ssm.CfnDocument(this, "DeployDoc", {
name: this.deployDocumentName,
documentType: "Command",
documentFormat: "YAML",
updateMethod: "NewVersion",
content: {
schemaVersion: "2.2",
description: `Roll the ${p} box to the latest published release (runs /opt/open-swe/bin/deploy.sh).`,
mainSteps: [
{
action: "aws:runShellScript",
name: "deploy",
inputs: {
// Fixed command - no parameters, so nothing untrusted is interpolated
// into the shell. The script itself reads /etc/open-swe/boot.env.
runCommand: ["bash /opt/open-swe/bin/deploy.sh"],
},
},
],
},
});
// Target group -> instance:80 (nginx). Health check hits nginx's /healthz
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
// (returns 200; the dashboard TG health path defined in open-swe.nginx.conf).
this.targetGroup = new elbv2.ApplicationTargetGroup(this, "Tg", {
vpc,
targetGroupName: `${p}-tg`,
port: 80,
protocol: elbv2.ApplicationProtocol.HTTP,
targetType: elbv2.TargetType.INSTANCE,
targets: [new elbTargets.InstanceTarget(this.instance)],
deregistrationDelay: cdk.Duration.seconds(15),
healthCheck: {
path: "/healthz",
healthyHttpCodes: "200",
interval: cdk.Duration.seconds(30),
timeout: cdk.Duration.seconds(5),
healthyThresholdCount: 2,
unhealthyThresholdCount: 3,
},
});
// Import the shared :443 listener (with its ALB SG) and ADD our two rules.
const albSg = ec2.SecurityGroup.fromSecurityGroupId(this, "AlbSg", SHARED.albSecurityGroupId, {
mutable: false,
});
const listener = elbv2.ApplicationListener.fromApplicationListenerAttributes(this, "HttpsListener", {
listenerArn: SHARED.httpsListenerArn,
securityGroup: albSg,
});
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// (1) Webhooks - accepted on EITHER host (integrations may target either), and
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
// MUST be below the on-prem path-only rule 5 (see ENV_NET note).
new elbv2.ApplicationListenerRule(this, "WebhooksRule", {
listener,
priority: net.webhookPriority,
conditions: [
elbv2.ListenerCondition.hostHeaders([net.dashboardHost, net.hooksHost]),
elbv2.ListenerCondition.pathPatterns(["/webhooks/*"]),
],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// (2) Dashboard SPA + /dashboard/api/ (OAuth) - DASHBOARD host ONLY. The hooks
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
// host intentionally serves nothing but /webhooks/* (rule 1), so the OAuth /
// dashboard surface stays single-origin (OSWE-T12-02). Non-webhook paths on the
// hooks host fall through to the on-prem default.
new elbv2.ApplicationListenerRule(this, "SiteRule", {
listener,
priority: net.sitePriority,
conditions: [elbv2.ListenerCondition.hostHeaders([net.dashboardHost])],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// Route53 ALIAS records -> the shared ALB, for both hostnames.
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
const zone = route53.HostedZone.fromHostedZoneAttributes(this, "PublicZone", {
hostedZoneId: SHARED.publicZoneId,
zoneName: SHARED.publicZoneName,
});
const albAlias: route53.IAliasRecordTarget = {
bind: () => ({
dnsName: SHARED.albDnsName,
hostedZoneId: SHARED.albCanonicalHostedZoneId,
}),
};
for (const [label, host] of [
["Dashboard", net.dashboardHost],
["Hooks", net.hooksHost],
] as const) {
new route53.ARecord(this, `${label}Alias`, {
zone,
recordName: host,
target: route53.RecordTarget.fromAlias(albAlias),
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
comment: `${p} ${label.toLowerCase()} -> shared seahaven-com ALB`,
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
});
}
// IaC-owned CloudWatch log groups at 30-day retention. Names mirror the
// CloudWatch-agent config (deploy/ami/templates/amazon-cloudwatch-agent.json);
// owning them here makes retention declarative rather than agent-set. Logs are
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
// not durable state -> DESTROY on stack delete.
feat: open-swe dev/prod compute + ALB ingress (T12) (#14) AppService construct wires the per-env EC2 box and its internet path. The seahaven-vpc and the internet-facing seahaven-com ALB are SHARED with the on-prem seahaven-site stack, so everything VPC/ALB/zone-side is IMPORTED and never owned/mutated; open-swe only ADDS its own resources. Per env (open-swe-stack.ts → AppService): - ARM64 EC2 box (t4g.medium dev / t4g.large prod) in private1 (us-east-1a, in-AZ NAT egress). requireImdsv2, gp3-encrypted root, deleteOnTermination (no RETAIN volume — replacement-tolerant; see ami-cache.ts). userDataCausesReplacement; user-data rendered from deploy/ami/user-data.sh. - Standalone instance SG: ingress ONLY from the shared ALB SG on :80; egress via NAT. The ALB SG is opened to the box via a STANDALONE CfnSecurityGroupEgress (the imported, on-prem-owned SG is never mutated). - Target group → instance:80 (nginx is sole ingress; LangGraph :2024 stays loopback). Health check GET /healthz. - Two rules on the imported :443 listener, both → the TG: * webhooks (priority 2 dev / 3 prod): host∈{openswe,hooks}-<env> AND /webhooks/* * site (priority 10 dev / 11 prod): host=openswe-<env> (dashboard SPA + api) Webhooks MUST sit below the on-prem host-agnostic /webhooks/* rule (priority 5) or it would steal every webhook — first-match-by-ascending-priority. - Route53 alias records (openswe[-dev] + hooks[-dev]) → shared ALB. - 4 CloudWatch log groups at 30-day retention (IaC-owned; mirrors CW-agent config). Security (/sh-security-review T12): iac-iam pass clean. Logic pass → 1 confirmed medium fixed (OSWE-T12-01: nginx 1MB default client_max_body_size would 413 large GitHub webhooks pre-signature-verification → set 25m on /webhooks/, 10m on /dashboard/api/); hooks host scoped to /webhooks/* only (OSWE-T12-02 hygiene); XFF-spoof candidate killed (no code trusts leftmost XFF). No confirmed critical/high. Synth-only; not deployed. AMI is the cdk.context.json placeholder until the baked open-swe-base-arm64 id is pinned pre-deploy. tsc/synth(dev+prod)/jest(16) clean. Next: T13 GPT-4.1 cross-review of the SG/listener diff before any deploy.
2026-06-26 16:10:23 -04:00
for (const suffix of ["app", "user-data", "nginx-access", "nginx-error"]) {
new logs.LogGroup(this, `Log-${suffix}`, {
logGroupName: `/open-swe/${env}/${suffix}`,
retention: logs.RetentionDays.ONE_MONTH,
removalPolicy: cdk.RemovalPolicy.DESTROY,
});
}
}
}