open-swe/infra/lib/constructs/ami-cache.ts

36 lines
1.7 KiB
TypeScript
Raw Normal View History

import * as ec2 from "aws-cdk-lib/aws-ec2";
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
import { REGION } from "../config";
/**
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
* The baked open-swe base AMI (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW agent
* + boot templates — NO secrets), produced by `deploy/ami/open-swe-base.pkr.hcl`.
* Pinned by EXACT id (not a name filter) so synth/deploy is fully offline and
* deterministic.
*
* Built 2026-06-26 from open-swe-base-arm64-20260626-203433.
*
* ── EBS / AMI replacement discipline (memory feedback_inline_ebs_volumes) ──
*
* Refresh DELIBERATELY: `cd deploy/ami && packer build open-swe-base.pkr.hcl`,
* then update this id. A new id → EC2 instance REPLACEMENT. Pinning by exact id
* (vs a `most_recent` name filter) is what prevents a routine deploy from silently
* swapping the AMI — the root cause of the file-share data-loss incidents
* (5/15, 5/27, 6/5).
*
* `userDataCausesReplacement: true` (AppService) is likewise DELIBERATE: user-data
* is provisioning-only and the box holds NO durable state (the langgraph store is
* in-memory, rebuilt every boot from S3 + Secrets Manager / SSM), so there is
* intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The design
* goal is replacement-TOLERANCE, not avoidance.
*
* Operational guard before ANY replacing deploy (AMI / userData / instance-type):
* snapshot the root volume AND wait `state=completed`, re-verify "no local-only
* durable state", and review the `cdk diff` replacement at PR time.
*/
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
export const BAKED_OPEN_SWE_AMI_ID = "ami-00080084502093021";
feat: stand up dev properly — assets bucket + artifact CD + baked AMI + on-box uv sync (T7+T19+T14) (#18) * feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3) Packer-build the custom base image and repoint AppService off the AL2023 placeholder onto it. deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real `packer build` (the config had only ever been `packer validate`'d at T8): - the file provisioner failed uploading the templates dir ('scp: …: Is a directory') — a trailing-slash contents-upload needs the dest dir to exist; added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the dest trailing slash. - the shell provisioner's custom execute_command omitted {{ .Vars }}, so the environment_vars never reached provision.sh (which runs under set -u and aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}. infra: - ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26 from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by exact id via MachineImage.genericLinux (offline, deterministic). Dropped the now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement discipline docs. - app-service.ts: machineImage → bakedOpenSweArm64(). - open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard). - cdk.context.json → {} (AMI is a static id pin; no context lookups remain). - README: Baked AMI + EBS-replacement-discipline section. tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI. NOTE: held — do NOT merge until the open-swe-dev secret values are populated (put-config.sh). The infra CD is live, so merging this to dev auto-deploys OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts → unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14). * fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021 Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects non-ASCII in the AMI Description attribute, so packer registered then DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error. Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available). Re-pinned BAKED_OPEN_SWE_AMI_ID. * fix(deploy): GitHub App + Slack required for prod only, not dev Per the migration decision: do NOT create/duplicate a separate dev GitHub App or Slack app — only prod owns the single shared app. So fetch-config.sh no longer hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/ CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET. Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env (boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is unchanged (prod still requires everything). * feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19) Make the dev/prod box deployable end-to-end: a real artifact pipeline and a re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy. Infra (T7): - assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access, SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent + abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput. - app-service.ts: open-swe-<env>-deploy SSM document that runs the baked /opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the baked open-swe-base-arm64 AMI (folds in the held #16). IAM (app deploy role — cross-review gated): - github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox. Boot/deploy (T19): - deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx. - user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB target is healthy even before the first release); deploy.sh is base64-rendered by CDK into user-data (a normal reviewable repo file, not a heredoc) and the first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy). CI (T7+T19): - build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite -> ui/.output/public -> spa.tar.gz), package the Python source via git archive (app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy. push dev -> dev (auto); push main -> prod (env "prod" approval gate). Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK, deploy.sh base64 round-trips exact. * harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01, the account-wide CDK cfn-exec residual already documented in config.ts; recorded in .security-review/suppressions.json with justification + flagged for the per-env bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified; these are the cheap defense-in-depth fixes worth taking regardless: - deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root never honors an archive's uid/mode → no setuid/foreign-owned file can land); and treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy failure (set -e stays loud once a release exists). - publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount), not CommandInvocations[0], so a partial failure across the brief 2-instance replacement window can't be reported as success. - instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app role's write scope) instead of the whole bucket. - package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz (defense in depth over .gitignore; scoped to data extensions so *_credentials.py source is not a false positive — verified against the real tree). Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz). shellcheck/tsc/jest(16) clean; both stacks synth offline. * fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard The instance-SG GroupDescription + ingress/egress rule descriptions carried an em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"), so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing from #14; same class as the AMI-description ASCII bug.) - app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the ingress/egress rule descriptions, and the Route53 comment. - test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup GroupDescription + rule descriptions are pure ASCII, so this fails the build instead of a deploy next time. jest 18/18; tsc clean. * fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`) The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule* descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`, which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII only" to the exact EC2 allowed charset so it catches `>` (and `<`) too. jest 18/18; tsc clean. * fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit The base64 deploy.sh embedded in user-data pushed the encoded boot script to 27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with "Encoded User data is limited to 25600 bytes". Strip full-line comments + blank lines from deploy.sh before base64-embedding it (repo file keeps comments; only the on-box copy is minified; the script is opaque base64 so user-data heredocs are unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded. jest 19/19; minified deploy.sh passes bash -n + shellcheck.
2026-06-26 18:49:09 -04:00
/** The baked open-swe base image, pinned by id (offline, deterministic). */
export function bakedOpenSweArm64(): ec2.IMachineImage {
return ec2.MachineImage.genericLinux({ [REGION]: BAKED_OPEN_SWE_AMI_ID });
}