* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
|
||
|---|---|---|
| .. | ||
| scripts | ||
| templates | ||
| deploy.sh | ||
| open-swe-base.pkr.hcl | ||
| README.md | ||
| user-data.sh | ||
Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(open-swe-dev / open-swe-prod) in the Open SWE → AWS migration. Builds an
ARM64 (Graviton) Ubuntu 24.04 LTS base AMI and provisions the box on first
boot with the stock langgraph dev runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo TODO.md ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, seahaven-vpc private subnet + NAT,
inbound only from the ALB SG. Runtime is stock langgraph dev (in-memory
store, --no-reload) + nginx + systemd. The box has no git auth — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), not on the box, so the old 8 GB-swapfile OOM hack is gone.
Files
| Path | Purpose |
|---|---|
open-swe-base.pkr.hcl |
Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
scripts/provision.sh |
Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
user-data.sh |
First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
templates/open-swe.service |
systemd unit TEMPLATE (@@tokens@@ rendered at boot). |
templates/open-swe.nginx.conf |
nginx site TEMPLATE (dashboard SPA + scoped /dashboard/api/ proxy). |
templates/amazon-cloudwatch-agent.json |
CW agent config TEMPLATE — 30-day log retention. |
deploy/seahaven/fetch-config.sh and deploy/seahaven/seed_store.sh are owned by
the parallel T10 work and ship inside the app artifact; this AMI wires them in
but does not author them (see "Integration contract" below).
Build the AMI
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
Builds in account 328440206208 / us-east-1. Source = latest Canonical Ubuntu
24.04 (Noble) arm64 AMI (source_ami_filter, owner 099720109477). Build host
is t4g.medium (ARM64). The output AMI is tagged:
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
Pinned versions live in the template variable defaults (uv_version,
python_version, node_major, the CW-agent / awscli URLs) and the
required_plugins block (amazon 1.3.6) — bump deliberately.
AMI → cdk.context.json pinning contract
The CDK stacks in /infra (owned by T3/T12) consume the AMI by id, pinned in the
committed infra/cdk.context.json — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached MachineImage.lookup resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
packer buildprints the new AMI id (and tags itopen-swe-base-arm64).- CDK looks the AMI up with
cachedInContext: true(e.g.MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })), which writes the resolved id intoinfra/cdk.context.json. infra/cdk.context.jsonis committed. From then on every synth/deploy uses the pinned id — no surprise replacement when a newer AMI exists.- To adopt a new AMI:
cdk context --reset <ami-lookup-key>(or edit the pinned value), commit the change, and review the cdk-diff — the PR will show "requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (project_open_swe_migration) per the
"memory updated for AMI id" build criterion.
userDataCausesReplacement rationale
user-data.sh is provisioning-only — it runs once at first boot and never
carries durable runtime config. CDK sets userDataCausesReplacement: true so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
fresh on every service start by fetch-config.sh (ExecStartPre) — changing a
secret or SSM value needs only a systemctl restart open-swe.service, not a
replacement.
EBS discipline (binding — feedback_inline_ebs_volumes)
The box holds no durable state of its own:
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs .env |
re-fetched at boot |
| app code + SPA | S3 open-swe-<env>-assets |
re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by seed_store.sh |
re-seeded at boot |
| logs | CloudWatch (30-day) — not a CFN resource in the stack | survive replacement |
→ No local-only durable state ⇒ no standalone RETAIN volume is needed. The root
volume is disposable; there is intentionally no inline data blockDevices to lose.
Even so, snapshot before any replacing deploy. Per the operational guard, before
merging/deploying any change that REPLACES the instance (userDataCausesReplacement,
AMI bump, instance-type change):
- Enumerate the instance's volumes and assert "no local-only durable state" (the table above is the checklist).
- Take an EBS snapshot of the root volume and WAIT for
state=completedbefore letting the deploy proceed. Keep it as insurance; delete after a grace period. - Confirm the CloudWatch log groups are not CFN-managed in the stack so history survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance criterion). This is the enforced version — not just an assertion in the runbook.
Integration contract (T10 — fetch-config.sh + seed_store.sh)
Both ship in the app artifact under deploy/seahaven/ and are wired into the unit:
fetch-config.sh(ExecStartPre, runs asopenswe): reads/etc/open-swe/boot.env(OPENSWE_ENV,AWS_REGION,SECRETS_PREFIX=open-swe-<env>,SSM_PREFIX=/open-swe-<env>,ENV_FILE=/run/open-swe/.env), pulls Secrets Manageropen-swe-<env>/*+ SSM/open-swe-<env>/*, and writes:/run/open-swe/.env(0600, tmpfs, secret-bearing app env incl. the multiline GitHub App PEM) — loaded by langgraph/dotenv via the${APP_DIR}/.envsymlink./run/open-swe/seed.env(0600, tmpfs, simpleOPENSWE_*vars only:OPENSWE_DEFAULT_REPO,OPENSWE_OWNER_LOGIN,OPENSWE_OWNER_EMAIL, model ids) — loaded by systemdEnvironmentFilesoseed_store.sh(ExecStartPost) has them.- It must fail-fast (non-zero exit) if any required value is missing, so the unit never starts half-configured.
seed_store.sh(ExecStartPost): existing script, reseedsteam_settings/defaultuser_mappings/<login>into the in-memory store after each start.
Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
cloud-init status --wait→done;/var/log/open-swe-user-data.logends with "user-data done" and shows the S3 pulls + service starts.systemctl is-active open-swe.service→active. (If it failed, checkExecStartPre/fetch-config.sh— fail-fast means missing config = failed unit.)- fetch-config fail-fast works:
/run/open-swe/.envexists, owneropenswe, mode0600, on tmpfs (findmnt /run/open-swe);seed.envpresent. curl -fsS http://127.0.0.1:2024/ok→200(raw LangGraph health).systemctl is-active nginx→active;curl -fsS http://127.0.0.1/healthz→200;curl -s http://127.0.0.1/threadsreturns the SPA shell, not JSON (proves the agent API is not proxied — the security boundary holds).seed_store: donein the journal / app.log (store reseeded).- CloudWatch: log groups
/open-swe/<env>/{app,user-data,nginx-access,nginx-error}exist with 30-day retention and are receiving events. - No swapfile (
swapon --showempty) — the on-box SPA build is gone. - From the ALB only: dashboard host serves the SPA;
hookshost reaches/webhooks/*on :2024 and nothing else (raw API paths hit the ALB default, not the box).
Assumptions
- Artifact bucket
open-swe-<env>-assets(T7), with objects${ARTIFACT_PREFIX}/app.tar.gz(Python app incl.deploy/seahaven/and a prebuilt arm64.venv) and${ARTIFACT_PREFIX}/spa.tar.gz(built SPA →/var/www/open-swe).ARTIFACT_PREFIXdefaults toreleases/latest; CDK renders the concrete value. - Instance role (defined in
/infra, least-privilege per T4/T12) grants:s3:GetObjectonopen-swe-<env>-assets/*;secretsmanager:GetSecretValueonopen-swe-<env>/*;ssm:GetParameter(s)/GetParametersByPathon/open-swe-<env>/*;logs:*for the CW agent log groups +cloudwatch:PutMetricData; SSM Session Manager (ssm:UpdateInstanceInformation,ssmmessages:*) for shell access. - CDK substitutes the
@@OPENSWE_ENV@@,@@ASSETS_BUCKET@@,@@SERVER_NAME@@,@@ARTIFACT_PREFIX@@tokens inuser-data.shwhen rendering the launch template. :2024binds0.0.0.0so the ALB hooks target group can reach/webhooks/*; it is reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is never internet-exposed — the ALB hooks rule is path-scoped to/webhooks/*.