open-swe/deploy/ami/README.md
Adam Moussa ce9e5fb51b
feat(deploy): EC2 AMI recipe (Packer) + cloud-init/user-data (PR#2) (#7)
PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64,
uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data
(userDataCausesReplacement rationale), systemd unit + nginx + CW templates.

Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0);
nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre
runs fetch-config as root (+) and passes the env arg; the app runs as the
unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
2026-06-26 15:06:36 -04:00

9.2 KiB

Open SWE base AMI (T8)

Packer recipe + first-boot user-data for the single EC2 instance per env (open-swe-dev / open-swe-prod) in the Open SWE → AWS migration. Builds an ARM64 (Graviton) Ubuntu 24.04 LTS base AMI and provisions the box on first boot with the stock langgraph dev runtime, nginx, and the CloudWatch agent.

The architecture is locked in the repo TODO.md ("Architecture (locked)"): ONE EC2 ARM64 (~t4g.large) instance per env, seahaven-vpc private subnet + NAT, inbound only from the ALB SG. Runtime is stock langgraph dev (in-memory store, --no-reload) + nginx + systemd. The box has no git auth — it pulls its deploy artifact from S3 via the instance role. The SPA build runs in GitHub Actions (T7), not on the box, so the old 8 GB-swapfile OOM hack is gone.

Files

Path Purpose
open-swe-base.pkr.hcl Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI.
scripts/provision.sh Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates.
user-data.sh First-boot provisioning (S3 artifact pull, render templates, CW agent, start services).
templates/open-swe.service systemd unit TEMPLATE (@@tokens@@ rendered at boot).
templates/open-swe.nginx.conf nginx site TEMPLATE (dashboard SPA + scoped /dashboard/api/ proxy).
templates/amazon-cloudwatch-agent.json CW agent config TEMPLATE — 30-day log retention.

deploy/seahaven/fetch-config.sh and deploy/seahaven/seed_store.sh are owned by the parallel T10 work and ship inside the app artifact; this AMI wires them in but does not author them (see "Integration contract" below).

Build the AMI

cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl

Builds in account 328440206208 / us-east-1. Source = latest Canonical Ubuntu 24.04 (Noble) arm64 AMI (source_ami_filter, owner 099720109477). Build host is t4g.medium (ARM64). The output AMI is tagged:

Name=open-swe-base-arm64   Purpose=open-swe-runtime-base   ManagedBy=packer

Pinned versions live in the template variable defaults (uv_version, python_version, node_major, the CW-agent / awscli URLs) and the required_plugins block (amazon 1.3.6) — bump deliberately.

AMI → cdk.context.json pinning contract

The CDK stacks in /infra (owned by T3/T12) consume the AMI by id, pinned in the committed infra/cdk.context.json — they never resolve "latest" at synth time. This is the EBS/AMI-fix discipline: an uncached MachineImage.lookup resolves a new AMI on every deploy and silently triggers instance replacement.

Contract (CDK side does the wiring; this is the handshake):

  1. packer build prints the new AMI id (and tags it open-swe-base-arm64).
  2. CDK looks the AMI up with cachedInContext: true (e.g. MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })), which writes the resolved id into infra/cdk.context.json.
  3. infra/cdk.context.json is committed. From then on every synth/deploy uses the pinned id — no surprise replacement when a newer AMI exists.
  4. To adopt a new AMI: cdk context --reset <ami-lookup-key> (or edit the pinned value), commit the change, and review the cdk-diff — the PR will show "requires replacement", which is the intended, visible signal.

Record the built AMI id in project memory (project_open_swe_migration) per the "memory updated for AMI id" build criterion.

userDataCausesReplacement rationale

user-data.sh is provisioning-only — it runs once at first boot and never carries durable runtime config. CDK sets userDataCausesReplacement: true so that any change to it is a deliberate, diff-visible instance replacement rather than a no-op edit that drifts from the running box. Durable runtime config is fetched fresh on every service start by fetch-config.sh (ExecStartPre) — changing a secret or SSM value needs only a systemctl restart open-swe.service, not a replacement.

EBS discipline (binding — feedback_inline_ebs_volumes)

The box holds no durable state of its own:

State Lives in On replacement
secrets / config Secrets Manager + SSM → tmpfs .env re-fetched at boot
app code + SPA S3 open-swe-<env>-assets re-pulled at boot
store (team_settings, user_mappings) reseeded by seed_store.sh re-seeded at boot
logs CloudWatch (30-day) — not a CFN resource in the stack survive replacement

→ No local-only durable state ⇒ no standalone RETAIN volume is needed. The root volume is disposable; there is intentionally no inline data blockDevices to lose.

Even so, snapshot before any replacing deploy. Per the operational guard, before merging/deploying any change that REPLACES the instance (userDataCausesReplacement, AMI bump, instance-type change):

  1. Enumerate the instance's volumes and assert "no local-only durable state" (the table above is the checklist).
  2. Take an EBS snapshot of the root volume and WAIT for state=completed before letting the deploy proceed. Keep it as insurance; delete after a grace period.
  3. Confirm the CloudWatch log groups are not CFN-managed in the stack so history survives; re-verify history after the new instance is healthy.

cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance criterion). This is the enforced version — not just an assertion in the runbook.

Integration contract (T10 — fetch-config.sh + seed_store.sh)

Both ship in the app artifact under deploy/seahaven/ and are wired into the unit:

  • fetch-config.sh (ExecStartPre, runs as openswe): reads /etc/open-swe/boot.env (OPENSWE_ENV, AWS_REGION, SECRETS_PREFIX=open-swe-<env>, SSM_PREFIX=/open-swe-<env>, ENV_FILE=/run/open-swe/.env), pulls Secrets Manager open-swe-<env>/* + SSM /open-swe-<env>/*, and writes:
    • /run/open-swe/.env (0600, tmpfs, secret-bearing app env incl. the multiline GitHub App PEM) — loaded by langgraph/dotenv via the ${APP_DIR}/.env symlink.
    • /run/open-swe/seed.env (0600, tmpfs, simple OPENSWE_* vars only: OPENSWE_DEFAULT_REPO, OPENSWE_OWNER_LOGIN, OPENSWE_OWNER_EMAIL, model ids) — loaded by systemd EnvironmentFile so seed_store.sh (ExecStartPost) has them.
    • It must fail-fast (non-zero exit) if any required value is missing, so the unit never starts half-configured.
  • seed_store.sh (ExecStartPost): existing script, reseeds team_settings/default
    • user_mappings/<login> into the in-memory store after each start.

Smoke-boot checklist (after first boot)

SSM Session Manager onto the instance (no public SSH — private subnet) and verify:

  • cloud-init status --wait → done; /var/log/open-swe-user-data.log ends with "user-data done" and shows the S3 pulls + service starts.
  • systemctl is-active open-swe.service → active. (If it failed, check ExecStartPre/fetch-config.sh — fail-fast means missing config = failed unit.)
  • fetch-config fail-fast works: /run/open-swe/.env exists, owner openswe, mode 0600, on tmpfs (findmnt /run/open-swe); seed.env present.
  • curl -fsS http://127.0.0.1:2024/ok → 200 (raw LangGraph health).
  • systemctl is-active nginx → active; curl -fsS http://127.0.0.1/healthz → 200; curl -s http://127.0.0.1/threads returns the SPA shell, not JSON (proves the agent API is not proxied — the security boundary holds).
  • seed_store: done in the journal / app.log (store reseeded).
  • CloudWatch: log groups /open-swe/<env>/{app,user-data,nginx-access,nginx-error} exist with 30-day retention and are receiving events.
  • No swapfile (swapon --show empty) — the on-box SPA build is gone.
  • From the ALB only: dashboard host serves the SPA; hooks host reaches /webhooks/* on :2024 and nothing else (raw API paths hit the ALB default, not the box).

Assumptions

  • Artifact bucket open-swe-<env>-assets (T7), with objects ${ARTIFACT_PREFIX}/app.tar.gz (Python app incl. deploy/seahaven/ and a prebuilt arm64 .venv) and ${ARTIFACT_PREFIX}/spa.tar.gz (built SPA → /var/www/open-swe). ARTIFACT_PREFIX defaults to releases/latest; CDK renders the concrete value.
  • Instance role (defined in /infra, least-privilege per T4/T12) grants: s3:GetObject on open-swe-<env>-assets/*; secretsmanager:GetSecretValue on open-swe-<env>/*; ssm:GetParameter(s)/GetParametersByPath on /open-swe-<env>/*; logs:* for the CW agent log groups + cloudwatch:PutMetricData; SSM Session Manager (ssm:UpdateInstanceInformation, ssmmessages:*) for shell access.
  • CDK substitutes the @@OPENSWE_ENV@@, @@ASSETS_BUCKET@@, @@SERVER_NAME@@, @@ARTIFACT_PREFIX@@ tokens in user-data.sh when rendering the launch template.
  • :2024 binds 0.0.0.0 so the ALB hooks target group can reach /webhooks/*; it is reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is never internet-exposed — the ALB hooks rule is path-scoped to /webhooks/*.