PR#2 of the AWS migration. deploy/ami/: Packer template (Ubuntu 24.04 arm64, uv+py3.12, nginx, awscli v2, CW agent; no swapfile), provisioning-only user-data (userDataCausesReplacement rationale), systemd unit + nginx + CW templates. Incorporates T5 /sh-security-review fixes: langgraph binds 127.0.0.1 (not 0.0.0.0); nginx is the sole ingress proxying only /dashboard/api/ + /webhooks/; ExecStartPre runs fetch-config as root (+) and passes the env arg; the app runs as the unprivileged openswe user reading an openswe-owned 0600 .env. packer validate clean.
9.2 KiB
Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(open-swe-dev / open-swe-prod) in the Open SWE → AWS migration. Builds an
ARM64 (Graviton) Ubuntu 24.04 LTS base AMI and provisions the box on first
boot with the stock langgraph dev runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo TODO.md ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, seahaven-vpc private subnet + NAT,
inbound only from the ALB SG. Runtime is stock langgraph dev (in-memory
store, --no-reload) + nginx + systemd. The box has no git auth — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), not on the box, so the old 8 GB-swapfile OOM hack is gone.
Files
| Path | Purpose |
|---|---|
open-swe-base.pkr.hcl |
Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
scripts/provision.sh |
Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
user-data.sh |
First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
templates/open-swe.service |
systemd unit TEMPLATE (@@tokens@@ rendered at boot). |
templates/open-swe.nginx.conf |
nginx site TEMPLATE (dashboard SPA + scoped /dashboard/api/ proxy). |
templates/amazon-cloudwatch-agent.json |
CW agent config TEMPLATE — 30-day log retention. |
deploy/seahaven/fetch-config.sh and deploy/seahaven/seed_store.sh are owned by
the parallel T10 work and ship inside the app artifact; this AMI wires them in
but does not author them (see "Integration contract" below).
Build the AMI
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
Builds in account 328440206208 / us-east-1. Source = latest Canonical Ubuntu
24.04 (Noble) arm64 AMI (source_ami_filter, owner 099720109477). Build host
is t4g.medium (ARM64). The output AMI is tagged:
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
Pinned versions live in the template variable defaults (uv_version,
python_version, node_major, the CW-agent / awscli URLs) and the
required_plugins block (amazon 1.3.6) — bump deliberately.
AMI → cdk.context.json pinning contract
The CDK stacks in /infra (owned by T3/T12) consume the AMI by id, pinned in the
committed infra/cdk.context.json — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached MachineImage.lookup resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
packer buildprints the new AMI id (and tags itopen-swe-base-arm64).- CDK looks the AMI up with
cachedInContext: true(e.g.MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })), which writes the resolved id intoinfra/cdk.context.json. infra/cdk.context.jsonis committed. From then on every synth/deploy uses the pinned id — no surprise replacement when a newer AMI exists.- To adopt a new AMI:
cdk context --reset <ami-lookup-key>(or edit the pinned value), commit the change, and review the cdk-diff — the PR will show "requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (project_open_swe_migration) per the
"memory updated for AMI id" build criterion.
userDataCausesReplacement rationale
user-data.sh is provisioning-only — it runs once at first boot and never
carries durable runtime config. CDK sets userDataCausesReplacement: true so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
fresh on every service start by fetch-config.sh (ExecStartPre) — changing a
secret or SSM value needs only a systemctl restart open-swe.service, not a
replacement.
EBS discipline (binding — feedback_inline_ebs_volumes)
The box holds no durable state of its own:
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs .env |
re-fetched at boot |
| app code + SPA | S3 open-swe-<env>-assets |
re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by seed_store.sh |
re-seeded at boot |
| logs | CloudWatch (30-day) — not a CFN resource in the stack | survive replacement |
→ No local-only durable state ⇒ no standalone RETAIN volume is needed. The root
volume is disposable; there is intentionally no inline data blockDevices to lose.
Even so, snapshot before any replacing deploy. Per the operational guard, before
merging/deploying any change that REPLACES the instance (userDataCausesReplacement,
AMI bump, instance-type change):
- Enumerate the instance's volumes and assert "no local-only durable state" (the table above is the checklist).
- Take an EBS snapshot of the root volume and WAIT for
state=completedbefore letting the deploy proceed. Keep it as insurance; delete after a grace period. - Confirm the CloudWatch log groups are not CFN-managed in the stack so history survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance criterion). This is the enforced version — not just an assertion in the runbook.
Integration contract (T10 — fetch-config.sh + seed_store.sh)
Both ship in the app artifact under deploy/seahaven/ and are wired into the unit:
fetch-config.sh(ExecStartPre, runs asopenswe): reads/etc/open-swe/boot.env(OPENSWE_ENV,AWS_REGION,SECRETS_PREFIX=open-swe-<env>,SSM_PREFIX=/open-swe-<env>,ENV_FILE=/run/open-swe/.env), pulls Secrets Manageropen-swe-<env>/*+ SSM/open-swe-<env>/*, and writes:/run/open-swe/.env(0600, tmpfs, secret-bearing app env incl. the multiline GitHub App PEM) — loaded by langgraph/dotenv via the${APP_DIR}/.envsymlink./run/open-swe/seed.env(0600, tmpfs, simpleOPENSWE_*vars only:OPENSWE_DEFAULT_REPO,OPENSWE_OWNER_LOGIN,OPENSWE_OWNER_EMAIL, model ids) — loaded by systemdEnvironmentFilesoseed_store.sh(ExecStartPost) has them.- It must fail-fast (non-zero exit) if any required value is missing, so the unit never starts half-configured.
seed_store.sh(ExecStartPost): existing script, reseedsteam_settings/defaultuser_mappings/<login>into the in-memory store after each start.
Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
cloud-init status --wait→done;/var/log/open-swe-user-data.logends with "user-data done" and shows the S3 pulls + service starts.systemctl is-active open-swe.service→active. (If it failed, checkExecStartPre/fetch-config.sh— fail-fast means missing config = failed unit.)- fetch-config fail-fast works:
/run/open-swe/.envexists, owneropenswe, mode0600, on tmpfs (findmnt /run/open-swe);seed.envpresent. curl -fsS http://127.0.0.1:2024/ok→200(raw LangGraph health).systemctl is-active nginx→active;curl -fsS http://127.0.0.1/healthz→200;curl -s http://127.0.0.1/threadsreturns the SPA shell, not JSON (proves the agent API is not proxied — the security boundary holds).seed_store: donein the journal / app.log (store reseeded).- CloudWatch: log groups
/open-swe/<env>/{app,user-data,nginx-access,nginx-error}exist with 30-day retention and are receiving events. - No swapfile (
swapon --showempty) — the on-box SPA build is gone. - From the ALB only: dashboard host serves the SPA;
hookshost reaches/webhooks/*on :2024 and nothing else (raw API paths hit the ALB default, not the box).
Assumptions
- Artifact bucket
open-swe-<env>-assets(T7), with objects${ARTIFACT_PREFIX}/app.tar.gz(Python app incl.deploy/seahaven/and a prebuilt arm64.venv) and${ARTIFACT_PREFIX}/spa.tar.gz(built SPA →/var/www/open-swe).ARTIFACT_PREFIXdefaults toreleases/latest; CDK renders the concrete value. - Instance role (defined in
/infra, least-privilege per T4/T12) grants:s3:GetObjectonopen-swe-<env>-assets/*;secretsmanager:GetSecretValueonopen-swe-<env>/*;ssm:GetParameter(s)/GetParametersByPathon/open-swe-<env>/*;logs:*for the CW agent log groups +cloudwatch:PutMetricData; SSM Session Manager (ssm:UpdateInstanceInformation,ssmmessages:*) for shell access. - CDK substitutes the
@@OPENSWE_ENV@@,@@ASSETS_BUCKET@@,@@SERVER_NAME@@,@@ARTIFACT_PREFIX@@tokens inuser-data.shwhen rendering the launch template. :2024binds0.0.0.0so the ALB hooks target group can reach/webhooks/*; it is reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is never internet-exposed — the ALB hooks rule is path-scoped to/webhooks/*.