# Open SWE base AMI (T8) Packer recipe + first-boot user-data for the single EC2 instance per env (`open-swe-dev` / `open-swe-prod`) in the Open SWE → AWS migration. Builds an **ARM64 (Graviton) Ubuntu 24.04 LTS** base AMI and provisions the box on first boot with the stock `langgraph dev` runtime, nginx, and the CloudWatch agent. The architecture is locked in the repo `TODO.md` ("Architecture (locked)"): ONE EC2 ARM64 (~t4g.large) instance per env, `seahaven-vpc` **private subnet + NAT**, inbound **only from the ALB SG**. Runtime is **stock `langgraph dev`** (in-memory store, `--no-reload`) + nginx + systemd. The box has **no git auth** — it pulls its deploy artifact from S3 via the instance role. The SPA build runs in GitHub Actions (T7), **not** on the box, so the old 8 GB-swapfile OOM hack is gone. ## Files | Path | Purpose | |---|---| | `open-swe-base.pkr.hcl` | Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. | | `scripts/provision.sh` | Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. | | `user-data.sh` | First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). | | `templates/open-swe.service` | systemd unit TEMPLATE (`@@tokens@@` rendered at boot). | | `templates/open-swe.nginx.conf` | nginx site TEMPLATE (dashboard SPA + scoped `/dashboard/api/` proxy). | | `templates/amazon-cloudwatch-agent.json` | CW agent config TEMPLATE — **30-day log retention**. | `deploy/seahaven/fetch-config.sh` and `deploy/seahaven/seed_store.sh` are owned by the parallel T10 work and ship **inside the app artifact**; this AMI wires them in but does not author them (see "Integration contract" below). ## Build the AMI ```bash cd deploy/ami packer init . packer fmt -check . packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl packer build open-swe-base.pkr.hcl ``` Builds in account **328440206208 / us-east-1**. Source = latest Canonical Ubuntu 24.04 (Noble) **arm64** AMI (`source_ami_filter`, owner `099720109477`). Build host is `t4g.medium` (ARM64). The output AMI is tagged: ``` Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer ``` Pinned versions live in the template `variable` defaults (`uv_version`, `python_version`, `node_major`, the CW-agent / awscli URLs) and the `required_plugins` block (`amazon` 1.3.6) — bump deliberately. ## AMI → `cdk.context.json` pinning contract The CDK stacks in `/infra` (owned by T3/T12) consume the AMI **by id, pinned in the committed `infra/cdk.context.json`** — they never resolve "latest" at synth time. This is the EBS/AMI-fix discipline: an uncached `MachineImage.lookup` resolves a new AMI on every deploy and silently triggers instance replacement. Contract (CDK side does the wiring; this is the handshake): 1. `packer build` prints the new AMI id (and tags it `open-swe-base-arm64`). 2. CDK looks the AMI up with **`cachedInContext: true`** (e.g. `MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })`), which writes the resolved id into `infra/cdk.context.json`. 3. **`infra/cdk.context.json` is committed.** From then on every synth/deploy uses the pinned id — no surprise replacement when a newer AMI exists. 4. To adopt a new AMI: `cdk context --reset ` (or edit the pinned value), commit the change, and review the cdk-diff — the PR will show "requires replacement", which is the intended, visible signal. Record the built AMI id in project memory (`project_open_swe_migration`) per the "memory updated for AMI id" build criterion. ## `userDataCausesReplacement` rationale `user-data.sh` is **provisioning-only** — it runs once at first boot and never carries durable runtime config. CDK sets **`userDataCausesReplacement: true`** so that any change to it is a deliberate, diff-visible instance replacement rather than a no-op edit that drifts from the running box. Durable runtime config is fetched **fresh on every service start** by `fetch-config.sh` (ExecStartPre) — changing a secret or SSM value needs only a `systemctl restart open-swe.service`, not a replacement. ## EBS discipline (binding — `feedback_inline_ebs_volumes`) **The box holds no durable state of its own:** | State | Lives in | On replacement | |---|---|---| | secrets / config | Secrets Manager + SSM → tmpfs `.env` | re-fetched at boot | | app code + SPA | S3 `open-swe--assets` | re-pulled at boot | | store (team_settings, user_mappings) | reseeded by `seed_store.sh` | re-seeded at boot | | logs | CloudWatch (30-day) — **not** a CFN resource in the stack | survive replacement | → **No local-only durable state ⇒ no standalone RETAIN volume is needed.** The root volume is disposable; there is intentionally no inline data `blockDevices` to lose. **Even so, snapshot before any replacing deploy.** Per the operational guard, before merging/deploying any change that REPLACES the instance (`userDataCausesReplacement`, AMI bump, instance-type change): 1. Enumerate the instance's volumes and assert **"no local-only durable state"** (the table above is the checklist). 2. Take an **EBS snapshot of the root volume and WAIT for `state=completed`** before letting the deploy proceed. Keep it as insurance; delete after a grace period. 3. Confirm the CloudWatch log groups are **not** CFN-managed in the stack so history survives; re-verify history after the new instance is healthy. cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance criterion). This is the enforced version — not just an assertion in the runbook. ## Integration contract (T10 — `fetch-config.sh` + `seed_store.sh`) Both ship in the app artifact under `deploy/seahaven/` and are wired into the unit: - **`fetch-config.sh`** (ExecStartPre, runs as `openswe`): reads `/etc/open-swe/boot.env` (`OPENSWE_ENV`, `AWS_REGION`, `SECRETS_PREFIX=open-swe-`, `SSM_PREFIX=/open-swe-`, `ENV_FILE=/run/open-swe/.env`), pulls Secrets Manager `open-swe-/*` + SSM `/open-swe-/*`, and writes: - `/run/open-swe/.env` (**0600, tmpfs**, secret-bearing app env incl. the multiline GitHub App PEM) — loaded by langgraph/dotenv via the `${APP_DIR}/.env` symlink. - `/run/open-swe/seed.env` (**0600, tmpfs**, simple `OPENSWE_*` vars only: `OPENSWE_DEFAULT_REPO`, `OPENSWE_OWNER_LOGIN`, `OPENSWE_OWNER_EMAIL`, model ids) — loaded by systemd `EnvironmentFile` so `seed_store.sh` (ExecStartPost) has them. - It must **fail-fast** (non-zero exit) if any required value is missing, so the unit never starts half-configured. - **`seed_store.sh`** (ExecStartPost): existing script, reseeds `team_settings/default` + `user_mappings/` into the in-memory store after each start. ## Smoke-boot checklist (after first boot) SSM Session Manager onto the instance (no public SSH — private subnet) and verify: - [ ] `cloud-init status --wait` → `done`; `/var/log/open-swe-user-data.log` ends with "user-data done" and shows the S3 pulls + service starts. - [ ] `systemctl is-active open-swe.service` → `active`. (If it failed, check `ExecStartPre`/`fetch-config.sh` — fail-fast means missing config = failed unit.) - [ ] **fetch-config fail-fast works:** `/run/open-swe/.env` exists, owner `openswe`, mode `0600`, on tmpfs (`findmnt /run/open-swe`); `seed.env` present. - [ ] `curl -fsS http://127.0.0.1:2024/ok` → `200` (raw LangGraph health). - [ ] `systemctl is-active nginx` → `active`; `curl -fsS http://127.0.0.1/healthz` → `200`; `curl -s http://127.0.0.1/threads` returns the SPA shell, **not** JSON (proves the agent API is not proxied — the security boundary holds). - [ ] `seed_store: done` in the journal / app.log (store reseeded). - [ ] CloudWatch: log groups `/open-swe//{app,user-data,nginx-access,nginx-error}` exist with **30-day** retention and are receiving events. - [ ] **No swapfile** (`swapon --show` empty) — the on-box SPA build is gone. - [ ] From the ALB only: dashboard host serves the SPA; `hooks` host reaches `/webhooks/*` on :2024 and nothing else (raw API paths hit the ALB default, not the box). ## Assumptions - **Artifact bucket** `open-swe--assets` (T7), with objects `${ARTIFACT_PREFIX}/app.tar.gz` (Python app incl. `deploy/seahaven/` and a prebuilt arm64 `.venv`) and `${ARTIFACT_PREFIX}/spa.tar.gz` (built SPA → `/var/www/open-swe`). `ARTIFACT_PREFIX` defaults to `releases/latest`; CDK renders the concrete value. - **Instance role** (defined in `/infra`, least-privilege per T4/T12) grants: `s3:GetObject` on `open-swe--assets/*`; `secretsmanager:GetSecretValue` on `open-swe-/*`; `ssm:GetParameter(s)`/`GetParametersByPath` on `/open-swe-/*`; `logs:*` for the CW agent log groups + `cloudwatch:PutMetricData`; SSM Session Manager (`ssm:UpdateInstanceInformation`, `ssmmessages:*`) for shell access. - **CDK substitutes** the `@@OPENSWE_ENV@@`, `@@ASSETS_BUCKET@@`, `@@SERVER_NAME@@`, `@@ARTIFACT_PREFIX@@` tokens in `user-data.sh` when rendering the launch template. - `:2024` binds `0.0.0.0` so the ALB hooks target group can reach `/webhooks/*`; it is reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is never internet-exposed — the ALB hooks rule is path-scoped to `/webhooks/*`.