mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 17:23:15 +00:00
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
141 lines
6.5 KiB
Bash
Executable file
141 lines
6.5 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Open SWE EC2 user-data — PROVISIONING-ONLY (runs once, at first boot).
|
|
#
|
|
# This is the rationale for `userDataCausesReplacement: true` in CDK: user-data
|
|
# does FIRST-BOOT provisioning, never durable runtime config. Editing it is a
|
|
# deliberate instance replacement. Durable runtime config is fetched fresh on
|
|
# every service start by deploy/seahaven/fetch-config.sh (ExecStartPre).
|
|
#
|
|
# The box holds NO durable state of its own:
|
|
# - secrets/config -> Secrets Manager + SSM, materialized to a tmpfs .env at boot
|
|
# - app artifact -> pulled from S3 (open-swe-<env>-assets) via the instance role
|
|
# - store state -> reseeded by seed_store.sh (ExecStartPost) on every start
|
|
# => there is no RETAIN volume to protect; replacement is tolerated. The EBS
|
|
# discipline (snapshot root + wait state=completed BEFORE any replacing deploy)
|
|
# is the safety net, not durable on-box state. See README "EBS discipline".
|
|
#
|
|
# Tokens (@@...@@) are substituted by CDK when it renders this script into the
|
|
# launch template. region is read from IMDSv2 as a fallback.
|
|
|
|
set -euo pipefail
|
|
exec > >(tee -a /var/log/open-swe-user-data.log) 2>&1
|
|
echo "==> open-swe user-data start $(date -u +%FT%TZ)"
|
|
|
|
# --- CDK-rendered values -----------------------------------------------------
|
|
OPENSWE_ENV="@@OPENSWE_ENV@@" # dev | prod
|
|
ASSETS_BUCKET="@@ASSETS_BUCKET@@" # open-swe-<env>-assets
|
|
SERVER_NAME="@@SERVER_NAME@@" # openswe[-dev].seahaven.com
|
|
ARTIFACT_PREFIX="@@ARTIFACT_PREFIX@@" # e.g. releases/latest
|
|
|
|
# --- fixed layout (must match provision.sh + templates) ----------------------
|
|
SERVICE_USER="openswe"
|
|
APP_DIR="/opt/open-swe/app"
|
|
VENV="${APP_DIR}/.venv"
|
|
WWW_ROOT="/var/www/open-swe"
|
|
TEMPLATE_DIR="/opt/open-swe/templates"
|
|
ENV_FILE="/run/open-swe/.env"
|
|
PORT="2024"
|
|
FETCH_CONFIG="${APP_DIR}/deploy/seahaven/fetch-config.sh"
|
|
SEED_STORE="${APP_DIR}/deploy/seahaven/seed_store.sh"
|
|
|
|
# region from IMDSv2
|
|
TOKEN="$(curl -fsS -X PUT "http://169.254.169.254/latest/api/token" \
|
|
-H "X-aws-ec2-metadata-token-ttl-seconds: 300" || true)"
|
|
AWS_REGION="$(curl -fsS -H "X-aws-ec2-metadata-token: ${TOKEN}" \
|
|
http://169.254.169.254/latest/meta-data/placement/region || echo us-east-1)"
|
|
export AWS_DEFAULT_REGION="$AWS_REGION"
|
|
echo "env=${OPENSWE_ENV} region=${AWS_REGION} bucket=${ASSETS_BUCKET} host=${SERVER_NAME}"
|
|
|
|
# --- boot.env: non-secret pointers fetch-config.sh reads ---------------------
|
|
install -d -o root -g root -m 0755 /etc/open-swe
|
|
cat >/etc/open-swe/boot.env <<EOF
|
|
OPENSWE_ENV=${OPENSWE_ENV}
|
|
AWS_REGION=${AWS_REGION}
|
|
ASSETS_BUCKET=${ASSETS_BUCKET}
|
|
ARTIFACT_PREFIX=${ARTIFACT_PREFIX}
|
|
ENV_FILE=${ENV_FILE}
|
|
SECRETS_PREFIX=open-swe-${OPENSWE_ENV}
|
|
SSM_PREFIX=/open-swe-${OPENSWE_ENV}
|
|
EOF
|
|
chmod 0644 /etc/open-swe/boot.env
|
|
|
|
# --- ensure the tmpfs for the materialized .env is mounted -------------------
|
|
# (baked into /etc/fstab by the AMI; mount it now in case it isn't yet.)
|
|
install -d -o root -g root -m 0755 /run/open-swe || true
|
|
mountpoint -q /run/open-swe || mount /run/open-swe || mount -t tmpfs \
|
|
-o rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m \
|
|
tmpfs /run/open-swe
|
|
|
|
# --- install the deploy script (single source of the app-deploy procedure) ---
|
|
# deploy.sh (deploy/ami/deploy.sh) pulls the release from S3, builds the venv with
|
|
# `uv sync`, and restarts the service. CDK base64-renders the file into the
|
|
# @@DEPLOY_SH_B64@@ token below so it is a normal reviewable repo file, not an
|
|
# inline heredoc. The `open-swe-<env>-deploy` SSM document runs this same script
|
|
# for every subsequent release.
|
|
echo "==> install /opt/open-swe/bin/deploy.sh"
|
|
install -d -o root -g root -m 0755 /opt/open-swe/bin
|
|
base64 -d >/opt/open-swe/bin/deploy.sh <<'DEPLOY_SH_B64'
|
|
@@DEPLOY_SH_B64@@
|
|
DEPLOY_SH_B64
|
|
chmod 0755 /opt/open-swe/bin/deploy.sh
|
|
|
|
# --- render + install the systemd unit ---------------------------------------
|
|
echo "==> install systemd unit"
|
|
sed \
|
|
-e "s|@@SERVICE_USER@@|${SERVICE_USER}|g" \
|
|
-e "s|@@APP_DIR@@|${APP_DIR}|g" \
|
|
-e "s|@@VENV@@|${VENV}|g" \
|
|
-e "s|@@PORT@@|${PORT}|g" \
|
|
-e "s|@@ENV_FILE@@|${ENV_FILE}|g" \
|
|
-e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
|
|
-e "s|@@FETCH_CONFIG@@|${FETCH_CONFIG}|g" \
|
|
-e "s|@@SEED_STORE@@|${SEED_STORE}|g" \
|
|
"${TEMPLATE_DIR}/open-swe.service" >/etc/systemd/system/open-swe.service
|
|
systemctl daemon-reload
|
|
|
|
# --- render + install the nginx site -----------------------------------------
|
|
echo "==> install nginx site"
|
|
sed \
|
|
-e "s|@@SERVER_NAME@@|${SERVER_NAME}|g" \
|
|
-e "s|@@WWW_ROOT@@|${WWW_ROOT}|g" \
|
|
-e "s|@@BACKEND_ADDR@@|127.0.0.1:${PORT}|g" \
|
|
"${TEMPLATE_DIR}/open-swe.nginx.conf" >/etc/nginx/sites-available/open-swe
|
|
ln -sfn /etc/nginx/sites-available/open-swe /etc/nginx/sites-enabled/open-swe
|
|
rm -f /etc/nginx/sites-enabled/default
|
|
nginx -t
|
|
|
|
# --- CloudWatch agent: 30-day log retention ----------------------------------
|
|
echo "==> configure CloudWatch agent (30-day retention)"
|
|
sed -e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
|
|
"${TEMPLATE_DIR}/amazon-cloudwatch-agent.json" \
|
|
>/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
|
|
/opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
|
|
-a fetch-config -m ec2 -s \
|
|
-c file:/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
|
|
|
|
# --- start nginx FIRST (the security boundary + health surface) --------------
|
|
# NOTE: intentionally NO swapfile here. The 8 GB-swapfile OOM hack existed only
|
|
# for the on-box Nitro SPA build, which now runs in GitHub Actions -> S3.
|
|
# nginx is brought up BEFORE the app is deployed so the ALB target-group health
|
|
# check (static `/healthz` -> 200) passes and the box is a healthy target even on
|
|
# the very first boot, before any release is published. open-swe.service is
|
|
# enabled (boot persistence) but STARTED by deploy.sh once the app is on disk.
|
|
echo "==> start nginx"
|
|
systemctl enable --now nginx
|
|
systemctl reload nginx
|
|
systemctl enable open-swe.service
|
|
|
|
# --- deploy the app (NON-FATAL on first boot) --------------------------------
|
|
# deploy.sh pulls the release, builds the venv, and starts open-swe.service. On a
|
|
# brand-new env no release exists yet, so this is allowed to fail WITHOUT aborting
|
|
# user-data: nginx is already up (healthy target), and the first `build-artifacts`
|
|
# run + `open-swe-<env>-deploy` SSM command will bring the app up. A failure here
|
|
# is logged, not fatal.
|
|
echo "==> initial app deploy (non-fatal if no release is published yet)"
|
|
if /opt/open-swe/bin/deploy.sh; then
|
|
echo "==> initial app deploy succeeded"
|
|
else
|
|
echo "==> no release yet (or deploy failed): open-swe.service deferred to the next SSM deploy"
|
|
fi
|
|
|
|
echo "==> open-swe user-data done $(date -u +%FT%TZ)"
|