chore: decommission self-hosted AWS LangGraph stack (#64)

* chore: decommission self-hosted AWS LangGraph stack

Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).

- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
  scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
  Vercel view; drop dead links to infra/ and deploy/seahaven

Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.

* chore: clean up dangling references left by the AWS decommission

Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:

- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
  rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
  a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
  infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)

The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
This commit is contained in:
Adam Moussa 2026-06-29 19:54:38 -04:00 • committed by GitHub
parent 430a1cdff9
commit 7f60324f0c
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
52 changed files with 15 additions and 9281 deletions

View file

@ -10,11 +10,10 @@ updates:
minor-and-patch:
update-types: ["minor", "patch"]
# JavaScript/TypeScript — CDK (/infra), Playwright (/tests/e2e), dashboard (/ui), root tooling
# JavaScript/TypeScript — Playwright (/tests/e2e), dashboard (/ui), root tooling
- package-ecosystem: "npm"
directories:
- "/"
- "/infra"
- "/tests/e2e"
- "/ui"
schedule:

View file

@ -1,52 +0,0 @@
#!/usr/bin/env bash
# Package the two release artifacts (run from the repo root by build-artifacts.yml):
#
# spa.tar.gz = CONTENTS of the built SPA dir (ui/.output/public/*), so it extracts
# straight into the nginx web root with _shell.html at the root.
# app.tar.gz = the Python source the box runs `uv sync` against. `git archive`
# gives a clean tree (no node_modules, no .venv, no local cruft);
# ui/ is intentionally excluded (it ships as spa.tar.gz).
set -euo pipefail
SPA_DIR="ui/.output/public"
[ -f "${SPA_DIR}/_shell.html" ] || {
echo "ERROR: SPA build output missing ${SPA_DIR}/_shell.html (did 'bun run build' run?)" >&2
exit 1
}
echo "==> spa.tar.gz from ${SPA_DIR}"
tar -C "${SPA_DIR}" -czf spa.tar.gz .
echo "==> app.tar.gz from source (git archive HEAD)"
git archive --format=tar.gz -o app.tar.gz HEAD \
agent deploy langgraph.json pyproject.toml uv.lock README.md
# List the archive ONCE into a variable. (`tar -tzf ... | grep -q ...` is unsafe
# under `set -o pipefail`: grep -q exits on first match, SIGPIPEs tar -> "write
# error" -> the pipeline reports non-zero even though grep succeeded, a false
# failure. Listing once avoids the pipe entirely.)
APP_LIST="$(tar -tzf app.tar.gz)"
# Sanity: the box's `uv sync --frozen` needs pyproject.toml + uv.lock at the root,
# and the package itself (agent/). Fail loudly here rather than on the box.
for required in pyproject.toml uv.lock agent/server.py langgraph.json; do
printf '%s\n' "${APP_LIST}" | grep -qx "${required}" || {
echo "ERROR: app.tar.gz is missing ${required}" >&2
exit 1
}
done
# Fail-closed secret guard: the source is git-archived wholesale, so reject the
# release if a secret-shaped FILE slipped into the tracked tree (defense in depth
# on top of .gitignore — the artifact lands on the box + in S3). Scoped to data
# extensions so credential-handling *source* (e.g. team_credentials.py) is not a
# false positive.
SECRET_RE='(^|/)(\.env(\..+)?|id_rsa|.*\.(pem|key|p12|pfx)|.*(secret|credential|password|token)s?\.(json|ya?ml|txt|env|ini|cfg))$'
if printf '%s\n' "${APP_LIST}" | grep -qiE "${SECRET_RE}"; then
echo "ERROR: app.tar.gz contains a secret-shaped file — refusing to publish:" >&2
printf '%s\n' "${APP_LIST}" | grep -iE "${SECRET_RE}" >&2
exit 1
fi
echo "==> artifacts:"
ls -la spa.tar.gz app.tar.gz

View file

@ -1,64 +0,0 @@
#!/usr/bin/env bash
# Publish the packaged artifacts to the env's S3 bucket and roll the box to them.
# Run by build-artifacts.yml AFTER aws creds are configured (env: ENV, BUCKET,
# DEPLOY_DOC). Each release is stored immutably under releases/<sha>/.
#
# releases/latest/ (what the box's deploy.sh pulls) is advanced TRANSACTIONALLY:
# it is pointed at the new release, the box is rolled, and ONLY on a successful
# roll is it kept — a failed roll reverts releases/latest/ to the prior release so
# a later box boot / replacement never self-deploys a release that failed to come
# up. releases/last-good/ (rollback fallback) is advanced only after success and
# means "last release whose deploy.sh brought the service up active" (deploy.sh
# gates on `systemctl is-active`), not merely "last uploaded".
#
# The fire/wait/gate against the box lives in roll-box.sh (shared with rollback.sh);
# the deploy is fired by TAG (project=open-swe,env=<env>), exactly what the app
# deploy role's tag-scoped ssm:SendCommand allows.
set -euo pipefail
: "${ENV:?}" "${BUCKET:?}" "${DEPLOY_DOC:?}"
SHA="${GITHUB_SHA:?}"
[ -f app.tar.gz ] && [ -f spa.tar.gz ] || { echo "ERROR: artifacts not built" >&2; exit 1; }
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Re-point releases/latest/ at the release stored under releases/<sha>/, recording
# the sha so the pointer is self-describing (used to revert on failure).
point_latest() {
local s="$1" f
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "s3://${BUCKET}/releases/${s}/${f}" "s3://${BUCKET}/releases/latest/${f}"
done
printf '%s\n' "${s}" | aws s3 cp - "s3://${BUCKET}/releases/latest/sha.txt"
}
# Which release does latest point at right now? (empty on the very first deploy.)
PREV_SHA="$(aws s3 cp "s3://${BUCKET}/releases/latest/sha.txt" - 2>/dev/null | tr -d '[:space:]' || true)"
echo "==> upload release ${SHA} to s3://${BUCKET}/releases/${SHA}/"
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "${f}" "s3://${BUCKET}/releases/${SHA}/${f}"
done
echo "==> point releases/latest/ -> ${SHA} (was ${PREV_SHA:-<none>})"
point_latest "${SHA}"
# Roll the box to releases/latest/. On a non-Success aggregate, roll-box.sh exits
# non-zero; revert latest to the prior release so no later boot pulls the bad one.
if ! ROLL_COMMENT="release ${SHA}" bash "${HERE}/roll-box.sh"; then
if [ -n "${PREV_SHA}" ]; then
echo "!! deploy failed — reverting releases/latest/ -> ${PREV_SHA}" >&2
point_latest "${PREV_SHA}"
else
echo "!! deploy failed on the FIRST release — leaving releases/latest/ = ${SHA} (no prior release to revert to)" >&2
fi
exit 1
fi
# Roll succeeded (service came up active) -> this release is now the known-good one.
echo "==> mark releases/last-good/ = ${SHA} (rollback fallback target)"
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "s3://${BUCKET}/releases/${SHA}/${f}" "s3://${BUCKET}/releases/last-good/${f}"
done
printf '%s\n' "${SHA}" | aws s3 cp - "s3://${BUCKET}/releases/last-good/sha.txt"
echo "==> ${ENV} rolled to release ${SHA} (now last-good)"

View file

@ -1,59 +0,0 @@
#!/usr/bin/env bash
# Fire the env's SSM deploy document (tag-targeted) and wait for it to finish,
# gating on the AGGREGATE command status. Shared by publish-and-deploy.sh (forward
# roll) and rollback.sh (backward roll) so the fire/wait/gate logic lives in ONE
# place. Requires ENV + DEPLOY_DOC in the environment and aws creds already set.
#
# Tag-targeting (project=open-swe,env=<env>) is exactly what the app deploy role's
# tag-scoped ssm:SendCommand allows — no ec2:DescribeInstances, no instance id.
set -euo pipefail
: "${ENV:?}" "${DEPLOY_DOC:?}"
COMMENT="${ROLL_COMMENT:-roll ${ENV}}"
echo "==> fire ${DEPLOY_DOC} via SSM (tag-targeted: project=open-swe, env=${ENV})"
CMD_ID="$(aws ssm send-command \
--document-name "${DEPLOY_DOC}" \
--targets "Key=tag:project,Values=open-swe" "Key=tag:env,Values=${ENV}" \
--comment "${COMMENT}" \
--query 'Command.CommandId' --output text)"
echo "command: ${CMD_ID}"
echo "==> wait for the deploy to finish"
IID=""
for _ in $(seq 1 60); do
sleep 10
IID="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].InstanceId' --output text 2>/dev/null || echo None)"
[ -z "${IID}" ] || [ "${IID}" = "None" ] && continue
STATUS="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].Status' --output text 2>/dev/null || echo Pending)"
case "${STATUS}" in
Success | Failed | Cancelled | TimedOut) break ;;
esac
done
if [ -z "${IID}" ] || [ "${IID}" = "None" ]; then
echo "ERROR: no box picked up the deploy command (is a running open-swe ${ENV} box registered with SSM?)" >&2
exit 1
fi
echo "==> deploy.sh output from ${IID}:"
echo "----- stdout -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardOutputContent' --output text || true
echo "----- stderr -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardErrorContent' --output text || true
# Gate on the AGGREGATE command status (Success only if EVERY targeted invocation
# succeeded), not CommandInvocations[0] — during a userDataCausesReplacement window
# two instances can briefly share the project/env tags, and a partial failure on the
# other instance must not be reported as success.
TARGETS="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].TargetCount' --output text 2>/dev/null || echo 1)"
[ "${TARGETS}" = "1" ] || echo "WARNING: deploy fanned out to ${TARGETS} instances (expected 1)"
AGG="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].Status' --output text 2>/dev/null || echo Failed)"
echo "==> aggregate deploy status: ${AGG} (across ${TARGETS} target(s))"
[ "${AGG}" = "Success" ] || { echo "ERROR: deploy did not succeed (${AGG})" >&2; exit 1; }

View file

@ -1,66 +0,0 @@
#!/usr/bin/env bash
# Roll an env BACK to a prior release: re-point releases/latest/ at a chosen release
# and re-fire the deploy. Run by rollback.yml after aws creds are configured.
#
# ENV dev | prod (required)
# BUCKET open-swe-<env>-assets (required)
# DEPLOY_DOC open-swe-<env>-deploy (required)
# TARGET_SHA release sha to restore; blank => releases/last-good/ (optional)
#
# releases/latest/ is moved TRANSACTIONALLY (same as the forward deploy): pointed at
# the target, the box rolled, and on a failed roll latest is reverted to whatever it
# was before the rollback attempt — so a failed rollback never leaves latest at a
# release the box could not bring up. releases/last-good/ is left untouched; advance
# it by running a forward deploy.
#
# Reuses the app deploy role's existing releases/* write + tag-scoped ssm:SendCommand
# — no new IAM. The fire/wait/gate is the SAME roll-box.sh the forward deploy uses.
set -euo pipefail
: "${ENV:?}" "${BUCKET:?}" "${DEPLOY_DOC:?}"
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Copy a release prefix into releases/latest/ AND record the resolved sha, so
# releases/latest/sha.txt is always a real commit a later revert can resolve.
point_latest() { # $1 = source prefix (releases/<sha>|releases/last-good); $2 = sha to record
local src="$1" sha="$2" f
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "s3://${BUCKET}/${src}/${f}" "s3://${BUCKET}/releases/latest/${f}"
done
printf '%s\n' "${sha}" | aws s3 cp - "s3://${BUCKET}/releases/latest/sha.txt"
}
if [ -n "${TARGET_SHA:-}" ]; then
SRC="releases/${TARGET_SHA}"
LABEL="${TARGET_SHA}"
RESOLVED_SHA="${TARGET_SHA}"
else
SRC="releases/last-good"
LABEL="last-good"
# resolve last-good's real sha so latest/sha.txt records a commit, not a label.
RESOLVED_SHA="$(aws s3 cp "s3://${BUCKET}/releases/last-good/sha.txt" - 2>/dev/null | tr -d '[:space:]' || true)"
RESOLVED_SHA="${RESOLVED_SHA:-last-good}"
fi
echo "==> rollback ${ENV} to ${LABEL} (s3://${BUCKET}/${SRC}/, sha=${RESOLVED_SHA})"
# Refuse to roll back to a release that is not fully present.
for f in app.tar.gz spa.tar.gz; do
aws s3 ls "s3://${BUCKET}/${SRC}/${f}" >/dev/null 2>&1 \
|| { echo "ERROR: ${SRC}/${f} not found in s3://${BUCKET} — cannot roll back to ${LABEL}" >&2; exit 1; }
done
# Capture what latest points at now, so a failed rollback can be reverted.
PREV_SHA="$(aws s3 cp "s3://${BUCKET}/releases/latest/sha.txt" - 2>/dev/null | tr -d '[:space:]' || true)"
echo "==> point releases/latest/ -> ${SRC} (was ${PREV_SHA:-<none>})"
point_latest "${SRC}" "${RESOLVED_SHA}"
if ! ROLL_COMMENT="rollback ${ENV} to ${LABEL}" bash "${HERE}/roll-box.sh"; then
if [ -n "${PREV_SHA}" ]; then
echo "!! rollback deploy failed — reverting releases/latest/ -> ${PREV_SHA}" >&2
point_latest "releases/${PREV_SHA}" "${PREV_SHA}"
fi
exit 1
fi
echo "==> ${ENV} rolled back to ${LABEL}"
echo "NOTE: releases/last-good/ is left unchanged; re-run a forward deploy to advance it."

View file

@ -1,123 +0,0 @@
name: Build & publish app artifacts
# T7 + T19 — build the release (SPA + Python source) and publish it to the per-env
# S3 artifact bucket, then roll the box to it.
#
# push to dev → publish to open-swe-dev-assets → deploy dev box (AUTO)
# push to main → publish to open-swe-prod-assets → deploy prod box (manual approval: env "prod")
#
# Two artifacts (the box's deploy.sh pulls both from releases/latest/):
# spa.tar.gz = the built dashboard SPA (vite -> ui/.output/public). Built HERE
# (not on the box) — the build is memory-heavy and the box is small.
# app.tar.gz = the Python source tree (NO ui/, NO .venv). The box runs
# `uv sync` to build a native-ARM64 venv at the real runtime path.
#
# Each release is uploaded under releases/<sha>/ (immutable, auditable) AND mirrored
# to releases/latest/ (what the box pulls). Then the open-swe-<env>-deploy SSM
# document is fired (tag-scoped to project=open-swe,env=<env>) to roll the box.
#
# OIDC subject alignment (matches the per-env app-role trust in infra/lib/config.ts):
# - publish-dev declares NO `environment:` → sub = repo:…:ref:refs/heads/dev
# - publish-prod declares `environment: prod` → sub = repo:…:environment:prod
# (also triggers the prod Environment's required-reviewer approval gate).
#
# Prerequisites:
# - repo variables AWS_DEPLOY_ROLE_APP_DEV / AWS_DEPLOY_ROLE_APP_PROD = the
# githubdeploy-open-swe-app-<env> role ARNs (open-swe-iam CfnOutputs).
# - the open-swe-<env> stack deployed (creates the bucket + the SSM deploy doc).
permissions:
contents: read
on:
push:
branches: [dev, main]
paths:
- "agent/**"
- "ui/**"
- "deploy/**"
- "langgraph.json"
- "pyproject.toml"
- "uv.lock"
- ".github/workflows/build-artifacts.yml"
- ".github/scripts/**"
workflow_dispatch:
concurrency:
# one publish+deploy per branch at a time; never cancel an in-flight release.
group: build-artifacts-${{ github.ref }}
cancel-in-progress: false
jobs:
publish-dev:
name: Publish + deploy (dev)
if: ${{ github.ref == 'refs/heads/dev' }}
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
id-token: write
contents: read
env:
ENV: dev
BUCKET: open-swe-dev-assets
DEPLOY_DOC: open-swe-dev-deploy
steps:
- uses: actions/checkout@v7
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- name: Build SPA (vite -> ui/.output/public)
working-directory: ui
# vite's bundle exceeds Node's default ~2 GB heap (the build that needed an
# 8 GB swapfile on-box); the runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"
run: |
bun install --frozen-lockfile
bun run build
- name: Package artifacts
run: bash .github/scripts/package-artifacts.sh
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_DEV }}
aws-region: us-east-1
- name: Publish to S3 + roll the box
run: bash .github/scripts/publish-and-deploy.sh
publish-prod:
name: Publish + deploy (prod)
if: ${{ github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 30
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam). Also
# makes the OIDC sub …:environment:prod (matches the prod app-role trust).
environment: prod
permissions:
id-token: write
contents: read
env:
ENV: prod
BUCKET: open-swe-prod-assets
DEPLOY_DOC: open-swe-prod-deploy
steps:
- uses: actions/checkout@v7
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- name: Build SPA (vite -> ui/.output/public)
working-directory: ui
# vite's bundle exceeds Node's default ~2 GB heap (the build that needed an
# 8 GB swapfile on-box); the runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"
run: |
bun install --frozen-lockfile
bun run build
- name: Package artifacts
run: bash .github/scripts/package-artifacts.sh
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_PROD }}
aws-region: us-east-1
- name: Publish to S3 + roll the box
run: bash .github/scripts/publish-and-deploy.sh

View file

@ -1,124 +0,0 @@
name: Infra CD
# Path-filtered CDK deploy for /infra, per env, OIDC-only (no static keys).
#
# push to dev → CI (tsc+jest+synth) → deploy OpenSweDevStack (AUTO, CI-green-gated)
# push to main → CI → deploy OpenSweProdStack (manual approval: env "prod")
#
# Why this is NOT the reusable cd-cdk.yaml: that workflow runs `cdk deploy --all`,
# which would deploy ALL THREE stacks (incl. the OTHER env + the shared IAM stack)
# from a single-env push — breaking the per-env dev/prod boundary. So we target one
# stack explicitly per env. (Infra CI still uses the reusable ci-typescript-cdk.)
#
# The shared IAM stack (open-swe-iam — owns BOTH envs' OIDC deploy roles) is
# intentionally NOT deployed here: it is a privileged, human-gated apply (T6), so a
# routine dev push can never alter prod's deploy role.
#
# OIDC subject alignment (must match the per-env trust in infra/lib/config.ts):
# - deploy-dev declares NO `environment:` → token sub = repo:…:ref:refs/heads/dev,
# which is exactly what githubdeploy-open-swe-infra-dev trusts.
# - deploy-prod declares `environment: prod` → token sub = repo:…:environment:prod,
# which githubdeploy-open-swe-infra-prod trusts AND which triggers the GitHub
# Environment's required-reviewer (manual approval) gate.
#
# Prerequisites (post-T6, when the roles exist):
# - repo variables AWS_DEPLOY_ROLE_INFRA_DEV / AWS_DEPLOY_ROLE_INFRA_PROD = the
# githubdeploy-open-swe-infra-<env> role ARNs (open-swe-iam CfnOutputs).
# - a GitHub Environment named "prod" with Adam as a required reviewer.
permissions:
contents: read
on:
push:
branches: [dev, main]
paths:
- "infra/**"
- ".github/workflows/cd-infra.yml"
workflow_dispatch:
concurrency:
# one infra deploy per branch at a time; never cancel an in-flight deploy.
group: cd-infra-${{ github.ref }}
cancel-in-progress: false
jobs:
# CI-green precondition — re-run tsc + jest + synth on the pushed commit before
# any deploy. A failure here blocks the deploy jobs (needs: ci).
ci:
name: Infra CI (pre-deploy)
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@main
with:
node-version: "24"
working-directory: infra
cache-dependency-path: infra/package-lock.json
run-typecheck: true
run-tests: true
run-cdk-synth: true
deploy-dev:
name: Deploy open-swe-dev
needs: ci
if: ${{ github.ref == 'refs/heads/dev' }}
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
id-token: write
contents: read
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v4
with:
node-version: "24"
cache: npm
cache-dependency-path: infra/package-lock.json
- name: Install deps
working-directory: infra
run: npm ci
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_INFRA_DEV }}
aws-region: us-east-1
- name: CDK deploy (dev only)
working-directory: infra
# --outputs-file lets us print the stack outputs from CDK's own result
# (the deploy role intentionally lacks cloudformation:DescribeStacks; CDK
# gets outputs via the bootstrap cfn-exec role it assumes, so no extra grant).
run: npx cdk deploy OpenSweDevStack --require-approval never --outputs-file cdk-outputs.json
- name: Stack outputs
working-directory: infra
run: cat cdk-outputs.json
deploy-prod:
name: Deploy open-swe-prod
needs: ci
if: ${{ github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 30
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam).
# Also makes the OIDC sub …:environment:prod (matches the prod role trust).
environment: prod
permissions:
id-token: write
contents: read
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v4
with:
node-version: "24"
cache: npm
cache-dependency-path: infra/package-lock.json
- name: Install deps
working-directory: infra
run: npm ci
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_INFRA_PROD }}
aws-region: us-east-1
- name: CDK deploy (prod only)
working-directory: infra
# See deploy-dev: --outputs-file avoids needing cloudformation:DescribeStacks.
run: npx cdk deploy OpenSweProdStack --require-approval never --outputs-file cdk-outputs.json
- name: Stack outputs
working-directory: infra
run: cat cdk-outputs.json

View file

@ -1,26 +0,0 @@
name: Infra CI
# Path-filtered CI for the /infra CDK app (TypeScript). The existing "CI"
# (ci.yml) covers the Python agent; this adds tsc + jest + cdk synth for /infra so
# infra changes are gated on a PR the same way. Runs only when /infra changes.
permissions:
contents: read
on:
pull_request:
paths:
- "infra/**"
- ".github/workflows/ci-infra.yml"
jobs:
infra-ci:
name: Infra CI (tsc + jest + synth)
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@main
with:
node-version: "24"
working-directory: infra
cache-dependency-path: infra/package-lock.json
run-typecheck: true
run-tests: true
run-cdk-synth: true

View file

@ -73,7 +73,7 @@ jobs:
- name: Run E2E
working-directory: tests/e2e
# Playwright's globalSetup runs the real `bun run build`, whose vite bundle
# exceeds Node's default ~2 GB heap (same OOM fixed in build-artifacts.yml).
# exceeds Node's default ~2 GB heap.
# The runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"

View file

@ -1,84 +0,0 @@
name: Rollback (re-point env to a prior release)
# Roll an env back to a previously published release without rebuilding. Re-points
# releases/latest/ at the chosen release and re-fires the open-swe-<env>-deploy SSM
# document — same fire/wait/gate path as a forward deploy (roll-box.sh).
#
# env=dev, sha blank → restore open-swe-dev-assets/releases/last-good/ (AUTO)
# env=prod, sha blank → restore open-swe-prod-assets/releases/last-good/ (manual
# approval: Environment "prod", same gate as a prod deploy)
# sha=<commit> → restore that exact releases/<sha>/ instead of last-good.
#
# No new IAM: reuses the githubdeploy-open-swe-app-<env> role's existing releases/*
# write + tag-scoped ssm:SendCommand. OIDC subject alignment matches build-artifacts:
# - rollback-dev declares NO `environment:` → sub = repo:…:ref:refs/heads/<branch>
# - rollback-prod declares `environment: prod` → sub = repo:…:environment:prod
permissions:
contents: read
on:
workflow_dispatch:
inputs:
env:
description: "Which environment to roll back"
required: true
type: choice
options: [dev, prod]
sha:
description: "Release SHA to restore (blank = releases/last-good)"
required: false
type: string
concurrency:
# never overlap a rollback with another rollback/deploy of the same env.
group: rollback-${{ inputs.env }}
cancel-in-progress: false
jobs:
rollback-dev:
name: Rollback (dev)
if: ${{ inputs.env == 'dev' }}
runs-on: ubuntu-latest
timeout-minutes: 20
permissions:
id-token: write
contents: read
env:
ENV: dev
BUCKET: open-swe-dev-assets
DEPLOY_DOC: open-swe-dev-deploy
TARGET_SHA: ${{ inputs.sha }}
steps:
- uses: actions/checkout@v7
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_DEV }}
aws-region: us-east-1
- name: Re-point releases/latest + redeploy
run: bash .github/scripts/rollback.sh
rollback-prod:
name: Rollback (prod)
if: ${{ inputs.env == 'prod' }}
runs-on: ubuntu-latest
timeout-minutes: 20
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam). Also
# makes the OIDC sub …:environment:prod (matches the prod app-role trust).
environment: prod
permissions:
id-token: write
contents: read
env:
ENV: prod
BUCKET: open-swe-prod-assets
DEPLOY_DOC: open-swe-prod-deploy
TARGET_SHA: ${{ inputs.sha }}
steps:
- uses: actions/checkout@v7
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_PROD }}
aws-region: us-east-1
- name: Re-point releases/latest + redeploy
run: bash .github/scripts/rollback.sh

View file

@ -10,16 +10,6 @@
"owner": "adam@seahavenind.com",
"added": "2026-06-29"
},
{
"id": "OSWE-IAC-SECRETS-LIST-01",
"title": "EC2 instance role grants BatchGetSecretValue on \"*\" (operation-level; secret-NAME existence enumeration account-wide)",
"file": "infra/lib/constructs/instance-role.ts",
"severity": "low",
"status": "confirmed",
"suppression_justification": "ACCEPTED LOW residual, metadata-only. secretsmanager:BatchGetSecretValue is a collection action that AWS cannot scope to a per-secret ARN, so it is granted on `*` (documented in commit 3c69dd9d and the construct comment). Secret VALUES remain strictly gated by the PREFIX-scoped GetSecretValue/DescribeSecret on secret:open-swe-<env>/* (checked per-secret even within the batch), so cross-env VALUE isolation is preserved; only a name EXISTENCE oracle remains, within Sea Haven's single-tenant account 328440206208. ListSecrets is intentionally NOT granted (so name FILTER enumeration AccessDenies). Confirmed by GPT-4.1 IAM cross-review (no BLOCK) and the iac-iam detector (one low residual, no critical/high).",
"owner": "adam@seahavenind.com",
"added": "2026-06-26"
},
{
"id": "gitleaks-generic-api-key-89",
"title": "Hardcoded credential flagged in encryption-roundtrip test fixture (CWE-798)",

View file

@ -153,18 +153,20 @@ This is an area where you can extend Open SWE for your org: add deterministic CI
## Deployment (Sea Haven fork)
This fork is **self-hosted on AWS** and live in production. Each env (`dev` /
`prod`) runs the stock `langgraph dev` server (all three graphs + the FastAPI
webapp) bound to loopback `127.0.0.1:2024` on a single ARM64 EC2 box, fronted by
nginx (the sole ingress) behind the shared `seahaven-com` ALB. CDK (`infra/`)
owns the per-env stacks; GitHub Actions handle CDK deploys (`cd-infra.yml`) and
app-artifact releases to S3 rolled onto the box via an SSM document
(`build-artifacts.yml`), with `dev` auto-deploying and `prod` gated behind a
manual GitHub Environment approval.
This fork runs on a **managed deployment**: the backend (all three graphs + the
FastAPI webapp) runs on
[LangGraph Cloud / Platform](https://langchain-ai.github.io/langgraph/cloud/),
and the `ui/` dashboard deploys to [Vercel](https://vercel.com/). Configuration
and secrets live in the LangGraph deployment config and Vercel environment
variables. Promotion from `dev` to `prod` (`main`) is handled by
[`.github/workflows/promote-to-main.yml`](.github/workflows/promote-to-main.yml).
**[`deploy/seahaven/DEPLOYMENT.md`](deploy/seahaven/DEPLOYMENT.md) is the canonical
deploy runbook** — full end-to-end pipeline, config seeding, promotion/rollback,
and live prod facts. CDK specifics live in [`infra/README.md`](infra/README.md).
See **[INSTALLATION.md § 10 "Production deployment"](INSTALLATION.md#10-production-deployment)**
for the full backend + dashboard setup.
> The earlier self-hosted AWS stack (CDK under `infra/`, an ARM64 EC2 box + nginx
> behind the shared ALB, and the `cd-infra` / `build-artifacts` release pipelines)
> was **decommissioned** in favor of the managed deployment above.
## License

View file

@ -1,167 +0,0 @@
# Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(`open-swe-dev` / `open-swe-prod`) in the Open SWE → AWS migration. Builds an
**ARM64 (Graviton) Ubuntu 24.04 LTS** base AMI and provisions the box on first
boot with the stock `langgraph dev` runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo `TODO.md` ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, `seahaven-vpc` **private subnet + NAT**,
inbound **only from the ALB SG**. Runtime is **stock `langgraph dev`** (in-memory
store, `--no-reload`) + nginx + systemd. The box has **no git auth** — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), **not** on the box, so the old 8 GB-swapfile OOM hack is gone.
## Files
| Path | Purpose |
|---|---|
| `open-swe-base.pkr.hcl` | Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
| `scripts/provision.sh` | Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
| `user-data.sh` | First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
| `templates/open-swe.service` | systemd unit TEMPLATE (`@@tokens@@` rendered at boot). |
| `templates/open-swe.nginx.conf` | nginx site TEMPLATE (dashboard SPA + scoped `/dashboard/api/` proxy). |
| `templates/amazon-cloudwatch-agent.json` | CW agent config TEMPLATE — **30-day log retention**. |
`deploy/seahaven/fetch-config.sh` and `deploy/seahaven/seed_store.sh` are owned by
the parallel T10 work and ship **inside the app artifact**; this AMI wires them in
but does not author them (see "Integration contract" below).
## Build the AMI
```bash
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
```
Builds in account **328440206208 / us-east-1**. Source = latest Canonical Ubuntu
24.04 (Noble) **arm64** AMI (`source_ami_filter`, owner `099720109477`). Build host
is `t4g.medium` (ARM64). The output AMI is tagged:
```
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
```
Pinned versions live in the template `variable` defaults (`uv_version`,
`python_version`, `node_major`, the CW-agent / awscli URLs) and the
`required_plugins` block (`amazon` 1.3.6) — bump deliberately.
## AMI → `cdk.context.json` pinning contract
The CDK stacks in `/infra` (owned by T3/T12) consume the AMI **by id, pinned in the
committed `infra/cdk.context.json`** — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached `MachineImage.lookup` resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
1. `packer build` prints the new AMI id (and tags it `open-swe-base-arm64`).
2. CDK looks the AMI up with **`cachedInContext: true`** (e.g.
`MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })`),
which writes the resolved id into `infra/cdk.context.json`.
3. **`infra/cdk.context.json` is committed.** From then on every synth/deploy uses
the pinned id — no surprise replacement when a newer AMI exists.
4. To adopt a new AMI: `cdk context --reset <ami-lookup-key>` (or edit the pinned
value), commit the change, and review the cdk-diff — the PR will show
"requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (`project_open_swe_migration`) per the
"memory updated for AMI id" build criterion.
## `userDataCausesReplacement` rationale
`user-data.sh` is **provisioning-only** — it runs once at first boot and never
carries durable runtime config. CDK sets **`userDataCausesReplacement: true`** so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
**fresh on every service start** by `fetch-config.sh` (ExecStartPre) — changing a
secret or SSM value needs only a `systemctl restart open-swe.service`, not a
replacement.
## EBS discipline (binding — `feedback_inline_ebs_volumes`)
**The box holds no durable state of its own:**
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs `.env` | re-fetched at boot |
| app code + SPA | S3 `open-swe-<env>-assets` | re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by `seed_store.sh` | re-seeded at boot |
| logs | CloudWatch (30-day) — **not** a CFN resource in the stack | survive replacement |
→ **No local-only durable state ⇒ no standalone RETAIN volume is needed.** The root
volume is disposable; there is intentionally no inline data `blockDevices` to lose.
**Even so, snapshot before any replacing deploy.** Per the operational guard, before
merging/deploying any change that REPLACES the instance (`userDataCausesReplacement`,
AMI bump, instance-type change):
1. Enumerate the instance's volumes and assert **"no local-only durable state"**
(the table above is the checklist).
2. Take an **EBS snapshot of the root volume and WAIT for `state=completed`** before
letting the deploy proceed. Keep it as insurance; delete after a grace period.
3. Confirm the CloudWatch log groups are **not** CFN-managed in the stack so history
survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance
criterion). This is the enforced version — not just an assertion in the runbook.
## Integration contract (T10 — `fetch-config.sh` + `seed_store.sh`)
Both ship in the app artifact under `deploy/seahaven/` and are wired into the unit:
- **`fetch-config.sh`** (ExecStartPre, runs as `openswe`): reads `/etc/open-swe/boot.env`
(`OPENSWE_ENV`, `AWS_REGION`, `SECRETS_PREFIX=open-swe-<env>`, `SSM_PREFIX=/open-swe-<env>`,
`ENV_FILE=/run/open-swe/.env`), pulls Secrets Manager `open-swe-<env>/*` + SSM
`/open-swe-<env>/*`, and writes:
- `/run/open-swe/.env` (**0600, tmpfs**, secret-bearing app env incl. the multiline
GitHub App PEM) — loaded by langgraph/dotenv via the `${APP_DIR}/.env` symlink.
- `/run/open-swe/seed.env` (**0600, tmpfs**, simple `OPENSWE_*` vars only:
`OPENSWE_DEFAULT_REPO`, `OPENSWE_OWNER_LOGIN`, `OPENSWE_OWNER_EMAIL`, model ids) —
loaded by systemd `EnvironmentFile` so `seed_store.sh` (ExecStartPost) has them.
- It must **fail-fast** (non-zero exit) if any required value is missing, so the
unit never starts half-configured.
- **`seed_store.sh`** (ExecStartPost): existing script, reseeds `team_settings/default`
+ `user_mappings/<login>` into the in-memory store after each start.
## Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
- [ ] `cloud-init status --wait` → `done`; `/var/log/open-swe-user-data.log` ends with
"user-data done" and shows the S3 pulls + service starts.
- [ ] `systemctl is-active open-swe.service` → `active`. (If it failed, check
`ExecStartPre`/`fetch-config.sh` — fail-fast means missing config = failed unit.)
- [ ] **fetch-config fail-fast works:** `/run/open-swe/.env` exists, owner `openswe`,
mode `0600`, on tmpfs (`findmnt /run/open-swe`); `seed.env` present.
- [ ] `curl -fsS http://127.0.0.1:2024/ok` → `200` (raw LangGraph health).
- [ ] `systemctl is-active nginx` → `active`; `curl -fsS http://127.0.0.1/healthz` →
`200`; `curl -s http://127.0.0.1/threads` returns the SPA shell, **not** JSON
(proves the agent API is not proxied — the security boundary holds).
- [ ] `seed_store: done` in the journal / app.log (store reseeded).
- [ ] CloudWatch: log groups `/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`
exist with **30-day** retention and are receiving events.
- [ ] **No swapfile** (`swapon --show` empty) — the on-box SPA build is gone.
- [ ] From the ALB only: dashboard host serves the SPA; `hooks` host reaches
`/webhooks/*` on :2024 and nothing else (raw API paths hit the ALB default, not
the box).
## Assumptions
- **Artifact bucket** `open-swe-<env>-assets` (T7), with objects
`${ARTIFACT_PREFIX}/app.tar.gz` (Python app incl. `deploy/seahaven/` and a prebuilt
arm64 `.venv`) and `${ARTIFACT_PREFIX}/spa.tar.gz` (built SPA → `/var/www/open-swe`).
`ARTIFACT_PREFIX` defaults to `releases/latest`; CDK renders the concrete value.
- **Instance role** (defined in `/infra`, least-privilege per T4/T12) grants:
`s3:GetObject` on `open-swe-<env>-assets/*`; `secretsmanager:GetSecretValue` on
`open-swe-<env>/*`; `ssm:GetParameter(s)`/`GetParametersByPath` on `/open-swe-<env>/*`;
`logs:*` for the CW agent log groups + `cloudwatch:PutMetricData`; SSM Session
Manager (`ssm:UpdateInstanceInformation`, `ssmmessages:*`) for shell access.
- **CDK substitutes** the `@@OPENSWE_ENV@@`, `@@ASSETS_BUCKET@@`, `@@SERVER_NAME@@`,
`@@ARTIFACT_PREFIX@@` tokens in `user-data.sh` when rendering the launch template.
- `:2024` binds `0.0.0.0` so the ALB hooks target group can reach `/webhooks/*`; it is
reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is
never internet-exposed — the ALB hooks rule is path-scoped to `/webhooks/*`.

View file

@ -1,102 +0,0 @@
#!/usr/bin/env bash
# Open SWE app deploy — RE-RUNNABLE (first boot + every subsequent release).
#
# Pulls the current release from S3 (open-swe-<env>-assets), builds the venv
# natively on the box, and restarts the service. This is the SINGLE source of the
# app-deploy procedure; it runs in two places:
#
# 1. first boot — user-data.sh decodes this script to /opt/open-swe/bin and
# calls it ONCE (non-fatal: if no release is published yet,
# nginx is already up and the box waits for the first deploy).
# 2. every release — the `open-swe-<env>-deploy` SSM document (CI fires it after
# uploading app.tar.gz / spa.tar.gz) runs this same script.
#
# It deploys CODE + STATIC ASSETS only. Secrets/config are NOT fetched here: the
# systemd unit's ExecStartPre=fetch-config.sh materializes the tmpfs .env on every
# (re)start, fail-fast — so `systemctl restart` below is what reloads config too.
#
# Contract:
# app.tar.gz = the Python source tree (pyproject.toml + uv.lock + agent/ +
# deploy/ + langgraph.json + README.md, NO ui/, NO .venv). The venv
# is built HERE with `uv sync` so it is native ARM64 and lives at
# the real runtime path (no cross-built / non-relocatable venv).
# spa.tar.gz = the built dashboard SPA (vite output: _shell.html + assets),
# extracted to the nginx web root.
set -euo pipefail
exec > >(tee -a /var/log/open-swe/deploy.log) 2>&1
echo "==> open-swe deploy start $(date -u +%FT%TZ)"
# Non-secret pointers written by user-data.sh (env, region, bucket, artifact prefix).
# shellcheck disable=SC1091
. /etc/open-swe/boot.env
export AWS_DEFAULT_REGION="${AWS_REGION:?boot.env missing AWS_REGION}"
: "${ASSETS_BUCKET:?boot.env missing ASSETS_BUCKET}"
: "${ARTIFACT_PREFIX:?boot.env missing ARTIFACT_PREFIX}"
# Fixed layout — must match provision.sh + user-data.sh + the templates.
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
WWW_ROOT="/var/www/open-swe"
ENV_FILE="/run/open-swe/.env"
UV_BIN="/usr/local/bin/uv"
UV_PYTHON_INSTALL_DIR="/opt/uv/python" # where provision.sh pre-installed py3.12
SERVICE_HOME="/opt/open-swe"
# Benign-vs-failure distinction: on a brand-new env no release is published yet.
# Treat "app.tar.gz absent in S3" as a benign no-op (exit 0) so first boot is not a
# scary failure; ONCE a release exists, any later step failing is loud (set -e).
if ! aws s3 ls "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" >/dev/null 2>&1; then
echo "==> no release published at s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/ yet — nothing to deploy"
exit 0
fi
echo "==> pull release from s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/"
tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" "${tmp}/app.tar.gz"
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/spa.tar.gz" "${tmp}/spa.tar.gz"
# Replace app source + SPA atomically-ish: clear the dirs (drops files removed in
# this release) then extract. The venv is rebuilt below, so wiping .venv too is
# fine — uv's cache (in the service home) makes the rebuild fast.
#
# Hardening: deploy.sh runs as root, so extract with --no-same-owner
# --no-same-permissions — files take root:root + umask perms (NOT the archive's
# uid/mode), so a tarball cannot land a setuid/setgid binary or a foreign-owned
# file; the chown -R below then hands the tree to the service user. (GNU tar also
# refuses `..`-escaping members by default.) Defense-in-depth: the only writer of
# this bucket is the CI OIDC app role, but the box never trusts the archive's
# ownership/mode regardless.
echo "==> install app source -> ${APP_DIR}"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$APP_DIR" "$WWW_ROOT"
find "$APP_DIR" -mindepth 1 -delete
find "$WWW_ROOT" -mindepth 1 -delete
tar --no-same-owner --no-same-permissions -xzf "${tmp}/app.tar.gz" -C "$APP_DIR"
tar --no-same-owner --no-same-permissions -xzf "${tmp}/spa.tar.gz" -C "$WWW_ROOT"
# langgraph reads ./.env from WorkingDirectory; point it at the tmpfs file the
# systemd ExecStartPre materializes.
ln -sfn "$ENV_FILE" "${APP_DIR}/.env"
chown -R "$SERVICE_USER":"$SERVICE_USER" "$APP_DIR" "$WWW_ROOT"
echo "==> build venv natively (uv sync --frozen --no-dev)"
# Run as the service user so the venv + uv cache are owned by it. Pin the
# pre-baked interpreter dir so uv never reaches out to download Python at deploy.
cd "$APP_DIR"
sudo -u "$SERVICE_USER" env \
HOME="$SERVICE_HOME" \
UV_PYTHON_INSTALL_DIR="$UV_PYTHON_INSTALL_DIR" \
UV_CACHE_DIR="${SERVICE_HOME}/.cache/uv" \
"$UV_BIN" sync --frozen --no-dev
echo "==> restart open-swe.service + reload nginx"
# ExecStartPre=fetch-config.sh fails-fast if secrets/config are missing, so a
# restart here surfaces a bad config as a failed unit (non-zero exit below).
systemctl restart open-swe.service
nginx -t && systemctl reload nginx
if systemctl is-active --quiet open-swe.service; then
echo "==> open-swe deploy OK $(date -u +%FT%TZ)"
else
echo "!! open-swe.service is not active after deploy (check fetch-config/secrets)"
exit 1
fi

View file

@ -1,143 +0,0 @@
# Open SWE base AMI — ARM64 (Graviton) Ubuntu 24.04 LTS.
#
# Builds the immutable base image for the single EC2 instance per env
# (open-swe-dev / open-swe-prod) in seahaven-vpc. The image bakes the runtime
# (uv + Python 3.12, nginx, awscli v2, CloudWatch agent) and the service-user /
# systemd / nginx TEMPLATES. It bakes NO secrets and NO env-specific values —
# those are materialized at first boot by user-data + deploy/seahaven/fetch-config.sh
# (Secrets Manager + SSM -> root-only tmpfs .env, fail-fast).
#
# Build: packer init . && packer build open-swe-base.pkr.hcl
# The resulting AMI id is pinned in infra/cdk.context.json (CDK cachedInContext:true);
# see README.md "AMI -> cdk.context.json pinning contract".
packer {
required_version = ">= 1.11.0, < 2.0.0"
required_plugins {
amazon = {
source = "github.com/hashicorp/amazon"
version = "1.3.6"
}
}
}
variable "aws_region" {
type = string
default = "us-east-1"
}
variable "instance_type" {
type = string
default = "t4g.medium" # ARM64 (Graviton) build host; runtime instances are ~t4g.large
}
variable "ami_name_prefix" {
type = string
default = "open-swe-base-arm64"
}
# Versions baked into the image. Pin and bump deliberately.
variable "python_version" {
type = string
default = "3.12"
}
variable "node_major" {
type = string
default = "24"
}
variable "uv_version" {
type = string
default = "0.11.24"
}
variable "cloudwatch_agent_deb_url" {
type = string
default = "https://amazoncloudwatch-agent.s3.amazonaws.com/ubuntu/arm64/latest/amazon-cloudwatch-agent.deb"
}
variable "awscli_zip_url" {
type = string
default = "https://awscli.amazonaws.com/awscli-exe-linux-aarch64.zip"
}
locals {
timestamp = formatdate("YYYYMMDD-hhmmss", timestamp())
}
# Latest Canonical Ubuntu 24.04 (Noble) arm64 server image.
source "amazon-ebs" "open-swe" {
region = var.aws_region
instance_type = var.instance_type
ssh_username = "ubuntu"
ami_name = "${var.ami_name_prefix}-${local.timestamp}"
# ASCII only — AWS rejects non-ASCII in the AMI Description attribute.
ami_description = "Open SWE base - Ubuntu 24.04 arm64 + uv/py3.12 + nginx + CW agent (templates only, no secrets)"
source_ami_filter {
filters = {
name = "ubuntu/images/hvm-ssd*/ubuntu-noble-24.04-arm64-server-*"
architecture = "arm64"
root-device-type = "ebs"
virtualization-type = "hvm"
}
owners = ["099720109477"] # Canonical
most_recent = true
}
# IMDSv2 required on the build host.
metadata_options {
http_endpoint = "enabled"
http_tokens = "required"
http_put_response_hop_limit = 1
}
# gp3 root, encrypted. Runtime root size is set by CDK; this is just the build host.
launch_block_device_mappings {
device_name = "/dev/sda1"
volume_size = 20
volume_type = "gp3"
encrypted = true
delete_on_termination = true
}
tags = {
Name = "open-swe-base-arm64"
Purpose = "open-swe-runtime-base"
ManagedBy = "packer"
}
}
build {
name = "open-swe-base"
sources = ["source.amazon-ebs.open-swe"]
# Stage the boot-time templates into the image. The destination dir must exist
# BEFORE a trailing-slash (contents-only) file upload — packer's file provisioner
# does not create it, and uploading the directory itself trips scp ("Is a
# directory"). So mkdir first, then upload the contents into it.
provisioner "shell" {
inline = ["mkdir -p /tmp/open-swe-templates"]
}
provisioner "file" {
source = "${path.root}/templates/"
destination = "/tmp/open-swe-templates"
}
provisioner "shell" {
environment_vars = [
"PYTHON_VERSION=${var.python_version}",
"NODE_MAJOR=${var.node_major}",
"UV_VERSION=${var.uv_version}",
"CLOUDWATCH_AGENT_DEB_URL=${var.cloudwatch_agent_deb_url}",
"AWSCLI_ZIP_URL=${var.awscli_zip_url}",
]
# {{ .Vars }} MUST be included or the environment_vars above never reach the
# script (provision.sh runs under `set -u` and fails on the first reference).
execute_command = "chmod +x {{ .Path }}; {{ .Vars }} sudo -E bash '{{ .Path }}'"
script = "${path.root}/scripts/provision.sh"
}
}

View file

@ -1,105 +0,0 @@
#!/usr/bin/env bash
# Packer provisioner for the Open SWE base AMI (ARM64 Ubuntu 24.04).
#
# Bakes the runtime + boot-time templates ONLY. No secrets, no env-specific
# values. Everything env-specific is materialized at first boot by user-data.sh
# + deploy/seahaven/fetch-config.sh.
set -euo pipefail
PYTHON_VERSION="${PYTHON_VERSION:-3.12}"
NODE_MAJOR="${NODE_MAJOR:-24}"
UV_VERSION="${UV_VERSION:-0.11.24}"
CLOUDWATCH_AGENT_DEB_URL="${CLOUDWATCH_AGENT_DEB_URL:?}"
AWSCLI_ZIP_URL="${AWSCLI_ZIP_URL:?}"
# Layout (must match user-data.sh and the templates).
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
SERVICE_HOME="/opt/open-swe"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
LOG_DIR="/var/log/open-swe"
UV_BIN="/usr/local/bin/uv"
export DEBIAN_FRONTEND=noninteractive
echo "==> apt base packages"
apt-get update -y
apt-get upgrade -y
apt-get install -y --no-install-recommends \
nginx jq curl unzip ca-certificates gnupg lsb-release \
build-essential pkg-config git acl
echo "==> awscli v2 (aarch64)"
tmp="$(mktemp -d)"
curl -fsSL "$AWSCLI_ZIP_URL" -o "$tmp/awscliv2.zip"
unzip -q "$tmp/awscliv2.zip" -d "$tmp"
"$tmp/aws/install" --update
rm -rf "$tmp"
aws --version
echo "==> CloudWatch agent (arm64)"
tmp="$(mktemp -d)"
curl -fsSL "$CLOUDWATCH_AGENT_DEB_URL" -o "$tmp/amazon-cloudwatch-agent.deb"
dpkg -i -E "$tmp/amazon-cloudwatch-agent.deb"
rm -rf "$tmp"
# Do NOT enable/start the agent during the build; user-data fetches its config
# (with env-specific log-group names + 30-day retention) and starts it at boot.
systemctl disable amazon-cloudwatch-agent.service || true
echo "==> uv ${UV_VERSION} + Python ${PYTHON_VERSION} (system-wide)"
export UV_INSTALL_DIR=/usr/local/bin
curl -fsSL "https://astral.sh/uv/${UV_VERSION}/install.sh" | env UV_NO_MODIFY_PATH=1 sh
"$UV_BIN" --version
# Pre-install the interpreter so the box never reaches out at boot to build a venv.
UV_PYTHON_INSTALL_DIR=/opt/uv/python "$UV_BIN" python install "$PYTHON_VERSION"
echo "==> node ${NODE_MAJOR} + bun (build-time UI tooling only; the SPA is built in CI)"
curl -fsSL "https://deb.nodesource.com/setup_${NODE_MAJOR}.x" | bash -
apt-get install -y --no-install-recommends nodejs
node --version
# bun installed system-wide; used only if any UI tooling must run on-box. The
# production SPA build runs in GitHub Actions -> S3 (no on-box build, no swapfile).
export BUN_INSTALL=/usr/local
curl -fsSL https://bun.sh/install | bash
/usr/local/bin/bun --version || true
echo "==> non-login service user '${SERVICE_USER}'"
if ! id "$SERVICE_USER" >/dev/null 2>&1; then
useradd --system --create-home --home-dir "$SERVICE_HOME" \
--shell /usr/sbin/nologin "$SERVICE_USER"
fi
echo "==> directories"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$SERVICE_HOME" "$APP_DIR"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$WWW_ROOT"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0750 "$LOG_DIR"
install -d -o root -g root -m 0755 "$TEMPLATE_DIR"
echo "==> stage boot-time templates into the image"
cp /tmp/open-swe-templates/* "$TEMPLATE_DIR/"
chown root:root "$TEMPLATE_DIR"/*
chmod 0644 "$TEMPLATE_DIR"/*
rm -rf /tmp/open-swe-templates
echo "==> tmpfs for the runtime .env (root/owner-only, noexec/nosuid/nodev)"
# /run is already tmpfs on Ubuntu; this is an explicit, deliberately-small mount
# scoped to the service user so the materialized .env never touches disk.
if ! grep -q '/run/open-swe' /etc/fstab; then
cat >>/etc/fstab <<EOF
tmpfs /run/open-swe tmpfs rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m 0 0
EOF
fi
echo "==> disable nginx default site (open-swe site is installed at boot)"
rm -f /etc/nginx/sites-enabled/default
systemctl enable nginx
echo "==> harden: no password auth, IMDSv2 already enforced by launch template"
# (sshd is not exposed publicly — instance is in a private subnet, SG inbound = ALB only.)
echo "==> clean apt caches"
apt-get clean
rm -rf /var/lib/apt/lists/*
echo "==> provision complete"

View file

@ -1,55 +0,0 @@
{
"agent": {
"metrics_collection_interval": 60,
"run_as_user": "root"
},
"metrics": {
"namespace": "open-swe/@@OPENSWE_ENV@@",
"append_dimensions": {
"InstanceId": "${aws:InstanceId}"
},
"metrics_collected": {
"mem": { "measurement": ["mem_used_percent"] },
"disk": {
"measurement": ["used_percent"],
"resources": ["/"]
}
}
},
"logs": {
"logs_collected": {
"files": {
"collect_list": [
{
"file_path": "/var/log/open-swe/app.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/app",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/open-swe-user-data.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/user-data",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/access.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-access",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/error.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-error",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
}
]
}
}
}
}

View file

@ -1,62 +0,0 @@
# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy.
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
#
# nginx is the SOLE ingress and security boundary (T5 OSWE-IAC-03): the backend
# binds 127.0.0.1:2024 and is NOT network-reachable. nginx proxies exactly two
# prefixes to it — /dashboard/api/* and /webhooks/* — and nothing else. The
# unauthenticated LangGraph agent API (/threads, /runs, /assistants, /store) is
# NEVER proxied; those paths return the SPA shell.
#
# Both ALB target groups (dashboard host + hooks host) point at this nginx :80,
# not at :2024 directly, so there is no path to the raw control plane even from
# inside the SG. Webhook signature verification still happens in the app (the raw
# body + GitHub/Slack/Linear signature headers are passed through unmodified).
server {
listen 80 default_server;
listen [::]:80 default_server;
server_name @@SERVER_NAME@@;
root @@WWW_ROOT@@;
index _shell.html;
# ALB target-group health check (dashboard TG).
location = /healthz { default_type text/plain; return 200 "ok\n"; }
# Dashboard API + OAuth callback -> backend webapp.
location /dashboard/api/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
client_max_body_size 10m;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 300s;
}
# Inbound webhooks (GitHub/Slack/Linear) -> backend webapp. Routed through
# nginx so :2024 stays loopback-only (T5 OSWE-IAC-03). The raw request body +
# signature headers pass through unmodified for in-app signature verification.
location /webhooks/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
# GitHub permits webhook payloads up to 25 MB; nginx's 1 MB default would
# 413 large push/PR events at the edge BEFORE in-app signature verification
# runs, silently dropping them (OSWE-T12-01). proxy_request_buffering off
# does not relax the size cap — set it explicitly.
client_max_body_size 25m;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_request_buffering off;
proxy_read_timeout 300s;
}
# Static assets + SPA shell fallback (client-side routing).
location / {
try_files $uri $uri/ /_shell.html;
}
}

View file

@ -1,54 +0,0 @@
# Open SWE — stock LangGraph dev server (3+ graphs + FastAPI webapp, :2024).
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
# In-memory runtime (--no-reload) + ExecStartPost reseed; no Aegra/Postgres.
#
# Boot contract:
# ExecStartPre = fetch-config.sh <env> -> runs as root (`+`) ONLY to materialize
# the SERVICE-USER-owned tmpfs .env (@@ENV_FILE@@) from Secrets
# Manager + SSM and chown it to @@SERVICE_USER@@, fail-fast (the
# unit does NOT start if config can't be fetched).
# ExecStart = langgraph dev (as @@SERVICE_USER@@, bound to 127.0.0.1 — nginx
# is the sole ingress; never binds 0.0.0.0).
# ExecStartPost = seed_store.sh <env> -> reseeds team_settings + user_mappings
# that the in-memory store loses on every restart.
[Unit]
Description=Open SWE stock LangGraph dev server (graphs + webapp, :@@PORT@@)
After=network-online.target
Wants=network-online.target
RequiresMountsFor=/run/open-swe
[Service]
Type=simple
User=@@SERVICE_USER@@
Group=@@SERVICE_USER@@
WorkingDirectory=@@APP_DIR@@
# No EnvironmentFile: the secret-bearing app .env (@@ENV_FILE@@) is loaded by
# langgraph/dotenv (so the multiline GitHub App PEM never hits systemd's env
# parser), and seed_store.sh reads the same .env directly (without sourcing it).
#
# ExecStartPre runs as root (`+`) so it can chown the tmpfs .env to the service
# user; the env arg (@@OPENSWE_ENV@@) selects the SSM/Secrets prefix (T5 BOOT-01).
ExecStartPre=+@@FETCH_CONFIG@@ @@OPENSWE_ENV@@
# Bind 127.0.0.1 only — nginx proxies dashboard + webhooks; :@@PORT@@ is never
# directly network-reachable (T5 OSWE-IAC-03).
ExecStart=@@VENV@@/bin/langgraph dev --host 127.0.0.1 --port @@PORT@@ --no-browser --no-reload
ExecStartPost=@@SEED_STORE@@ @@OPENSWE_ENV@@
# App logs to a file CloudWatch collects (30-day retention set in the CW config).
StandardOutput=append:/var/log/open-swe/app.log
StandardError=append:/var/log/open-swe/app.log
Restart=on-failure
RestartSec=5
TimeoutStartSec=180
# Hardening — the box holds no durable state of its own.
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
PrivateTmp=true
ReadWritePaths=/var/log/open-swe /var/www/open-swe /run/open-swe @@APP_DIR@@
[Install]
WantedBy=multi-user.target

View file

@ -1,145 +0,0 @@
#!/usr/bin/env bash
# Open SWE EC2 user-data — PROVISIONING-ONLY (runs once, at first boot).
#
# This is the rationale for `userDataCausesReplacement: true` in CDK: user-data
# does FIRST-BOOT provisioning, never durable runtime config. Editing it is a
# deliberate instance replacement. Durable runtime config is fetched fresh on
# every service start by deploy/seahaven/fetch-config.sh (ExecStartPre).
#
# The box holds NO durable state of its own:
# - secrets/config -> Secrets Manager + SSM, materialized to a tmpfs .env at boot
# - app artifact -> pulled from S3 (open-swe-<env>-assets) via the instance role
# - store state -> reseeded by seed_store.sh (ExecStartPost) on every start
# => there is no RETAIN volume to protect; replacement is tolerated. The EBS
# discipline (snapshot root + wait state=completed BEFORE any replacing deploy)
# is the safety net, not durable on-box state. See README "EBS discipline".
#
# Tokens (@@...@@) are substituted by CDK when it renders this script into the
# launch template. region is read from IMDSv2 as a fallback.
set -euo pipefail
exec > >(tee -a /var/log/open-swe-user-data.log) 2>&1
echo "==> open-swe user-data start $(date -u +%FT%TZ)"
# --- CDK-rendered values -----------------------------------------------------
# NOTE: CDK-substituted tokens use %%...%% (rendered by app-service.ts), DISTINCT
# from the @@...@@ tokens this script seds into the baked systemd/nginx templates.
# The two MUST NOT share a delimiter: a shared @@OPENSWE_ENV@@ / @@SERVER_NAME@@
# let CDK clobber the sed PATTERN, leaving the unit's token unsubstituted.
OPENSWE_ENV="%%OPENSWE_ENV%%" # dev | prod
ASSETS_BUCKET="%%ASSETS_BUCKET%%" # open-swe-<env>-assets
SERVER_NAME="%%SERVER_NAME%%" # openswe[-dev].seahaven.com
ARTIFACT_PREFIX="%%ARTIFACT_PREFIX%%" # e.g. releases/latest
# --- fixed layout (must match provision.sh + templates) ----------------------
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
VENV="${APP_DIR}/.venv"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
ENV_FILE="/run/open-swe/.env"
PORT="2024"
FETCH_CONFIG="${APP_DIR}/deploy/seahaven/fetch-config.sh"
SEED_STORE="${APP_DIR}/deploy/seahaven/seed_store.sh"
# region from IMDSv2
TOKEN="$(curl -fsS -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 300" || true)"
AWS_REGION="$(curl -fsS -H "X-aws-ec2-metadata-token: ${TOKEN}" \
http://169.254.169.254/latest/meta-data/placement/region || echo us-east-1)"
export AWS_DEFAULT_REGION="$AWS_REGION"
echo "env=${OPENSWE_ENV} region=${AWS_REGION} bucket=${ASSETS_BUCKET} host=${SERVER_NAME}"
# --- boot.env: non-secret pointers fetch-config.sh reads ---------------------
install -d -o root -g root -m 0755 /etc/open-swe
cat >/etc/open-swe/boot.env <<EOF
OPENSWE_ENV=${OPENSWE_ENV}
AWS_REGION=${AWS_REGION}
ASSETS_BUCKET=${ASSETS_BUCKET}
ARTIFACT_PREFIX=${ARTIFACT_PREFIX}
ENV_FILE=${ENV_FILE}
SECRETS_PREFIX=open-swe-${OPENSWE_ENV}
SSM_PREFIX=/open-swe-${OPENSWE_ENV}
EOF
chmod 0644 /etc/open-swe/boot.env
# --- ensure the tmpfs for the materialized .env is mounted -------------------
# (baked into /etc/fstab by the AMI; mount it now in case it isn't yet.)
install -d -o root -g root -m 0755 /run/open-swe || true
mountpoint -q /run/open-swe || mount /run/open-swe || mount -t tmpfs \
-o rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m \
tmpfs /run/open-swe
# --- install the deploy script (single source of the app-deploy procedure) ---
# deploy.sh (deploy/ami/deploy.sh) pulls the release from S3, builds the venv with
# `uv sync`, and restarts the service. CDK base64-renders the file into the
# %%DEPLOY_SH_B64%% token below so it is a normal reviewable repo file, not an
# inline heredoc. The `open-swe-<env>-deploy` SSM document runs this same script
# for every subsequent release.
echo "==> install /opt/open-swe/bin/deploy.sh"
install -d -o root -g root -m 0755 /opt/open-swe/bin
base64 -d >/opt/open-swe/bin/deploy.sh <<'DEPLOY_SH_B64'
%%DEPLOY_SH_B64%%
DEPLOY_SH_B64
chmod 0755 /opt/open-swe/bin/deploy.sh
# --- render + install the systemd unit ---------------------------------------
echo "==> install systemd unit"
sed \
-e "s|@@SERVICE_USER@@|${SERVICE_USER}|g" \
-e "s|@@APP_DIR@@|${APP_DIR}|g" \
-e "s|@@VENV@@|${VENV}|g" \
-e "s|@@PORT@@|${PORT}|g" \
-e "s|@@ENV_FILE@@|${ENV_FILE}|g" \
-e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
-e "s|@@FETCH_CONFIG@@|${FETCH_CONFIG}|g" \
-e "s|@@SEED_STORE@@|${SEED_STORE}|g" \
"${TEMPLATE_DIR}/open-swe.service" >/etc/systemd/system/open-swe.service
systemctl daemon-reload
# --- render + install the nginx site -----------------------------------------
echo "==> install nginx site"
sed \
-e "s|@@SERVER_NAME@@|${SERVER_NAME}|g" \
-e "s|@@WWW_ROOT@@|${WWW_ROOT}|g" \
-e "s|@@BACKEND_ADDR@@|127.0.0.1:${PORT}|g" \
"${TEMPLATE_DIR}/open-swe.nginx.conf" >/etc/nginx/sites-available/open-swe
ln -sfn /etc/nginx/sites-available/open-swe /etc/nginx/sites-enabled/open-swe
rm -f /etc/nginx/sites-enabled/default
nginx -t
# --- CloudWatch agent: 30-day log retention ----------------------------------
echo "==> configure CloudWatch agent (30-day retention)"
sed -e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
"${TEMPLATE_DIR}/amazon-cloudwatch-agent.json" \
>/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
/opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
-a fetch-config -m ec2 -s \
-c file:/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
# --- start nginx FIRST (the security boundary + health surface) --------------
# NOTE: intentionally NO swapfile here. The 8 GB-swapfile OOM hack existed only
# for the on-box Nitro SPA build, which now runs in GitHub Actions -> S3.
# nginx is brought up BEFORE the app is deployed so the ALB target-group health
# check (static `/healthz` -> 200) passes and the box is a healthy target even on
# the very first boot, before any release is published. open-swe.service is
# enabled (boot persistence) but STARTED by deploy.sh once the app is on disk.
echo "==> start nginx"
systemctl enable --now nginx
systemctl reload nginx
systemctl enable open-swe.service
# --- deploy the app (NON-FATAL on first boot) --------------------------------
# deploy.sh pulls the release, builds the venv, and starts open-swe.service. On a
# brand-new env no release exists yet, so this is allowed to fail WITHOUT aborting
# user-data: nginx is already up (healthy target), and the first `build-artifacts`
# run + `open-swe-<env>-deploy` SSM command will bring the app up. A failure here
# is logged, not fatal.
echo "==> initial app deploy (non-fatal if no release is published yet)"
if /opt/open-swe/bin/deploy.sh; then
echo "==> initial app deploy succeeded"
else
echo "==> no release yet (or deploy failed): open-swe.service deferred to the next SSM deploy"
fi
echo "==> open-swe user-data done $(date -u +%FT%TZ)"

View file

@ -1,280 +0,0 @@
# Sea Haven — Open SWE deployment runbook
How this fork is deployed at Sea Haven. **PROD is LIVE as of 2026-06-29.** The
runtime is the **stock LangGraph dev server** (not the Aegra path — see
[Aegra](#aegra-deferred)), running on a self-hosted ARM64 EC2 box behind the
shared `seahaven-com` ALB.
Account `328440206208`, region `us-east-1`. Internal addresses, ARNs, snapshot
ids, and account-scoped values that are sensitive are shown as `<PLACEHOLDERS>`;
the real values live in the private IT docs (Confluence "AWS Architecture Map",
id 1540098) and in AWS — **do not commit them to this public fork.**
This is the canonical deploy runbook. The CDK details live in
[`infra/README.md`](../../infra/README.md); the rotation procedure in
[`ROTATION.md`](ROTATION.md).
## Live prod facts (2026-06-29)
| | |
|---|---|
| Dashboard | `https://openswe.seahaven.com` |
| Webhooks | `https://hooks.seahaven.com/webhooks/*` |
| Ingress | shared internet-facing ALB `app/seahaven-com` → target group `open-swe-prod-tg` → EC2 `i-08a729e50779c4b07` (`t4g.large`, ARM64) on nginx `:80` |
| Backend | `langgraph dev` bound to `127.0.0.1:2024` (loopback only); nginx is the sole ingress |
| CDK stacks | `open-swe-iam` (OIDC roles) · `open-swe-dev` · `open-swe-prod` |
Dev mirrors prod with `-dev` hosts (`openswe-dev.seahaven.com` /
`hooks-dev.seahaven.com`), a `t4g.medium` box, and no GitHub-App/Slack/webhook
integration (it is a deployment-validation env, not a live-triggered agent).
> The retired on-prem `*.seahavenind.com` ALB routing and DNS were removed on
> 2026-06-29; prod is now live exclusively on `*.seahaven.com`.
## Hosting model
```
GitHub / Slack ──▶ hooks.seahaven.com ──┐
│ (shared ALB :443, host+path rules)
Browser ─────────▶ openswe.seahaven.com ──┤
▼
ALB app/seahaven-com ──▶ open-swe-prod-tg ──▶ EC2 box :80 (nginx)
├─ nginx — SPA + scoped proxy
│ /dashboard/api/* and /webhooks/*
└─ langgraph dev 127.0.0.1:2024
└─▶ LangSmith cloud sandbox (build/git/PR)
```
- A **single** VPC and a **single** internet-facing ALB (`app/seahaven-com`) are
shared with the on-prem `seahaven-site` stack. open-swe **imports** the VPC,
ALB SG, `:443` listener, and `seahaven.com` zone — it never owns/mutates them;
it only adds its own instance SG, a standalone ALB-egress rule, two listener
rules, a target group, and Route53 aliases.
- The EC2 box is in a **private** subnet (us-east-1a, same AZ as the single NAT
for in-AZ egress). It is reachable **only** from the shared ALB SG on `:80`.
- **nginx is the security boundary.** It serves the static dashboard SPA and
proxies exactly two prefixes to `:2024` — `/dashboard/api/*` and `/webhooks/*`.
The unauthenticated LangGraph API (`/threads`, `/runs`, `/assistants`,
`/store`) is never proxied; those paths return the SPA shell. `:2024` is
loopback-only and never network-reachable, even inside the SG.
- **Webhooks** ride listener rules below the on-prem host-agnostic `/webhooks/*`
rule (priority 2 dev / 3 prod, host-scoped to the open-swe hosts) so they reach
the open-swe box and never steal an on-prem host's webhooks.
The box holds **no durable state of its own**: secrets/config are materialized to
a tmpfs `.env` at boot, the app artifact is pulled from S3, and the in-memory
LangGraph store is re-seeded on every start. Replacement is tolerated; there is no
RETAIN volume.
---
## Deploy pipeline (end to end)
Two independent CD lanes, both OIDC-only (no static keys), both with a manual
approval gate on prod via the GitHub **`prod` Environment** (required reviewer:
Adam). The `environment: prod` declaration both fires the approval gate and makes
the OIDC subject `…:environment:prod`, which is the only subject the prod deploy
roles trust — so a dev-branch token can never reach prod.
### (a) Infra CD — `cd-infra.yml`
Deploys the CDK stacks. Path-filtered to `infra/**`.
```
push to dev → Infra CI (tsc + jest + cdk synth) → cdk deploy OpenSweDevStack (AUTO, CI-green-gated)
push to main → Infra CI → cdk deploy OpenSweProdStack (manual approval: env "prod")
```
- Roles: `githubdeploy-open-swe-infra-{dev,prod}` (in the `open-swe-iam` stack;
set as repo variables `AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}`).
- It targets **one stack explicitly per env** (`cdk deploy OpenSweDevStack` /
`OpenSweProdStack`), not `cdk deploy --all`, so a single-env push can never
deploy the other env or the shared IAM stack.
- The shared `open-swe-iam` stack (owns both envs' OIDC deploy roles) is **not**
deployed by CD — it is a privileged, human-gated apply.
**Stack order on a clean account:** `open-swe-iam` first (creates the OIDC roles;
set the repo deploy-role variables and configure the `prod` Environment reviewer
from its outputs), then `open-swe-dev`, then `open-swe-prod`.
### (b) Seed the config store — `put-config.sh <env>`
Run **after** `cdk deploy open-swe-<env>` and **before** the box first boots. CDK
creates the value-less Secrets Manager shells (`open-swe-<env>/<VAR>`) and the
IaC-managed SSM params (`/open-swe-<env>/<VAR>`); `put-config.sh` populates the
secret values plus the out-of-band SSM params that cannot live in IaC.
```bash
deploy/seahaven/put-config.sh <dev|prod> # set each value inline, via OPENSWE_PUT_<VAR>, or from a vault
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
```
`put-config.sh` ships `<FILL>` placeholders only — **no real secret values are
committed**. It does not touch the IaC-managed SSM params (CDK owns those).
**13 prod boot-required vars** — `fetch-config.sh` fail-fasts (refuses to write a
partial `.env`, the unit does not start) if any are missing/empty:
- **9 secrets** (Secrets Manager `open-swe-prod/<VAR>`): `DASHBOARD_JWT_SECRET`,
`TOKEN_ENCRYPTION_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`LANGSMITH_API_KEY_PROD`, `GITHUB_APP_PRIVATE_KEY`, `GITHUB_APP_CLIENT_SECRET`,
`GITHUB_WEBHOOK_SECRET`, `SLACK_SIGNING_SECRET`.
- **4 SSM params** (`/open-swe-prod/<VAR>`): `DEFAULT_SANDBOX_SNAPSHOT_ID`,
`GITHUB_APP_ID`, `GITHUB_APP_INSTALLATION_ID`, `GITHUB_APP_CLIENT_ID`.
(`ANTHROPIC_API_KEY` + `OPENAI_API_KEY` are required because that is the seeded
cross-family pair; the active set follows `REQUIRED_PROVIDER_KEYS`. The
`LANGSMITH_API_KEY_PROD` + `DEFAULT_SANDBOX_SNAPSHOT_ID` pair is required because
`SANDBOX_TYPE=langsmith`.) Dev boots without the GitHub-App / Slack / webhook
secrets — it has no such integration.
`fetch-config.sh` runs as an `ExecStartPre=+` hook (root, only long enough to
write the `openswe`-owned `0600` tmpfs `.env`), reads all `/open-swe-<env>/*` SSM
params + all `open-swe-<env>/*` secrets via the instance role, and forces
`DEFAULT_REPO_OWNER` away from the upstream `langchain-ai` org.
### (c) App artifact deploy — `build-artifacts.yml`
Builds the release and rolls the box. Path-filtered to `agent/**`, `ui/**`,
`deploy/**`, `langgraph.json`, `pyproject.toml`, `uv.lock`.
```
push to dev → build SPA + package → open-swe-dev-assets/releases/ → SSM open-swe-dev-deploy (AUTO)
push to main → build SPA + package → open-swe-prod-assets/releases/ → SSM open-swe-prod-deploy (manual approval: env "prod")
```
1. The dashboard SPA is built **on the runner** (`bun run build` → vite →
`ui/.output/public`) — the box is small, so the memory-heavy build runs in CI.
2. `package-artifacts.sh` produces two tarballs: `spa.tar.gz` (built SPA) and
`app.tar.gz` (Python source tree — no `ui/`, no `.venv`).
3. Both are uploaded to S3 `open-swe-<env>-assets` under `releases/<sha>/`
(immutable, auditable) and mirrored to `releases/latest/` (what the box pulls).
4. CI fires the `open-swe-<env>-deploy` SSM document (tag-scoped to
`project=open-swe,env=<env>`), which runs `/opt/open-swe/bin/deploy.sh` on the
box: pull the release from S3, build a native-ARM64 venv with
`uv sync --frozen --no-dev`, extract the SPA to the nginx web root,
`systemctl restart open-swe.service`, reload nginx, then gate on
`systemctl is-active --quiet open-swe.service` (a non-active unit exits the
deploy non-zero).
Roles: `githubdeploy-open-swe-app-{dev,prod}` (repo variables
`AWS_DEPLOY_ROLE_APP_{DEV,PROD}`) — tag-scoped `ssm:SendCommand` on the deploy
document only (not the generic `AWS-RunShellScript`) + write to the env's S3
bucket.
Secrets/config are **not** fetched by `deploy.sh`; the `systemctl restart`'s
`ExecStartPre=fetch-config.sh` re-materializes the `.env` on every restart, so a
bad config surfaces as a failed unit.
### (d) dev → main promotion + rollback
**Promotion — `promote-dev-to-prod.yml`** (nightly cron `0 8 * * *` + manual
dispatch): mints a GitHub App installation token (a bypass actor on the `main`
ruleset), gates on **every check-run on the dev HEAD commit being completed and
passing**, then **fast-forward-only** pushes `dev` → `main`. A diverged `main`
fails loudly rather than force-updating. The push to `main` is what triggers the
prod lanes of `cd-infra.yml` / `build-artifacts.yml` (each still behind the `prod`
Environment approval). Re-gating via a PR on `main` would be redundant since the
commit already passed every check on dev.
**Rollback — `rollback.yml`** (manual dispatch, `env` + optional `sha`): re-points
`releases/latest/` at a prior release and re-fires the `open-swe-<env>-deploy` SSM
document — same fire/wait/gate path as a forward deploy, no rebuild.
```
env=dev, sha blank → restore open-swe-dev-assets/releases/last-good/ (AUTO)
env=prod, sha blank → restore open-swe-prod-assets/releases/last-good/ (manual approval: env "prod")
sha=<commit> → restore that exact releases/<sha>/ instead
```
It reuses the existing `githubdeploy-open-swe-app-<env>` role (no new IAM).
---
## On-box layout (reference)
| Path | What |
|---|---|
| `open-swe.service` (systemd) | `langgraph dev --host 127.0.0.1 --port 2024 --no-browser --no-reload` as the unprivileged `openswe` user. In-memory runtime. |
| `fetch-config.sh` | `ExecStartPre=+` — materializes the tmpfs `.env` from Secrets Manager + SSM, fail-fast. |
| `seed_store.sh` | `ExecStartPost` — re-seeds `team_settings/default` + `user_mappings` (the in-memory store loses them on every restart). |
| nginx | SPA from `/var/www/open-swe`, proxy `/dashboard/api/` + `/webhooks/` → `127.0.0.1:2024`, `/healthz` → 200. |
| `deploy.sh` | the release procedure run on first boot (non-fatal) and by every SSM deploy. |
| CloudWatch logs | `/open-swe/<env>/{app,user-data,nginx-access,nginx-error}` at 30-day retention. |
The live systemd unit + nginx site are the **AMI templates**
(`deploy/ami/templates/open-swe.service`, `open-swe.nginx.conf`), rendered at
first boot by `deploy/ami/user-data.sh`. The AMI is the baked
`open-swe-base-arm64` image (Ubuntu 24.04 + uv/py3.12 + nginx + CW agent), pinned
by exact id in `infra/lib/constructs/ami-cache.ts`. There is intentionally **no**
on-box swapfile — the OOM-prone SPA build now runs in CI, not on the box.
> `deploy/seahaven/{nginx/openswe.conf,systemd/open-swe.service}` are the
> **retired on-prem VM** variants (run as `adam` from a home dir, bound `0.0.0.0`,
> Postgres-backed). They are kept only for on-prem-contrast reference and are not
> used by the AWS deployment.
### Models
Model selection is **store-driven**, not env. The `team_settings/default` store
doc wins (then per-user profile, then per-thread); `LLM_MODEL_ID` is only a
seed-time fallback. Defaults seeded by `seed_store.sh`:
- builder: `bedrock_converse:us.anthropic.claude-opus-4-8` (effort `high`)
- reviewer: `bedrock_converse:us.anthropic.claude-opus-4-8` (effort `high`) — set
`SEED_REVIEWER_MODEL` (or change it in the UI) to a Fireworks model if you want a
cross-family reviewer. Only ids present in `SUPPORTED_MODELS`
(`agent/dashboard/options.py`) are valid; OpenAI/Google models were removed in the
Bedrock/Fireworks migration.
- the `analyzer` graph is hardcoded to the code default and ignores team settings.
## Triggering
Mention **`@openswe`** (or `@open-swe` / `@seahaven-openswe`) in a GitHub issue or
PR comment, a Linear comment, or a Slack thread. The commenter must have a
`user_mappings` entry (seeded by `seed_store.sh` from `CONFIGURED_ADMINS` /
`SEED_USER_MAPPINGS`) or the run is skipped.
Live integration endpoints (set in each provider's app config):
| Integration | URL |
|---|---|
| GitHub webhook | `https://hooks.seahaven.com/webhooks/github` |
| Slack events | `https://hooks.seahaven.com/webhooks/slack` (+ `/webhooks/slack/interactivity`) |
| Linear webhook | `https://hooks.seahaven.com/webhooks/linear` |
| GitHub OAuth callback | `https://openswe.seahaven.com/dashboard/api/auth/callback` |
---
## Troubleshooting
**RETAIN secret-shell orphan on stack re-create.** The Secrets Manager shells use
`DeletionPolicy: Retain` + a fixed `open-swe-<env>/<VAR>` name. If a stack's first
create rolls back (or on a teardown/rebuild, a secret logical-id refactor, or
standing up a new env), the empty shells survive and keep their global names, so
every later create fails `AlreadyExists` — and a plain `delete-secret` does not
free the name (it stays reserved for the 7–30 day recovery window). Before
re-creating the stack, **force-delete the empty orphans** (only shells with no
value version — never a populated secret). Hit on prod 2026-06-29 (PR #51 deploy
failure). Full recovery command + rationale:
[`infra/README.md`](../../infra/README.md) (PR #52).
**`langgraph dev` won't start after a deploy.** `fetch-config.sh` fail-fasts on a
missing/empty required var and prints the offending variable **names** (never
values) to the unit journal. Confirm the 13 prod boot-required vars are populated
(`put-config.sh prod`), then `systemctl restart open-swe.service`.
**ALB target unhealthy.** The TG health check is `GET /healthz` on nginx `:80`
(static 200). nginx starts before the app on first boot, so an unhealthy target
usually means the box can't reach the ALB SG on `:80` (the standalone ALB-egress
rule) rather than an app fault.
---
## Aegra (deferred)
`aegra/aegra.json` + `aegra/aegra_entry.py` are the self-hosted-runtime
alternative (Apache-2.0, avoids the LangGraph-Platform Elastic license). Not
active on the stock deployment. To use: place both at the repo root, run
`aegra serve` (:2026), and point `LANGGRAPH_URL` at `:2026`. Aegra gives a
Postgres-backed durable store/checkpointer, which removes the need for
`seed_store.sh` and survives restarts (paused HITL interrupts persist).
</content>

View file

@ -1,92 +0,0 @@
# Sea Haven — Open SWE secret & config rotation
How secrets and config reach the running app, and how to rotate either one.
## How values flow at boot
```
AWS Secrets Manager open-swe-<env>/* ─┐
AWS SSM Param Store /open-swe-<env>/* ─┤── fetch-config.sh ──▶ tmpfs /run/open-swe/.env (root:root 0600)
│ (systemd ExecStartPre=+, EC2 role) │
▼ ▼
FAIL-FAST if a app symlink <APP_DIR>/.env
required var is empty python-dotenv reads at import
```
The app reads `.env` **once, at import**. There is no hot-reload of secrets.
Therefore the rotation contract is always the same two steps:
> **Rotation = (1) update the value in Secrets Manager / SSM, then (2) restart the
> service** so `fetch-config.sh` re-materializes the `.env`.
```bash
# after updating a secret/param in AWS:
sudo systemctl restart open-swe.service
# ExecStartPre=+ -> fetch-config.sh re-pulls + rewrites the tmpfs .env (fail-fast)
# ExecStartPost -> seed_store.sh re-seeds the in-memory store (team_settings + user_mappings)
```
There is **no zero-downtime path for most secrets** on the stock in-memory
runtime — a restart is required and it also wipes the in-memory store (re-seeded
by `seed_store.sh` automatically). The one secret built for zero-downtime overlap
is `TOKEN_ENCRYPTION_KEY` (see below), but even it needs the restart to load the
new key list.
## Rotating a secret (Secrets Manager)
```bash
ENV=prod # or dev
NAME=DASHBOARD_JWT_SECRET
aws secretsmanager put-secret-value \
--secret-id "open-swe-${ENV}/${NAME}" \
--secret-string 'NEW_VALUE' \
--region us-east-1
sudo systemctl restart open-swe.service # on the box
```
(`update-secret`/`put-secret-value` both create a new version; the boot hook
always reads `AWSCURRENT`.)
## Rotating a config param (SSM)
```bash
aws ssm put-parameter --overwrite \
--name "/open-swe-${ENV}/DASHBOARD_BASE_URL" \
--type String --value 'https://openswe.seahaven.com' \
--region us-east-1
sudo systemctl restart open-swe.service
```
## Per-secret rotation notes
| Secret | Rotation notes |
|---|---|
| **TOKEN_ENCRYPTION_KEY** | Fernet key(s). Supports a **comma/newline-separated list** (`agent/encryption.py`) for zero-downtime key rotation: prepend the NEW key, keep the OLD key(s) in the list. New data is encrypted with the first key; old data still decrypts with the trailing keys. After all encrypted-at-rest tokens (per-user GitHub OAuth tokens in thread metadata) have been re-encrypted/expired, drop the old key. Store the list as one secret value; `fetch-config.sh` writes it verbatim. **Never** rotate to a single new key in one step or every existing encrypted token becomes undecryptable. |
| **GITHUB_APP_PRIVATE_KEY** | Multiline PEM. Generate a new private key in the GitHub App settings (you may have **two active keys** during overlap), put the new PEM into the secret, restart, verify install-token minting + a webhook delivery, then delete the old key in GitHub. `fetch-config.sh` writes the PEM as a double-quoted multiline value (python-dotenv-safe); paste the full `-----BEGIN…-----END-----` block including newlines. |
| **GITHUB_WEBHOOK_SECRET** | Webhook HMAC. GitHub allows only **one** webhook secret per App, so this is a brief-break rotation: update the secret in AWS **and** the GitHub App webhook config, restart. Deliveries signed with the old secret during the gap will 401 (GitHub auto-redelivers). Required in **prod** (fail-fast). |
| **SLACK_SIGNING_SECRET** | Slack request-signature secret. Rotate in the Slack app config and AWS together, restart. Required in **prod** (fail-fast). A stale value silently 401s `url_verification`/events until restart (known gotcha). |
| **LINEAR_WEBHOOK_SECRET** | Linear webhook signature. Required in prod only when the Linear integration is wired (`LINEAR_API_KEY` present). Rotate in Linear + AWS together, restart. |
| **SLACK_CLIENT_SECRET / GITHUB_APP_CLIENT_SECRET** | OAuth client secrets (dashboard login / Slack OAuth). Rotate in the provider console + AWS, restart. Existing dashboard sessions are JWT-signed by `DASHBOARD_JWT_SECRET`, not these, so they survive. |
| **DASHBOARD_JWT_SECRET** | Signs dashboard session cookies. Rotating **invalidates all active sessions** (users re-login). Hard-required (RuntimeError if empty). No overlap list — single value. |
| **Model provider keys** (`FIREWORKS_API_KEY`; legacy `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` / `GOOGLE_API_KEY` / `GROQ_API_KEY`) | Standard API-key rotation: issue new key, update AWS, restart, revoke old. Bedrock (the seeded Claude builder + reviewer) authenticates via the **host IAM role — no API key**; the only fail-fast-required provider key is now `FIREWORKS_API_KEY` (fallback / subagents / any non-Claude model). Keep `REQUIRED_PROVIDER_KEYS` in sync if you change the seeded models. |
| **LANGSMITH_API_KEY_PROD** | LangSmith key powering the `langsmith` sandbox (the only provider with working in-sandbox git/gh auth). Required when `SANDBOX_TYPE=langsmith` (fail-fast). Rotate in LangSmith + AWS, restart; existing sandboxes keep their already-injected proxy token until recycled. |
| **SLACK_BOT_TOKEN / LINEAR_API_KEY / GITHUB_PAT / EXA_API_KEY / DAYTONA_API_KEY / RUNLOOP_API_KEY / CORRIDOR_* / USER_ID_API_KEY_MAP / X_SERVICE_AUTH_JWT_SECRET** | Plain API-token rotation: update AWS, restart, revoke old at the provider. None support an overlap list. |
## Fail-fast safety
`fetch-config.sh` refuses to write the `.env` (exit 1) if any required var is
empty after a rotation — so a botched rotation (e.g. an empty `put-secret-value`)
stops the service at `ExecStartPre` instead of starting it with a partial `.env`.
The missing variable **names** are printed to the journal (values never are):
```bash
sudo journalctl -u open-swe.service -b | grep fetch-config
```
Required set enforced: `DASHBOARD_JWT_SECRET`, `TOKEN_ENCRYPTION_KEY`,
`GITHUB_APP_ID`, `GITHUB_APP_PRIVATE_KEY`, `GITHUB_APP_INSTALLATION_ID`,
`GITHUB_APP_CLIENT_ID`, `GITHUB_APP_CLIENT_SECRET`, the active provider keys
(`REQUIRED_PROVIDER_KEYS`, default `ANTHROPIC_API_KEY,OPENAI_API_KEY`), the
sandbox key for `SANDBOX_TYPE` (`langsmith` ⇒ `LANGSMITH_API_KEY_PROD` +
`DEFAULT_SANDBOX_SNAPSHOT_ID`), and — in **prod** — `GITHUB_WEBHOOK_SECRET`,
`SLACK_SIGNING_SECRET` (plus `LINEAR_WEBHOOK_SECRET` when Linear is wired).

View file

@ -1,11 +0,0 @@
{
"dependencies": ["."],
"graphs": {
"agent": "./aegra_entry.py:agent_graph",
"reviewer": "./aegra_entry.py:reviewer_graph",
"analyzer": "./aegra_entry.py:analyzer_graph"
},
"http": {
"app": "./aegra_entry.py:webapp_app"
}
}

View file

@ -1,19 +0,0 @@
"""Aegra entrypoint for Open SWE graphs.
Aegra loads graph files standalone via importlib.spec_from_file_location, which
gives them a synthetic module name and breaks agent/*.py's package-relative
imports (e.g. `from .dashboard.admin import ...`). Re-exporting the graphs here
through the installed ``agent`` package (absolute imports) restores correct
``__package__`` resolution, so the relative imports inside the agent modules work.
"""
from agent.analyzer import traced_analyzer as analyzer_graph
from agent.reviewer import traced_reviewer_agent as reviewer_graph
from agent.server import traced_agent as agent_graph
# Open SWE's FastAPI webapp (GitHub/Slack webhooks + dashboard API), mounted by
# Aegra via the "http" key in aegra.json. Absolute import for the same reason as
# the graphs above (relative imports break under Aegra's standalone file loader).
from agent.webapp import app as webapp_app
__all__ = ["agent_graph", "reviewer_graph", "analyzer_graph", "webapp_app"]

View file

@ -1,363 +0,0 @@
#!/usr/bin/env bash
# fetch-config.sh — AWS-sourced boot hook that materializes the app's .env.
#
# The stock `langgraph dev` runtime + the Open SWE app read a plain `.env` from
# the app working directory (python-dotenv). On the AWS lift-and-shift we do NOT
# commit a .env; instead every non-sensitive value lives in SSM Parameter Store
# (`/open-swe-<env>/*`) and every secret lives in AWS Secrets Manager
# (`open-swe-<env>/*`). This hook is run by systemd BEFORE the service starts; it
# pulls both sources via the EC2 instance role (no static keys), assembles a
# single .env on a tmpfs, and writes it owned by the unprivileged service user
# `chmod 600` (T5 SC-01: the privileged pre-hook materializes the secret; the app
# itself then runs as that NON-root service user, not root).
#
# It is intentionally FAIL-FAST: if any required secret/param is missing or empty
# it prints the offending variable NAMES (never values) and exits 1, so the
# service never starts with a partial .env.
#
# ---------------------------------------------------------------------------
# Naming contract (source of truth: T9 env/secret/config inventory)
# SSM /open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# Secrets open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# i.e. the last path segment IS the literal environment-variable name. This is a
# deliberate (documented) deviation from the handbook's kebab-case value-name
# example (`my-stack/slack-signing`): a .env materializer needs a lossless,
# unambiguous round-trip from store key -> env var, and the env var name is the
# only key that guarantees that. The `open-swe-<env>` stack prefix still follows
# kebab-case per naming-conventions.md.
# ---------------------------------------------------------------------------
#
# Wiring into systemd (AWS EC2 variant):
# The unit runs as the unprivileged service user (User=openswe). ONLY the
# ExecStartPre pre-hook runs as root (the `+` prefix) so it can pull from AWS,
# write the tmpfs .env, and chown it to the service user. The app (ExecStart)
# and the seeder (ExecStartPost) then run as openswe and read the openswe-owned
# 0600 .env — the agent never runs as root (T5 SC-01). Pass the env as the
# positional arg (T5 BOOT-01):
#
# [Service]
# User=openswe
# Group=openswe
# Environment=ENV_DIR=/run/open-swe SERVICE_USER=openswe
# # ExecStartPre runs as root (+) so it can chown the .env to the service user.
# ExecStartPre=+/opt/open-swe/deploy/seahaven/fetch-config.sh prod
# ExecStart=/opt/open-swe/.venv/bin/langgraph dev --host 127.0.0.1 --port 2024 \
# --no-browser --no-reload
# ExecStartPost=/opt/open-swe/deploy/seahaven/seed_store.sh prod
#
# tmpfs: /run is already a tmpfs on systemd hosts, so ENV_DIR=/run/open-swe is
# tmpfs-backed by default (the .env never touches disk). Set RUN_DEDICATED_TMPFS=1
# to mount a private tmpfs at ENV_DIR instead. The app's CWD `.env` is a symlink
# into ENV_DIR (created idempotently below), so python-dotenv finds it unchanged.
#
# Idempotent, re-runnable on every (re)start. No secret is ever echoed.
set -euo pipefail
umask 077
# --- Inputs ------------------------------------------------------------------
ENV="${1:-${OPENSWE_ENV:-}}"
case "$ENV" in
dev | prod) ;;
*)
echo "fetch-config: ENV must be 'dev' or 'prod' (got '${ENV:-<empty>}')" >&2
echo "usage: fetch-config.sh <dev|prod> (or set OPENSWE_ENV)" >&2
exit 2
;;
esac
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
SSM_PREFIX="/open-swe-${ENV}/"
SECRET_PREFIX="open-swe-${ENV}/"
ENV_DIR="${ENV_DIR:-/run/open-swe}" # tmpfs-backed (/run) by default
ENV_FILE="${ENV_DIR}/.env"
APP_DIR="${APP_DIR:-/opt/open-swe}" # where the app + its CWD .env live
APP_ENV_LINK="${APP_DIR}/.env" # symlink -> ENV_FILE
# The unprivileged service user that runs the app and OWNS the .env (T5 SC-01).
# fetch-config runs as root (ExecStartPre=+) only to chown the secret to it.
SERVICE_USER="${SERVICE_USER:-openswe}"
SERVICE_GROUP="${SERVICE_GROUP:-${SERVICE_USER}}"
# Sea Haven owner guard inputs (applied after the store is read, below). The owner
# normally comes from SSM /open-swe-<env>/DEFAULT_REPO_OWNER; OPENSWE_REPO_OWNER is an
# explicit operator override that wins over the store. FORBIDDEN = the upstream org
# the fork must never target; SAFE = the fallback when the resolved owner is blank or
# forbidden. (Upper/lower + whitespace are normalized before the guard check.)
OPENSWE_REPO_OWNER="${OPENSWE_REPO_OWNER:-}"
FORBIDDEN_REPO_OWNER="langchain-ai"
# Fallback org when the resolved owner is blank/forbidden — PER-ENV (mirrors the
# iacManagedSsm owner) so a dev box can NEVER fall back into the real Sea Haven org;
# it stays isolated in its own dev org. Defends the blank/upstream cases in-env.
case "$ENV" in
dev) SAFE_REPO_OWNER="seahaven-open-swe-dev" ;;
*) SAFE_REPO_OWNER="Sea-Haven-Industries" ;;
esac
for bin in aws jq; do
command -v "$bin" >/dev/null 2>&1 || { echo "fetch-config: '$bin' not found on PATH" >&2; exit 3; }
done
log() { echo "fetch-config[$ENV]: $*"; } # NAMES/counts only — never values
b64d() { base64 --decode; } # GNU coreutils on the EC2 host
# Accept a store key into VARS iff it is a valid env-var identifier and not a
# duplicate. Rejects non-identifier names (T5 SH-INJ-002 / set -e DoS hardening)
# and flat-namespace collisions (T5 SSM-05). $3 = source label for logs.
accept_var() {
local key="$1" value="$2" src="$3"
if ! [[ "$key" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
log "WARNING: skipping ${src} key with non-identifier name (rejected)"
return 0
fi
if [ -n "${VARS[$key]+set}" ]; then
echo "fetch-config[$ENV]: FAIL-FAST — duplicate key '${key}' from ${src} (flat-namespace collision)" >&2
exit 1
fi
VARS["$key"]="$value"
}
# --- tmpfs ------------------------------------------------------------------
mkdir -p "$ENV_DIR"
# Owned by the service user so the unprivileged app can traverse it (T5 SC-01).
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_DIR" 2>/dev/null || true
chmod 700 "$ENV_DIR"
if [ "${RUN_DEDICATED_TMPFS:-0}" = "1" ] && ! mountpoint -q "$ENV_DIR"; then
mount -t tmpfs -o nosuid,nodev,noexec,mode=0700,size=4m tmpfs "$ENV_DIR"
log "mounted dedicated tmpfs at $ENV_DIR"
fi
# --- Collect values into an associative array --------------------------------
declare -A VARS=()
# 1) SSM Parameter Store (non-sensitive config). NOT --recursive: the contract is
# a FLAT namespace /open-swe-<env>/<VAR>, so a non-recursive list returns exactly
# those keys and cannot collapse two nested paths onto one name (T5 SSM-05). aws
# CLI v2 auto-paginates NextToken.
log "reading SSM params under ${SSM_PREFIX} ..."
ssm_json="$(
aws ssm get-parameters-by-path \
--path "$SSM_PREFIX" \
--with-decryption \
--region "$REGION" \
--no-cli-pager \
--output json
)"
# Records are base64-encoded (name<TAB>value) so values with spaces/newlines/tabs
# survive the line-based read intact.
ssm_count=0
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}" # strip /open-swe-<env>/ prefix
[ -n "$key" ] || continue
accept_var "$key" "$value" "SSM"
ssm_count=$((ssm_count + 1))
done < <(jq -r '.Parameters[] | (.Name|@base64) + "\t" + (.Value|@base64)' <<<"$ssm_json")
log "loaded ${ssm_count} config param(s) from SSM"
# 2) Secrets Manager (sensitive values). We request secrets by EXPLICIT id
# (`batch-get-secret-value --secret-id-list ...`) rather than a name-prefix
# `--filters` collection scan (OSWE-IAC-SECRETS-LIST-01). Two wins:
# (a) Least privilege — an explicit id list lets the instance role scope
# BatchGetSecretValue to the per-secret ARN prefix and DROP the account-wide
# `secretsmanager:ListSecrets` grant that a filtered scan unavoidably forces
# (ListSecrets has no resource-level scoping). A filtered batch call also only
# authorizes against `*`; an id-list call authorizes per-secret ARN.
# (b) Deterministic set — the value set is the fixed SECRET_VARS shells created by
# the CDK ConfigStore, so we no longer depend on a list scan returning every
# page. Value-less shells and ids with no current value come back in the
# response `.Errors[]` (ResourceNotFound), never in `.SecretValues[]`, so a
# genuinely-missing REQUIRED secret is still caught by the FAIL-FAST check
# below — an absent optional secret is simply skipped.
# `--secret-id-list` is capped at 20 ids per call, so we chunk it. `--no-cli-pager`
# disables only the OUTPUT pager. Each record is base64(name<TAB>value).
#
# SOURCE OF TRUTH for this list: infra/lib/constructs/config-store.ts `SECRET_VARS`.
# Keep the two in lockstep — a new secret shell created there must be added here or it
# will never be fetched into the .env.
SECRET_VARS=(
ANTHROPIC_API_KEY CORRIDOR_API_TOKEN CORRIDOR_MCP_TOKEN CORRIDOR_TOKEN
DASHBOARD_JWT_SECRET DAYTONA_API_KEY EXA_API_KEY FIREWORKS_API_KEY
GITHUB_APP_CLIENT_SECRET GITHUB_APP_PRIVATE_KEY GITHUB_PAT GITHUB_WEBHOOK_SECRET
JUDGE_ANTHROPIC_API_KEY LANGSMITH_API_KEY
LANGSMITH_API_KEY_PROD LANGCHAIN_API_KEY LINEAR_API_KEY LINEAR_WEBHOOK_SECRET
RUNLOOP_API_KEY SLACK_BOT_TOKEN SLACK_CLIENT_SECRET
SLACK_SIGNING_SECRET TOKEN_ENCRYPTION_KEY USER_ID_API_KEY_MAP X_SERVICE_AUTH_JWT_SECRET
)
batch_get_secrets_tsv() {
local -a ids=()
local v
for v in "${SECRET_VARS[@]}"; do ids+=("${SECRET_PREFIX}${v}"); done
local i page
local -a chunk
for ((i = 0; i < ${#ids[@]}; i += 20)); do
chunk=("${ids[@]:i:20}")
# Capture the response into a variable FIRST so a non-zero `aws` exit (throttle,
# AccessDenied, KMS DecryptionFailure) aborts under set -e instead of being
# silently swallowed — then we'd FAIL-FAST below as "missing secret" with a wrong
# root cause. (Value-less / absent shells come back in .Errors[], not .SecretValues[].)
page="$(aws secretsmanager batch-get-secret-value \
--secret-id-list "${chunk[@]}" \
--region "$REGION" --no-cli-pager --output json)"
printf '%s' "$page" \
| jq -r '.SecretValues[] | select(.SecretString != null) | (.Name|@base64) + "\t" + (.SecretString|@base64)'
done
}
log "reading secrets under ${SECRET_PREFIX} ..."
secret_count=0
# Capture into a variable (NOT `done < <(...)` process substitution) so a non-zero
# exit from batch_get_secrets_tsv propagates under set -e — process substitution hides
# the producer's exit status from the parent shell, which would let a failed AWS call
# fall through to a misleading "missing required var" FAIL-FAST. Mirrors the SSM read.
secrets_tsv="$(batch_get_secrets_tsv)"
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
case "$name" in
"${SECRET_PREFIX}"*) ;; # defensive: exact-prefix only
*) continue ;;
esac
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}"
[ -n "$key" ] || continue
accept_var "$key" "$value" "Secrets"
secret_count=$((secret_count + 1))
done <<<"$secrets_tsv"
log "loaded ${secret_count} secret(s) from Secrets Manager"
# --- Sea Haven DEFAULT_REPO_OWNER guard --------------------------------------
# OSWE-OWNER-04 (revised for multi-org): HONOR the configured owner — the
# OPENSWE_REPO_OWNER env override if set, else the store value — so per-env orgs
# work (dev = seahaven-open-swe-dev, prod = Sea-Haven-Industries). But GUARD the two
# values that must NEVER reach the agent: blank, and the upstream 'langchain-ai' org
# (the fork's origin). Either falls back to the Sea Haven org so a stale/blank/mis-set
# value can never point the agent upstream. Comparison is case- and whitespace-
# insensitive. The POSITIVE org allowlist is enforced by the app (ALLOWED_GITHUB_ORGS).
resolved_owner="${OPENSWE_REPO_OWNER:-${VARS[DEFAULT_REPO_OWNER]:-}}"
# Normalize for the guard CHECK ONLY (the original value is what gets stored when
# allowed): lowercase, strip whitespace, take the FIRST path segment so a value like
# 'langchain-ai/open-swe' still trips the guard, and drop dots (GitHub owners contain
# none) so 'langchain-ai.' can't slip past. Homoglyph/unicode variants are out of scope
# here — the owner comes from admin-written SSM/IaC, not attacker-controlled input.
norm_owner="$(printf '%s' "$resolved_owner" | tr '[:upper:]' '[:lower:]' | tr -d '[:space:]')"
norm_owner="${norm_owner%%/*}"
norm_owner="${norm_owner//./}"
case "$norm_owner" in
"" | "$FORBIDDEN_REPO_OWNER")
log "WARNING: DEFAULT_REPO_OWNER ('${resolved_owner:-<blank>}') is blank or the upstream org -> forcing '${SAFE_REPO_OWNER}'"
resolved_owner="$SAFE_REPO_OWNER"
;;
esac
VARS[DEFAULT_REPO_OWNER]="$resolved_owner"
# --- FAIL-FAST: required vars -------------------------------------------------
# Hard-required regardless of mode:
required=(
DASHBOARD_JWT_SECRET # RuntimeError on startup if missing (oauth.py)
TOKEN_ENCRYPTION_KEY # Fernet key(s); decrypts per-user GitHub tokens
)
# NOTE: the GitHub App is NOT created/duplicated for dev — only prod owns the
# (single, shared) GitHub App + Slack app. So the GitHub App quintet + Slack +
# webhook-signing secrets are required for PROD only (see the prod block below).
# Dev boots without them: it has no GitHub-App/Slack/webhook integration — it is a
# deployment-validation env (boot/health/boundary), not a live-triggered agent.
# Active model-provider key(s): model selection is store-driven (team_settings),
# so fetch-config cannot infer it from .env. Bedrock (Claude builder + reviewer)
# authenticates via the host IAM role — no API key; Fireworks (fallback / subagents /
# any non-Claude model) needs its key. Override with a comma list if the active models change.
IFS=',' read -r -a provider_keys <<<"${REQUIRED_PROVIDER_KEYS:-FIREWORKS_API_KEY}"
for k in "${provider_keys[@]}"; do
k="${k//[[:space:]]/}"
[ -n "$k" ] && required+=("$k")
done
# Sandbox provider key(s) — depends on SANDBOX_TYPE (default langsmith).
sandbox_type="${VARS[SANDBOX_TYPE]:-langsmith}"
case "$sandbox_type" in
langsmith) required+=(LANGSMITH_API_KEY_PROD DEFAULT_SANDBOX_SNAPSHOT_ID) ;;
daytona) required+=(DAYTONA_API_KEY) ;;
runloop) required+=(RUNLOOP_API_KEY) ;;
modal | local) ;; # no key required
*) log "WARNING: unknown SANDBOX_TYPE='${sandbox_type}' — not enforcing a sandbox key" ;;
esac
# Prod-only: the GitHub App (installation-token minting + dashboard OAuth) and the
# webhook-signing secrets. Dev has no GitHub/Slack app, so none of these are
# required there; prod owns the single shared app and must have all of them.
if [ "$ENV" = "prod" ]; then
required+=(
GITHUB_APP_ID # GitHub App trio (installation-token minting) ...
GITHUB_APP_PRIVATE_KEY # ... multiline PEM ...
GITHUB_APP_INSTALLATION_ID # ... used by utils/github_app.py
GITHUB_APP_CLIENT_ID # dashboard OAuth login
GITHUB_APP_CLIENT_SECRET # dashboard OAuth login
GITHUB_WEBHOOK_SECRET # webhook signature verification
SLACK_SIGNING_SECRET # Slack webhook signature verification
)
if [ -n "${VARS[LINEAR_API_KEY]:-}" ] && [ "${OPENSWE_REQUIRE_LINEAR:-1}" = "1" ]; then
required+=(LINEAR_WEBHOOK_SECRET)
fi
fi
missing=()
for k in "${required[@]}"; do
[ -n "${VARS[$k]:-}" ] || missing+=("$k")
done
# de-dup the names for a clean report
if [ "${#missing[@]}" -gt 0 ]; then
mapfile -t missing < <(printf '%s\n' "${missing[@]}" | sort -u)
echo "fetch-config[$ENV]: FAIL-FAST — ${#missing[@]} required var(s) missing/empty:" >&2
printf ' - %s\n' "${missing[@]}" >&2
echo "fetch-config[$ENV]: refusing to write a partial .env; service will not start." >&2
exit 1
fi
# --- Write the .env atomically (root-only on tmpfs) --------------------------
# python-dotenv reads double-quoted values (incl. multiline PEMs). Its decoder
# unescapes ONLY backslash and double-quote (\\ -> \, \" -> "); it does NOT honor
# \$ or \` escapes, so escaping those would leave a spurious backslash. Escape
# exactly backslash then double-quote — real newlines stay literal (multiline OK).
# (Caveat: python-dotenv interpolates a literal `${VAR}` substring; the secret
# domain here — base64/hex/PEM keys — never contains one, so no extra guard.)
emit_var() {
local name="$1" value="$2" esc
esc="${value//\\/\\\\}"
esc="${esc//\"/\\\"}"
printf '%s="%s"\n' "$name" "$esc"
}
tmp="$(mktemp "${ENV_DIR}/.env.XXXXXX")"
chmod 600 "$tmp"
{
printf '# Generated by fetch-config.sh for env=%s at %s — DO NOT EDIT.\n' \
"$ENV" "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
printf '# Source: SSM /open-swe-%s/* + Secrets Manager open-swe-%s/*\n\n' "$ENV" "$ENV"
for k in $(printf '%s\n' "${!VARS[@]}" | sort); do
emit_var "$k" "${VARS[$k]}"
done
} >"$tmp"
mv -f "$tmp" "$ENV_FILE"
# Owned by the unprivileged service user (T5 SC-01) so the app reads it without
# running as root. fetch-config itself runs as root (ExecStartPre=+) to chown.
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_FILE"
chmod 600 "$ENV_FILE"
# Point the app's CWD .env at the tmpfs file (idempotent).
if [ "$APP_ENV_LINK" != "$ENV_FILE" ]; then
if [ -L "$APP_ENV_LINK" ] || [ ! -e "$APP_ENV_LINK" ]; then
ln -sfn "$ENV_FILE" "$APP_ENV_LINK"
elif [ "$(readlink -f "$APP_ENV_LINK" 2>/dev/null || true)" != "$(readlink -f "$ENV_FILE")" ]; then
log "WARNING: ${APP_ENV_LINK} exists and is not a symlink to ${ENV_FILE} — leaving it untouched"
fi
fi
total=$((ssm_count + secret_count))
log "wrote ${ENV_FILE} (${total} vars, sandbox=${sandbox_type}) — ${SERVICE_USER}:${SERVICE_GROUP} 0600"

View file

@ -1,36 +0,0 @@
# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy.
# RETIRED on-prem VM variant — kept for on-prem-contrast reference only. The LIVE
# AWS nginx site is the AMI template deploy/ami/templates/open-swe.nginx.conf
# (rendered from an @@SERVER_NAME@@ token at first boot). See DEPLOYMENT.md.
#
# nginx is the security boundary: ONLY /dashboard/api/* reaches the backend;
# the unauthenticated LangGraph agent API (/threads,/runs,/assistants,/store) is NOT proxied.
server {
listen 80 default_server;
listen [::]:80 default_server;
server_name openswe.seahaven.com;
root /var/www/openswe;
index _shell.html;
# ALB health check
location = /healthz { default_type text/plain; return 200 "ok\n"; }
# Dashboard API + OAuth callback -> backend webapp on :2024 (the ONLY proxied path)
location /dashboard/api/ {
proxy_pass http://127.0.0.1:2024;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 300s;
}
# Static assets + SPA shell fallback (client-side routing)
location / {
try_files $uri $uri/ /_shell.html;
}
}

View file

@ -1,152 +0,0 @@
#!/usr/bin/env bash
# put-config.sh — out-of-band populator for the open-swe config store.
#
# T11 (CDK) creates the RESOURCE SHELLS:
# - 28 Secrets Manager secrets open-swe-<env>/<VAR> (value-LESS shells)
# - the IaC-managed SSM params /open-swe-<env>/<VAR> (real values, owned in CDK)
# This script sets the values that CANNOT live in IaC — every secret value, plus
# the out-of-band SSM params (operationally-variable / env-specific-unknown). Run
# it AFTER `cdk deploy open-swe-<env>` and BEFORE the EC2/T12 box first boots, so
# fetch-config.sh finds every REQUIRED var populated and never writes a partial .env.
#
# The secret list below mirrors config-store.ts SECRET_VARS (28 names — the
# inventory's "29" double-counted JUDGE_ANTHROPIC_BASE_URL, which is config
# (SSM), not a secret). Keep the two lists in lockstep.
#
# SAFETY:
# - NO real secret values live in this file — every secret is a <FILL> placeholder.
# Replace <FILL...> inline at run time, pipe from a vault, or export the
# matching OPENSWE_PUT_<VAR> env var; NEVER commit real values.
# - It does NOT touch the IaC-managed SSM params (SANDBOX_TYPE, DEFAULT_REPO_OWNER,
# ALLOWED_GITHUB_ORGS, DEFAULT_REPO_NAME, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID)
# — CDK owns those; setting them here would cause drift.
# - Secrets go to Secrets Manager; the box's instance role grants read on
# open-swe-<env>/* and /open-swe-<env>/* (no kms:Decrypt — AWS-managed keys).
#
# Idempotent: put-secret-value adds a new AWSCURRENT version; put-parameter
# --overwrite updates in place.
set -euo pipefail
ENV="${1:-}"
case "$ENV" in
dev | prod) ;;
*)
echo "usage: put-config.sh <dev|prod>" >&2
exit 2
;;
esac
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
SECRET_PREFIX="open-swe-${ENV}/"
SSM_PREFIX="/open-swe-${ENV}/"
command -v aws >/dev/null 2>&1 || { echo "put-config: 'aws' not found on PATH" >&2; exit 3; }
# --- helpers -----------------------------------------------------------------
# put_secret VAR : set the value of an EXISTING secret shell open-swe-<env>/VAR.
# Resolution order for the value: $OPENSWE_PUT_<VAR> env var, else the literal
# <FILL> placeholder (which aborts so an unset secret is never silently shipped).
put_secret() {
local var="$1"
local override_name="OPENSWE_PUT_${var}"
local value="${!override_name:-<FILL>}"
if [ "$value" = "<FILL>" ]; then
echo "put-config[$ENV]: SKIP secret ${var} (no value; set ${override_name} or edit inline)" >&2
return 0
fi
aws secretsmanager put-secret-value \
--secret-id "${SECRET_PREFIX}${var}" \
--secret-string "$value" \
--region "$REGION" \
--no-cli-pager >/dev/null
echo "put-config[$ENV]: set secret ${var}"
}
# put_param VAR [TYPE] : create/update an out-of-band SSM param /open-swe-<env>/VAR.
# TYPE defaults to String; pass SecureString for anything sensitive-but-not-a-secret.
put_param() {
local var="$1" type="${2:-String}"
local override_name="OPENSWE_PUT_${var}"
local value="${!override_name:-<FILL>}"
if [ "$value" = "<FILL>" ]; then
echo "put-config[$ENV]: SKIP param ${var} (no value; set ${override_name} or edit inline)" >&2
return 0
fi
aws ssm put-parameter \
--name "${SSM_PREFIX}${var}" \
--value "$value" \
--type "$type" \
--overwrite \
--region "$REGION" \
--no-cli-pager >/dev/null
echo "put-config[$ENV]: set param ${var} (${type})"
}
# --- 1) Secrets (open-swe-<env>/<VAR>) — 28 shells from config-store.ts -------
# REQUIRED at boot (fetch-config fail-fast): DASHBOARD_JWT_SECRET,
# TOKEN_ENCRYPTION_KEY, GITHUB_APP_PRIVATE_KEY/CLIENT_SECRET, the active provider
# key(s) (ANTHROPIC_API_KEY + OPENAI_API_KEY by default), LANGSMITH_API_KEY_PROD,
# and (prod) GITHUB_WEBHOOK_SECRET + SLACK_SIGNING_SECRET.
put_secret ANTHROPIC_API_KEY # optional — eval judge only (JUDGE_ANTHROPIC_API_KEY fallback); Bedrock builder/reviewer use the host IAM role
put_secret DASHBOARD_JWT_SECRET # REQUIRED — dashboard session JWT signing
put_secret TOKEN_ENCRYPTION_KEY # REQUIRED — Fernet key(s) for GH-token crypto
put_secret GITHUB_APP_PRIVATE_KEY # REQUIRED — GitHub App PEM (multiline; quote it)
put_secret GITHUB_APP_CLIENT_SECRET # REQUIRED — dashboard OAuth login
put_secret LANGSMITH_API_KEY_PROD # REQUIRED (langsmith sandbox) — prod key
put_secret GITHUB_WEBHOOK_SECRET # prod-REQUIRED — GitHub webhook signature
put_secret SLACK_SIGNING_SECRET # prod-REQUIRED — Slack webhook signature
put_secret LINEAR_WEBHOOK_SECRET # required only when Linear is wired
# Optional / conditional secrets — set the ones this deployment actually uses.
put_secret CORRIDOR_API_TOKEN # optional — Corridor MCP
put_secret CORRIDOR_MCP_TOKEN # optional — Corridor MCP (alt name)
put_secret CORRIDOR_TOKEN # optional — Corridor MCP (alt name)
put_secret DAYTONA_API_KEY # only if SANDBOX_TYPE=daytona
put_secret EXA_API_KEY # optional — Exa web search
put_secret FIREWORKS_API_KEY # active non-Claude key (fallback/subagents); Bedrock uses the host IAM role
put_secret GITHUB_PAT # optional — PAT fallback
put_secret JUDGE_ANTHROPIC_API_KEY # optional — eval judge (falls back to ANTHROPIC)
put_secret LANGSMITH_API_KEY # optional — LangSmith (dev)
put_secret LANGCHAIN_API_KEY # optional — LangSmith alt name
put_secret LINEAR_API_KEY # optional — Linear API
put_secret RUNLOOP_API_KEY # only if SANDBOX_TYPE=runloop
put_secret SLACK_BOT_TOKEN # optional — Slack bot token
put_secret SLACK_CLIENT_SECRET # optional — Slack OAuth
put_secret USER_ID_API_KEY_MAP # optional — JSON map user id -> API key (per-user auth)
put_secret X_SERVICE_AUTH_JWT_SECRET # optional — service-auth JWT
# --- 2) Out-of-band SSM params (/open-swe-<env>/<VAR>) ------------------------
# These are NOT created by CDK (operationally-variable / env-specific-unknown).
# put_param creates them on first run.
put_param DEFAULT_SANDBOX_SNAPSHOT_ID # REQUIRED (langsmith) — changes every rebuild
put_param GITHUB_APP_ID # REQUIRED — GitHub App numeric id
put_param GITHUB_APP_INSTALLATION_ID # REQUIRED — GitHub App installation id
put_param GITHUB_APP_CLIENT_ID # REQUIRED — dashboard OAuth client id
put_param GITHUB_OAUTH_PROVIDER_ID # GitHub OAuth provider id
put_param LANGSMITH_TENANT_ID_PROD # LangSmith prod tenant id
put_param LANGSMITH_URL_PROD # LangSmith prod URL
put_param LANGSMITH_ENDPOINT # LangSmith API endpoint
put_param LANGSMITH_ENDPOINT_PROD # LangSmith prod API endpoint
put_param LANGSMITH_HOST_API_URL # LangSmith host API URL
put_param LANGGRAPH_URL # LangGraph server URL
put_param LANGGRAPH_URL_PROD # LangGraph server URL (prod)
put_param LANGCHAIN_REVISION_ID # LangChain revision id
put_param SLACK_CLIENT_ID # Slack OAuth client id
put_param SLACK_TEAM_ID # Slack workspace/team id
put_param SLACK_BOT_USER_ID # Slack bot user id
put_param SLACK_BOT_USERNAME # Slack bot username
put_param SLACK_REPO_OWNER # Slack default repo owner
put_param SLACK_REPO_NAME # Slack default repo name
put_param CONFIGURED_ADMINS # dashboard admin GitHub logins (comma list)
put_param OBSERVABILITY_AUTHORIZED_EMAILS # observability allowlist (comma list)
put_param PUBLIC_REPO_ORG_GATE # public-repo trigger gate (org name; empty=off)
put_param ALLOWED_GITHUB_REPOS # extra repo allowlist (comma list)
put_param LLM_FALLBACK_MODEL_ID # optional — model fallback id
put_param DATADOG_MCP_TOOLSETS # optional — Datadog MCP toolsets
put_param NOTION_MCP_CLIENT_NAME # optional — Notion MCP client name
put_param API_STANDARDS_SKILL_HANDLE # optional — API standards skill handle
put_param REPO_SNAPSHOT_BASE_IMAGE # optional — repo snapshot base image
put_param REPO_SNAPSHOT_BUILD_TIMEOUT_SECONDS # optional — snapshot build timeout
put_param REPO_SNAPSHOT_STALE_BUILD_SECONDS # optional — snapshot stale threshold
echo "put-config[$ENV]: done. Verify with fetch-config.sh ${ENV} before first boot."

View file

@ -1,184 +0,0 @@
#!/usr/bin/env bash
# Seed the LangGraph store after a (re)start.
#
# The stock `langgraph dev` server uses an IN-MEMORY store, so anything written
# to it (team model settings, user mappings) is lost on every restart. This
# script idempotently re-PUTs that state and is wired as a systemd
# ExecStartPost on the open-swe.service unit so it runs after each start.
# IT MUST RE-RUN ON EVERY RESTART — the in-memory store starts empty each boot.
#
# Replace this with Postgres-backed durability (Aegra / `langgraph up`) to make
# the store survive restarts and drop this script.
#
# AWS-env-aware: pass the env as $1 (dev|prod). Seed values (default repo, model
# ids, user mappings) come from the fetch-config-materialized .env.
#
# SECURITY (T5 /sh-security-review):
# - SH-INJ-001: this script NEVER `source`s the .env. python-dotenv and bash
# have incompatible escaping, and a config value like `$(cmd)` would execute
# when sourced. We extract the few single-line seed keys with a non-eval
# reader (read_env) instead.
# - SH-INJ-003: the store PUT bodies are built with `jq --arg`, so values are
# always JSON-encoded (no string interpolation into a JSON heredoc).
# - SH-INJ-004: BASE is pinned to loopback — never derived from store/SSM
# config (LANGGRAPH_URL) — so a tampered value can't redirect the PUTs.
# - SC-02: not sourcing the .env means secrets are never exported into this
# script's (or curl's) environment.
#
# Seed values read from the materialized .env (set in SSM /open-swe-<env>/*):
# DEFAULT_REPO_OWNER / DEFAULT_REPO_NAME -> team_settings default_repo
# LLM_MODEL_ID -> default builder model (fallback)
# SEED_AGENT_MODEL / SEED_AGENT_EFFORT -> builder model + effort (optional)
# SEED_REVIEWER_MODEL / SEED_REVIEWER_EFFORT -> reviewer model + effort (optional)
# SEED_USER_MAPPINGS -> "login:email,login:email" (optional)
# CONFIGURED_ADMINS -> "login,email" fallback for the mapping
# Legacy OPENSWE_* process-env overrides are still honored (highest precedence).
set -euo pipefail
ENV="${1:-${OPENSWE_ENV:-}}"
case "$ENV" in
dev | prod | "") ;; # empty allowed: pure-env / on-prem backward-compat mode
*)
echo "seed_store: ENV must be 'dev' or 'prod' (got '$ENV')" >&2
exit 2
;;
esac
ENV_DIR="${ENV_DIR:-/run/open-swe}"
ENV_FILE="${ENV_FILE:-${ENV_DIR}/.env}"
# read_env KEY -> prints the value of a SINGLE-LINE `KEY="..."` entry from the
# materialized .env WITHOUT shell evaluation (SH-INJ-001 fix). Seed keys are
# simple single-line values; multiline secrets (e.g. the PEM) are never read
# here. Returns empty if the key is absent/unreadable.
read_env() {
local key="$1" line
[ -r "$ENV_FILE" ] || return 0
line="$(grep -m1 -- "^${key}=" "$ENV_FILE" 2>/dev/null || true)"
[ -n "$line" ] || return 0
line="${line#*=}"
# strip one layer of surrounding double quotes (python-dotenv double-quoted form)
if [ "${line#\"}" != "$line" ]; then line="${line%\"}"; line="${line#\"}"; fi
# reverse python-dotenv double-quote escaping (only \" and \\ are escaped)
line="${line//\\\"/\"}"; line="${line//\\\\/\\}"
printf '%s' "$line"
}
# Process-env override (legacy/on-prem) -> .env value -> default.
pick() { # pick DEFAULT OVERRIDE_VALUE FILE_KEY...
local def="$1" override="$2"; shift 2
if [ -n "$override" ]; then printf '%s' "$override"; return; fi
local k v
for k in "$@"; do v="$(read_env "$k")"; [ -n "$v" ] && { printf '%s' "$v"; return; }; done
printf '%s' "$def"
}
# BASE is loopback-pinned (SH-INJ-004): this on-box seeder only talks to the
# local server; OPENSWE_PORT may override the port but never the host.
BASE="http://127.0.0.1:${OPENSWE_PORT:-2024}"
AGENT_MODEL="$(pick 'bedrock_converse:us.anthropic.claude-opus-4-8' "${OPENSWE_AGENT_MODEL:-}" SEED_AGENT_MODEL LLM_MODEL_ID)"
AGENT_EFFORT="$(pick 'high' "${OPENSWE_AGENT_EFFORT:-}" SEED_AGENT_EFFORT)"
REVIEWER_MODEL="$(pick 'bedrock_converse:us.anthropic.claude-opus-4-8' "${OPENSWE_REVIEWER_MODEL:-}" SEED_REVIEWER_MODEL)"
REVIEWER_EFFORT="$(pick 'high' "${OPENSWE_REVIEWER_EFFORT:-}" SEED_REVIEWER_EFFORT)"
# default_repo = owner/name from AWS config (DEFAULT_REPO_OWNER is hard-pinned
# away from upstream by fetch-config.sh).
REPO_OWNER="$(pick '' '' DEFAULT_REPO_OWNER)"
REPO_NAME="$(pick '' '' DEFAULT_REPO_NAME)"
if [ -n "${OPENSWE_DEFAULT_REPO:-}" ]; then
DEFAULT_REPO="$OPENSWE_DEFAULT_REPO"
elif [ -n "$REPO_OWNER" ] && [ -n "$REPO_NAME" ]; then
DEFAULT_REPO="${REPO_OWNER}/${REPO_NAME}"
else
echo "seed_store: set OPENSWE_DEFAULT_REPO=owner/repo (or DEFAULT_REPO_OWNER + DEFAULT_REPO_NAME in .env)" >&2
exit 1
fi
# user_mappings: explicit SEED_USER_MAPPINGS ("login:email,..."), then legacy
# OPENSWE_OWNER_LOGIN/EMAIL, then parse CONFIGURED_ADMINS ("login,email").
SEED_MAP="$(pick '' "${SEED_USER_MAPPINGS:-}" SEED_USER_MAPPINGS)"
ADMINS="$(pick '' "${CONFIGURED_ADMINS:-}" CONFIGURED_ADMINS)"
declare -a MAPPINGS=()
if [ -n "$SEED_MAP" ]; then
IFS=',' read -r -a _pairs <<<"$SEED_MAP"
for p in "${_pairs[@]}"; do
p="${p//[[:space:]]/}"
[ -n "$p" ] && MAPPINGS+=("$p")
done
elif [ -n "${OPENSWE_OWNER_LOGIN:-}" ] && [ -n "${OPENSWE_OWNER_EMAIL:-}" ]; then
MAPPINGS+=("${OPENSWE_OWNER_LOGIN}:${OPENSWE_OWNER_EMAIL}")
elif [ -n "$ADMINS" ]; then
_login="" _email=""
IFS=',' read -r -a _toks <<<"$ADMINS"
for t in "${_toks[@]}"; do
t="${t//[[:space:]]/}"
[ -z "$t" ] && continue
case "$t" in
*@*) [ -z "$_email" ] && _email="$t" ;;
*) [ -z "$_login" ] && _login="$t" ;;
esac
done
[ -n "$_login" ] && [ -n "$_email" ] && MAPPINGS+=("${_login}:${_email}")
fi
# No user mapping is NON-FATAL (OSWE-SEED-03 precedent: never fail the unit into a
# restart loop over a seeding gap — same as the server-not-ready path below). The
# server itself is healthy; an unseeded user_mappings table only means the @openswe
# trigger won't resolve a commenter, which a deployment-validation env (e.g. dev)
# does not need. team_settings is still seeded. Set SEED_USER_MAPPINGS (or
# CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL) to seed the mapping when wanted.
if [ "${#MAPPINGS[@]}" -eq 0 ]; then
echo "seed_store: no user mapping resolved — skipping user_mappings seed (set SEED_USER_MAPPINGS or OPENSWE_OWNER_LOGIN/EMAIL to enable)" >&2
fi
NOW="$(date -u +%Y-%m-%dT%H:%M:%S+00:00)"
# Wait for the server to accept requests (up to ~60s). Authoritative (OSWE-SEED-03):
# if it never comes up, log and exit 0 — do NOT fail the unit into a restart loop.
READY=0
for _ in $(seq 1 30); do
if [ "$(curl -s -o /dev/null -w '%{http_code}' "$BASE/ok" || true)" = "200" ]; then READY=1; break; fi
sleep 2
done
if [ "$READY" -ne 1 ]; then
echo "seed_store: server not ready at $BASE after ~60s; skipping seed (will reseed on next restart)" >&2
exit 0
fi
# 1) team_settings/default — JSON built with jq --arg (SH-INJ-003 fix).
team_body="$(jq -n \
--arg am "$AGENT_MODEL" --arg ae "$AGENT_EFFORT" \
--arg rm "$REVIEWER_MODEL" --arg re "$REVIEWER_EFFORT" \
--arg repo "$DEFAULT_REPO" --arg now "$NOW" \
'{namespace:["team_settings"],key:"default",value:{
review_draft_prs:false, pr_summaries:true, review_trace_links:true,
org_guidelines:null,
default_agent_model:$am, default_agent_reasoning_effort:$ae,
default_agent_subagent_model:$am, default_agent_subagent_reasoning_effort:$ae,
default_repo:$repo,
default_reviewer_model:$rm, default_reviewer_reasoning_effort:$re,
default_reviewer_subagent_model:$rm, default_reviewer_subagent_reasoning_effort:$re,
default_grouping_model:null, default_grouping_reasoning_effort:null,
default_chat_model:null, default_chat_reasoning_effort:null,
updated_at:$now}}')"
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$team_body" >/dev/null \
|| echo "seed_store: WARN team_settings PUT failed (will reseed next restart)" >&2
# 2) user_mappings/<login> — required, or the @openswe trigger ignores the commenter.
for pair in "${MAPPINGS[@]}"; do
login="${pair%%:*}"
email="${pair#*:}"
if [ -z "$login" ] || [ -z "$email" ] || [ "$login" = "$pair" ]; then
echo "seed_store: skipping malformed mapping '$pair' (want login:email)" >&2
continue
fi
map_body="$(jq -n --arg login "$login" --arg email "$email" --arg now "$NOW" \
'{namespace:["user_mappings"],key:$login,value:{
github_login:$login, work_email:$email, slack_user_id:null,
source:"slack_oauth", status:"active", created_at:$now, updated_at:$now}}')"
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$map_body" >/dev/null \
|| echo "seed_store: WARN user_mapping PUT failed for $login" >&2
done
echo "seed_store: done at $NOW (env=${ENV:-none}, repo=$DEFAULT_REPO, mappings=${#MAPPINGS[@]})"

View file

@ -1,17 +0,0 @@
[Unit]
Description=Open SWE stock LangGraph dev server (graphs + webapp, :2024)
After=network-online.target postgresql.service
Wants=network-online.target
[Service]
Type=simple
User=adam
WorkingDirectory=/home/adam/open-swe
ExecStart=/home/adam/open-swe/.venv/bin/langgraph dev --host 0.0.0.0 --port 2024 --no-browser --no-reload
ExecStartPost=/home/adam/open-swe/seed_store.sh
Restart=on-failure
RestartSec=5
TimeoutStartSec=120
[Install]
WantedBy=multi-user.target

21
infra/.gitignore vendored
View file

@ -1,21 +0,0 @@
# CDK / build output
cdk.out/
*.js
*.d.ts
*.js.map
# ...but jest.config.js is hand-authored config, not build output — keep it.
!jest.config.js
# deps
node_modules/
# env
.env
# coverage
coverage/
# NOTE: cdk.context.json IS committed on purpose (pins the AMI / lookups so
# deploys are reproducible and don't implicitly pick up a newer AMI — see
# lib/constructs/ami-cache.ts and feedback_inline_ebs_volumes).
!cdk.context.json

View file

@ -1,348 +0,0 @@
# open-swe infra (CDK TypeScript)
AWS infrastructure for the Open SWE → AWS migration. **Synth-only at this stage —
nothing here is deployed yet.** All IAM is applied only after the Phase-1 security
gate (T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review`) clears (T6).
## Layout
```
infra/
├── bin/
│ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
│ ├── aspects/
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
│ └── constructs/
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│ └── ami-cache.ts # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore
```
## Stacks
| Stack name (kebab) | Construct | Contents |
|---|---|---|
| `open-swe-iam` | `OpenSweIamStack` | Account-level shared GitHub OIDC deploy roles (singletons). |
| `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role`, config store (T11), and the EC2 box + ALB ingress (T12, `AppService`). |
| `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role`, config store, and the EC2 box + ALB ingress. |
Account `328440206208`, region `us-east-1`. Stack names are set explicitly so CDK
never defaults to PascalCase; resource names follow `open-swe-<env>-*`.
> The two env stacks (`open-swe-dev` / `open-swe-prod`) are the required pair. The
> shared OIDC deploy roles are account-wide singletons (one `RoleName` each), so
> they live in their own dedicated `open-swe-iam` stack rather than being
> duplicated across the env stacks — and that stack deploys first (see ordering).
## IAM roles defined (unapplied)
- **`githubdeploy-open-swe-infra`** — GitHub OIDC role for CDK/CFN infra deploys.
Trust scoped to `repo:Sea-Haven-Industries/open-swe` on the `main`/`dev`
branches only. Permission is the org-standard CDK pattern: `sts:AssumeRole` on
the CDK bootstrap roles (`cdk-hnb659fds-*`) — the real CFN/IAM/resource scope
lives in the bootstrap `cfn-exec-role`, not in this role.
- **`githubdeploy-open-swe-app`** — GitHub OIDC role for app deploys. Tag-scoped
`ssm:SendCommand` (instances tagged `project=open-swe` + `env in {dev,prod}`) +
read-only access to the `open-swe-<env>-assets` S3 artifact buckets.
- **`open-swe-<env>-instance-role`** — EC2 instance role, least-privilege: read
`open-swe-<env>-assets` (S3), read `/open-swe-<env>/*` (SSM), read
`open-swe-<env>/*` (Secrets Manager), put `/open-swe/<env>/*` CloudWatch Logs,
plus `AmazonSSMManagedInstanceCore` for SSM agent registration. No admin.
The GitHub OIDC provider already exists account-wide (created for seahaven-site);
it is referenced by ARN, never re-created.
## CDK bootstrap qualifiers — per-env deploy isolation (B-1 / OSWE-IAC-01)
Each env's infra deploy role may assume **only its own bootstrap qualifier's**
roles, so a dev-branch token can never assume the bootstrap roles whose admin
`cfn-exec-role` deploys prod (closing the cross-env escalation that bypassed
prod's Environment approval gate). Mapping lives in `config.ts:bootstrapQualifier`:
| Env | Qualifier | Toolkit stack | Infra role assumes |
|---|---|---|---|
| dev | `oswedev` | `CDKToolkit-oswedev` | `cdk-oswedev-*` |
| prod | `hnb659fds` (default) | `CDKToolkit` | `cdk-hnb659fds-*` |
The dev stack synthesizes with `DefaultStackSynthesizer({ qualifier: "oswedev" })`
(`bin/app.ts`); prod uses the default. Bootstrap a new env qualifier with:
```bash
npx cdk bootstrap --qualifier <qual> --toolkit-stack-name CDKToolkit-<qual> \
--cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess \
aws://328440206208/us-east-1
```
**Deploy order matters** when changing an env's qualifier: bootstrap the new
qualifier and deploy the env stack onto it **before** re-scoping that env's infra
role in `open-swe-iam` — otherwise a pipeline deploy with the re-scoped role would
fail to assume the not-yet-targeted bootstrap roles.
## Kebab-case naming Aspect
`KebabNamingAspect` (applied app-wide in `bin/app.ts`) fails synth via
`Annotations.addError` when a stack name or an explicit physical resource name
(`RoleName`, `BucketName`, …) is not kebab-case. Path-style names (Secrets
Manager `a/b`, SSM `/a/b`, log groups `/aws/.../x`) are validated per `/`-segment.
CDK logical construct ids are intentionally NOT validated (they are conventionally
PascalCase). Covered by `test/kebab-naming-aspect.test.ts`.
## Config store (Secrets Manager + SSM shells — T11)
`ConfigStore` (`lib/constructs/config-store.ts`, one per env from `OpenSweStack`)
renders the resource shells the boot hook `deploy/seahaven/fetch-config.sh` reads.
The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var
names as the last path segment — `open-swe-<env>/<VAR>` for secrets,
`/open-swe-<env>/<VAR>` (FLAT) for config — because fetch-config strips the prefix
and exports that segment verbatim.
Three buckets:
1. **Secret shells (Secrets Manager) — 27 secrets.** Created value-LESS (an L1
`CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which
CloudFormation creates as an empty secret). The real value is set **out-of-band**
(`put-config.sh`) — CDK never owns it, so a later `cdk deploy` can never clobber
it. `UpdateReplacePolicy/DeletionPolicy: Retain` so a teardown can't destroy
operator-set secret material. AWS-managed key (no CMK — matches the instance
role, which omits `kms:Decrypt`).
> The T9 header says "29 secrets" but its table enumerates **27** distinct VAR
> names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29.
> **Confirm** the 27-vs-29 count (code-only candidates not in the table:
> `USER_ID_API_KEY_MAP`, `JUDGE_ANTHROPIC_BASE_URL`).
> ⚠️ **`Retain` + fixed name orphans these shells on a failed FIRST create.**
> If the stack's initial create fails and rolls back, `Retain` keeps the shells
> instead of deleting them. The stack is then gone, but the secrets survive,
> still holding the global `open-swe-<env>/<VAR>` names — so every later create
> fails with `AlreadyExists`. A plain `delete-secret` does **not** clear it
> (the name stays reserved for the 7–30 day recovery window). This bites on a
> **teardown/rebuild, a secret logical-id change/refactor, or standing up a new
> env** — never on routine updates of an already-created stack. **Recovery —**
> before re-creating the stack, force-delete the *empty* orphans so the names
> free immediately:
> ```bash
> aws secretsmanager list-secrets --region us-east-1 \
> --filters Key=name,Values=open-swe-<env>/ \
> --query 'SecretList[].Name' --output text | tr '\t' '\n' | while read -r n; do
> aws secretsmanager delete-secret --secret-id "$n" \
> --region us-east-1 --force-delete-without-recovery
> done
> ```
> Force-delete only shells with **no value version** — a populated secret holds
> real operator material. (Hit on prod 2026-06-29; see PR #51's deploy failure.)
2. **IaC-managed SSM config — 8 params, real values owned in code:**
| Param | dev | prod |
|---|---|---|
| `SANDBOX_TYPE` | `langsmith` | `langsmith` |
| `DEFAULT_REPO_OWNER` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `DEFAULT_REPO_NAME` | `open-swe-pilot` *(confirm)* | `open-swe-pilot` *(confirm)* |
| `DASHBOARD_BASE_URL` | `https://openswe-dev.seahaven.com` *(confirm host)* | `https://openswe.seahaven.com` *(confirm host)* |
| `DASHBOARD_API_BASE_URL` | same as base | same as base |
| `DASHBOARD_ALLOWED_ORIGINS` | same as base | same as base |
| `LLM_MODEL_ID` | `bedrock_converse:us.anthropic.claude-opus-4-8` | `bedrock_converse:us.anthropic.claude-opus-4-8` |
3. **Out-of-band SSM config — NOT created by CDK.** Operationally-variable or
env-specific-unknown values listed in `OUT_OF_BAND_SSM` and populated by
`put-config.sh`. The keystone is `DEFAULT_SANDBOX_SNAPSHOT_ID` (changes on every
snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the
GitHub App ids, Slack ids, and LangSmith tenant/urls.
### Kebab-Aspect deviation
`KebabNamingAspect` exempts `AWS::SecretsManager::Secret` and `AWS::SSM::Parameter`
from the kebab check (see the `KEBAB_EXEMPT_RESOURCE_TYPES` set) — the UPPER_SNAKE
env-var segment is a required, documented deviation for a lossless store→env
round-trip. Every other explicitly-named resource is still validated. Covered by a
dedicated case in `test/kebab-naming-aspect.test.ts`.
### Deploy ordering (values BEFORE the box boots)
The shells are synth-able now (T11). Population is out-of-band and happens **after**
`cdk deploy open-swe-<env>` but **before** the EC2/T12 box first boots:
```bash
cdk deploy open-swe-<env> # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod> # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
```
`put-config.sh` ships `<FILL>` placeholders only (no real secret values committed);
provide each value inline, via `OPENSWE_PUT_<VAR>` env vars, or from a vault. It does
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
## Compute + ingress (`AppService` — T12)
`AppService` (`lib/constructs/app-service.ts`, one per env from `OpenSweStack`)
builds the box and its path to the internet. A **single** internet-facing ALB
(`app/seahaven-com`) and a **single** VPC are shared with the on-prem
`seahaven-site` stack, so open-swe **imports** the VPC, the ALB security group
(`sg-0b0301deed193258a`), the `:443` listener, and the public `seahaven.com`
zone — and never owns/mutates them. It **adds**:
- **One ARM64 EC2 box** (`open-swe-<env>-box`, `t4g.medium` dev / `t4g.large`
prod) in **private1 (us-east-1a)** — same AZ as the single NAT for in-AZ egress.
`requireImdsv2`, gp3 **encrypted** root, `deleteOnTermination` (no RETAIN
volume — see below). `userDataCausesReplacement: true`; user-data is rendered
from `deploy/ami/user-data.sh`.
- **A standalone instance SG** reachable **only** from the shared ALB SG on `:80`
(nginx). Egress open (NAT). The ALB SG is opened to the box via a **standalone
`CfnSecurityGroupEgress`** so the imported (on-prem-owned) SG is never mutated.
- **A target group → instance `:80`** (nginx is the sole ingress; the LangGraph
control plane stays on loopback `:2024`). Health check `GET /healthz`.
- **Two listener rules** on the imported `:443` listener, both → the same TG:
- **Webhooks** (priority **2** dev / **3** prod): `host ∈ {openswe-<env>, hooks-<env>}.seahaven.com` **AND** path `/webhooks/*`.
- **Site** (priority **10** dev / **11** prod): `host = openswe-<env>.seahaven.com` (dashboard SPA + `/dashboard/api/`).
- **Route53 alias records** `openswe[-dev]` + `hooks[-dev]` → the shared ALB.
- **Four CloudWatch log groups** (`/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`) at **30-day** retention (IaC-owned; mirrors the CW-agent config).
### Listener-rule ordering (load-bearing)
The shared listener already has a **host-agnostic** `/webhooks/*` PATH rule at
**priority 5** (on-prem). ALB rules are first-match by ascending priority, so the
open-swe webhook rule **must** sit below 5 or every `…/webhooks/*` request (any
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
a host condition, so it does **not** steal the on-prem hosts' webhooks. The
dashboard "site" rule carries no path that collides with rule 5, so it sits at
10/11.
**Cross-stack coordination (T13 review).** The `seahaven-site` (on-prem) and
`open-swe` stacks both add resources to the *same imported* listener and ALB SG.
This is safe: each stack owns only the resources it declares (its own logical
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
and the standalone `CfnSecurityGroupEgress` never mutates the shared SG's own
definition (the pattern on-prem itself uses). The one shared namespace that needs
care is **listener-rule priority** (globally unique per listener; a collision is
a fail-*safe* deploy error, not silent drift). Ownership — keep disjoint:
`seahaven-site` = **4-7 + default**; `open-swe` = **2, 3, 10, 11**. open-swe's
webhook rules are host-scoped to its own `*.seahaven.com` hosts, so they never
match an on-prem `seahavenind.com` host.
### Security review (T5/T12 `/sh-security-review`)
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
One **confirmed medium** (OSWE-T12-01: nginx's 1 MB default `client_max_body_size`
would 413 large GitHub webhooks before in-app signature verification) is fixed in
`open-swe.nginx.conf` (`25m` on `/webhooks/`, `10m` on `/dashboard/api/`). The
hooks hostname is scoped to `/webhooks/*` only (OSWE-T12-02 hygiene). An
X-Forwarded-For spoof candidate was **killed** — no code trusts the leftmost XFF.
No confirmed critical/high; no block.
## Baked AMI + EBS-replacement discipline
`bakedOpenSweArm64()` (in `lib/constructs/ami-cache.ts`) pins the custom
**open-swe-base-arm64** image by EXACT id (`BAKED_OPEN_SWE_AMI_ID`) via
`MachineImage.genericLinux({ "us-east-1": "<ami-id>" })` — no SSM lookup, so synth
and deploy are fully offline/deterministic. The image is built by
`deploy/ami/open-swe-base.pkr.hcl` (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW
agent + boot templates, **no secrets**); the box's `user-data.sh` assumes that
baked layout (`/opt/open-swe`, `openswe` user, nginx, CW agent).
Pinning by exact id (vs a `most_recent` name filter) is what prevents a routine
deploy from silently swapping the AMI → **EC2 instance replacement** (the
file-share data-loss root cause — memory `feedback_inline_ebs_volumes`).
- `userDataCausesReplacement: true` is **deliberate** — user-data is
provisioning-only and carries no durable state.
- **No durable state on the box → no RETAIN volume.** The in-memory langgraph
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The goal is
replacement-*tolerance*, not avoidance.
- **Snapshot-before-replace** still applies operationally: before any replacing
deploy snapshot the root volume and wait `state=completed`, and re-verify "no
local-only durable state" first.
Refresh the AMI deliberately:
```bash
cd deploy/ami && packer build open-swe-base.pkr.hcl # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack # WILL show "requires replacement"
```
> `cdk.context.json` is `{}` — nothing is resolved via context anymore (the AMI is
> a static id pin), so synth makes no live AWS call.
## Commands
```bash
npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test # jest — naming Aspect
```
## CI/CD (T18 — `.github/workflows/ci-infra.yml` + `cd-infra.yml`)
Path-filtered, OIDC-only (no static keys). The Python agent keeps its own
`ci.yml` ("CI"); these two add the `/infra` half.
| Workflow | Trigger | Does |
|---|---|---|
| `ci-infra.yml` | PR touching `infra/**` | `tsc` + `jest` + `cdk synth` (reusable `ci-typescript-cdk.yaml`). |
| `cd-infra.yml` | push to `dev`/`main` touching `infra/**`, or dispatch | CI (pre-deploy) → per-env `cdk deploy`. |
`cd-infra.yml` flow:
- **push to `dev`** → CI green → **auto** `cdk deploy OpenSweDevStack` (assumes
`githubdeploy-open-swe-infra-dev`; the job declares **no** `environment:`, so the
OIDC subject is `…:ref:refs/heads/dev` — matching that role's trust).
- **push to `main`** → CI green → `cdk deploy OpenSweProdStack` behind the
**`prod` GitHub Environment** (required reviewer = Adam). The `environment: prod`
declaration both fires the manual-approval gate and makes the OIDC subject
`…:environment:prod` — matching `githubdeploy-open-swe-infra-prod`'s trust.
**Why not the reusable `cd-cdk.yaml`:** it runs `cdk deploy --all`, which from a
single-env push would deploy the *other* env + the shared IAM stack — breaking the
per-env boundary. So CD targets one stack explicitly per env. The shared
`open-swe-iam` stack is **not** deployed by CD (privileged, human-gated — T6).
**Gating note:** infra CI is enforced at the *deploy* boundary (`cd-infra`'s
`deploy-*` jobs `needs: ci`), not as a branch-protection required check —
path-filtering a *required* check would deadlock app-only PRs (a skipped required
check never satisfies). Making `Infra CI` a required check later needs a
skip-aware shim or dropping its path filter.
**Prerequisites (set post-T6, when the roles exist):**
- repo **variables** `AWS_DEPLOY_ROLE_INFRA_DEV` / `AWS_DEPLOY_ROLE_INFRA_PROD`
= the `githubdeploy-open-swe-infra-<env>` role ARNs (`open-swe-iam` outputs).
- a GitHub **Environment** named `prod` with Adam as a required reviewer.
> App-side CD (CI → S3 artifact → SSM deploy via `githubdeploy-open-swe-app-<env>`)
> is **T19**, not here.
## Deploy ordering (when the gate clears — NOT yet)
1. **`open-swe-iam` first** — apply the IAM stack (T6, human-gated), then set the
repo `AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}` variables from its role-ARN outputs and
configure the `prod` Environment reviewer (BLOCK#3).
2. **Security gate** — T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review` on
the synth; resolve every confirmed critical/high.
3. **IAM applied** (T6) — only after the gate.
4. Env stacks: first `open-swe-dev` (T14, manual validate), then CD auto-deploys
dev on push; `open-swe-prod` (T21) behind the `prod` Environment approval.
## Version policy
`aws-cdk-lib` is pinned EXACT (`2.260.0`) — no `^`/`~`. Dependabot keeps it
current; CI (`npm ci` + `cdk synth`) + dependency review gate each bump. See
`aws-infrastructure.md` "CDK Version Policy" and memory
`feedback_cdk_lib_bundled_deps`.

View file

@ -1,43 +0,0 @@
#!/usr/bin/env node
import "source-map-support/register";
import * as cdk from "aws-cdk-lib";
import { ACCOUNT, REGION, bootstrapQualifier } from "../lib/config";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
import { KebabNamingAspect } from "../lib/aspects/kebab-naming-aspect";
const app = new cdk.App();
const env = { account: ACCOUNT, region: REGION };
// Account-level shared OIDC deploy roles (singletons). Deployed FIRST.
new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env,
});
// The two env stacks — explicit kebab-case stackName (never let CDK default to
// PascalCase), env-parameterised so resources are `open-swe-<env>-*`.
//
// B-1 / OSWE-IAC-01: dev synthesizes against its OWN bootstrap qualifier
// (`oswedev`), so it deploys via the cdk-oswedev-* roles the dev infra role is
// scoped to — and NOT the default hnb659fds bootstrap roles that deploy prod.
// Prod stays on the default qualifier (no synthesizer override).
new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env,
envName: "dev",
synthesizer: new cdk.DefaultStackSynthesizer({
qualifier: bootstrapQualifier("dev"),
}),
});
new OpenSweStack(app, "OpenSweProdStack", {
stackName: "open-swe-prod",
env,
envName: "prod",
});
// Fail synth on any non-kebab-case explicit resource/stack name.
cdk.Aspects.of(app).add(new KebabNamingAspect());
app.synth();

View file

@ -1 +0,0 @@
{}

View file

@ -1,21 +0,0 @@
{
"app": "npx ts-node --prefer-ts-exts bin/app.ts",
"watch": {
"include": ["**"],
"exclude": [
"README.md",
"cdk*.json",
"**/*.d.ts",
"**/*.js",
"tsconfig.json",
"package*.json",
"node_modules",
"cdk.out"
]
},
"context": {
"@aws-cdk/aws-lambda:recognizeLayerVersion": true,
"@aws-cdk/core:checkSecretUsage": true,
"@aws-cdk/core:target-partitions": ["aws"]
}
}

View file

@ -1,9 +0,0 @@
module.exports = {
testEnvironment: "node",
roots: ["<rootDir>/test"],
testMatch: ["**/*.test.ts"],
preset: "ts-jest",
transform: {
"^.+\\.tsx?$": ["ts-jest", { tsconfig: "tsconfig.json" }],
},
};

View file

@ -1,119 +0,0 @@
import { Annotations, CfnResource, IAspect, Stack, Token } from "aws-cdk-lib";
import { IConstruct } from "constructs";
/**
* One "/"-delimited segment must be lower kebab-case: `a-b-c`, digits allowed.
*/
const KEBAB_SEGMENT = /^[a-z0-9]+(-[a-z0-9]+)*$/;
/**
* CloudFormation property keys that carry an *explicit physical name*. The
* codegen'd L1 stores these either camelCased (`roleName`) or CFN-cased
* (`RoleName`) depending on the construct, so the aspect matches keys
* case-insensitively.
*
* We deliberately validate physical NAMES + the stack name only — not CDK
* logical construct ids (those are conventionally PascalCase, e.g.
* `InfraDeployRole`, and validating them would be wrong).
*/
const NAME_PROPERTY_KEYS = [
"RoleName",
"BucketName",
"FunctionName",
"TableName",
"LogGroupName",
"QueueName",
"TopicName",
"SecretName",
"StreamName",
"RepositoryName",
"DBInstanceIdentifier",
"DBClusterIdentifier",
"StateMachineName",
"RuleName",
"UserPoolName",
];
// NOTE: `PolicyName` is intentionally NOT checked — CDK auto-generates inline
// `DefaultPolicy` names (e.g. "InstanceRoleDefaultPolicyF15F...") from the
// logical id; those are not explicit, user-controlled physical names and are
// outside the naming convention's scope.
const NAME_KEYS_LC = new Set(NAME_PROPERTY_KEYS.map((k) => k.toLowerCase()));
/**
* Resource types whose physical NAME is a REQUIRED deviation from kebab-case:
* the open-swe config store names secrets `open-swe-<env>/<ENV_VAR_NAME>` and SSM
* params `/open-swe-<env>/<ENV_VAR_NAME>`, where the last segment is the LITERAL
* UPPER_SNAKE environment-variable name. The boot hook
* (deploy/seahaven/fetch-config.sh) strips the prefix and exports that segment
* verbatim, so a lossless store→env round-trip needs the exact env-var name —
* it cannot be kebab-cased. These two resource types are therefore exempt; every
* OTHER explicitly-named resource is still validated. (The `open-swe-<env>`
* prefix is code-generated from `prefix(env)` and is always kebab-case.)
*/
const KEBAB_EXEMPT_RESOURCE_TYPES = new Set([
"AWS::SecretsManager::Secret",
"AWS::SSM::Parameter",
]);
/**
* `true` when every non-empty "/"-delimited segment is kebab-case.
*
* Path-style names are tolerated so the same check works for Secrets Manager
* (`open-swe-dev/foo`), SSM params (`/open-swe-dev/foo`) and log groups
* (`/open-swe/dev/agent`): each segment is validated independently, and a
* leading slash (empty first segment) is ignored.
*/
export function isKebabCase(value: string): boolean {
return value
.split("/")
.filter((seg) => seg.length > 0)
.every((seg) => KEBAB_SEGMENT.test(seg));
}
/**
* Aspect that FAILS synth (`Annotations.addError`) when an explicitly-named
* resource — or a stack name — is not kebab-case. Enforces the org naming
* convention (naming-conventions.md) deterministically at synth time so a
* non-conforming name can never reach a deploy. Wired in bin/app.ts via
* `Aspects.of(app).add(new KebabNamingAspect())`.
*/
export class KebabNamingAspect implements IAspect {
public visit(node: IConstruct): void {
if (node instanceof Stack) {
const name = node.stackName;
if (!Token.isUnresolved(name) && !isKebabCase(name)) {
Annotations.of(node).addError(
`Stack name "${name}" is not kebab-case (open-swe naming convention).`,
);
}
return;
}
if (node instanceof CfnResource) {
// The config store's Secret/Parameter names carry the literal UPPER_SNAKE
// env-var name per the fetch-config naming contract — a required deviation.
if (KEBAB_EXEMPT_RESOURCE_TYPES.has(node.cfnResourceType)) {
return;
}
// `_cfnProperties` is the props as set on the L1; resolve to collapse any
// intrinsic tokens (refs/getatt) so only literal strings are checked.
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const raw = (node as any)._cfnProperties ?? {};
const resolved = Stack.of(node).resolve(raw) ?? {};
for (const [key, value] of Object.entries(resolved)) {
if (
NAME_KEYS_LC.has(key.toLowerCase()) &&
typeof value === "string" &&
!Token.isUnresolved(value) &&
!isKebabCase(value)
) {
Annotations.of(node).addError(
`Resource "${node.node.path}" property ${key}="${value}" is not kebab-case ` +
`(open-swe naming convention).`,
);
}
}
}
}
}

View file

@ -1,59 +0,0 @@
/**
* Shared, non-sensitive constants for the open-swe infra app.
* Account / region are locked per the migration spec (TODO.md "Architecture (locked)").
*/
export const ACCOUNT = "328440206208";
export const REGION = "us-east-1";
export const GITHUB_ORG = "Sea-Haven-Industries";
export const GITHUB_REPO = "open-swe";
export type EnvName = "dev" | "prod";
/** `open-swe-dev` / `open-swe-prod` — kebab-case stack + resource prefix. */
export const prefix = (env: EnvName): string => `open-swe-${env}`;
/**
* Per-ENV GitHub OIDC trust subject for the deploy roles (T5 OSWE-IAC-01/02 fix:
* the dev/prod boundary is enforced in the IAM trust, not by convention).
*
* - `dev` → the `dev` integration branch ref (auto-deploy on push to dev).
* - `prod` → the **GitHub `prod` Environment** subject. A workflow can only mint
* a token with sub `…:environment:prod` by declaring `environment: prod`,
* which triggers the Environment's manual-approval gate (Adam, T18). So the
* prod approval is now expressed at the IAM layer: a dev-branch token can
* never assume a prod deploy role.
*
* Each env gets its OWN infra + app role (githubdeploy-open-swe-{infra,app}-<env>)
* so a dev token cannot reach prod. Exact subject → StringEquals (no `*`).
*
* Cross-env deploy isolation is enforced at the bootstrap layer too — see
* `bootstrapQualifier`: dev runs on its own qualifier so the dev infra role
* cannot assume the bootstrap roles that deploy prod.
*/
export const oidcSubject = (env: EnvName): string =>
env === "prod"
? `repo:${GITHUB_ORG}/${GITHUB_REPO}:environment:prod`
: `repo:${GITHUB_ORG}/${GITHUB_REPO}:ref:refs/heads/dev`;
/**
* Per-env CDK bootstrap qualifier (B-1 / OSWE-IAC-01 fix). Dev runs on its OWN
* qualifier `oswedev` (bootstrapped into the `CDKToolkit-oswedev` stack), so the
* dev infra deploy role only assumes `cdk-oswedev-*` and can NO LONGER assume the
* default `cdk-hnb659fds-*` set whose admin `cfn-exec-role` deploys prod. Prod
* stays on the default qualifier. This closes the cross-env escalation where a
* dev-branch token could `cdk deploy open-swe-prod` via the shared bootstrap
* roles, bypassing prod's Environment approval gate.
*/
export const DEFAULT_BOOTSTRAP_QUALIFIER = "hnb659fds";
export const bootstrapQualifier = (env: EnvName): string =>
env === "dev" ? "oswedev" : DEFAULT_BOOTSTRAP_QUALIFIER;
/**
* The GitHub Actions OIDC provider already exists account-wide (created for
* seahaven-site; see .github/oidc-deploy-roles.yaml `CreateOIDCProvider=false`).
* Reference it by ARN — never create a duplicate `AWS::IAM::OIDCProvider`
* (CloudFormation rejects a second provider for the same URL).
*/
export const GITHUB_OIDC_PROVIDER_ARN = `arn:aws:iam::${ACCOUNT}:oidc-provider/token.actions.githubusercontent.com`;

View file

@ -1,35 +0,0 @@
import * as ec2 from "aws-cdk-lib/aws-ec2";
import { REGION } from "../config";
/**
* The baked open-swe base AMI (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW agent
* + boot templates — NO secrets), produced by `deploy/ami/open-swe-base.pkr.hcl`.
* Pinned by EXACT id (not a name filter) so synth/deploy is fully offline and
* deterministic.
*
* Built 2026-06-26 from open-swe-base-arm64-20260626-203433.
*
* ── EBS / AMI replacement discipline (memory feedback_inline_ebs_volumes) ──
*
* Refresh DELIBERATELY: `cd deploy/ami && packer build open-swe-base.pkr.hcl`,
* then update this id. A new id → EC2 instance REPLACEMENT. Pinning by exact id
* (vs a `most_recent` name filter) is what prevents a routine deploy from silently
* swapping the AMI — the root cause of the file-share data-loss incidents
* (5/15, 5/27, 6/5).
*
* `userDataCausesReplacement: true` (AppService) is likewise DELIBERATE: user-data
* is provisioning-only and the box holds NO durable state (the langgraph store is
* in-memory, rebuilt every boot from S3 + Secrets Manager / SSM), so there is
* intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The design
* goal is replacement-TOLERANCE, not avoidance.
*
* Operational guard before ANY replacing deploy (AMI / userData / instance-type):
* snapshot the root volume AND wait `state=completed`, re-verify "no local-only
* durable state", and review the `cdk diff` replacement at PR time.
*/
export const BAKED_OPEN_SWE_AMI_ID = "ami-00080084502093021";
/** The baked open-swe base image, pinned by id (offline, deterministic). */
export function bakedOpenSweArm64(): ec2.IMachineImage {
return ec2.MachineImage.genericLinux({ [REGION]: BAKED_OPEN_SWE_AMI_ID });
}

View file

@ -1,368 +0,0 @@
import * as fs from "fs";
import * as path from "path";
import * as cdk from "aws-cdk-lib";
import * as ec2 from "aws-cdk-lib/aws-ec2";
import * as elbv2 from "aws-cdk-lib/aws-elasticloadbalancingv2";
import * as elbTargets from "aws-cdk-lib/aws-elasticloadbalancingv2-targets";
import * as logs from "aws-cdk-lib/aws-logs";
import * as route53 from "aws-cdk-lib/aws-route53";
import * as iam from "aws-cdk-lib/aws-iam";
import * as ssm from "aws-cdk-lib/aws-ssm";
import { Construct, IConstruct } from "constructs";
import { EnvName, prefix } from "../config";
import { bakedOpenSweArm64 } from "./ami-cache";
/**
* Shared seahaven-vpc + internet-facing ALB facts (read-only recon 2026-06-26;
* scratchpad/T12-infra-facts.md). A SINGLE VPC and a SINGLE shared ALB front
* both the on-prem `seahaven-site` stack and open-swe. We IMPORT every one of
* these and NEVER own them - open-swe only ADDS its own instance SG, a standalone
* ALB-egress rule, listener rules, a target group, and DNS records.
*
* ── Cross-stack coordination on the SHARED listener + ALB SG (T13 review) ──
* Two CDK stacks (seahaven-site, open-swe) add resources to the same imported
* `:443` listener and ALB SG. This is safe because each stack owns ONLY the
* resources it declares (its own logical ids): an on-prem `cdk deploy` computes a
* changeset over its own template and cannot delete rules/egress it never
* declared. The standalone-egress pattern is what on-prem itself uses
* (sgr-0c57812752a3bca13), so it does not mutate the shared SG's own definition.
*
* The ONE shared namespace that REQUIRES coordination is listener-rule PRIORITY
* (globally unique per listener; a collision is a fail-SAFE deploy error, not
* silent drift). Ownership map - keep these disjoint when editing either stack:
* - seahaven-site (on-prem): priorities 4-7 + default.
* - open-swe: priorities 2, 3 (webhooks) and 10, 11 (site).
* open-swe's webhook rules are HOST-scoped to its own *.seahaven.com hosts, so
* they never match (let alone "steal") any seahavenind.com / on-prem host.
*/
const SHARED = {
vpcId: "vpc-0d3d4b67bd0cf8a68",
availabilityZones: ["us-east-1a", "us-east-1b"],
// Private subnets host the EC2 box. The single NAT gateway lives in 1a, so the
// box is pinned to private1 (1a) for in-AZ NAT egress (no cross-AZ data $).
privateSubnetIds: ["subnet-04e38c507e96f1926", "subnet-0a0b4fc6f296dfba5"],
instanceSubnetId: "subnet-04e38c507e96f1926",
instanceAz: "us-east-1a",
albDnsName: "seahaven-com-1856441924.us-east-1.elb.amazonaws.com",
albCanonicalHostedZoneId: "Z35SXDOTRQ7X7K",
albSecurityGroupId: "sg-0b0301deed193258a",
httpsListenerArn:
"arn:aws:elasticloadbalancing:us-east-1:328440206208:listener/app/seahaven-com/222c3257354ab559/bab8bcf0da0e2927",
publicZoneId: "Z06652411XKH89KTZD3XA",
publicZoneName: "seahaven.com",
} as const;
/**
* Per-env public hostnames, listener-rule priorities, and instance size.
*
* ── Listener-rule ordering hazard (load-bearing) ──
* The shared listener already has a HOST-AGNOSTIC `/webhooks/*` PATH rule at
* priority 5 (the on-prem seahaven-site stack owns it). ALB rules are first-match
* by ASCENDING priority, so a `…/webhooks/*` request to our host would match
* rule 5 (priority 5) and be forwarded to the on-prem target BEFORE any host rule
* at 10+. Therefore our webhook rule MUST sit below priority 5. The catch-all
* "site" rule (dashboard SPA + /dashboard/api/) carries no path that collides
* with rule 5, so it can sit at any free higher number (10/11). Free priorities
* confirmed by recon: 1-3 and 8+ (4=forgejo, 5=/webhooks/*, 6/7=seahavenind).
*/
const ENV_NET: Record<
EnvName,
{
dashboardHost: string;
hooksHost: string;
webhookPriority: number;
sitePriority: number;
instanceType: string;
}
> = {
dev: {
dashboardHost: "openswe-dev.seahaven.com",
hooksHost: "hooks-dev.seahaven.com",
webhookPriority: 2,
sitePriority: 10,
instanceType: "t4g.medium",
},
prod: {
dashboardHost: "openswe.seahaven.com",
hooksHost: "hooks.seahaven.com",
webhookPriority: 3,
sitePriority: 11,
instanceType: "t4g.large",
},
};
export interface AppServiceProps {
readonly envName: EnvName;
/** Least-privilege EC2 instance role (per-env; from InstanceRole). */
readonly instanceRole: iam.IRole;
/** S3 artifact key prefix the box pulls app.tar.gz / spa.tar.gz from. */
readonly artifactPrefix?: string;
}
/**
* The open-swe compute + ingress wiring for one env (T12):
* - one ARM64 EC2 box in private1 (1a), replacement-tolerant (no RETAIN volume),
* - a standalone instance SG reachable ONLY from the shared ALB SG on :80,
* - a target group -> instance:80 (nginx is the sole ingress; :2024 stays loopback),
* - two listener rules on the imported :443 listener (webhooks below the on-prem
* path rule; site catch-all above it), both -> the same TG,
* - Route53 alias records for both hostnames -> the shared ALB,
* - IaC-owned CloudWatch log groups at 30-day retention.
*
* Everything ALB/VPC/zone-side is IMPORTED. Synth is offline: the AMI is the
* cdk.context.json-pinned AL2023 ARM64 placeholder until the baked
* open-swe-base-arm64 id is pinned before the first real deploy.
*/
export class AppService extends Construct {
public readonly instance: ec2.Instance;
public readonly targetGroup: elbv2.ApplicationTargetGroup;
/** Name of the SSM document CI fires to roll the box to the latest release. */
public readonly deployDocumentName: string;
constructor(scope: Construct, id: string, props: AppServiceProps) {
super(scope, id);
const env = props.envName;
const p = prefix(env);
const net = ENV_NET[env];
const artifactPrefix = props.artifactPrefix ?? "releases/latest";
// Import the shared VPC with explicit attributes (no fromLookup -> offline synth).
const vpc = ec2.Vpc.fromVpcAttributes(this, "Vpc", {
vpcId: SHARED.vpcId,
availabilityZones: [...SHARED.availabilityZones],
privateSubnetIds: [...SHARED.privateSubnetIds],
});
// Standalone instance SG. Egress open (NAT path); ingress only from the ALB SG.
const instanceSg = new ec2.SecurityGroup(this, "InstanceSg", {
vpc,
securityGroupName: `${p}-instance-sg`,
description: `${p} instance SG - ingress only from the shared ALB SG on :80; egress via NAT.`,
allowAllOutbound: true,
});
instanceSg.addIngressRule(
ec2.Peer.securityGroupId(SHARED.albSecurityGroupId),
ec2.Port.tcp(80),
`${p}: shared ALB SG to nginx :80`,
);
// Open the IMPORTED ALB SG to our instance via a STANDALONE egress rule, so we
// never mutate the ALB SG's own (on-prem-owned) definition.
new ec2.CfnSecurityGroupEgress(this, "AlbToInstanceEgress", {
groupId: SHARED.albSecurityGroupId,
ipProtocol: "tcp",
fromPort: 80,
toPort: 80,
destinationSecurityGroupId: instanceSg.securityGroupId,
description: `${p}: ALB to instance nginx :80`,
});
// The app-deploy procedure (deploy/ami/deploy.sh) is a normal reviewable repo
// file; CDK base64-encodes it (single line - no `$`/regex-special chars in the
// base64 alphabet) and renders it into user-data's @@DEPLOY_SH_B64@@ token, so
// user-data writes it verbatim to /opt/open-swe/bin/deploy.sh at first boot.
// The same file is run by the open-swe-<env>-deploy SSM document on every
// release - a single source of truth for "pull release, build venv, restart".
const deployShPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "deploy.sh");
// Minify before embedding: strip full-line comments + blank lines (keep the
// shebang) so the base64 fits EC2's 25.6 KB user-data limit. The repo file
// keeps its comments; only the on-box copy is minified. deploy.sh becomes
// opaque base64 here, so this never affects user-data's heredoc parsing.
const deployShMin = fs
.readFileSync(deployShPath, "utf8")
.split("\n")
.filter((line, i) => i === 0 || (!/^\s*#/.test(line) && line.trim() !== ""))
.join("\n");
const deployShB64 = Buffer.from(deployShMin, "utf8").toString("base64");
// Render the provisioning script's @@tokens@@ into the instance user-data.
// userDataCausesReplacement makes a bootstrap change roll a fresh box (the box
// holds no durable state - see ami-cache.ts / user-data.sh). Editing deploy.sh
// therefore also rolls the box (its base64 is embedded here) - acceptable: the
// box is replacement-tolerant, and ongoing releases never touch user-data.
const userDataPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "user-data.sh");
const userData = ec2.UserData.custom(
fs
.readFileSync(userDataPath, "utf8")
// %%...%% tokens are CDK-substituted here; they are DELIBERATELY a
// different delimiter from the @@...@@ tokens user-data.sh seds into the
// baked systemd/nginx templates, so CDK can never clobber a sed pattern
// (a shared @@OPENSWE_ENV@@/@@SERVER_NAME@@ left the unit unsubstituted).
.replace(/%%OPENSWE_ENV%%/g, env)
.replace(/%%ASSETS_BUCKET%%/g, `${p}-assets`)
.replace(/%%SERVER_NAME%%/g, net.dashboardHost)
.replace(/%%ARTIFACT_PREFIX%%/g, artifactPrefix)
.replace(/%%DEPLOY_SH_B64%%/g, deployShB64),
);
this.instance = new ec2.Instance(this, "Instance", {
vpc,
vpcSubnets: {
subnets: [
ec2.Subnet.fromSubnetAttributes(this, "InstanceSubnet", {
subnetId: SHARED.instanceSubnetId,
availabilityZone: SHARED.instanceAz,
}),
],
},
instanceType: new ec2.InstanceType(net.instanceType),
// The baked open-swe base AMI (deploy/ami packer build) - ARM64 Ubuntu 24.04
// with the /opt/open-swe layout, openswe user, nginx, and CW agent that
// user-data.sh assumes. Pinned by exact id (see ami-cache.ts); refresh by
// rebuilding and updating BAKED_OPEN_SWE_AMI_ID.
machineImage: bakedOpenSweArm64(),
role: props.instanceRole,
securityGroup: instanceSg,
userData,
userDataCausesReplacement: true,
requireImdsv2: true,
instanceName: `${p}-box`,
blockDevices: [
{
deviceName: "/dev/xvda",
// gp3 encrypted root; deleteOnTermination (no durable on-box state ->
// intentionally NO standalone RETAIN volume; see ami-cache.ts).
volume: ec2.BlockDeviceVolume.ebs(30, {
volumeType: ec2.EbsDeviceVolumeType.GP3,
encrypted: true,
deleteOnTermination: true,
}),
},
],
});
// `requireImdsv2: true` makes CDK auto-create a launch template, and it names
// that LT from the construct id ("Instance" -> "InstanceLaunchTemplate") with NO
// env qualifier — so OpenSweDevStack and OpenSweProdStack both want the identical
// LT name and the second env to deploy fails with
// InvalidLaunchTemplateName.AlreadyExistsException (prod rollback, 2026-06-29).
// Force a per-env LT name. Done via an aspect because the LT is created at synth
// time by the requireImdsv2 handling, not in this constructor.
cdk.Aspects.of(this.instance).add({
visit(node: IConstruct) {
if (node instanceof ec2.CfnLaunchTemplate) {
node.launchTemplateName = `${p}-lt`;
}
// The instance references the LT BY NAME, so the reference must be renamed in
// lockstep (preserve the GetAtt version) or CFN can't find the template.
if (node instanceof ec2.CfnInstance && node.launchTemplate) {
const spec = node.launchTemplate as ec2.CfnInstance.LaunchTemplateSpecificationProperty;
node.launchTemplate = { ...spec, launchTemplateName: `${p}-lt` };
}
},
});
// SSM deploy document (open-swe-<env>-deploy): runs the baked
// /opt/open-swe/bin/deploy.sh to pull the latest release + restart. CI fires it
// (tag-scoped to project=open-swe,env=<env>) after uploading a release, so the
// app deploy role needs SendCommand ONLY on this document - NOT on the generic
// AWS-RunShellScript (closes the T4 BLOCK#3 arbitrary-shell timebox).
this.deployDocumentName = `${p}-deploy`;
new ssm.CfnDocument(this, "DeployDoc", {
name: this.deployDocumentName,
documentType: "Command",
documentFormat: "YAML",
updateMethod: "NewVersion",
content: {
schemaVersion: "2.2",
description: `Roll the ${p} box to the latest published release (runs /opt/open-swe/bin/deploy.sh).`,
mainSteps: [
{
action: "aws:runShellScript",
name: "deploy",
inputs: {
// Fixed command - no parameters, so nothing untrusted is interpolated
// into the shell. The script itself reads /etc/open-swe/boot.env.
runCommand: ["bash /opt/open-swe/bin/deploy.sh"],
},
},
],
},
});
// Target group -> instance:80 (nginx). Health check hits nginx's /healthz
// (returns 200; the dashboard TG health path defined in open-swe.nginx.conf).
this.targetGroup = new elbv2.ApplicationTargetGroup(this, "Tg", {
vpc,
targetGroupName: `${p}-tg`,
port: 80,
protocol: elbv2.ApplicationProtocol.HTTP,
targetType: elbv2.TargetType.INSTANCE,
targets: [new elbTargets.InstanceTarget(this.instance)],
deregistrationDelay: cdk.Duration.seconds(15),
healthCheck: {
path: "/healthz",
healthyHttpCodes: "200",
interval: cdk.Duration.seconds(30),
timeout: cdk.Duration.seconds(5),
healthyThresholdCount: 2,
unhealthyThresholdCount: 3,
},
});
// Import the shared :443 listener (with its ALB SG) and ADD our two rules.
const albSg = ec2.SecurityGroup.fromSecurityGroupId(this, "AlbSg", SHARED.albSecurityGroupId, {
mutable: false,
});
const listener = elbv2.ApplicationListener.fromApplicationListenerAttributes(this, "HttpsListener", {
listenerArn: SHARED.httpsListenerArn,
securityGroup: albSg,
});
// (1) Webhooks - accepted on EITHER host (integrations may target either), and
// MUST be below the on-prem path-only rule 5 (see ENV_NET note).
new elbv2.ApplicationListenerRule(this, "WebhooksRule", {
listener,
priority: net.webhookPriority,
conditions: [
elbv2.ListenerCondition.hostHeaders([net.dashboardHost, net.hooksHost]),
elbv2.ListenerCondition.pathPatterns(["/webhooks/*"]),
],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
// (2) Dashboard SPA + /dashboard/api/ (OAuth) - DASHBOARD host ONLY. The hooks
// host intentionally serves nothing but /webhooks/* (rule 1), so the OAuth /
// dashboard surface stays single-origin (OSWE-T12-02). Non-webhook paths on the
// hooks host fall through to the on-prem default.
new elbv2.ApplicationListenerRule(this, "SiteRule", {
listener,
priority: net.sitePriority,
conditions: [elbv2.ListenerCondition.hostHeaders([net.dashboardHost])],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
// Route53 ALIAS records -> the shared ALB, for both hostnames.
const zone = route53.HostedZone.fromHostedZoneAttributes(this, "PublicZone", {
hostedZoneId: SHARED.publicZoneId,
zoneName: SHARED.publicZoneName,
});
const albAlias: route53.IAliasRecordTarget = {
bind: () => ({
dnsName: SHARED.albDnsName,
hostedZoneId: SHARED.albCanonicalHostedZoneId,
}),
};
for (const [label, host] of [
["Dashboard", net.dashboardHost],
["Hooks", net.hooksHost],
] as const) {
new route53.ARecord(this, `${label}Alias`, {
zone,
recordName: host,
target: route53.RecordTarget.fromAlias(albAlias),
comment: `${p} ${label.toLowerCase()} -> shared seahaven-com ALB`,
});
}
// IaC-owned CloudWatch log groups at 30-day retention. Names mirror the
// CloudWatch-agent config (deploy/ami/templates/amazon-cloudwatch-agent.json);
// owning them here makes retention declarative rather than agent-set. Logs are
// not durable state -> DESTROY on stack delete.
for (const suffix of ["app", "user-data", "nginx-access", "nginx-error"]) {
new logs.LogGroup(this, `Log-${suffix}`, {
logGroupName: `/open-swe/${env}/${suffix}`,
retention: logs.RetentionDays.ONE_MONTH,
removalPolicy: cdk.RemovalPolicy.DESTROY,
});
}
}
}

View file

@ -1,58 +0,0 @@
import * as cdk from "aws-cdk-lib";
import * as s3 from "aws-cdk-lib/aws-s3";
import { Construct } from "constructs";
import { EnvName, prefix } from "../config";
/**
* The per-env S3 artifact bucket (`open-swe-<env>-assets`) the box pulls its
* release from (T7). CI builds the SPA + packages the app source and uploads
* `app.tar.gz` / `spa.tar.gz` under `releases/<sha>/` + `releases/latest/`
* (`build-artifacts.yml`, via the `githubdeploy-open-swe-app-<env>` OIDC role);
* the box pulls `releases/latest/*` at boot / on deploy via its instance role.
*
* The bucket holds ONLY build artifacts — no secrets (those live in Secrets
* Manager + SSM), no durable runtime state (the langgraph store is in-memory and
* rebuilt every boot). It is therefore safe to treat as reproducible-from-CI, but
* we RETAIN it on stack delete so an accidental `cdk destroy` cannot strand the
* box with no artifact to pull on its next replacement.
*
* Security posture (locked, reviewed in T7):
* - `BLOCK_ALL` public access (this is an internal artifact store; ALB/nginx is
* the only public surface — never S3 directly).
* - SSE-S3 encryption at rest + `enforceSSL` (deny any non-TLS request).
* - versioned, so a bad release can be rolled back to the previous object
* version (the last-good-artifact story in T19); a lifecycle rule expires
* NONcurrent versions after 30 days so history does not grow unbounded.
* - aborts incomplete multipart uploads after 7 days (cost hygiene).
*
* The name is the load-bearing contract: `instance-role.ts` (read), the app
* deploy role in `github-deploy-roles.ts` (write), and `user-data.sh` /
* `deploy.sh` (`@@ASSETS_BUCKET@@`) all reference `open-swe-<env>-assets` by
* literal name, so it is set explicitly here rather than auto-generated.
*/
export class AssetsBucket extends Construct {
public readonly bucket: s3.Bucket;
constructor(scope: Construct, id: string, envName: EnvName) {
super(scope, id);
const p = prefix(envName);
this.bucket = new s3.Bucket(this, "Bucket", {
bucketName: `${p}-assets`,
blockPublicAccess: s3.BlockPublicAccess.BLOCK_ALL,
encryption: s3.BucketEncryption.S3_MANAGED,
enforceSSL: true,
versioned: true,
// Artifacts are reproducible from CI, but RETAIN protects against an
// accidental stack delete leaving the box with nothing to pull (see above).
removalPolicy: cdk.RemovalPolicy.RETAIN,
lifecycleRules: [
{
id: "expire-noncurrent-artifact-versions",
noncurrentVersionExpiration: cdk.Duration.days(30),
abortIncompleteMultipartUploadAfter: cdk.Duration.days(7),
},
],
});
}
}

View file

@ -1,265 +0,0 @@
import * as cdk from "aws-cdk-lib";
import * as secretsmanager from "aws-cdk-lib/aws-secretsmanager";
import * as ssm from "aws-cdk-lib/aws-ssm";
import { Construct } from "constructs";
import { EnvName } from "../config";
/**
* Config / secret "shells" for the boot hook (`deploy/seahaven/fetch-config.sh`).
*
* The naming contract (source of truth: the T9 env/secret/config inventory + the
* fetch-config header) is LITERAL env-var names as the last path segment:
*
* Secrets open-swe-<env>/<ENV_VAR_NAME> (AWS Secrets Manager)
* Config /open-swe-<env>/<ENV_VAR_NAME> (AWS SSM Parameter Store, FLAT)
*
* fetch-config reads secrets with `batch-get-secret-value --filters
* Key=name,Values=open-swe-<env>/` and config with `get-parameters-by-path
* --path /open-swe-<env>/` (NON-recursive), then strips the prefix so the last
* segment IS the exported variable name. So these resources MUST carry the
* UPPER_SNAKE env-var name verbatim — which is why both resource types are
* exempted from the kebab-naming Aspect (see aspects/kebab-naming-aspect.ts).
*
* Three buckets:
*
* 1. SECRETS_SHELLS — the 29 secrets. Created as value-LESS shells (an L1
* `CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which
* CloudFormation creates as an empty secret with no version). The real value
* is set out-of-band via `deploy/seahaven/put-config.sh` (put-secret-value)
* BEFORE the box boots. Because CDK never owns the value, a later
* `cdk deploy` can never clobber the operator-set value. AWS-managed key
* (alias/aws/secretsmanager) — no CMK, matching the instance role which
* deliberately omits kms:Decrypt.
*
* 2. IAC_MANAGED_SSM — stable / derivable config. Real values are owned here in
* IaC (one StringParameter each) so they are reproducible and reviewed.
*
* 3. Out-of-band SSM (NOT created here) — operationally-variable or
* env-specific-unknown config (e.g. DEFAULT_SANDBOX_SNAPSHOT_ID, which
* changes on every snapshot rebuild and would be clobbered by a deploy if it
* were CDK-managed; GitHub App ids; Slack ids; LangSmith tenant/urls). These
* are listed in OUT_OF_BAND_SSM purely for documentation and are set by
* `put-config.sh`, never by CDK.
*/
/** The 29 Secrets Manager secret VAR names (T9 inventory SECRETS table). */
export const SECRET_VARS: readonly string[] = [
"ANTHROPIC_API_KEY",
"CORRIDOR_API_TOKEN",
"CORRIDOR_MCP_TOKEN",
"CORRIDOR_TOKEN",
"DASHBOARD_JWT_SECRET",
"DAYTONA_API_KEY",
"EXA_API_KEY",
"FIREWORKS_API_KEY",
"GITHUB_APP_CLIENT_SECRET",
"GITHUB_APP_PRIVATE_KEY",
"GITHUB_PAT",
"GITHUB_WEBHOOK_SECRET",
"JUDGE_ANTHROPIC_API_KEY",
"LANGSMITH_API_KEY",
"LANGSMITH_API_KEY_PROD",
"LANGCHAIN_API_KEY",
"LINEAR_API_KEY",
"LINEAR_WEBHOOK_SECRET",
"RUNLOOP_API_KEY",
"SLACK_BOT_TOKEN",
"SLACK_CLIENT_SECRET",
"SLACK_SIGNING_SECRET",
"TOKEN_ENCRYPTION_KEY",
"USER_ID_API_KEY_MAP",
"X_SERVICE_AUTH_JWT_SECRET",
] as const;
// 25 secret shells (OPENAI_API_KEY / GOOGLE_API_KEY / GROQ_API_KEY removed in the
// Bedrock/Fireworks migration — those providers are dropped from SUPPORTED_MODELS and
// their keys revoked + secret objects deleted). The T9 inventory header said "29" vs
// 27 enumerated; reconciled
// (Adam confirm 2026-06-26): USER_ID_API_KEY_MAP (maps user ids -> API keys; flagged
// sensitive by the T5 security review) is a SECRET and is included here.
// JUDGE_ANTHROPIC_BASE_URL is a URL (non-sensitive config, eval-only) -> SSM/default,
// NOT a secret. So the inventory's "29" was effectively a miscount.
/** Short, value-free descriptions for the secret shells (no secret material). */
const SECRET_DESCRIPTIONS: Record<string, string> = {
ANTHROPIC_API_KEY: "Claude LLM API key (primary builder provider).",
CORRIDOR_API_TOKEN: "Corridor MCP token (optional).",
CORRIDOR_MCP_TOKEN: "Corridor MCP token alt name (optional).",
CORRIDOR_TOKEN: "Corridor MCP token alt name (optional).",
DASHBOARD_JWT_SECRET: "JWT signing secret for dashboard session cookies (REQUIRED).",
DAYTONA_API_KEY: "Daytona sandbox key (only if SANDBOX_TYPE=daytona).",
EXA_API_KEY: "Exa web-search key (optional).",
FIREWORKS_API_KEY: "Fireworks LLM key (only if a fireworks: model is used).",
GITHUB_APP_CLIENT_SECRET: "GitHub App OAuth client secret (dashboard login).",
GITHUB_APP_PRIVATE_KEY: "GitHub App private key PEM (installation-token minting).",
GITHUB_PAT: "GitHub PAT fallback (optional).",
GITHUB_WEBHOOK_SECRET: "GitHub webhook signature secret (prod-required).",
JUDGE_ANTHROPIC_API_KEY: "Eval judge key (optional; falls back to ANTHROPIC_API_KEY).",
LANGSMITH_API_KEY: "LangSmith key (dev).",
LANGSMITH_API_KEY_PROD: "LangSmith key (prod / deployed sandbox).",
LANGCHAIN_API_KEY: "LangSmith key alt name (fallback).",
LINEAR_API_KEY: "Linear API key (optional).",
LINEAR_WEBHOOK_SECRET: "Linear webhook signature secret (required when Linear is wired).",
RUNLOOP_API_KEY: "Runloop sandbox key (only if SANDBOX_TYPE=runloop).",
SLACK_BOT_TOKEN: "Slack bot token (optional).",
SLACK_CLIENT_SECRET: "Slack OAuth client secret.",
SLACK_SIGNING_SECRET: "Slack webhook signing secret (prod-required).",
TOKEN_ENCRYPTION_KEY: "Fernet key(s) for per-user GitHub-token encryption (REQUIRED).",
USER_ID_API_KEY_MAP: "JSON map of user id -> API key for per-user auth (optional, sensitive).",
X_SERVICE_AUTH_JWT_SECRET: "Service-auth JWT secret (optional).",
};
/**
* IaC-managed SSM config: stable / derivable values owned in code, per env.
* Values are functions of envName so dev/prod render correct hosts.
*
* Anything operationally-variable or env-specific-unknown is deliberately NOT
* here — see OUT_OF_BAND_SSM.
*/
export function iacManagedSsm(env: EnvName): Record<string, string> {
// Public host = seahaven.com (the migration's new AWS public face; confirmed by
// recon: seahaven.com Route53 zone + *.seahaven.com ACM cert are live on the ALB
// — distinct from the on-prem seahavenind.com). dev = openswe-dev, prod = openswe.
const host = `https://openswe${env === "dev" ? "-dev" : ""}.seahaven.com`;
// Repo targeting is PER-ENV: dev drives the disposable sandbox repo in the
// dedicated seahaven-open-swe-dev org (isolates dev-agent activity from the real
// Sea Haven org); prod stays on the Sea Haven org pilot repo. fetch-config's owner
// GUARD honors this value (rejecting only blank / the upstream langchain-ai org).
const repo =
env === "dev"
? { owner: "seahaven-open-swe-dev", name: "openswe-dev-sandbox" }
: { owner: "Sea-Haven-Industries", name: "open-swe-pilot" };
const managed: Record<string, string> = {
// Sandbox provider — plan keeps stock langsmith (T9). Stable.
SANDBOX_TYPE: "langsmith",
DEFAULT_REPO_OWNER: repo.owner,
// Repo-owner allowlist (comma list) — the env's org. The app gates triggers on this.
ALLOWED_GITHUB_ORGS: repo.owner,
DEFAULT_REPO_NAME: repo.name,
// Dashboard URLs — derived from the public host.
DASHBOARD_BASE_URL: host,
DASHBOARD_API_BASE_URL: host,
DASHBOARD_ALLOWED_ORIGINS: host,
// Primary builder model. seed_store.sh's `pick` precedence is
// OPENSWE_AGENT_MODEL > SEED_AGENT_MODEL > LLM_MODEL_ID > script default, so this
// SSM value overrides the seed-script default — it MUST be a supported id. Post
// Bedrock/Fireworks migration the only Bedrock-Claude id is the inference profile;
// `anthropic:claude-opus-4-8` was removed from SUPPORTED_MODELS.
LLM_MODEL_ID: "bedrock_converse:us.anthropic.claude-opus-4-8",
};
// Dev e2e smoke: seed the owner's user_mapping so an @openswe comment from the
// triggering GitHub login resolves (an unmapped commenter is silently skipped).
// Dev-only — prod seeds its mappings via its own operator config. login:email.
if (env === "dev") {
managed.SEED_USER_MAPPINGS = "amoussa1229:adam@seahavenind.com";
}
return managed;
}
/**
* Out-of-band SSM config: NOT created by CDK. Listed for documentation and for
* `put-config.sh` to populate before the box boots. Each MUST stay out of IaC
* because its value is operationally-variable or env-specific and unknown at
* synth time — making it CDK-managed would either clobber the operator value on
* the next deploy (e.g. DEFAULT_SANDBOX_SNAPSHOT_ID) or hardcode a secret-ish id.
*/
export const OUT_OF_BAND_SSM: readonly string[] = [
// Sandbox snapshot id — changes on EVERY snapshot rebuild. MUST NOT be
// CDK-managed or a deploy clobbers it. Required for SANDBOX_TYPE=langsmith.
"DEFAULT_SANDBOX_SNAPSHOT_ID",
// GitHub App identifiers — set when the per-env GitHub App is created.
"GITHUB_APP_ID",
"GITHUB_APP_CLIENT_ID",
"GITHUB_APP_INSTALLATION_ID",
"GITHUB_OAUTH_PROVIDER_ID",
// LangSmith deployment coordinates (prod tenant/urls/endpoints).
"LANGSMITH_TENANT_ID_PROD",
"LANGSMITH_URL_PROD",
"LANGSMITH_ENDPOINT",
"LANGSMITH_ENDPOINT_PROD",
"LANGSMITH_HOST_API_URL",
"LANGGRAPH_URL",
"LANGGRAPH_URL_PROD",
"LANGCHAIN_REVISION_ID",
// Slack workspace ids — set after the Slack app is installed.
"SLACK_CLIENT_ID",
"SLACK_TEAM_ID",
"SLACK_BOT_USER_ID",
"SLACK_BOT_USERNAME",
"SLACK_REPO_OWNER",
"SLACK_REPO_NAME",
// Access / observability allowlists — operator-curated.
"CONFIGURED_ADMINS",
"OBSERVABILITY_AUTHORIZED_EMAILS",
"PUBLIC_REPO_ORG_GATE",
"ALLOWED_GITHUB_REPOS",
// Optional integrations + tuning knobs (left to code defaults unless set).
"LLM_FALLBACK_MODEL_ID",
"DATADOG_MCP_TOOLSETS",
"NOTION_MCP_CLIENT_NAME",
"API_STANDARDS_SKILL_HANDLE",
"REPO_SNAPSHOT_BASE_IMAGE",
"REPO_SNAPSHOT_BUILD_TIMEOUT_SECONDS",
"REPO_SNAPSHOT_STALE_BUILD_SECONDS",
] as const;
export interface ConfigStoreProps {
readonly envName: EnvName;
}
/**
* Per-env Secrets Manager + SSM Parameter Store shells the boot hook reads.
* Instantiated from OpenSweStack. Synth-able now (T11); values populated
* out-of-band BEFORE the EC2/T12 deploy. See infra/README.md "Config store".
*/
export class ConfigStore extends Construct {
public readonly secrets: secretsmanager.CfnSecret[] = [];
public readonly params: ssm.StringParameter[] = [];
constructor(scope: Construct, id: string, props: ConfigStoreProps) {
super(scope, id);
const env = props.envName;
// --- 1) Secret shells (value-LESS; populated out-of-band) ----------------
for (const varName of SECRET_VARS) {
const secret = new secretsmanager.CfnSecret(this, `Secret-${varName}`, {
name: `open-swe-${env}/${varName}`,
description: SECRET_DESCRIPTIONS[varName] ?? `open-swe ${varName}`,
// Deliberately NO secretString / generateSecretString: CloudFormation
// creates an empty secret, so the out-of-band value is never clobbered.
});
// RETAIN: a stack teardown must not destroy operator-set secret material.
//
// GOTCHA — RETAIN + fixed name orphans these shells on a FAILED FIRST
// CREATE. If the stack's initial create fails and rolls back, RETAIN keeps
// the shells instead of deleting them; the stack is then gone but the
// secrets survive, still holding the global `open-swe-<env>/<VAR>` names.
// Every later create then fails with `AlreadyExists` (and a normal
// delete-secret keeps the name reserved for the 7–30 day recovery window,
// so it does NOT clear the deadlock). Recovery: before re-creating the
// stack, force-delete the orphans so the names free immediately, e.g.
// aws secretsmanager list-secrets --filters Key=name,Values=open-swe-<env>/ \
// --query 'SecretList[].Name' --output text | tr '\t' '\n' | while read n; do
// aws secretsmanager delete-secret --secret-id "$n" \
// --force-delete-without-recovery; done
// Only force-delete shells that are EMPTY (no value version) — a populated
// secret holds real operator material. This bites on teardown/rebuild, a
// secret logical-id change/refactor, or standing up a new env — NOT on
// routine updates of an already-created stack. (Hit on prod 2026-06-29.)
secret.applyRemovalPolicy(cdk.RemovalPolicy.RETAIN);
this.secrets.push(secret);
}
// --- 2) IaC-managed SSM config (real, derivable values) ------------------
const managed = iacManagedSsm(env);
for (const [varName, value] of Object.entries(managed)) {
this.params.push(
new ssm.StringParameter(this, `Param-${varName}`, {
parameterName: `/open-swe-${env}/${varName}`,
stringValue: value,
description: `IaC-managed open-swe ${varName} (${env}).`,
tier: ssm.ParameterTier.STANDARD,
}),
);
}
}
}

View file

@ -1,173 +0,0 @@
import * as iam from "aws-cdk-lib/aws-iam";
import { Construct } from "constructs";
import {
ACCOUNT,
EnvName,
GITHUB_OIDC_PROVIDER_ARN,
REGION,
bootstrapQualifier,
oidcSubject,
} from "../config";
/**
* Per-ENV GitHub Actions OIDC deploy roles. Created ONCE per env in the
* dedicated `open-swe-iam` stack. Two roles per env, per the locked architecture's
* "dual OIDC roles":
*
* - githubdeploy-open-swe-infra-<env> → CFN/IAM (CDK) deploys of that env's stack
* - githubdeploy-open-swe-app-<env> → app deploys (env-tag-scoped SSM + S3 read)
*
* T5 OSWE-IAC-01/02 fix: roles are split per env and the trust subject is
* env-scoped (dev = dev branch ref; prod = the GitHub `prod` Environment subject,
* so the manual-approval gate is IAM-enforced). A dev-branch token therefore
* cannot SendCommand to the prod box nor assume a prod deploy role.
*
* Reviewed at T4 (GPT-4.1 IAM cross-review) + T5 (/sh-security-review) and
* deployed FIRST (BLOCK#3 "OIDC-role-first" ordering) before any other infra or
* secrets CI step.
*/
export class GithubDeployRoles extends Construct {
public readonly infraRole: iam.Role;
public readonly appRole: iam.Role;
constructor(scope: Construct, id: string, envName: EnvName) {
super(scope, id);
// The provider already exists account-wide — reference, never re-create.
const provider = iam.OpenIdConnectProvider.fromOpenIdConnectProviderArn(
this,
"GithubOidcProvider",
GITHUB_OIDC_PROVIDER_ARN,
);
// T4 BLOCK#2 + T5 IAC-01/02: exact env-scoped subject via StringEquals (no
// StringLike, no `*`). prod = environment:prod (manual-approval gate),
// dev = the dev branch ref.
const trust = new iam.WebIdentityPrincipal(provider.openIdConnectProviderArn, {
StringEquals: {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": oidcSubject(envName),
},
});
// ---- githubdeploy-open-swe-infra-<env> --------------------------------
this.infraRole = new iam.Role(this, "InfraDeployRole", {
roleName: `githubdeploy-open-swe-infra-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for CDK deploys of the open-swe-${envName} infra stack (assumes CDK bootstrap roles).`,
});
// Org-standard CDK deploy pattern (mirrors githubdeploy-seahaven-account-
// baseline / -forgejo / -apm-wo-analysis): the deploy role only needs to
// assume the CDK bootstrap roles. The actual CloudFormation + IAM + resource
// permissions are exercised by the bootstrap `cfn-exec-role`, whose scope is
// owned by the CDKToolkit stack — NOT granted directly here.
//
// B-1 / OSWE-IAC-01 fix: scope the assume to THIS env's bootstrap qualifier.
// Dev uses `oswedev` (its own CDKToolkit-oswedev bootstrap), prod uses the
// default `hnb659fds`. The dev infra role can therefore no longer assume the
// bootstrap roles whose admin cfn-exec-role deploys prod — closing the prior
// cross-env escalation (a dev-branch token could `cdk deploy open-swe-prod`
// via the shared account-wide bootstrap roles, bypassing prod's Environment
// approval gate). The qualifier wildcard still matches only the handful of
// roles `cdk bootstrap` creates for that qualifier.
this.infraRole.addToPolicy(
new iam.PolicyStatement({
sid: "AssumeCdkBootstrapRoles",
actions: ["sts:AssumeRole"],
resources: [`arn:aws:iam::${ACCOUNT}:role/cdk-${bootstrapQualifier(envName)}-*`],
}),
);
// ---- githubdeploy-open-swe-app-<env> ----------------------------------
this.appRole = new iam.Role(this, "AppDeployRole", {
roleName: `githubdeploy-open-swe-app-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for open-swe-${envName} app deploys: env-tag-scoped ssm:SendCommand + read of the ${envName} S3 artifact bucket.`,
});
// T5 OSWE-IAC-01 fix: SendCommand only to instances tagged project=open-swe
// AND env=<this env> (a SINGLE value, not {dev,prod}). The dev app role can
// never command the prod box and vice versa — env isolation in IAM.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandTagScoped",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ec2:${REGION}:${ACCOUNT}:instance/*`],
conditions: {
StringEquals: {
"ssm:resourceTag/project": "open-swe",
"ssm:resourceTag/env": envName,
},
},
}),
);
// SendCommand also has to reference the command document. Scope to this env's
// open-swe deploy document ONLY.
// T4 BLOCK#3 (CLOSED at T19): GPT-4.1 flagged AWS-RunShellScript as an
// arbitrary-shell escalation path. The dedicated `open-swe-${envName}-deploy`
// SSM document (app-service.ts) now runs the fixed, parameter-less command
// `bash /opt/open-swe/bin/deploy.sh`, so AWS-RunShellScript is dropped here:
// this role can run ONLY that one document, and only on its own env's box
// (tag-scoped by the SsmSendCommandTagScoped statement above).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandDocuments",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:document/open-swe-${envName}-deploy`],
}),
);
// Poll command results. These read actions do not support resource-level
// scoping, so `*` is required by the API (T4 FIX: API limitation, documented).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmReadCommandStatus",
actions: [
"ssm:GetCommandInvocation",
"ssm:ListCommands",
"ssm:ListCommandInvocations",
],
resources: ["*"],
}),
);
// Read+WRITE access to THIS env's artifact bucket only (T19): the
// build-artifacts workflow uploads app.tar.gz / spa.tar.gz under releases/*,
// then fires the deploy document so the box pulls them via its instance role.
// Object actions are scoped to releases/* (the only prefix CI writes), and to
// THIS env's bucket — a dev token can never write the prod bucket. No
// bucket-level mutation (no PutBucket*/Delete bucket) — that stays with CDK.
// GetObject + PutObject (S3-to-S3 copy = Get source + Put dest) is all the
// publish/rollback path uses; s3:DeleteObject is deliberately NOT granted so a
// CI token cannot erase an immutable release or the releases/last-good rollback
// fallback (lifecycle expiry handles old-version cleanup, not CI).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ReadWriteArtifactObjects",
actions: ["s3:GetObject", "s3:PutObject"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets/releases/*`],
}),
);
// ListBucket is constrained to the releases/ prefix (F-1/IAC-04): the
// publish/rollback scripts only ever list under releases/, so a leaked CI
// token cannot enumerate anything else in the bucket. GetBucketLocation
// carries no s3:prefix, so it stays a separate, unconditioned statement.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ListArtifactBucket",
actions: ["s3:ListBucket"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
conditions: { StringLike: { "s3:prefix": ["releases/*"] } },
}),
);
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "GetArtifactBucketLocation",
actions: ["s3:GetBucketLocation"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
}),
);
}
}

View file

@ -1,160 +0,0 @@
import * as iam from "aws-cdk-lib/aws-iam";
import { Construct } from "constructs";
import { ACCOUNT, EnvName, REGION, prefix } from "../config";
/**
* Least-privilege EC2 instance role for the open-swe box (one per env).
*
* Grants exactly what the boot/runtime flow needs and NOTHING ELSE — no admin,
* no `*` resources except where the AWS action genuinely has no resource-level
* scoping. Per-env so the dev box can never read prod secrets/config and vice
* versa. Reviewed at T4 (GPT-4.1 IAM cross-review) / T5 (/sh-security-review)
* before it is ever deployed (T6).
*/
export class InstanceRole extends Construct {
public readonly role: iam.Role;
constructor(scope: Construct, id: string, env: EnvName) {
super(scope, id);
const p = prefix(env);
this.role = new iam.Role(this, "Role", {
roleName: `${p}-instance-role`,
assumedBy: new iam.ServicePrincipal("ec2.amazonaws.com"),
description: `EC2 instance role for the ${p} open-swe box (least-privilege).`,
});
// AWS-managed: lets the SSM agent register the instance and RECEIVE the
// app-deploy `ssm:SendCommand` from githubdeploy-open-swe-app. This is the
// standard Session-Manager / RunCommand grant and is the only managed
// policy on the role. DELIBERATE — flag for T4 confirmation.
this.role.addManagedPolicy(
iam.ManagedPolicy.fromAwsManagedPolicyName("AmazonSSMManagedInstanceCore"),
);
// Read the build artifact from the env's S3 asset bucket (deploy = pull).
// Scoped to releases/* — the only prefix CI writes and the box pulls — so a
// compromised box (or stolen IMDS creds) cannot read anything else that might
// ever land in the bucket (least-privilege; mirrors the app role's write scope).
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadArtifactObjects",
actions: ["s3:GetObject"],
resources: [`arn:aws:s3:::${p}-assets/releases/*`],
}),
);
// ListBucket is constrained to the releases/ prefix (F-1/IAC-04) — the box
// only ever lists release artifacts, so a compromised box cannot enumerate
// any other object that might land in the bucket. GetBucketLocation has no
// s3:prefix in its request context, so it stays a separate, unconditioned
// statement (the condition would otherwise AccessDeny it).
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ListArtifactBucket",
actions: ["s3:ListBucket"],
resources: [`arn:aws:s3:::${p}-assets`],
conditions: { StringLike: { "s3:prefix": ["releases/*"] } },
}),
);
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "GetArtifactBucketLocation",
actions: ["s3:GetBucketLocation"],
resources: [`arn:aws:s3:::${p}-assets`],
}),
);
// Read non-sensitive config from SSM Parameter Store under /open-swe-<env>/*.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadSsmConfig",
actions: ["ssm:GetParameter", "ssm:GetParameters", "ssm:GetParametersByPath"],
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:parameter/${p}/*`],
}),
);
// VALUE access — Secrets Manager under open-swe-<env>/*. Secret ARNs carry a
// random 6-char suffix, hence the trailing `*`. This is the statement that
// actually gates which secret VALUES the box can read: prefix-scoped, so the
// dev box can never read prod secret values (and vice versa). GetSecretValue is
// checked per-secret even when the value is returned via the batch call below.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadSecretValues",
actions: ["secretsmanager:GetSecretValue", "secretsmanager:DescribeSecret"],
resources: [`arn:aws:secretsmanager:${REGION}:${ACCOUNT}:secret:${p}/*`],
}),
);
// BatchGetSecretValue MUST be granted on `*` — it is a collection action that
// AWS authorizes against the account, NOT the per-secret ARN, REGARDLESS of
// whether the caller uses `--filters` or `--secret-id-list`. A prefix-scoped
// BatchGetSecretValue AccessDenies the whole call ("no identity-based policy
// allows the secretsmanager:BatchGetSecretValue action") — VERIFIED on the live
// dev box 2026-06-29 (the OSWE-IAC-SECRETS-LIST-01 attempt to prefix-scope it
// crash-looped the box once the prior broad grant's eventual-consistency lapsed).
// This `*` does NOT widen VALUE access: a secret value is only returned when the
// prefix-scoped GetSecretValue above also allows it, so cross-env value isolation
// holds. The win that DID survive: fetch-config uses `--secret-id-list` (explicit
// names, no name filter), so `secretsmanager:ListSecrets` is NOT needed and is
// intentionally omitted — the box cannot enumerate secret names account-wide.
// F-2 (accepted residual): because the grant is `*`, a caller naming a secret
// in ANOTHER env's prefix learns whether that name EXISTS (an existence oracle
// via the per-secret AccessDenied-vs-not signal) even though the VALUE stays
// gated by the prefix-scoped GetSecretValue above. Accepted within Sea Haven's
// single-tenant account 328440206208 — cross-env VALUE isolation is preserved.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "BatchGetSecretValues",
actions: ["secretsmanager:BatchGetSecretValue"],
resources: ["*"],
}),
);
// NOTE (T11): SSM SecureString + Secrets Manager here are assumed to use the
// AWS-managed keys (alias/aws/ssm, alias/aws/secretsmanager) for which the
// service grants Decrypt implicitly — so NO kms:Decrypt is granted. If T11
// moves these to a customer CMK, add a scoped `kms:Decrypt` on that key ARN
// ONLY (not `*`).
// Invoke the Bedrock Claude model. DEFAULT_MODEL_ID is
// `bedrock_converse:us.anthropic.claude-opus-4-8`, and the model runs in the
// LangGraph server PROCESS on this box (not in the sandbox), so the EC2
// instance role is the calling principal. The `us.` cross-region inference
// profile fans out to us-east-1 / us-east-2 / us-west-2, and Bedrock authorizes
// InvokeModel against BOTH the inference-profile ARN AND the underlying
// foundation-model ARN in each routed region — all four resources are required
// or the call AccessDenies. Scoped to opus-4-8 ONLY (least-privilege): adding a
// new Bedrock model to SUPPORTED_MODELS means extending this resource list.
// IAM change — flag for T4 (GPT-4.1 IAM cross-review) / T5 (/sh-security-review).
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "InvokeBedrockClaude",
actions: ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
resources: [
`arn:aws:bedrock:${REGION}:${ACCOUNT}:inference-profile/us.anthropic.claude-opus-4-8`,
"arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-opus-4-8",
"arn:aws:bedrock:us-east-2::foundation-model/anthropic.claude-opus-4-8",
"arn:aws:bedrock:us-west-2::foundation-model/anthropic.claude-opus-4-8",
],
}),
);
// Ship application logs to CloudWatch Logs under /open-swe/<env>/*.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "PutAppLogs",
actions: [
"logs:CreateLogGroup",
"logs:CreateLogStream",
"logs:PutLogEvents",
"logs:DescribeLogStreams",
],
resources: [
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*`,
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*:*`,
],
}),
);
}
}

View file

@ -1,38 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { GithubDeployRoles } from "./constructs/github-deploy-roles";
/**
* Account-level IAM stack: the per-ENV GitHub OIDC deploy roles
* (githubdeploy-open-swe-{infra,app}-{dev,prod} — four roles).
*
* T5 OSWE-IAC-01/02 fix: roles are split per env with env-scoped OIDC trust, so
* a dev-branch token cannot reach prod (prod roles require the GitHub
* `prod` Environment manual-approval gate). They live in this dedicated stack
* rather than the env stacks because IAM roles are global and this stack ships
* FIRST (TODO.md BLOCK#3): the infra OIDC roles + the repo deploy-role-ARN
* secrets must exist before any infra/secrets CI step. Synth-only until the
* Phase-1 security gate (T4 + T5) clears (T6).
*/
export class OpenSweIamStack extends cdk.Stack {
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
super(scope, id, props);
const dev = new GithubDeployRoles(this, "DeployRolesDev", "dev");
const prod = new GithubDeployRoles(this, "DeployRolesProd", "prod");
cdk.Tags.of(this).add("project", "open-swe");
cdk.Tags.of(this).add("ManagedBy", "cdk");
const out = (id: string, role: { roleName?: string }, env: string, kind: string) =>
new cdk.CfnOutput(this, id, {
value: `arn:aws:iam::${this.account}:role/${role.roleName}`,
description: `OIDC role ARN for ${env} ${kind} deploys — set as the ${env} deploy-role secret.`,
});
out("InfraDeployRoleDevArn", dev.infraRole, "dev", "infra (CDK)");
out("AppDeployRoleDevArn", dev.appRole, "dev", "app (tag-scoped SSM + S3)");
out("InfraDeployRoleProdArn", prod.infraRole, "prod", "infra (CDK)");
out("AppDeployRoleProdArn", prod.appRole, "prod", "app (tag-scoped SSM + S3)");
}
}

View file

@ -1,90 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { EnvName, prefix } from "./config";
import { AppService } from "./constructs/app-service";
import { AssetsBucket } from "./constructs/assets-bucket";
import { ConfigStore } from "./constructs/config-store";
import { InstanceRole } from "./constructs/instance-role";
import { BAKED_OPEN_SWE_AMI_ID } from "./constructs/ami-cache";
export interface OpenSweStackProps extends cdk.StackProps {
/** open-swe environment — drives the `open-swe-<env>-*` resource naming. */
readonly envName: EnvName;
}
/**
* Per-env open-swe stack (`open-swe-dev` / `open-swe-prod`). Resource names are
* prefixed `open-swe-<env>-*`.
*
* Composes: the per-env least-privilege instance role (T6), the Secrets/SSM
* config store (T11), and the compute + ingress wiring (T12, AppService — EC2
* box, instance SG, target group, imported-listener rules, Route53 aliases,
* 30-day log groups). The shared VPC and ALB are imported, never owned. Synth is
* offline (AMI is the cdk.context.json-pinned placeholder until T12-deploy).
*/
export class OpenSweStack extends cdk.Stack {
public readonly instanceRole: InstanceRole;
public readonly configStore: ConfigStore;
public readonly assetsBucket: AssetsBucket;
public readonly appService: AppService;
constructor(scope: Construct, id: string, props: OpenSweStackProps) {
super(scope, id, props);
const envName = props.envName;
const p = prefix(envName);
cdk.Tags.of(this).add("project", "open-swe");
cdk.Tags.of(this).add("env", envName);
cdk.Tags.of(this).add("ManagedBy", "cdk");
// Per-env least-privilege EC2 instance role (open-swe-<env>-instance-role).
this.instanceRole = new InstanceRole(this, "Instance", envName);
// Secrets Manager + SSM Parameter Store shells the boot hook reads
// (deploy/seahaven/fetch-config.sh). Secret shells are value-less and
// populated out-of-band; IaC-managed SSM params carry real derivable values.
// The instance role already grants read on open-swe-<env>/* + /open-swe-<env>/*.
this.configStore = new ConfigStore(this, "Config", { envName });
// T7: the S3 artifact bucket (open-swe-<env>-assets) CI uploads releases to
// and the box pulls app.tar.gz / spa.tar.gz from. The instance role already
// grants read on it by name; the app deploy role grants write.
this.assetsBucket = new AssetsBucket(this, "Assets", envName);
// Surface the baked open-swe base AMI id the box runs on (pinned by id in
// ami-cache.ts; refreshed by a deliberate packer rebuild → replacement).
new cdk.CfnOutput(this, "BakedAmiId", {
value: BAKED_OPEN_SWE_AMI_ID,
description: "Baked open-swe-base-arm64 AMI id consumed by the EC2 instance.",
});
// T12: compute + ingress. Imports the shared seahaven-vpc + ALB and adds the
// env's EC2 box, instance SG, target group, listener rules, DNS, log groups.
this.appService = new AppService(this, "App", {
envName,
instanceRole: this.instanceRole.role,
});
new cdk.CfnOutput(this, "InstanceRoleArn", {
value: this.instanceRole.role.roleArn,
description: `${p} EC2 instance role ARN.`,
});
new cdk.CfnOutput(this, "InstanceId", {
value: this.appService.instance.instanceId,
description: `${p} EC2 instance id.`,
});
new cdk.CfnOutput(this, "TargetGroupArn", {
value: this.appService.targetGroup.targetGroupArn,
description: `${p} ALB target group ARN (→ instance:80 nginx).`,
});
new cdk.CfnOutput(this, "AssetsBucketName", {
value: this.assetsBucket.bucket.bucketName,
description: `${p} S3 artifact bucket (CI uploads releases; box pulls).`,
});
new cdk.CfnOutput(this, "DeployDocumentName", {
value: this.appService.deployDocumentName,
description: `${p} SSM document that rolls the box to the latest release.`,
});
}
}

4487
infra/package-lock.json generated

File diff suppressed because it is too large Load diff

View file

@ -1,31 +0,0 @@
{
"name": "open-swe-infra",
"version": "1.0.0",
"description": "Open SWE AWS infrastructure (CDK TypeScript) — open-swe-dev / open-swe-prod stacks + shared OIDC deploy roles.",
"private": true,
"bin": {
"open-swe-infra": "bin/app.js"
},
"scripts": {
"build": "tsc",
"cdk": "cdk",
"synth": "cdk synth",
"diff": "cdk diff",
"test": "jest"
},
"devDependencies": {
"@types/jest": "^29.5.14",
"@types/node": "^24.0.0",
"@types/source-map-support": "^0.5.10",
"aws-cdk": "^2.1029.0",
"jest": "^29.7.0",
"source-map-support": "^0.5.21",
"ts-jest": "^29.2.5",
"ts-node": "^10.9.2",
"typescript": "~5.6.3"
},
"dependencies": {
"aws-cdk-lib": "2.260.0",
"constructs": "^10.0.0"
}
}

View file

@ -1,87 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Template } from "aws-cdk-lib/assertions";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
/**
* Several AWS APIs reject non-ASCII in fields that `cdk synth` happily emits and
* `tsc` happily compiles — so a stray em-dash/arrow only blows up at DEPLOY time
* (e.g. EC2 SecurityGroup GroupDescription: "Character sets beyond ASCII are not
* supported"). This has bitten us twice (the AMI Description, then the instance-SG
* description). This test fails the build at synth time instead.
*
* Scope: the EC2 fields with a documented ASCII/restricted-charset constraint —
* SecurityGroup GroupDescription and ingress/egress rule descriptions. (CloudFormation
* Output descriptions + Route53 comments accept UTF-8, so they are not asserted.)
*
* EC2 rule descriptions are stricter than ASCII: the allowed set is
* `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*` — note it EXCLUDES `<` and `>`, which is why a
* naive em-dash -> "->" replacement still fails at deploy. We assert that exact set.
*/
// Characters NOT in the EC2 description allowed set.
const DISALLOWED = /[^a-zA-Z0-9. _:/()#,@[\]+=&;{}!$*-]/;
function synthDev(): Record<string, { Type: string; Properties?: Record<string, unknown> }> {
const app = new cdk.App();
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
return Template.fromStack(dev).toJSON().Resources;
}
describe("ASCII-only EC2 description fields", () => {
const resources = synthDev();
it("SecurityGroup GroupDescription is ASCII", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::SecurityGroup") continue;
const desc = (r.Properties?.GroupDescription as string) ?? "";
expect(DISALLOWED.test(desc) ? `${id}: ${desc}` : "ascii").toBe("ascii");
}
});
it("SecurityGroup ingress/egress rule descriptions are ASCII", () => {
for (const [id, r] of Object.entries(resources)) {
const props = r.Properties ?? {};
const groups: Array<{ Description?: string }> = [];
if (Array.isArray(props.SecurityGroupIngress)) groups.push(...props.SecurityGroupIngress);
if (Array.isArray(props.SecurityGroupEgress)) groups.push(...props.SecurityGroupEgress);
// Standalone AWS::EC2::SecurityGroupEgress / ...Ingress resources.
if (r.Type === "AWS::EC2::SecurityGroupEgress" || r.Type === "AWS::EC2::SecurityGroupIngress") {
groups.push(props as { Description?: string });
}
for (const rule of groups) {
const desc = rule.Description ?? "";
expect(DISALLOWED.test(desc) ? `${id}: ${desc}` : "ascii").toBe("ascii");
}
}
});
// EC2 caps base64-encoded user-data at 25600 bytes; CDK + tsc don't check it, so
// an oversized boot script (e.g. an embedded deploy.sh) only fails at deploy.
it("EC2 user-data fits the 25600-byte encoded limit", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::Instance") continue;
const ud = (r.Properties?.UserData as { "Fn::Base64"?: string }) ?? {};
const script = typeof ud["Fn::Base64"] === "string" ? ud["Fn::Base64"] : "";
const encoded = Buffer.from(script, "utf8").toString("base64").length;
expect(`${id}: ${encoded} bytes`).toBe(encoded < 25600 ? `${id}: ${encoded} bytes` : "OVER 25600");
}
});
// CDK substitutes %%...%% tokens in user-data at synth. Any %%TOKEN%% left in the
// rendered script means a token wasn't wired in app-service.ts (the @@...@@ tokens
// are intentional — user-data seds those into the baked templates at boot).
it("user-data has no unresolved %%CDK%% tokens", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::Instance") continue;
const ud = (r.Properties?.UserData as { "Fn::Base64"?: string }) ?? {};
const script = typeof ud["Fn::Base64"] === "string" ? ud["Fn::Base64"] : "";
const leftover = script.match(/%%[A-Z0-9_]+%%/g) ?? [];
expect(`${id}: ${leftover.join(",")}`).toBe(`${id}: `);
}
});
});

View file

@ -1,55 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Match, Template } from "aws-cdk-lib/assertions";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { bootstrapQualifier } from "../lib/config";
const ENV = { account: "328440206208", region: "us-east-1" };
// B-1 / OSWE-IAC-01: each env's infra deploy role may assume ONLY its own
// bootstrap qualifier's roles. Dev runs on `oswedev`, so a dev-branch token can
// no longer assume the default `hnb659fds` bootstrap roles whose admin
// cfn-exec-role deploys prod. Prod stays on the default qualifier.
describe("Per-env CDK bootstrap qualifier isolation (B-1/OSWE-IAC-01)", () => {
it("maps dev -> oswedev and prod -> hnb659fds", () => {
expect(bootstrapQualifier("dev")).toBe("oswedev");
expect(bootstrapQualifier("prod")).toBe("hnb659fds");
});
it("dev infra deploy role assumes only cdk-oswedev-* bootstrap roles", () => {
const app = new cdk.App();
const stack = new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env: ENV,
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "AssumeCdkBootstrapRoles",
Action: "sts:AssumeRole",
Resource: "arn:aws:iam::328440206208:role/cdk-oswedev-*",
}),
]),
}),
});
});
it("prod infra deploy role stays on the default cdk-hnb659fds-* bootstrap roles", () => {
const app = new cdk.App();
const stack = new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env: ENV,
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "AssumeCdkBootstrapRoles",
Action: "sts:AssumeRole",
Resource: "arn:aws:iam::328440206208:role/cdk-hnb659fds-*",
}),
]),
}),
});
});
});

View file

@ -1,109 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Annotations, Match, Template } from "aws-cdk-lib/assertions";
import * as iam from "aws-cdk-lib/aws-iam";
import { KebabNamingAspect, isKebabCase } from "../lib/aspects/kebab-naming-aspect";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
describe("isKebabCase", () => {
it.each([
"open-swe-dev",
"open-swe-prod-instance-role",
"githubdeploy-open-swe-infra",
"open-swe-dev/slack-signing", // Secrets Manager path
"/open-swe-dev/feature-flag", // SSM param path
"/open-swe/dev/agent", // log group path
"abc123",
])("accepts conforming name %s", (name) => {
expect(isKebabCase(name)).toBe(true);
});
it.each([
"OpenSweDev",
"open_swe_dev",
"openSweDev",
"Open-Swe-Dev",
"open-swe-dev/SlackSigning",
])("rejects non-conforming name %s", (name) => {
expect(isKebabCase(name)).toBe(false);
});
});
describe("KebabNamingAspect", () => {
it("passes the real app stacks (no errors)", () => {
const app = new cdk.App();
cdk.Aspects.of(app).add(new KebabNamingAspect());
new OpenSweIamStack(app, "OpenSweIamStack", { stackName: "open-swe-iam", env: ENV });
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
const prod = new OpenSweStack(app, "OpenSweProdStack", {
stackName: "open-swe-prod",
env: ENV,
envName: "prod",
});
for (const s of [dev, prod]) {
Annotations.fromStack(s).hasNoError("*", Match.anyValue());
}
});
it("exempts Secrets Manager + SSM names that carry the literal env-var segment", () => {
// The config store names a secret open-swe-dev/ANTHROPIC_API_KEY and a param
// /open-swe-dev/SANDBOX_TYPE — the UPPER_SNAKE last segment is a REQUIRED
// deviation from kebab (fetch-config naming contract). Must NOT be flagged.
const app = new cdk.App();
cdk.Aspects.of(app).add(new KebabNamingAspect());
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
const tpl = Template.fromStack(dev);
// Sanity: the shells actually render with the literal env-var names.
tpl.hasResourceProperties("AWS::SecretsManager::Secret", {
Name: "open-swe-dev/ANTHROPIC_API_KEY",
});
tpl.hasResourceProperties("AWS::SSM::Parameter", {
Name: "/open-swe-dev/SANDBOX_TYPE",
Value: "langsmith",
});
Annotations.fromStack(dev).hasNoError("*", Match.anyValue());
});
it("flags a deliberately non-kebab-case resource name", () => {
const app = new cdk.App();
const stack = new cdk.Stack(app, "ConformingStackId", { stackName: "open-swe-test", env: ENV });
cdk.Aspects.of(stack).add(new KebabNamingAspect());
// Deliberately bad physical name — must be flagged.
new iam.Role(stack, "BadlyNamedRole", {
roleName: "OpenSweBadRole",
assumedBy: new iam.ServicePrincipal("ec2.amazonaws.com"),
});
Annotations.fromStack(stack).hasError(
"*",
Match.stringLikeRegexp("not kebab-case"),
);
});
it("flags a deliberately non-kebab-case stack name", () => {
const app = new cdk.App();
// PascalCase stackName — the convention CDK defaults to and that we forbid.
const stack = new cdk.Stack(app, "BadStack", { stackName: "OpenSweBadStack", env: ENV });
cdk.Aspects.of(stack).add(new KebabNamingAspect());
Annotations.fromStack(stack).hasError(
"*",
Match.stringLikeRegexp("Stack name .* is not kebab-case"),
);
});
});

View file

@ -1,71 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Match, Template } from "aws-cdk-lib/assertions";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
// F-1 / IAC-04: s3:ListBucket must be constrained to the releases/ prefix so a
// compromised box / leaked CI token cannot enumerate the rest of the bucket.
const RELEASES_PREFIX_CONDITION = { StringLike: { "s3:prefix": ["releases/*"] } };
describe("S3 ListBucket prefix scoping (F-1/IAC-04)", () => {
it("instance role ListBucket is constrained to releases/*", () => {
const app = new cdk.App();
const stack = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "ListArtifactBucket",
Action: "s3:ListBucket",
Condition: RELEASES_PREFIX_CONDITION,
}),
]),
}),
});
});
it("github deploy app role ListBucket is constrained to releases/*", () => {
const app = new cdk.App();
const stack = new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env: ENV,
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "ListArtifactBucket",
Action: "s3:ListBucket",
Condition: RELEASES_PREFIX_CONDITION,
}),
]),
}),
});
});
it("GetBucketLocation stays a separate, unconditioned statement", () => {
const app = new cdk.App();
const stack = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "GetArtifactBucketLocation",
Action: "s3:GetBucketLocation",
Condition: Match.absent(),
}),
]),
}),
});
});
});

View file

@ -1,24 +0,0 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "commonjs",
"lib": ["ES2022"],
"types": ["node", "jest"],
"declaration": true,
"strict": true,
"noImplicitAny": true,
"strictNullChecks": true,
"noImplicitReturns": true,
"noFallthroughCasesInSwitch": true,
"inlineSourceMap": true,
"inlineSources": true,
"strictPropertyInitialization": false,
"outDir": "./cdk.out",
"rootDir": ".",
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"resolveJsonModule": true,
"esModuleInterop": true
},
"exclude": ["node_modules", "cdk.out"]
}