Compare commits

...

3 commits

Author SHA1 Message Date
Adam Moussa
a30ce2ab40
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint

Prepare the dashboard backend for the managed LangGraph Cloud + Vercel
runtime, where the API is HTTPS and cross-site from the UI.

- OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to
  https:// in _api_base_url() so GitHub stops rejecting login with
  "redirect_uri not associated with this application". _cookie_security()
  now treats a schemeless (managed) value as Secure; SameSite=None too,
  consistent with the coerced scheme.
- OAuth state cookie (#3): document that osw_oauth_state is host-only by
  design (a Domain cookie is unsafe across *.vercel.app, a public suffix),
  so login must always start on the stable alias to avoid "oauth state
  mismatch". Operational contract; no behavioral change.
- Admin user mappings (#4): add POST /admin/user-mappings so an admin can
  set the github_login -> work_email link from the dashboard instead of a
  raw Store write. New "admin" MappingSource provenance value.

* fix(webapp): refresh user-mapping cache on GitHub webhook paths

On managed LangGraph Cloud the backend runs multiple replicas, so the
per-process GitHub<->work-email mapping cache can be stale on the replica
handling a webhook (a mapping created on another replica is invisible
until refresh). process_github_pr_comment and process_github_issue now
refresh the cache from the durable Store before resolving the author's
email, matching the existing Slack mention path (process_slack_mention).

* perf(webapp): defer deepagents import to speed custom-app cold start

The custom FastAPI app (agent.webapp:app, the langgraph.json http.app)
pulled deepagents -> langchain_anthropic -> anthropic into its import
graph via dashboard.routes, only to build skill/chat seed files. Defer
those create_file_data imports into the functions that use them. Removes
deepagents/langchain_anthropic/anthropic from app import entirely and
roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger
cold-start saving since native anthropic init is skipped). Behavior
identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.)

* feat(ui): set work_email user mappings from the admin dashboard

Add an "Add / update" form to the admin User mappings section and the
adminUpsertUserMapping API client method, wiring the new
POST /admin/user-mappings endpoint. Admins can now create or update a
github_login -> work_email mapping directly instead of waiting for the
user to self-connect Slack.

* docs: document managed LangGraph Cloud + Vercel deployment

- INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL,
  DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty
  VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login
  and vercel.json stable-deployment-URL requirements, multi-replica cache
  note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting.
  Refresh the langgraph.json snippet to all six graphs.
- README: reframe deployment around the managed migration; link the plan.
- deploy/MIGRATION.md: import the self-hosted -> managed migration plan.
2026-06-29 19:58:38 -04:00
Adam Moussa
7f60324f0c
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack

Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).

- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
  scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
  Vercel view; drop dead links to infra/ and deploy/seahaven

Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.

* chore: clean up dangling references left by the AWS decommission

Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:

- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
  rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
  a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
  infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)

The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
2026-06-29 19:54:38 -04:00
Adam Moussa
430a1cdff9
ci: re-home prod promotion into a gated promote-to-main workflow (PR1: managed-LGC migration) (#63)
* ci: re-home prod promotion into a gated promote-to-main workflow

Migrate the prod-deploy gate to managed LangGraph Cloud (git-connected to
`main`) + Vercel. Under managed, a push to `main` auto-deploys prod, so the
dev -> main fast-forward IS the prod deploy trigger -- the bespoke AWS CD step
is obsolete and already gone from this workflow.

Re-home `promote-dev-to-prod.yml` -> `promote-to-main.yml`:
- gate the promote job on the `prod` GitHub Environment (required reviewer
  amoussa1229), restoring the manual prod-approval that the retired AWS CD job
  used to carry;
- drop the nightly auto-promote cron -- a scheduled auto-promotion conflicts
  with a manual approval gate now that the push deploys prod; promotion is
  workflow_dispatch only;
- keep the dev-HEAD-fully-green precondition and the seahaven-promotion App
  fast-forward push (sole non-admin bypass actor on `main` ruleset 18238334).

Update the companion check-dev-green.sh filename reference.

* ci: refuse promote-to-main dispatch from any ref other than dev

Defense-in-depth atop the already-pinned `ref: dev` checkout: workflow_dispatch
runs the workflow definition from the launched ref, so reject a non-dev dispatch
before the App token is minted. Surfaced by the GPT-4.1 cross-review of #63.
2026-06-29 19:49:21 -04:00
63 changed files with 580 additions and 9300 deletions

View file

@ -10,11 +10,10 @@ updates:
minor-and-patch:
update-types: ["minor", "patch"]
# JavaScript/TypeScript — CDK (/infra), Playwright (/tests/e2e), dashboard (/ui), root tooling
# JavaScript/TypeScript — Playwright (/tests/e2e), dashboard (/ui), root tooling
- package-ecosystem: "npm"
directories:
- "/"
- "/infra"
- "/tests/e2e"
- "/ui"
schedule:

View file

@ -4,7 +4,7 @@
# Reads check-runs on stdin — one
# name<US>status<US>conclusion<US>details_url
# per line, fields separated by ASCII Unit Separator (0x1F) — so it is unit-testable
# WITHOUT GitHub. promote-dev-to-prod.yml pipes the live `gh api .../check-runs`
# WITHOUT GitHub. promote-to-main.yml pipes the live `gh api .../check-runs`
# output in. 0x1F (not TAB) is used deliberately: TAB is IFS-whitespace, so an empty
# conclusion (every in_progress check has a null conclusion) would collapse and shift
# the columns — which would make the promote run fail to exclude itself. 0x1F is

View file

@ -1,52 +0,0 @@
#!/usr/bin/env bash
# Package the two release artifacts (run from the repo root by build-artifacts.yml):
#
# spa.tar.gz = CONTENTS of the built SPA dir (ui/.output/public/*), so it extracts
# straight into the nginx web root with _shell.html at the root.
# app.tar.gz = the Python source the box runs `uv sync` against. `git archive`
# gives a clean tree (no node_modules, no .venv, no local cruft);
# ui/ is intentionally excluded (it ships as spa.tar.gz).
set -euo pipefail
SPA_DIR="ui/.output/public"
[ -f "${SPA_DIR}/_shell.html" ] || {
echo "ERROR: SPA build output missing ${SPA_DIR}/_shell.html (did 'bun run build' run?)" >&2
exit 1
}
echo "==> spa.tar.gz from ${SPA_DIR}"
tar -C "${SPA_DIR}" -czf spa.tar.gz .
echo "==> app.tar.gz from source (git archive HEAD)"
git archive --format=tar.gz -o app.tar.gz HEAD \
agent deploy langgraph.json pyproject.toml uv.lock README.md
# List the archive ONCE into a variable. (`tar -tzf ... | grep -q ...` is unsafe
# under `set -o pipefail`: grep -q exits on first match, SIGPIPEs tar -> "write
# error" -> the pipeline reports non-zero even though grep succeeded, a false
# failure. Listing once avoids the pipe entirely.)
APP_LIST="$(tar -tzf app.tar.gz)"
# Sanity: the box's `uv sync --frozen` needs pyproject.toml + uv.lock at the root,
# and the package itself (agent/). Fail loudly here rather than on the box.
for required in pyproject.toml uv.lock agent/server.py langgraph.json; do
printf '%s\n' "${APP_LIST}" | grep -qx "${required}" || {
echo "ERROR: app.tar.gz is missing ${required}" >&2
exit 1
}
done
# Fail-closed secret guard: the source is git-archived wholesale, so reject the
# release if a secret-shaped FILE slipped into the tracked tree (defense in depth
# on top of .gitignore — the artifact lands on the box + in S3). Scoped to data
# extensions so credential-handling *source* (e.g. team_credentials.py) is not a
# false positive.
SECRET_RE='(^|/)(\.env(\..+)?|id_rsa|.*\.(pem|key|p12|pfx)|.*(secret|credential|password|token)s?\.(json|ya?ml|txt|env|ini|cfg))$'
if printf '%s\n' "${APP_LIST}" | grep -qiE "${SECRET_RE}"; then
echo "ERROR: app.tar.gz contains a secret-shaped file — refusing to publish:" >&2
printf '%s\n' "${APP_LIST}" | grep -iE "${SECRET_RE}" >&2
exit 1
fi
echo "==> artifacts:"
ls -la spa.tar.gz app.tar.gz

View file

@ -1,64 +0,0 @@
#!/usr/bin/env bash
# Publish the packaged artifacts to the env's S3 bucket and roll the box to them.
# Run by build-artifacts.yml AFTER aws creds are configured (env: ENV, BUCKET,
# DEPLOY_DOC). Each release is stored immutably under releases/<sha>/.
#
# releases/latest/ (what the box's deploy.sh pulls) is advanced TRANSACTIONALLY:
# it is pointed at the new release, the box is rolled, and ONLY on a successful
# roll is it kept — a failed roll reverts releases/latest/ to the prior release so
# a later box boot / replacement never self-deploys a release that failed to come
# up. releases/last-good/ (rollback fallback) is advanced only after success and
# means "last release whose deploy.sh brought the service up active" (deploy.sh
# gates on `systemctl is-active`), not merely "last uploaded".
#
# The fire/wait/gate against the box lives in roll-box.sh (shared with rollback.sh);
# the deploy is fired by TAG (project=open-swe,env=<env>), exactly what the app
# deploy role's tag-scoped ssm:SendCommand allows.
set -euo pipefail
: "${ENV:?}" "${BUCKET:?}" "${DEPLOY_DOC:?}"
SHA="${GITHUB_SHA:?}"
[ -f app.tar.gz ] && [ -f spa.tar.gz ] || { echo "ERROR: artifacts not built" >&2; exit 1; }
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Re-point releases/latest/ at the release stored under releases/<sha>/, recording
# the sha so the pointer is self-describing (used to revert on failure).
point_latest() {
local s="$1" f
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "s3://${BUCKET}/releases/${s}/${f}" "s3://${BUCKET}/releases/latest/${f}"
done
printf '%s\n' "${s}" | aws s3 cp - "s3://${BUCKET}/releases/latest/sha.txt"
}
# Which release does latest point at right now? (empty on the very first deploy.)
PREV_SHA="$(aws s3 cp "s3://${BUCKET}/releases/latest/sha.txt" - 2>/dev/null | tr -d '[:space:]' || true)"
echo "==> upload release ${SHA} to s3://${BUCKET}/releases/${SHA}/"
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "${f}" "s3://${BUCKET}/releases/${SHA}/${f}"
done
echo "==> point releases/latest/ -> ${SHA} (was ${PREV_SHA:-<none>})"
point_latest "${SHA}"
# Roll the box to releases/latest/. On a non-Success aggregate, roll-box.sh exits
# non-zero; revert latest to the prior release so no later boot pulls the bad one.
if ! ROLL_COMMENT="release ${SHA}" bash "${HERE}/roll-box.sh"; then
if [ -n "${PREV_SHA}" ]; then
echo "!! deploy failed — reverting releases/latest/ -> ${PREV_SHA}" >&2
point_latest "${PREV_SHA}"
else
echo "!! deploy failed on the FIRST release — leaving releases/latest/ = ${SHA} (no prior release to revert to)" >&2
fi
exit 1
fi
# Roll succeeded (service came up active) -> this release is now the known-good one.
echo "==> mark releases/last-good/ = ${SHA} (rollback fallback target)"
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "s3://${BUCKET}/releases/${SHA}/${f}" "s3://${BUCKET}/releases/last-good/${f}"
done
printf '%s\n' "${SHA}" | aws s3 cp - "s3://${BUCKET}/releases/last-good/sha.txt"
echo "==> ${ENV} rolled to release ${SHA} (now last-good)"

View file

@ -1,59 +0,0 @@
#!/usr/bin/env bash
# Fire the env's SSM deploy document (tag-targeted) and wait for it to finish,
# gating on the AGGREGATE command status. Shared by publish-and-deploy.sh (forward
# roll) and rollback.sh (backward roll) so the fire/wait/gate logic lives in ONE
# place. Requires ENV + DEPLOY_DOC in the environment and aws creds already set.
#
# Tag-targeting (project=open-swe,env=<env>) is exactly what the app deploy role's
# tag-scoped ssm:SendCommand allows — no ec2:DescribeInstances, no instance id.
set -euo pipefail
: "${ENV:?}" "${DEPLOY_DOC:?}"
COMMENT="${ROLL_COMMENT:-roll ${ENV}}"
echo "==> fire ${DEPLOY_DOC} via SSM (tag-targeted: project=open-swe, env=${ENV})"
CMD_ID="$(aws ssm send-command \
--document-name "${DEPLOY_DOC}" \
--targets "Key=tag:project,Values=open-swe" "Key=tag:env,Values=${ENV}" \
--comment "${COMMENT}" \
--query 'Command.CommandId' --output text)"
echo "command: ${CMD_ID}"
echo "==> wait for the deploy to finish"
IID=""
for _ in $(seq 1 60); do
sleep 10
IID="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].InstanceId' --output text 2>/dev/null || echo None)"
[ -z "${IID}" ] || [ "${IID}" = "None" ] && continue
STATUS="$(aws ssm list-command-invocations --command-id "${CMD_ID}" \
--query 'CommandInvocations[0].Status' --output text 2>/dev/null || echo Pending)"
case "${STATUS}" in
Success | Failed | Cancelled | TimedOut) break ;;
esac
done
if [ -z "${IID}" ] || [ "${IID}" = "None" ]; then
echo "ERROR: no box picked up the deploy command (is a running open-swe ${ENV} box registered with SSM?)" >&2
exit 1
fi
echo "==> deploy.sh output from ${IID}:"
echo "----- stdout -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardOutputContent' --output text || true
echo "----- stderr -----"
aws ssm get-command-invocation --command-id "${CMD_ID}" --instance-id "${IID}" \
--query 'StandardErrorContent' --output text || true
# Gate on the AGGREGATE command status (Success only if EVERY targeted invocation
# succeeded), not CommandInvocations[0] — during a userDataCausesReplacement window
# two instances can briefly share the project/env tags, and a partial failure on the
# other instance must not be reported as success.
TARGETS="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].TargetCount' --output text 2>/dev/null || echo 1)"
[ "${TARGETS}" = "1" ] || echo "WARNING: deploy fanned out to ${TARGETS} instances (expected 1)"
AGG="$(aws ssm list-commands --command-id "${CMD_ID}" \
--query 'Commands[0].Status' --output text 2>/dev/null || echo Failed)"
echo "==> aggregate deploy status: ${AGG} (across ${TARGETS} target(s))"
[ "${AGG}" = "Success" ] || { echo "ERROR: deploy did not succeed (${AGG})" >&2; exit 1; }

View file

@ -1,66 +0,0 @@
#!/usr/bin/env bash
# Roll an env BACK to a prior release: re-point releases/latest/ at a chosen release
# and re-fire the deploy. Run by rollback.yml after aws creds are configured.
#
# ENV dev | prod (required)
# BUCKET open-swe-<env>-assets (required)
# DEPLOY_DOC open-swe-<env>-deploy (required)
# TARGET_SHA release sha to restore; blank => releases/last-good/ (optional)
#
# releases/latest/ is moved TRANSACTIONALLY (same as the forward deploy): pointed at
# the target, the box rolled, and on a failed roll latest is reverted to whatever it
# was before the rollback attempt — so a failed rollback never leaves latest at a
# release the box could not bring up. releases/last-good/ is left untouched; advance
# it by running a forward deploy.
#
# Reuses the app deploy role's existing releases/* write + tag-scoped ssm:SendCommand
# — no new IAM. The fire/wait/gate is the SAME roll-box.sh the forward deploy uses.
set -euo pipefail
: "${ENV:?}" "${BUCKET:?}" "${DEPLOY_DOC:?}"
HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# Copy a release prefix into releases/latest/ AND record the resolved sha, so
# releases/latest/sha.txt is always a real commit a later revert can resolve.
point_latest() { # $1 = source prefix (releases/<sha>|releases/last-good); $2 = sha to record
local src="$1" sha="$2" f
for f in app.tar.gz spa.tar.gz; do
aws s3 cp "s3://${BUCKET}/${src}/${f}" "s3://${BUCKET}/releases/latest/${f}"
done
printf '%s\n' "${sha}" | aws s3 cp - "s3://${BUCKET}/releases/latest/sha.txt"
}
if [ -n "${TARGET_SHA:-}" ]; then
SRC="releases/${TARGET_SHA}"
LABEL="${TARGET_SHA}"
RESOLVED_SHA="${TARGET_SHA}"
else
SRC="releases/last-good"
LABEL="last-good"
# resolve last-good's real sha so latest/sha.txt records a commit, not a label.
RESOLVED_SHA="$(aws s3 cp "s3://${BUCKET}/releases/last-good/sha.txt" - 2>/dev/null | tr -d '[:space:]' || true)"
RESOLVED_SHA="${RESOLVED_SHA:-last-good}"
fi
echo "==> rollback ${ENV} to ${LABEL} (s3://${BUCKET}/${SRC}/, sha=${RESOLVED_SHA})"
# Refuse to roll back to a release that is not fully present.
for f in app.tar.gz spa.tar.gz; do
aws s3 ls "s3://${BUCKET}/${SRC}/${f}" >/dev/null 2>&1 \
|| { echo "ERROR: ${SRC}/${f} not found in s3://${BUCKET} — cannot roll back to ${LABEL}" >&2; exit 1; }
done
# Capture what latest points at now, so a failed rollback can be reverted.
PREV_SHA="$(aws s3 cp "s3://${BUCKET}/releases/latest/sha.txt" - 2>/dev/null | tr -d '[:space:]' || true)"
echo "==> point releases/latest/ -> ${SRC} (was ${PREV_SHA:-<none>})"
point_latest "${SRC}" "${RESOLVED_SHA}"
if ! ROLL_COMMENT="rollback ${ENV} to ${LABEL}" bash "${HERE}/roll-box.sh"; then
if [ -n "${PREV_SHA}" ]; then
echo "!! rollback deploy failed — reverting releases/latest/ -> ${PREV_SHA}" >&2
point_latest "releases/${PREV_SHA}" "${PREV_SHA}"
fi
exit 1
fi
echo "==> ${ENV} rolled back to ${LABEL}"
echo "NOTE: releases/last-good/ is left unchanged; re-run a forward deploy to advance it."

View file

@ -1,123 +0,0 @@
name: Build & publish app artifacts
# T7 + T19 — build the release (SPA + Python source) and publish it to the per-env
# S3 artifact bucket, then roll the box to it.
#
# push to dev → publish to open-swe-dev-assets → deploy dev box (AUTO)
# push to main → publish to open-swe-prod-assets → deploy prod box (manual approval: env "prod")
#
# Two artifacts (the box's deploy.sh pulls both from releases/latest/):
# spa.tar.gz = the built dashboard SPA (vite -> ui/.output/public). Built HERE
# (not on the box) — the build is memory-heavy and the box is small.
# app.tar.gz = the Python source tree (NO ui/, NO .venv). The box runs
# `uv sync` to build a native-ARM64 venv at the real runtime path.
#
# Each release is uploaded under releases/<sha>/ (immutable, auditable) AND mirrored
# to releases/latest/ (what the box pulls). Then the open-swe-<env>-deploy SSM
# document is fired (tag-scoped to project=open-swe,env=<env>) to roll the box.
#
# OIDC subject alignment (matches the per-env app-role trust in infra/lib/config.ts):
# - publish-dev declares NO `environment:` → sub = repo:…:ref:refs/heads/dev
# - publish-prod declares `environment: prod` → sub = repo:…:environment:prod
# (also triggers the prod Environment's required-reviewer approval gate).
#
# Prerequisites:
# - repo variables AWS_DEPLOY_ROLE_APP_DEV / AWS_DEPLOY_ROLE_APP_PROD = the
# githubdeploy-open-swe-app-<env> role ARNs (open-swe-iam CfnOutputs).
# - the open-swe-<env> stack deployed (creates the bucket + the SSM deploy doc).
permissions:
contents: read
on:
push:
branches: [dev, main]
paths:
- "agent/**"
- "ui/**"
- "deploy/**"
- "langgraph.json"
- "pyproject.toml"
- "uv.lock"
- ".github/workflows/build-artifacts.yml"
- ".github/scripts/**"
workflow_dispatch:
concurrency:
# one publish+deploy per branch at a time; never cancel an in-flight release.
group: build-artifacts-${{ github.ref }}
cancel-in-progress: false
jobs:
publish-dev:
name: Publish + deploy (dev)
if: ${{ github.ref == 'refs/heads/dev' }}
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
id-token: write
contents: read
env:
ENV: dev
BUCKET: open-swe-dev-assets
DEPLOY_DOC: open-swe-dev-deploy
steps:
- uses: actions/checkout@v7
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- name: Build SPA (vite -> ui/.output/public)
working-directory: ui
# vite's bundle exceeds Node's default ~2 GB heap (the build that needed an
# 8 GB swapfile on-box); the runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"
run: |
bun install --frozen-lockfile
bun run build
- name: Package artifacts
run: bash .github/scripts/package-artifacts.sh
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_DEV }}
aws-region: us-east-1
- name: Publish to S3 + roll the box
run: bash .github/scripts/publish-and-deploy.sh
publish-prod:
name: Publish + deploy (prod)
if: ${{ github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 30
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam). Also
# makes the OIDC sub …:environment:prod (matches the prod app-role trust).
environment: prod
permissions:
id-token: write
contents: read
env:
ENV: prod
BUCKET: open-swe-prod-assets
DEPLOY_DOC: open-swe-prod-deploy
steps:
- uses: actions/checkout@v7
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- name: Build SPA (vite -> ui/.output/public)
working-directory: ui
# vite's bundle exceeds Node's default ~2 GB heap (the build that needed an
# 8 GB swapfile on-box); the runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"
run: |
bun install --frozen-lockfile
bun run build
- name: Package artifacts
run: bash .github/scripts/package-artifacts.sh
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_PROD }}
aws-region: us-east-1
- name: Publish to S3 + roll the box
run: bash .github/scripts/publish-and-deploy.sh

View file

@ -1,124 +0,0 @@
name: Infra CD
# Path-filtered CDK deploy for /infra, per env, OIDC-only (no static keys).
#
# push to dev → CI (tsc+jest+synth) → deploy OpenSweDevStack (AUTO, CI-green-gated)
# push to main → CI → deploy OpenSweProdStack (manual approval: env "prod")
#
# Why this is NOT the reusable cd-cdk.yaml: that workflow runs `cdk deploy --all`,
# which would deploy ALL THREE stacks (incl. the OTHER env + the shared IAM stack)
# from a single-env push — breaking the per-env dev/prod boundary. So we target one
# stack explicitly per env. (Infra CI still uses the reusable ci-typescript-cdk.)
#
# The shared IAM stack (open-swe-iam — owns BOTH envs' OIDC deploy roles) is
# intentionally NOT deployed here: it is a privileged, human-gated apply (T6), so a
# routine dev push can never alter prod's deploy role.
#
# OIDC subject alignment (must match the per-env trust in infra/lib/config.ts):
# - deploy-dev declares NO `environment:` → token sub = repo:…:ref:refs/heads/dev,
# which is exactly what githubdeploy-open-swe-infra-dev trusts.
# - deploy-prod declares `environment: prod` → token sub = repo:…:environment:prod,
# which githubdeploy-open-swe-infra-prod trusts AND which triggers the GitHub
# Environment's required-reviewer (manual approval) gate.
#
# Prerequisites (post-T6, when the roles exist):
# - repo variables AWS_DEPLOY_ROLE_INFRA_DEV / AWS_DEPLOY_ROLE_INFRA_PROD = the
# githubdeploy-open-swe-infra-<env> role ARNs (open-swe-iam CfnOutputs).
# - a GitHub Environment named "prod" with Adam as a required reviewer.
permissions:
contents: read
on:
push:
branches: [dev, main]
paths:
- "infra/**"
- ".github/workflows/cd-infra.yml"
workflow_dispatch:
concurrency:
# one infra deploy per branch at a time; never cancel an in-flight deploy.
group: cd-infra-${{ github.ref }}
cancel-in-progress: false
jobs:
# CI-green precondition — re-run tsc + jest + synth on the pushed commit before
# any deploy. A failure here blocks the deploy jobs (needs: ci).
ci:
name: Infra CI (pre-deploy)
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@main
with:
node-version: "24"
working-directory: infra
cache-dependency-path: infra/package-lock.json
run-typecheck: true
run-tests: true
run-cdk-synth: true
deploy-dev:
name: Deploy open-swe-dev
needs: ci
if: ${{ github.ref == 'refs/heads/dev' }}
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
id-token: write
contents: read
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v4
with:
node-version: "24"
cache: npm
cache-dependency-path: infra/package-lock.json
- name: Install deps
working-directory: infra
run: npm ci
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_INFRA_DEV }}
aws-region: us-east-1
- name: CDK deploy (dev only)
working-directory: infra
# --outputs-file lets us print the stack outputs from CDK's own result
# (the deploy role intentionally lacks cloudformation:DescribeStacks; CDK
# gets outputs via the bootstrap cfn-exec role it assumes, so no extra grant).
run: npx cdk deploy OpenSweDevStack --require-approval never --outputs-file cdk-outputs.json
- name: Stack outputs
working-directory: infra
run: cat cdk-outputs.json
deploy-prod:
name: Deploy open-swe-prod
needs: ci
if: ${{ github.ref == 'refs/heads/main' }}
runs-on: ubuntu-latest
timeout-minutes: 30
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam).
# Also makes the OIDC sub …:environment:prod (matches the prod role trust).
environment: prod
permissions:
id-token: write
contents: read
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v4
with:
node-version: "24"
cache: npm
cache-dependency-path: infra/package-lock.json
- name: Install deps
working-directory: infra
run: npm ci
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_INFRA_PROD }}
aws-region: us-east-1
- name: CDK deploy (prod only)
working-directory: infra
# See deploy-dev: --outputs-file avoids needing cloudformation:DescribeStacks.
run: npx cdk deploy OpenSweProdStack --require-approval never --outputs-file cdk-outputs.json
- name: Stack outputs
working-directory: infra
run: cat cdk-outputs.json

View file

@ -1,26 +0,0 @@
name: Infra CI
# Path-filtered CI for the /infra CDK app (TypeScript). The existing "CI"
# (ci.yml) covers the Python agent; this adds tsc + jest + cdk synth for /infra so
# infra changes are gated on a PR the same way. Runs only when /infra changes.
permissions:
contents: read
on:
pull_request:
paths:
- "infra/**"
- ".github/workflows/ci-infra.yml"
jobs:
infra-ci:
name: Infra CI (tsc + jest + synth)
uses: Sea-Haven-Industries/.github/.github/workflows/ci-typescript-cdk.yaml@main
with:
node-version: "24"
working-directory: infra
cache-dependency-path: infra/package-lock.json
run-typecheck: true
run-tests: true
run-cdk-synth: true

View file

@ -73,7 +73,7 @@ jobs:
- name: Run E2E
working-directory: tests/e2e
# Playwright's globalSetup runs the real `bun run build`, whose vite bundle
# exceeds Node's default ~2 GB heap (same OOM fixed in build-artifacts.yml).
# exceeds Node's default ~2 GB heap.
# The runner has ~16 GB, so lift the heap cap.
env:
NODE_OPTIONS: "--max-old-space-size=8192"

View file

@ -1,11 +1,21 @@
name: Promote dev to main (prod)
# Gated dev -> main promotion (the PROD deploy gate under managed LangGraph Cloud).
#
# PROD now runs on managed LangGraph Cloud (git-connected to `main`) + Vercel (UI).
# The platform AUTO-DEPLOYS prod on every push to `main`, so a fast-forward of `main`
# to `dev` IS the prod deploy trigger -- there is no separate AWS deploy step anymore
# (the bespoke S3/SSM/packer CD is retired). This workflow therefore carries the whole
# prod-promotion gate:
# 1. manual dispatch only (no scheduled auto-promote -- prod ships when a human asks),
# 2. the `prod` GitHub Environment approval (required reviewer: amoussa1229),
# 3. the dev-HEAD-fully-green precondition (check-dev-green.sh),
# 4. a fast-forward-only push of `main` -> `dev` via the seahaven-promotion App, the
# sole non-admin bypass actor on the `main` ruleset (id 18238334).
name: Promote to main (prod)
permissions:
contents: write
on:
schedule:
- cron: "0 8 * * *"
workflow_dispatch:
concurrency:
@ -15,10 +25,29 @@ concurrency:
jobs:
promote:
runs-on: ubuntu-latest
# The manual-approval gate. The `prod` Environment's required reviewer
# (amoussa1229) must approve before this job runs -- and because the FF push
# below auto-deploys managed prod, that approval IS the prod-deploy approval.
environment: prod
permissions:
contents: write
checks: read
steps:
# Defense-in-depth: the checkout below pins `ref: dev`, so this workflow only
# ever promotes dev's HEAD. But workflow_dispatch runs the workflow DEFINITION
# from whichever ref it was launched on, so a branch that edited this file could
# otherwise reach the privileged steps. Refuse any dispatch not from `dev`,
# before the App token is minted. (The `prod` Environment approval still gates
# everything after this regardless.)
- name: Guard — only promote from dev
env:
DISPATCH_REF: ${{ github.ref_name }}
run: |
if [ "${DISPATCH_REF}" != "dev" ]; then
echo "::error::promote-to-main must be dispatched from 'dev' (got '${DISPATCH_REF}')."
exit 1
fi
echo "Dispatch ref OK: ${DISPATCH_REF}"
# Mint a GitHub App installation token for the protected-branch push below.
# A plain ref push by github-actions[bot] is REJECTED by the `main` ruleset
# (PRs required + a required status check; the default token is not a bypass
@ -56,10 +85,11 @@ jobs:
-q '.check_runs[] | [.name, .status, (.conclusion // ""), (.details_url // "")] | join("\u001f")' \
| bash .github/scripts/check-dev-green.sh
- name: Fast-forward main (PROD) to dev
# main is the production branch. A direct ref push is normally rejected by the
# `main` ruleset (PRs required), so this succeeds only because the App minted
# above is a bypass actor. The push is fast-forward-only, so a diverged main
# fails loudly rather than force-updating.
# main is the production branch: managed LangGraph Cloud is git-connected to it
# and auto-deploys prod on every push, so THIS push is the prod deploy trigger.
# A direct ref push is normally rejected by the `main` ruleset (PRs required),
# so it succeeds only because the App minted above is a bypass actor. The push
# is fast-forward-only, so a diverged main fails loudly rather than force-updating.
#
# SECURITY: this is the LAST step on purpose. The App token persists as the
# git credential after checkout — do not add steps after this push that run

View file

@ -1,84 +0,0 @@
name: Rollback (re-point env to a prior release)
# Roll an env back to a previously published release without rebuilding. Re-points
# releases/latest/ at the chosen release and re-fires the open-swe-<env>-deploy SSM
# document — same fire/wait/gate path as a forward deploy (roll-box.sh).
#
# env=dev, sha blank → restore open-swe-dev-assets/releases/last-good/ (AUTO)
# env=prod, sha blank → restore open-swe-prod-assets/releases/last-good/ (manual
# approval: Environment "prod", same gate as a prod deploy)
# sha=<commit> → restore that exact releases/<sha>/ instead of last-good.
#
# No new IAM: reuses the githubdeploy-open-swe-app-<env> role's existing releases/*
# write + tag-scoped ssm:SendCommand. OIDC subject alignment matches build-artifacts:
# - rollback-dev declares NO `environment:` → sub = repo:…:ref:refs/heads/<branch>
# - rollback-prod declares `environment: prod` → sub = repo:…:environment:prod
permissions:
contents: read
on:
workflow_dispatch:
inputs:
env:
description: "Which environment to roll back"
required: true
type: choice
options: [dev, prod]
sha:
description: "Release SHA to restore (blank = releases/last-good)"
required: false
type: string
concurrency:
# never overlap a rollback with another rollback/deploy of the same env.
group: rollback-${{ inputs.env }}
cancel-in-progress: false
jobs:
rollback-dev:
name: Rollback (dev)
if: ${{ inputs.env == 'dev' }}
runs-on: ubuntu-latest
timeout-minutes: 20
permissions:
id-token: write
contents: read
env:
ENV: dev
BUCKET: open-swe-dev-assets
DEPLOY_DOC: open-swe-dev-deploy
TARGET_SHA: ${{ inputs.sha }}
steps:
- uses: actions/checkout@v7
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_DEV }}
aws-region: us-east-1
- name: Re-point releases/latest + redeploy
run: bash .github/scripts/rollback.sh
rollback-prod:
name: Rollback (prod)
if: ${{ inputs.env == 'prod' }}
runs-on: ubuntu-latest
timeout-minutes: 20
# Manual-approval gate: the "prod" Environment requires a reviewer (Adam). Also
# makes the OIDC sub …:environment:prod (matches the prod app-role trust).
environment: prod
permissions:
id-token: write
contents: read
env:
ENV: prod
BUCKET: open-swe-prod-assets
DEPLOY_DOC: open-swe-prod-deploy
TARGET_SHA: ${{ inputs.sha }}
steps:
- uses: actions/checkout@v7
- uses: aws-actions/configure-aws-credentials@v6
with:
role-to-assume: ${{ vars.AWS_DEPLOY_ROLE_APP_PROD }}
aws-region: us-east-1
- name: Re-point releases/latest + redeploy
run: bash .github/scripts/rollback.sh

View file

@ -10,16 +10,6 @@
"owner": "adam@seahavenind.com",
"added": "2026-06-29"
},
{
"id": "OSWE-IAC-SECRETS-LIST-01",
"title": "EC2 instance role grants BatchGetSecretValue on \"*\" (operation-level; secret-NAME existence enumeration account-wide)",
"file": "infra/lib/constructs/instance-role.ts",
"severity": "low",
"status": "confirmed",
"suppression_justification": "ACCEPTED LOW residual, metadata-only. secretsmanager:BatchGetSecretValue is a collection action that AWS cannot scope to a per-secret ARN, so it is granted on `*` (documented in commit 3c69dd9d and the construct comment). Secret VALUES remain strictly gated by the PREFIX-scoped GetSecretValue/DescribeSecret on secret:open-swe-<env>/* (checked per-secret even within the batch), so cross-env VALUE isolation is preserved; only a name EXISTENCE oracle remains, within Sea Haven's single-tenant account 328440206208. ListSecrets is intentionally NOT granted (so name FILTER enumeration AccessDenies). Confirmed by GPT-4.1 IAM cross-review (no BLOCK) and the iac-iam detector (one low residual, no critical/high).",
"owner": "adam@seahavenind.com",
"added": "2026-06-26"
},
{
"id": "gitleaks-generic-api-key-89",
"title": "Hardcoded credential flagged in encryption-roundtrip test fixture (CWE-798)",

View file

@ -638,14 +638,17 @@ Production runs the backend and dashboard separately.
3. Set all environment variables from step 6 in the deployment config. Set `DASHBOARD_BASE_URL` and `LANGGRAPH_URL` to your production URLs (all `https://`). Set `DASHBOARD_API_BASE_URL` to the URL browsers use for dashboard API requests and OAuth callbacks: either the backend URL for direct cross-origin calls, or the dashboard/Vercel URL when a same-origin rewrite proxies `/dashboard/api/*`.
4. Update your webhook URLs (Linear, Slack, GitHub App) and the GitHub App / Slack OAuth callback URLs to your production URLs (replace the ngrok / localhost values). The dashboard GitHub App callback must be `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback`.
The `langgraph.json` at the project root defines the three graphs and the HTTP app:
The `langgraph.json` at the project root defines the graphs and the custom HTTP app:
```json
{
"graphs": {
"agent": "agent.server:traced_agent",
"reviewer": "agent.reviewer:traced_reviewer_agent",
"analyzer": "agent.analyzer:traced_analyzer"
"analyzer": "agent.analyzer:traced_analyzer",
"chat": "agent.chat:traced_chat_agent",
"scheduler": "agent.scheduler:get_scheduler",
"ci_monitor": "agent.ci_monitor:get_ci_monitor"
},
"http": {
"app": "agent.webapp:app"
@ -657,6 +660,23 @@ The `langgraph.json` at the project root defines the three graphs and the HTTP a
Alternatively, you can run the dashboard as a direct cross-origin client: set `VITE_DASHBOARD_API_BASE_URL` to the hosted backend origin, set `DASHBOARD_API_BASE_URL` to that same backend origin, and include the dashboard origin in `DASHBOARD_ALLOWED_ORIGINS`.
### Managed LangGraph Cloud + Vercel — the production env triad
The production runtime is **managed LangGraph Cloud** (backend) + **Vercel** (UI), not a self-hosted `langgraph dev` process. Get these three right, in the recommended same-origin setup:
| Variable | Where | Value | Why |
|---|---|---|---|
| `LANGGRAPH_URL` | backend (LangGraph Cloud) | the deployment's own `https://…us.langgraph.app` URL | The server-side `langgraph_client()` (thread sidebar + run creation) connects here. It defaults to `http://localhost:2024` for local dev; **on managed it must be set** or every dashboard/thread call `ConnectError`s. |
| `DASHBOARD_BASE_URL` | backend | the dashboard origin, `https://…` | Where token-free settings links point. |
| `DASHBOARD_API_BASE_URL` | backend | the **same dashboard/Vercel origin**, `https://…` | Builds the OAuth `redirect_uri` and drives the session-cookie `Secure; SameSite=None` flags. **Must include the `https://` scheme** — a schemeless value yields a schemeless `redirect_uri` and GitHub rejects login with "redirect_uri not associated with this application." (The server now coerces a schemeless value to `https://` as a backstop, but set it explicitly.) |
| `VITE_DASHBOARD_API_BASE_URL` | UI build (Vercel) | **empty** | Leave empty so the SPA calls relative `/dashboard/api/*` same-origin; `ui/vercel.json` rewrites those to the backend so the `osw_session` cookie rides along. |
`ui/vercel.json`'s rewrite `destination` must be the **stable LangGraph deployment URL** for the target environment (e.g. `https://<deployment>-<hash>.us.langgraph.app`) — a per-deployment alias that is stable across revisions. Never point it at an immutable per-deploy / preview URL. Vercel does not interpolate env vars into `vercel.json` rewrites, so this hostname is committed per Vercel project; set the dev project's rewrite to the dev deployment and the prod project's to the prod deployment.
**Always start dashboard login on the stable alias / custom domain**, not on an immutable per-deploy Vercel URL. The OAuth state cookie (`osw_oauth_state`) is intentionally host-only (a `Domain` cookie is unsafe across `*.vercel.app`, which is on the public-suffix list). If login starts on one host and GitHub redirects back to another, the cookie is dropped and the callback fails with "oauth state mismatch."
On managed the backend runs as **multiple replicas**. In-process caches that must stay coherent across replicas (notably the GitHub⇄work-email user mapping) are refreshed from the durable Store on the webhook hot paths (Slack and GitHub) before they are read.
## Troubleshooting
### Webhook not receiving events
@ -675,7 +695,8 @@ Alternatively, you can run the dashboard as a direct cross-origin client: set `V
### Dashboard login fails or won't stay logged in
- `500 GITHUB_APP_CLIENT_ID not configured` (or client secret): set `GITHUB_APP_CLIENT_ID` / `GITHUB_APP_CLIENT_SECRET` (step 3c) and `DASHBOARD_JWT_SECRET`.
- OAuth `redirect_uri` mismatch: the GitHub App must list `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` as a callback URL (step 3b). Locally that's `http://localhost:2024/dashboard/api/auth/callback`.
- OAuth `redirect_uri` mismatch / "redirect_uri not associated with this application": the GitHub App must list `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` as a callback URL (step 3b). Locally that's `http://localhost:2024/dashboard/api/auth/callback`. On managed, a common cause is a **schemeless** `DASHBOARD_API_BASE_URL` (e.g. `dash.example.com` instead of `https://dash.example.com`), which produces a schemeless `redirect_uri`; always set the full `https://` URL.
- "oauth state mismatch" after the GitHub redirect: login was started on a different host than the one GitHub redirected back to (typically an immutable per-deploy Vercel URL vs the stable alias). The `osw_oauth_state` cookie is host-only by design, so it isn't presented to the other host. Always start login on the stable alias / custom domain.
- Login redirects but the session doesn't stick: this is almost always a cookie problem. Locally, keep `DASHBOARD_API_BASE_URL` on `http://` (so cookies are `SameSite=Lax`); in prod use `https://` for both API and frontend and add the frontend origin to `DASHBOARD_ALLOWED_ORIGINS`.
- Login rejected with an org error: `ALLOWED_GITHUB_ORGS` gates dashboard login (and requires the App's Organization → Members: Read-only permission). See step 5.
- Admin pages 403: add your GitHub login or email to `CONFIGURED_ADMINS`.

View file

@ -153,18 +153,20 @@ This is an area where you can extend Open SWE for your org: add deterministic CI
## Deployment (Sea Haven fork)
This fork is **self-hosted on AWS** and live in production. Each env (`dev` /
`prod`) runs the stock `langgraph dev` server (all three graphs + the FastAPI
webapp) bound to loopback `127.0.0.1:2024` on a single ARM64 EC2 box, fronted by
nginx (the sole ingress) behind the shared `seahaven-com` ALB. CDK (`infra/`)
owns the per-env stacks; GitHub Actions handle CDK deploys (`cd-infra.yml`) and
app-artifact releases to S3 rolled onto the box via an SSM document
(`build-artifacts.yml`), with `dev` auto-deploying and `prod` gated behind a
manual GitHub Environment approval.
This fork runs on a **managed deployment**: the backend (all three graphs + the
FastAPI webapp) runs on
[LangGraph Cloud / Platform](https://langchain-ai.github.io/langgraph/cloud/),
and the `ui/` dashboard deploys to [Vercel](https://vercel.com/). Configuration
and secrets live in the LangGraph deployment config and Vercel environment
variables. Promotion from `dev` to `prod` (`main`) is handled by
[`.github/workflows/promote-to-main.yml`](.github/workflows/promote-to-main.yml).
**[`deploy/seahaven/DEPLOYMENT.md`](deploy/seahaven/DEPLOYMENT.md) is the canonical
deploy runbook** — full end-to-end pipeline, config seeding, promotion/rollback,
and live prod facts. CDK specifics live in [`infra/README.md`](infra/README.md).
See **[INSTALLATION.md § 10 "Production deployment"](INSTALLATION.md#10-production-deployment)**
for the full backend + dashboard setup.
> The earlier self-hosted AWS stack (CDK under `infra/`, an ARM64 EC2 box + nginx
> behind the shared ALB, and the `cd-infra` / `build-artifacts` release pipelines)
> was **decommissioned** in favor of the managed deployment above.
## License

View file

@ -17,7 +17,6 @@ from datetime import UTC, datetime
from typing import Any
import httpx
from deepagents.backends.utils import create_file_data
from fastapi import HTTPException
from ..reviewer_diff import fetch_pr_diff
@ -223,6 +222,12 @@ async def _build_pr_context(
Accepts an already-fetched ``review`` to avoid re-fetching it when the caller
has just read it to decide whether a reseed is needed.
"""
# Deferred import: deepagents pulls langchain_anthropic / anthropic (~0.7s)
# into the import graph. It is otherwise dragged into the custom FastAPI
# app's cold start via dashboard.routes; only needed when seeding chat
# files, so import it lazily here.
from deepagents.backends.utils import create_file_data
if review is None:
review = await get_review(owner, repo, pr_number)
findings = review.get("findings") if isinstance(review.get("findings"), list) else []

View file

@ -236,6 +236,13 @@ def _api_base_url() -> str:
v = os.environ.get("DASHBOARD_API_BASE_URL", "").rstrip("/")
if not v:
raise HTTPException(500, "DASHBOARD_API_BASE_URL not configured")
if not v.startswith(("http://", "https://")):
# A schemeless value (e.g. "open-swe-prod.us.langgraph.app") produces a
# schemeless OAuth redirect_uri, which GitHub rejects with
# "redirect_uri not associated with this application". Managed
# deployments are served over HTTPS, so default to https:// when the
# operator omitted the scheme.
v = f"https://{v}"
return v
@ -255,9 +262,14 @@ def _cookie_security() -> tuple[bool, str]:
rejected and the frontend/API are same-site, so fall back to
``SameSite=Lax`` without ``Secure``.
"""
if os.environ.get("DASHBOARD_API_BASE_URL", "").startswith("https://"):
return True, "none"
return False, "lax"
api = os.environ.get("DASHBOARD_API_BASE_URL", "")
# Only an explicit ``http://`` (local dev) or an unconfigured value falls
# back to the insecure same-site cookie. ``https://`` *and* a schemeless
# managed host (which ``_api_base_url`` coerces to https) are cross-site
# over TLS and must use ``Secure; SameSite=None``.
if not api or api.startswith("http://"):
return False, "lax"
return True, "none"
def _set_session_cookie(response: Response, jwt_token: str) -> None:
@ -277,6 +289,15 @@ def _set_state_cookie(response: Response, nonce: str) -> None:
# SameSite=Lax so GitHub's top-level redirect back to /auth/callback
# still presents this cookie; the cookie is single-purpose and lives
# only for the duration of one OAuth round-trip.
#
# This cookie is intentionally host-only (no Domain attribute): scoping it
# to a shared parent domain is not safe across Vercel's immutable per-deploy
# hostnames (``*.vercel.app`` is on the public-suffix list, so a Domain
# cookie there is rejected). The operational contract is therefore to
# *always start login on the stable alias / custom domain* so the host that
# sets this cookie is the same host GitHub redirects back to. Starting the
# flow on an immutable per-deploy URL and finishing on the alias (or vice
# versa) drops the cookie and surfaces as "oauth state mismatch".
secure, _ = _cookie_security()
response.set_cookie(
key=STATE_COOKIE_NAME,
@ -803,6 +824,33 @@ async def admin_list_user_mappings(
}
class UserMappingUpsert(BaseModel):
github_login: str
work_email: str
slack_user_id: str | None = None
@router.post("/admin/user-mappings")
async def admin_upsert_user_mapping(
body: UserMappingUpsert,
_admin: dict[str, Any] = _ADMIN_DEP,
) -> dict[str, Any]:
"""Create or update a GitHub↔work-email mapping from the admin dashboard.
Lets an admin set the ``work_email`` link directly instead of waiting for
the user to self-connect Slack (or doing a raw Store write).
"""
try:
return await upsert_mapping(
github_login=body.github_login,
work_email=body.work_email,
slack_user_id=body.slack_user_id or None,
source="admin",
)
except ValueError as e:
raise HTTPException(400, str(e)) from e
@router.delete("/admin/user-mappings/{github_login}")
async def admin_delete_user_mapping(
github_login: str,

View file

@ -35,7 +35,7 @@ logger = logging.getLogger(__name__)
USER_MAPPINGS_NAMESPACE: list[str] = ["user_mappings"]
MappingSource = Literal["slack_oauth"]
MappingSource = Literal["slack_oauth", "admin"]
MappingStatus = Literal["active", "pending"]

View file

@ -17,8 +17,6 @@ from __future__ import annotations
from pathlib import Path
from typing import Any
from deepagents.backends.utils import create_file_data
SKILLS_DIR = Path(__file__).resolve().parent.parent / "skills"
SKILLS_ROUTE = "/skills/"
@ -41,6 +39,13 @@ def build_skill_files() -> dict[str, Any]:
can serve them. Keys omit the ``/skills`` prefix (stripped by the composite
route); values are ``FileData`` v2 entries.
"""
# Deferred import: deepagents (and its langchain_anthropic / anthropic
# transitive deps) is heavy (~0.7s) and is otherwise pulled into the custom
# FastAPI app's import chain via dashboard.routes, slowing cold start. It's
# only needed when a skill bundle is actually built (analyzer launch), so
# import it lazily here.
from deepagents.backends.utils import create_file_data
files: dict[str, Any] = {}
for skill in ANALYZER_MODES.values():
skill_md = SKILLS_DIR / skill / "SKILL.md"

View file

@ -3049,6 +3049,16 @@ async def process_github_pr_comment(payload: dict[str, Any], event_type: str) ->
else:
logger.warning("Failed to persist branch_name metadata for thread %s", thread_id)
# Refresh the per-process user-mapping cache from the Store before
# resolving the author's email. On a multi-replica managed deployment this
# replica's cache may be stale (a mapping created on another replica is not
# otherwise visible), which would drop a legitimately-mapped user. Mirrors
# the Slack mention path (process_slack_mention).
try:
await refresh_user_mapping_cache()
except Exception: # noqa: BLE001
logger.debug("Could not refresh user mapping cache for GitHub PR comment", exc_info=True)
email = await email_for_login(github_login) or ""
if email:
github_token = await _get_or_resolve_thread_github_token(thread_id, email, repo=repo_config)
@ -3328,6 +3338,14 @@ async def process_github_issue(payload: dict[str, Any], event_type: str) -> None
logger.warning("Missing GitHub issue id/number, skipping")
return
# Refresh the per-process user-mapping cache from the Store before
# resolving the author's email (multi-replica staleness; mirrors the Slack
# mention path in process_slack_mention).
try:
await refresh_user_mapping_cache()
except Exception: # noqa: BLE001
logger.debug("Could not refresh user mapping cache for GitHub issue", exc_info=True)
email = await email_for_login(github_login) or ""
if not email:
logger.warning("No email mapping for GitHub user '%s', skipping", github_login)

361
deploy/MIGRATION.md Normal file
View file

@ -0,0 +1,361 @@
# Open SWE — Migration Plan: Self-Hosted AWS → Managed LangGraph Cloud + Vercel
**Repo:** `Sea-Haven-Industries/open-swe` (private) · **AWS:** 328440206208 / us-east-1
**Author:** Adam Moussa · **Date:** 2026-06-29 · **Status:** DRAFT — owes a `/sh-plan-review` before prod cutover (Phase C gate)
---
## 1. Executive Summary
### Decision (settled — not re-litigated here)
Move the Open SWE deployment off the bespoke self-hosted AWS stack (stock `langgraph dev`, in-memory store, no durability — issue **#9**) onto the **managed runtime**:
- **Backend → LangGraph Cloud** (a.k.a. "LangSmith Deployment"): git-connected, auto-builds a revision on push, zero-downtime, revision rollback, durable Postgres-backed store/checkpointer.
- **UI → Vercel**: `ui/` SPA, atomic deploys, instant rollback, per-PR previews (net-new capability).
### Why
1. **Dissolves the licensing blocker.** Self-hosted `langgraph up` on the LangSmith **Plus** key is "Self-Hosted Lite" (1M node-exec/yr cap, Elastic License 2.0, production legally ambiguous). Managed **IS** the licensed product with **no node cap** — runs billed flat at $0.005, nodes not charged.
2. **Deletes the entire durable-runtime build** (RDS/Redis/Docker/AMI) **and most of the bespoke AWS CD** (S3 release pipeline, SSM deploy docs, packer, CDK box/ALB stacks).
3. **It's the upstream-canonical deployment** and the repo is **already wired for it**: `ui/vercel.json` has the same-origin `/dashboard/api/*` rewrite; `langgraph.json` is cloud-format (6 graphs + `http.app`).
### Cost
≈ **$160/mo incremental** for one always-on prod deployment:
- $0.0036/min prod uptime ≈ **$155/mo** (always-on)
- + **$0.005/run**
- + traces pay-as-you-go above 10k/mo
- The **$39/mo Plus seat is already paid** (not incremental)
- **Dev deployment is FREE** (1 included on Plus, preemptible)
A self-hosted `langgraph up` + RDS plan (already `/sh-plan-review`'d to APPROVE-after-revision this session) is the documented **FALLBACK** if managed is ever rejected (see §11).
---
## 2. Architecture: Before → After
### Before (self-hosted AWS — LIVE as of 2026-06-29)
```
GitHub/Slack/Linear ──webhook──▶ hooks.seahaven.com ─┐
Browser (dashboard) ─────────────▶ openswe.seahaven.com ─┤
▼
shared seahaven-com ALB (:443, host+path rules)
▼
EC2 (prod i-08a729e50779c4b07 t4g.large ARM64, private subnet)
nginx :80 ──proxy /dashboard/api + /webhooks──▶ langgraph dev :2024 (loopback)
in-memory store (reseeded by seed_store.sh ExecStartPost)
▼
Config: Secrets Manager open-swe-prod/* + SSM /open-swe-prod/*
Release: GitHub Actions → S3 open-swe-prod-assets/releases/* → SSM doc deploy.sh
IaC: CDK OpenSweIamStack + OpenSweDevStack + OpenSweProdStack
Sandbox: LangSmith cloud (DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub proxy)
```
### After (managed)
```
GitHub/Slack/Linear ──webhook──▶ *.langgraph.app (or hooks.seahaven.com CNAME → TODO §10)
Browser (dashboard) ─────────────▶ open-swe-dashboard.vercel.app (stable alias / custom domain)
│ same-origin rewrite /dashboard/api/* (ui/vercel.json)
▼
LangGraph Cloud "Deployment" (managed, git-connected to `main` for prod / `dev` for dev)
├─ serves the 6 graphs (agent, reviewer, analyzer, chat, scheduler, ci_monitor)
├─ serves the custom http.app (agent.webapp:app = webhooks + dashboard API + OAuth)
├─ durable Postgres store + checkpointer (issue #9 SOLVED) — autoscaled 1→10 replicas
└─ env/secrets in the Deployment config (NOT Secrets Manager) — Adam accepted deviation
▼
Sandbox: LangSmith cloud (UNCHANGED — DEFAULT_SANDBOX_SNAPSHOT_ID + GitHub-App proxy)
Auth: GitHub App seahaven-openswe (UNCHANGED — App 4146115 / Install 142615168)
```
**What changes shape:** runtime host (EC2 → managed PaaS), durability (in-memory → managed Postgres), CD (bespoke S3/SSM/packer → git-connected auto-build), config home (Secrets Manager/SSM → Deployment+Vercel env). **What stays:** the GitHub App, the LangSmith sandbox plane, CI (lint/format/unit/Playwright), the app code itself.
---
## 3. Phased Plan
### Phase A — Dev spike (MOSTLY DONE)
Goal: prove managed serves our custom app + durability, at $0, before committing prod $.
**Proven this session:**
- ✅ Dev backend deployed to LangGraph Cloud, connected to branch `dev`:
`https://open-swe-dev-hosted-e76c2b0e8a7955fe8ad3110a7a54e5d0.us.langgraph.app`
(Bedrock + a Fireworks key set; AWS creds for Bedrock deferred — see §10 open decision).
- ✅ UI deployed to Vercel — team `sea-haven`, project `open-swe-dashboard`, `https://open-swe-dashboard.vercel.app`; `ui/vercel.json` rewrite repointed at the dev deployment.
- ✅ Managed serves the custom `http.app` (dashboard API + webhooks) — **no platform auth gate** in front of our routes (webhook returns 401 sig-enforced, so signatures still govern).
- ✅ Vercel same-origin rewrite → backend works.
- ✅ GitHub OAuth dashboard login end-to-end.
- ✅ **DURABILITY** — team-default model survived a revision redeploy (this is issue **#9**'s core acceptance goal, proven on managed).
**Remaining Phase A items (the GitHub `@openswe` run-trigger loop):**
- [ ] A1. Get an `@openswe` GitHub comment to dispatch a run end-to-end on dev. Blocked by the user-mapping + cache gotchas (see fixes #4 and #5 in §5).
- [ ] A2. Create the owner user-mapping in the managed Store (namespace `["user_mappings"]`, key = lowercased login `amoussa1229`, value `{github_login, work_email}`) — the Admin UI can't set `work_email` yet (fix #4).
- [ ] A3. Verify a freshly-added mapping is seen without a redeploy across managed's multi-replica autoscaling (fix #5).
- [ ] A4. Confirm `LANGGRAPH_URL` is set to the deployment URL on the dev deployment so `langgraph_client()` calls resolve (fix #1 — config now, code later).
### Phase B — Harden (the code fixes + env codification)
Land the 6 code fixes (§5), codify env/config, and resolve the Bedrock-auth decision **before** standing up prod.
- [ ] B1. Land code fixes #1–#6 (§5) as a PR into `dev` (or split into focused PRs). Re-run `make lint` + `make test`.
- [ ] B2. **Codify the env contract for managed.** Update `ui/vercel.json` rewrite to the *prod* deployment URL (Phase C) but keep dev pointing at dev. Add a documented LangGraph-Cloud env list to the repo (see §4) — NOT secrets, just the variable inventory + which are excluded.
- [ ] B3. **Resolve the Bedrock-auth decision** (§10 open decision): static AWS keys for a Bedrock-scoped IAM user in the deployment env, OR run the agent on Fireworks. This intersects the #62 model work.
- [x] B4. ~~Merge #62 first~~ — **DONE / N/A.** #62 (Bedrock + Fireworks) **IS** merged to `dev` and is what the dev deployment runs — verified on `origin/dev` (`options.py` → `DEFAULT_MODEL_ID = "bedrock_converse:us.anthropic.claude-opus-4-8"`) and confirmed by the live deployment's `git_ref_sha: a4ed19ba…` (= the #62 merge commit). The contrary note came from a STALE local checkout; no merge action needed.
- [ ] B5. Audit ALL in-process caches for the single-process → multi-replica assumption (generalization of fix #5; see §5).
- [ ] B6. Measure + reduce custom-app import time (fix #6) so the deployment isn't flagged unhealthy / slow to scale.
### Phase C — Prod deployment
- [ ] C1. Create a **prod LangGraph Cloud deployment** tracking branch `main` (the durable autoscaled 1→10 tier, not the free Dev tier). Record its `*.langgraph.app` URL.
- [ ] C2. Create the **prod Vercel project/target** (or promote the existing `open-swe-dashboard` to production); set its env (same-origin mode: `VITE_DASHBOARD_API_BASE_URL` empty).
- [ ] C3. Set the **prod env triad** (§4) on the prod deployment + Vercel:
- `LANGGRAPH_URL` = the prod `*.langgraph.app` URL
- `DASHBOARD_BASE_URL` + `DASHBOARD_API_BASE_URL` = the prod Vercel origin (with `https://` scheme — fix #2)
- [ ] C4. **Establish the prod approval gate** (§6): git-connected auto-deploy needs an explicit gate. Preserve the prod manual-approval that exists today (the GitHub `prod` Environment reviewer). Mechanism = protected `main` ruleset + the platform's "require manual promotion to production" if available (§10 open decision).
- [ ] C5. **Repoint webhooks + OAuth URLs** to prod:
- GitHub App `seahaven-openswe` webhook URL → prod backend webhook URL (`*.langgraph.app/webhooks/github` or `hooks.seahaven.com` — §10 custom-domain decision)
- Slack Event Subscriptions + Interactivity Request URLs → prod
- Linear webhook URL → prod (if/when wired)
- GitHub App OAuth callback → `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` (the Vercel prod origin; fixes #2, #3)
- [ ] C6. **Custom-domain decision** (§10): dashboard custom domain solved by Vercel; webhooks `hooks.seahaven.com` either re-point to `*.langgraph.app` directly or front via CNAME. Managed issues `*.langgraph.app` URLs ONLY (no documented custom-domain support on the backend).
### Phase D — Cutover + verify
- [ ] D1. Run the issue **#9 acceptance criteria on PROD**: durable store survives a revision redeploy; team defaults / user mappings persist; a paused/in-flight run survives a redeploy (the durability win self-host never had).
- [ ] D2. End-to-end prod smoke: `@openswe` GitHub comment → run dispatch → sandbox → draft PR → reply in source channel.
- [ ] D3. Dashboard prod smoke: GitHub OAuth login on the stable Vercel alias (fix #3); admin pages reachable for `CONFIGURED_ADMINS`.
- [ ] D4. Webhook sig-enforcement smoke: unsigned POST to each `/webhooks/*` returns 401; signed returns 200.
- [ ] D5. **Soak** for an agreed window (recommend 3–7 days) with the AWS prod stack still standing as the rollback target (§11) before any teardown.
### Phase E — AWS decommission (AFTER soak)
Only after managed prod is proven + soaked. See §7 for the precise retire-vs-keep list. Phased teardown:
- [ ] E1. Export the live ALB listener-rule / target-group config for `open-swe-prod` and `open-swe-dev` to JSON (the export is the rollback source — these rules were partly hand-built).
- [ ] E2. Disable the bespoke CD (build-artifacts / cd-infra workflows) so nothing re-deploys the EC2 boxes.
- [ ] E3. `cdk destroy OpenSweProdStack` then `OpenSweDevStack` (mind the **RETAIN** secret shells — they survive and hold the global `open-swe-<env>/*` names; force-delete only EMPTY shells once env is fully migrated; keep populated ones transitionally as the env source — §7).
- [ ] E4. Remove ALB rules/target groups, S3 assets buckets, SSM deploy docs, and the per-env CDK bootstrap qualifier `oswedev` (`CDKToolkit-oswedev`) — once nothing references them.
- [ ] E5. Decide DNS: repoint `openswe.seahaven.com` → Vercel; resolve `hooks.seahaven.com` → `*.langgraph.app` (or CNAME).
- [ ] E6. Leave `OpenSweIamStack` until last — the promotion App + any retained roles may still be referenced.
### Phase F — Docs / memory
- [ ] F1. Rework Confluence "AWS Architecture Map" (page **1540098**) + the open-swe child page (**26116098**): the RDS/EC2/ALB subgraph is replaced by an **external-services view** (LangGraph Cloud + Vercel + the retained GitHub App + LangSmith sandbox).
- [ ] F2. Update `README.md` + `INSTALLATION.md` (§10 is already canonical-correct; align Sea Haven specifics).
- [ ] F3. Update project memories `project_open_swe_migration.md` + `reference_open_swe_deployment.md` to mark the managed cutover and the AWS teardown.
---
## 4. Env / Config Reference
### The prod triad (per INSTALLATION.md §10, lines 630–658)
| Var | Prod value | Notes |
|---|---|---|
| `LANGGRAPH_URL` | the deployment URL (`https://...langgraph.app`) | **NOT** localhost. Drives `thread_ops.langgraph_url()` (fix #1). |
| `DASHBOARD_BASE_URL` | the Vercel origin | same-origin rewrite mode |
| `DASHBOARD_API_BASE_URL` | the Vercel origin, **with `https://` scheme** | scheme required or OAuth `redirect_uri` is schemeless and GitHub rejects (fix #2) |
| `VITE_DASHBOARD_API_BASE_URL` | **empty** | same-origin mode; UI calls relative `/dashboard/api/*`, Vercel rewrites to backend |
GitHub App dashboard OAuth callback = `<DASHBOARD_API_BASE_URL>/dashboard/api/auth/callback` (the Vercel prod origin).
### Where secrets/env live now
**LangGraph Cloud Deployment config + Vercel env** — NOT AWS Secrets Manager. This **deviates from the Sea Haven `secrets-and-config.md` "Secrets Manager for all sensitive" handbook rule** — **Adam ACCEPTED this deviation** (managed has no instance role / no fetch-config boot hook; the platform's own secret store is the mechanism).
### Sourcing env from the existing AWS stack (transitional)
Pull from the existing `open-swe-dev` Secrets Manager + SSM via a documented CLI dump:
- `aws secretsmanager list-secrets --filters Key=name,Values=open-swe-dev/` then per-name `get-secret-value` — **NOT** `batch-get-secret-value` (its pagination silently drops values past page 1; this bit the box twice — see memory).
- SSM: `aws ssm get-parameters-by-path --path /open-swe-dev/`.
### EXCLUDE when copying to managed
| Category | Vars |
|---|---|
| Box-/self-host-specific | `LANGGRAPH_URL` (set fresh to deployment URL), `LANGGRAPH_URL_PROD`, `LANGSMITH_ENDPOINT`, `LANGSMITH_ENDPOINT_PROD`, `LANGSMITH_URL_PROD`, `LANGSMITH_TENANT_ID_PROD`, `LANGSMITH_HOST_API_URL`, `LANGCHAIN_REVISION_ID` |
| Dropped providers (PR #62) | `OPENAI_API_KEY`, `GOOGLE_API_KEY`, `GROQ_API_KEY` (these 6 Secrets Manager shells deleted 2026-06-29, final purge 2026-07-06) |
| Unused sandbox providers | `DAYTONA_API_KEY`, `RUNLOOP_API_KEY` |
| Eval-only | `JUDGE_ANTHROPIC_API_KEY` (kept on AWS for `evals/reviewer/judge.py`, not needed in the runtime deployment) |
### KEY mappings on managed
- `LANGSMITH_API_KEY` = the **`LANGSMITH_API_KEY_PROD`** value (and `LANGCHAIN_API_KEY` = same).
- The full secret inventory is `SECRET_VARS` in `infra/lib/constructs/config-store.ts:46` (28 shells) — use it as the checklist of what to carry, minus the EXCLUDE rows above.
### Models (post-#62) + Bedrock auth
Post-#62 `SUPPORTED_MODELS` = `bedrock_converse:us.anthropic.claude-opus-4-8` (**DEFAULT**) + 3 Fireworks models; **no direct `anthropic:` option**.
✅ **#62 is merged to `dev` and deployed** — verified on `origin/dev` (`options.py` → `bedrock_converse` default) and the live deployment `git_ref_sha a4ed19ba`. (The pre-#62 values appear only on a stale LOCAL checkout — ignore.)
**Bedrock on managed has NO EC2 instance role** (the #62 IAM design used the EC2 instance role as the Bedrock principal — that breaks on managed). Options (OPEN DECISION, §10):
- (a) Static `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION` for a **Bedrock-scoped IAM user** in the deployment env, or
- (b) Run the agent on **Fireworks** and avoid Bedrock entirely on managed.
---
## 5. Required Code Fixes (the 6 gotchas)
These are self-host → managed assumption breaks discovered on the spike. Each gets a config workaround now and a code fix to land in Phase B.
### Fix #1 — `langgraph_url()` localhost fallback
**File:** `agent/utils/thread_ops.py:38-41`
```python
def langgraph_url() -> str:
return os.environ.get("LANGGRAPH_URL") or os.environ.get(
"LANGGRAPH_URL_PROD", "http://localhost:2024"
)
```
On managed, every `langgraph_client()` call (thread sidebar, run creation, the same pattern in `agent/dashboard/user_mappings.py:43` `get_client()` with no URL) `ConnectError`s unless `LANGGRAPH_URL` is set to the deployment URL.
- **Config fix (now):** set `LANGGRAPH_URL` = the deployment URL on every deployment (Phase A4 / C3).
- **Code fix (Phase B):** default to the in-process / deployment URL rather than `http://localhost:2024` when running inside a managed deployment (e.g. honor the platform's own URL env, or fail loud instead of silently hitting localhost).
### Fix #2 — OAuth `DASHBOARD_API_BASE_URL` must include scheme
**Where:** OAuth redirect construction (dashboard `oauth.py` / `routes.py`), driven by `DASHBOARD_API_BASE_URL`.
A schemeless value yields a schemeless `redirect_uri` → GitHub rejects with "redirect_uri not associated."
- **Fix:** always set `DASHBOARD_API_BASE_URL` **with `https://`** (Phase C3). Code hardening (Phase B): assert/normalize a scheme at startup and fail loud if missing.
### Fix #3 — OAuth state cookie is host-only
**Cookie:** `osw_oauth_state` (host-only). Login must START on the SAME host as `DASHBOARD_API_BASE_URL` — the **stable Vercel alias**, never the immutable per-deploy URL — or you get "oauth state mismatch."
- **Fix:** pin `DASHBOARD_API_BASE_URL` + the login entry point to the stable alias / custom domain (Phase C3, D3). Document this so a future per-deploy preview URL isn't used for login.
### Fix #4 — Admin "User mappings" UI can't set `work_email`
**Files:** the Admin → User mappings UI (`ui/`) + `agent/dashboard/user_mappings.py` (`upsert_mapping` at line 257 takes `work_email` but the UI form doesn't expose it).
Mappings can't be fully created from the dashboard → must write the Store directly: namespace `["user_mappings"]`, key = **lowercased login**, value `{github_login, work_email}` (matching `_index_record` at `user_mappings.py:73-84`, which keys `_by_login` on `login.lower()`).
- **Fix (Phase B):** add the `work_email` field to the Admin mappings form so mappings are fully creatable from the UI.
### Fix #5 — GitHub webhook path doesn't refresh the user-mapping cache (multi-replica break)
**Files:** `agent/webapp.py:3052` (GitHub path) vs `agent/webapp.py:1091` (Slack path).
The Slack path refreshes before lookup:
```python
# agent/webapp.py:1089-1093 (Slack)
await refresh_user_mapping_cache()
...
```
The GitHub path does **not** — it calls `email = await email_for_login(github_login)` (`webapp.py:3052`, again at `:3331`) cold. The cache (`user_mappings.py` `_ensure_cache_loaded`, line 197) is **one-shot per process** (`_cache_loaded` flag). On self-host single-process this was fine; on managed's **multi-replica autoscaling**, a freshly-added mapping isn't seen by a replica whose cache loaded earlier — until restart.
- **Fix:** refresh-before-lookup on the GitHub path (mirror the Slack path), or add a TTL / cross-replica invalidation to the cache.
- **Generalize (Phase B5):** audit ALL in-process caches for the single-process → multi-replica assumption — `SANDBOX_BACKENDS` dict (`agent/utils/sandbox_state.py`), `_THREAD_RUN_LOCKS` (`thread_ops.py:18`), `_by_login`/`_by_email`/`_by_slack_id` (`user_mappings.py:67-69`). Sandbox affinity is already thread-keyed + persisted in thread metadata (`sandbox_id`), so it's the cache/lock state that needs the multi-replica review.
### Fix #6 — Slow custom-app import (~8s startup)
**Symptom:** "exceeded expected startup time" → risks the deployment being marked unhealthy / slow to scale out.
- **Fix (Phase B6):** lazy imports / reduce import-time work in `agent/webapp.py` and the graph factories. Profile with `FF_PROFILE_IMPORTS` (the import-profiling flag) to find the heavy modules.
---
## 6. CD / Ops Changes
### Retires (bespoke pipeline)
- GitHub Actions **build → S3 releases → SSM doc → `deploy.sh` → health-gate** (`.github/scripts/`, `deploy/ami/deploy.sh`)
- `roll-box.sh` / `publish-and-deploy.sh` / `rollback.sh`
- AMI baking (packer `deploy/ami/open-swe-base.pkr.hcl`) + `cdk.context.json` AMI pin
- `cd-infra.yml` CDK deploys + OIDC bootstrap qualifiers
- `fetch-config.sh` / `seed_store.sh` (boot-time config materialization + Store reseed)
### Replaced by git-connected PaaS
- **LangGraph Cloud** auto-builds a revision on push (first-party zero-downtime + revision rollback).
- **Vercel** builds / atomic-deploys / instant-rollback + per-PR previews (**net-new** capability the AWS stack never had).
### Stays
- **CI** (lint / format / unit / Playwright E2E) — still matters and still gates merges. `make lint`, `make test`.
### Gate re-homing
- The **dev → main promotion gate** (`check-dev-green.sh` + protected-`main` ruleset 18238334) **re-homes**: the **dev deployment tracks `dev`**, the **prod deployment tracks `main`**, so the gate governs exactly what reaches prod.
- Today's promotion machinery: `promote-dev-to-prod.yml` + the `seahaven-promotion` GitHub App (app_id 4170147, in ruleset 18238334 bypass_actors) FF-pushes `dev → main`. Under managed, a push to `main` is what triggers the prod build — so the promotion gate IS the prod deploy gate.
### Preserve the PROD manual-approval gate (LOAD-BEARING)
Today the `prod` GitHub Environment (required reviewer `amoussa1229`) is the manual approval. Git-connected auto-deploy removes the CD job that consulted that Environment, so the approval must be re-established explicitly:
- **Mechanism options (OPEN DECISION §10):** protected `main` (only the promotion App can FF-push, and that push is itself the gate) AND/OR the platform's "require manual promotion to production" if LangGraph Cloud exposes it.
- **Norm shift:** deploy-then-merge → **merge-to-deploy**. Lean on the dev deployment + Vercel previews to verify *before* promoting `dev → main`. (This inverts the Sea Haven handbook deploy-then-merge default — call it out in the README + handbook note.)
---
## 7. AWS Decommission — Retire vs Keep
### RETIRE (becomes vestigial once managed prod is live + soaked)
| Item | Path / resource |
|---|---|
| AMI bake + cloud-init | `deploy/ami/` (`open-swe-base.pkr.hcl`, `provision.sh`, `user-data.sh`, `deploy.sh`, `templates/`) |
| Self-host boot/config scripts | `deploy/seahaven/fetch-config.sh`, `seed_store.sh`, `put-config.sh`, `nginx/openswe.conf`, `systemd/open-swe.service`, `aegra/` (already deferred) |
| CDK box/ALB/AMI stacks | `infra/lib/open-swe-stack.ts`, `constructs/app-service.ts`, `assets-bucket.ts`, `ami-cache.ts`, `instance-role.ts`, `github-deploy-roles.ts` |
| EC2 instances | prod `i-08a729e50779c4b07` (t4g.large), dev `i-0a3bb8e0ddd36c29b` |
| S3 release buckets | `open-swe-dev-assets`, `open-swe-prod-assets` |
| SSM deploy docs | `open-swe-dev-deploy`, `open-swe-prod-deploy` |
| ALB plumbing | `open-swe-prod-tg` / `open-swe-dev-tg` target groups + the `seahaven-com` ALB host/path listener rules for openswe/hooks |
| Per-env bootstrap qualifier | dev `oswedev` (`CDKToolkit-oswedev`) |
| RDS / Redis durable plan | **never built** — fully dropped (managed provides durability) |
| Bespoke CD workflows | `cd-infra.yml`, `build-artifacts.yml`, `roll-box.sh`/`publish-and-deploy.sh`/`rollback.sh` |
### KEEP
| Item | Why |
|---|---|
| **GitHub App `seahaven-openswe`** (App 4146115 / Install 142615168) | unchanged auth + webhook source; only the webhook/OAuth URLs repoint |
| **LangSmith sandbox setup** incl. `DEFAULT_SANDBOX_SNAPSHOT_ID` (`dc36e509-d4d2-4efc-8a4e-61f74f3446ec`) | the only sandbox with working in-sandbox git/gh auth (`_configure_github_proxy`); managed default `SANDBOX_TYPE=langsmith` uses it |
| **Secrets Manager `open-swe-{dev,prod}/*` values** | the **source you copy env FROM**, at least transitionally — do NOT force-delete populated shells until managed env is fully cut over and verified |
| **DNS** | repoint `openswe.seahaven.com` → Vercel; decide `hooks.seahaven.com` → `*.langgraph.app` or a CNAME (§10) |
| **`seahaven-promotion` App** + main/dev rulesets (18238334 / 18238542) | still govern `dev → main` promotion = the prod deploy gate |
| **`evals/` + `JUDGE_ANTHROPIC_API_KEY`** | eval pipeline unaffected (separate from runtime) |
**Phased teardown rule:** remove AWS infra ONLY after managed prod is proven + soaked (Phase D5). The exported ALB JSON (E1) is the rollback source for the hand-built listener rules.
---
## 8. Cost
| Line item | Monthly |
|---|---|
| Prod deployment uptime ($0.0036/min, always-on) | ≈ $155 |
| Runs ($0.005/run) | usage-based |
| Traces above 10k/mo | pay-as-you-go |
| LangSmith **Plus** seat ($39) | already paid (not incremental) |
| Dev deployment | **$0** (1 free on Plus, preemptible) |
| **Incremental total** | **≈ $160/mo** |
Compare to the retired AWS stack: 2× EC2 (t4g.large prod + t4g.medium/large dev), 2× S3 buckets, ALB share, NAT egress, plus the **never-built** RDS/Redis durable tier the self-host path would have added. Managed nets out roughly cost-neutral-to-cheaper while *adding* durability + previews. **TODO: confirm current AWS run-rate for an apples-to-apples delta** (not verified in repo).
---
## 9. Sea Haven Gates & Obligations
- [ ] **`/sh-plan-review`** (GPT-4.1 `cross_reviewer`) on THIS plan **BEFORE prod cutover** (Phase C gate). ⚠️ `run.py` misroutes reviewer-framed prompts to the no-op `done` route — call `models.get_cross_reviewer()` directly (load orchestrator `.env`, `ChatOpenAI("gpt-4.1")`) per `feedback_orchestrator_usage`.
- [ ] **`/sh-security-review`** on the sensitive surface: this migration touches **auth/webhook signature verification** (the custom `http.app` is now publicly reachable on `*.langgraph.app` with no platform gate in front) and **secrets handling** (secrets move from Secrets Manager to the Deployment/Vercel env). Required, not opt-in — resolve confirmed critical/high before cutover.
- [ ] **GPT-4.1 cross-family review** if the Bedrock-auth fix changes IAM (new Bedrock-scoped IAM user + policy = IAM change → mandatory cross-review).
- [ ] **Confluence** — rework "AWS Architecture Map" (**1540098**) + open-swe child page (**26116098**): replace the RDS/EC2/ALB subgraph with an external-services view (LangGraph Cloud + Vercel + retained GitHub App + LangSmith sandbox). Do it in the same conversation as the cutover, not as a follow-up.
- [ ] **README + INSTALLATION** updated (INSTALLATION §10 already canonical; align Sea Haven specifics + the merge-to-deploy norm shift).
- [ ] **Memories** — update `project_open_swe_migration.md` + `reference_open_swe_deployment.md`.
---
## 10. Decisions — RESOLVED (2026-06-29, Adam)
1. **Bedrock auth on managed → STATIC AWS KEYS.** Create a **Bedrock-scoped IAM user**; set `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` / `AWS_REGION=us-east-1` in the deployment env; reuse the #62 instance-role Bedrock policy (`bedrock:InvokeModel[WithResponseStream]` on the `us.anthropic.claude-opus-4-8` inference-profile ARN + per-region foundation-model ARNs in us-east-1/2 + us-west-2). Caveat: long-lived keys → rotation reminder + keep the policy tightly scoped. **New IAM user+policy = mandatory GPT-4.1 IAM cross-review + `/sh-security-review` before merge.**
2. **Webhook domain → RAW `*.langgraph.app`** (the "easier" path — no CNAME/DNS/TLS). Point GitHub/Slack/Linear webhook URLs straight at the deployment. Dashboard gets a **Vercel custom domain** (e.g. `openswe.seahaven.com`) since that's trivial on Vercel.
3. **Prod approval gate → GATED PROMOTE-TO-MAIN WORKFLOW.** Re-home `promote-dev-to-prod.yml`: gate the promote job on the **`prod` GitHub Environment reviewer (`amoussa1229`)**; on approval it FF-pushes `dev → main`; the push triggers the managed prod build. Preserves today's manual-approval semantics — approval is on the promote job, the merge to `main` is the deploy trigger.
4. **Cost/trace monitoring → YES, in LangSmith** (usage/trace budget + alert in the LangSmith workspace).
5. **Dev/prod home → CONFIRMED:** prod = LangChain managed (LangGraph Cloud) + Vercel UI; dev = managed **free** Dev tier (same runtime as prod).
6. ~~#62 merge status~~ — RESOLVED (merged + deployed).
## 10b. Execution approach — 2-PR SPLIT (Adam's call, 2026-06-29, post-plan-review)
The plan-review BLOCKed a literal single PR and recommended a 3-PR minimum (isolate cache fix #5 + the promotion-gate workflow). Adam chose a **2-PR split** — isolating the prod-deploy-gate change (highest risk to prod), bundling the rest:
- **PR1 — promotion-gate workflow + ruleset.** Re-home `promote-dev-to-prod.yml` to the gated promote-to-main (decision 3), and lock `main` so ONLY the promote App can push (BLOCK 3). Isolated because it changes *how prod deploys* — must be independently reviewable/revertible.
- **PR2 — the 6 code fixes (§5) + `ui/vercel.json` prod repoint + README/INSTALLATION.** (Accepted deviation from the reviewer's 3-PR rec: the multi-replica cache fix #5 stays bundled in PR2 rather than isolated — Adam's call.)
- **NOT in either PR (rollback safety):** AWS **infra-code deletion** (`deploy/ami`, `deploy/seahaven`, `infra/` self-host stack) + `cdk destroy` stay a **POST-SOAK cleanup (Phase E)**.
- **Not PR content (operational, sequenced around the merges):** LangGraph Cloud prod deployment + env, Vercel prod + custom domain, webhook/OAuth repoint, Bedrock IAM user, LangSmith budget, Confluence/memory.
## 10c. Plan-review resolution (GPT-4.1 cross_reviewer, 2026-06-29) — verdict REQUEST CHANGES, all BLOCKs addressed
- **BLOCK 1 (single PR)** → addressed via the **2-PR split** above (conscious deviation: cache fix #5 not isolated — Adam accepted).
- **BLOCK 2 (raw LangGraph API public?)** → **EMPIRICALLY RESOLVED.** Unauthenticated probes of the dev deployment returned **403 "Missing authentication headers"** for `/threads/search`, `/assistants/search`, `/store/items`; `/ok`=200, `/dashboard/api/me`=401. The platform gates the raw control-plane API. **Keep as a pre-prod gate (Phase D4):** re-probe the PROD deployment URL before cutover.
- **BLOCK 3 (prod gate enforcement)** → PR1 must **lock `main`** so only the promote App can push (verify ruleset 18238334: no direct-push path, PR-required, promote App is the sole FF bypass). Any non-gated push to `main` would auto-deploy prod.
- **BLOCK 4 (secrets posture)** → add to Phase B/E: **rotate** all secrets after migrating them to the platform env stores and BEFORE deleting from AWS; **document who can read/write** the LangGraph Cloud + Vercel env (access audit); delete AWS shells only post-cutover; no dual-homed/stale secrets; record in memory.
- **FIX (rollback integrity)** → Phase E checklist: "no IaC/DNS/config the rollback needs is altered in PR1/PR2 or during cutover."
- **FIX (gates resolved-before-merge)** → make explicit: do NOT merge PR1/PR2 until the GPT-4.1 IAM cross-review (Bedrock IAM user) + `/sh-security-review` (auth/webhook/secrets surface) are **resolved with no critical/high**; Confluence (1540098 + 26116098) + memory updated in the cutover conversation.
- **NITs/QUESTIONs** → carried into §8/§10 TODOs (cost delta, `FF_PROFILE_IMPORTS`, platform feature availability, key-rotation owner, env-store access logging, cache race-review under autoscaling, atomic webhook repoint).
---
## 11. Rollback Story
**Pre-cutover:** the self-host `langgraph up` + RDS plan (RDS+Redis+Docker; `/sh-plan-review`'d to APPROVE-after-revision this session) is the documented **fallback** if managed is rejected before cutover. Its durable design wins (env-scoped RDS physical names; `Credentials.fromGeneratedSecret({secretName:"open-swe-<env>/rds-credentials"})` under the existing instance-role secret prefix → no new IAM; DESTROY-on-rollback RDS-managed secret; derive `DATABASE_URI` in fetch-config) are captured in the project memory.
**Post-cutover (managed is live, AWS still standing during soak):** if managed prod fails, **rollback = repoint webhooks + DNS back** to the AWS stack:
- GitHub App / Slack / Linear webhook URLs → `hooks.seahaven.com`
- `openswe.seahaven.com` DNS → the `seahaven-com` ALB (restore from the E1 export)
- `ui/vercel.json` rewrite → the AWS dashboard origin (or stop using Vercel)
- Do NOT run Phase E teardown until the soak passes — the AWS stack IS the rollback target.
**Revision-level rollback (within managed):** LangGraph Cloud revision rollback (backend) + Vercel instant rollback (UI) cover bad deploys without leaving the platform.
---
## TODOs flagged (couldn't verify in repo)
- ~~#62 merge state~~ — RESOLVED: merged to `dev` + deployed (the draft read a stale local checkout; `origin/dev` and the deployment `git_ref_sha a4ed19ba` confirm Bedrock+Fireworks).
- **Current AWS run-rate** for the cost delta (§8) — not derivable from repo.
- **LangGraph Cloud custom-domain + "manual promotion to production" feature availability** (§10.2, §10.3) — platform features, verify in the LangGraph Cloud console/docs at execution time.
- **`FF_PROFILE_IMPORTS`** exact flag name/usage (§5 fix #6) — referenced from session context; confirm the flag exists in `agent/` before relying on it.
- Exact webhook path on `*.langgraph.app` (whether the custom `http.app` mounts at root so `/webhooks/github` is reachable as-is) — proven reachable on the dev spike; re-verify the exact path on prod.

View file

@ -1,167 +0,0 @@
# Open SWE base AMI (T8)
Packer recipe + first-boot user-data for the single EC2 instance per env
(`open-swe-dev` / `open-swe-prod`) in the Open SWE → AWS migration. Builds an
**ARM64 (Graviton) Ubuntu 24.04 LTS** base AMI and provisions the box on first
boot with the stock `langgraph dev` runtime, nginx, and the CloudWatch agent.
The architecture is locked in the repo `TODO.md` ("Architecture (locked)"): ONE
EC2 ARM64 (~t4g.large) instance per env, `seahaven-vpc` **private subnet + NAT**,
inbound **only from the ALB SG**. Runtime is **stock `langgraph dev`** (in-memory
store, `--no-reload`) + nginx + systemd. The box has **no git auth** — it pulls
its deploy artifact from S3 via the instance role. The SPA build runs in GitHub
Actions (T7), **not** on the box, so the old 8 GB-swapfile OOM hack is gone.
## Files
| Path | Purpose |
|---|---|
| `open-swe-base.pkr.hcl` | Packer template (HCL2). Latest Canonical 24.04 arm64 source → base AMI. |
| `scripts/provision.sh` | Packer provisioner. System packages, uv+py3.12, node+bun, service user, stages templates. |
| `user-data.sh` | First-boot provisioning (S3 artifact pull, render templates, CW agent, start services). |
| `templates/open-swe.service` | systemd unit TEMPLATE (`@@tokens@@` rendered at boot). |
| `templates/open-swe.nginx.conf` | nginx site TEMPLATE (dashboard SPA + scoped `/dashboard/api/` proxy). |
| `templates/amazon-cloudwatch-agent.json` | CW agent config TEMPLATE — **30-day log retention**. |
`deploy/seahaven/fetch-config.sh` and `deploy/seahaven/seed_store.sh` are owned by
the parallel T10 work and ship **inside the app artifact**; this AMI wires them in
but does not author them (see "Integration contract" below).
## Build the AMI
```bash
cd deploy/ami
packer init .
packer fmt -check .
packer validate -var aws_region=us-east-1 open-swe-base.pkr.hcl
packer build open-swe-base.pkr.hcl
```
Builds in account **328440206208 / us-east-1**. Source = latest Canonical Ubuntu
24.04 (Noble) **arm64** AMI (`source_ami_filter`, owner `099720109477`). Build host
is `t4g.medium` (ARM64). The output AMI is tagged:
```
Name=open-swe-base-arm64 Purpose=open-swe-runtime-base ManagedBy=packer
```
Pinned versions live in the template `variable` defaults (`uv_version`,
`python_version`, `node_major`, the CW-agent / awscli URLs) and the
`required_plugins` block (`amazon` 1.3.6) — bump deliberately.
## AMI → `cdk.context.json` pinning contract
The CDK stacks in `/infra` (owned by T3/T12) consume the AMI **by id, pinned in the
committed `infra/cdk.context.json`** — they never resolve "latest" at synth time.
This is the EBS/AMI-fix discipline: an uncached `MachineImage.lookup` resolves a new
AMI on every deploy and silently triggers instance replacement.
Contract (CDK side does the wiring; this is the handshake):
1. `packer build` prints the new AMI id (and tags it `open-swe-base-arm64`).
2. CDK looks the AMI up with **`cachedInContext: true`** (e.g.
`MachineImage.lookup({ name: "open-swe-base-arm64-*", owners: ["328440206208"], cachedInContext: true })`),
which writes the resolved id into `infra/cdk.context.json`.
3. **`infra/cdk.context.json` is committed.** From then on every synth/deploy uses
the pinned id — no surprise replacement when a newer AMI exists.
4. To adopt a new AMI: `cdk context --reset <ami-lookup-key>` (or edit the pinned
value), commit the change, and review the cdk-diff — the PR will show
"requires replacement", which is the intended, visible signal.
Record the built AMI id in project memory (`project_open_swe_migration`) per the
"memory updated for AMI id" build criterion.
## `userDataCausesReplacement` rationale
`user-data.sh` is **provisioning-only** — it runs once at first boot and never
carries durable runtime config. CDK sets **`userDataCausesReplacement: true`** so
that any change to it is a deliberate, diff-visible instance replacement rather than
a no-op edit that drifts from the running box. Durable runtime config is fetched
**fresh on every service start** by `fetch-config.sh` (ExecStartPre) — changing a
secret or SSM value needs only a `systemctl restart open-swe.service`, not a
replacement.
## EBS discipline (binding — `feedback_inline_ebs_volumes`)
**The box holds no durable state of its own:**
| State | Lives in | On replacement |
|---|---|---|
| secrets / config | Secrets Manager + SSM → tmpfs `.env` | re-fetched at boot |
| app code + SPA | S3 `open-swe-<env>-assets` | re-pulled at boot |
| store (team_settings, user_mappings) | reseeded by `seed_store.sh` | re-seeded at boot |
| logs | CloudWatch (30-day) — **not** a CFN resource in the stack | survive replacement |
→ **No local-only durable state ⇒ no standalone RETAIN volume is needed.** The root
volume is disposable; there is intentionally no inline data `blockDevices` to lose.
**Even so, snapshot before any replacing deploy.** Per the operational guard, before
merging/deploying any change that REPLACES the instance (`userDataCausesReplacement`,
AMI bump, instance-type change):
1. Enumerate the instance's volumes and assert **"no local-only durable state"**
(the table above is the checklist).
2. Take an **EBS snapshot of the root volume and WAIT for `state=completed`** before
letting the deploy proceed. Keep it as insurance; delete after a grace period.
3. Confirm the CloudWatch log groups are **not** CFN-managed in the stack so history
survives; re-verify history after the new instance is healthy.
cdk-diff-on-PR must flag "requires replacement" at review time (T8/T12 acceptance
criterion). This is the enforced version — not just an assertion in the runbook.
## Integration contract (T10 — `fetch-config.sh` + `seed_store.sh`)
Both ship in the app artifact under `deploy/seahaven/` and are wired into the unit:
- **`fetch-config.sh`** (ExecStartPre, runs as `openswe`): reads `/etc/open-swe/boot.env`
(`OPENSWE_ENV`, `AWS_REGION`, `SECRETS_PREFIX=open-swe-<env>`, `SSM_PREFIX=/open-swe-<env>`,
`ENV_FILE=/run/open-swe/.env`), pulls Secrets Manager `open-swe-<env>/*` + SSM
`/open-swe-<env>/*`, and writes:
- `/run/open-swe/.env` (**0600, tmpfs**, secret-bearing app env incl. the multiline
GitHub App PEM) — loaded by langgraph/dotenv via the `${APP_DIR}/.env` symlink.
- `/run/open-swe/seed.env` (**0600, tmpfs**, simple `OPENSWE_*` vars only:
`OPENSWE_DEFAULT_REPO`, `OPENSWE_OWNER_LOGIN`, `OPENSWE_OWNER_EMAIL`, model ids) —
loaded by systemd `EnvironmentFile` so `seed_store.sh` (ExecStartPost) has them.
- It must **fail-fast** (non-zero exit) if any required value is missing, so the
unit never starts half-configured.
- **`seed_store.sh`** (ExecStartPost): existing script, reseeds `team_settings/default`
+ `user_mappings/<login>` into the in-memory store after each start.
## Smoke-boot checklist (after first boot)
SSM Session Manager onto the instance (no public SSH — private subnet) and verify:
- [ ] `cloud-init status --wait` → `done`; `/var/log/open-swe-user-data.log` ends with
"user-data done" and shows the S3 pulls + service starts.
- [ ] `systemctl is-active open-swe.service` → `active`. (If it failed, check
`ExecStartPre`/`fetch-config.sh` — fail-fast means missing config = failed unit.)
- [ ] **fetch-config fail-fast works:** `/run/open-swe/.env` exists, owner `openswe`,
mode `0600`, on tmpfs (`findmnt /run/open-swe`); `seed.env` present.
- [ ] `curl -fsS http://127.0.0.1:2024/ok` → `200` (raw LangGraph health).
- [ ] `systemctl is-active nginx` → `active`; `curl -fsS http://127.0.0.1/healthz` →
`200`; `curl -s http://127.0.0.1/threads` returns the SPA shell, **not** JSON
(proves the agent API is not proxied — the security boundary holds).
- [ ] `seed_store: done` in the journal / app.log (store reseeded).
- [ ] CloudWatch: log groups `/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`
exist with **30-day** retention and are receiving events.
- [ ] **No swapfile** (`swapon --show` empty) — the on-box SPA build is gone.
- [ ] From the ALB only: dashboard host serves the SPA; `hooks` host reaches
`/webhooks/*` on :2024 and nothing else (raw API paths hit the ALB default, not
the box).
## Assumptions
- **Artifact bucket** `open-swe-<env>-assets` (T7), with objects
`${ARTIFACT_PREFIX}/app.tar.gz` (Python app incl. `deploy/seahaven/` and a prebuilt
arm64 `.venv`) and `${ARTIFACT_PREFIX}/spa.tar.gz` (built SPA → `/var/www/open-swe`).
`ARTIFACT_PREFIX` defaults to `releases/latest`; CDK renders the concrete value.
- **Instance role** (defined in `/infra`, least-privilege per T4/T12) grants:
`s3:GetObject` on `open-swe-<env>-assets/*`; `secretsmanager:GetSecretValue` on
`open-swe-<env>/*`; `ssm:GetParameter(s)`/`GetParametersByPath` on `/open-swe-<env>/*`;
`logs:*` for the CW agent log groups + `cloudwatch:PutMetricData`; SSM Session
Manager (`ssm:UpdateInstanceInformation`, `ssmmessages:*`) for shell access.
- **CDK substitutes** the `@@OPENSWE_ENV@@`, `@@ASSETS_BUCKET@@`, `@@SERVER_NAME@@`,
`@@ARTIFACT_PREFIX@@` tokens in `user-data.sh` when rendering the launch template.
- `:2024` binds `0.0.0.0` so the ALB hooks target group can reach `/webhooks/*`; it is
reachable only from the ALB SG (private subnet, SG-scoped inbound). The raw API is
never internet-exposed — the ALB hooks rule is path-scoped to `/webhooks/*`.

View file

@ -1,102 +0,0 @@
#!/usr/bin/env bash
# Open SWE app deploy — RE-RUNNABLE (first boot + every subsequent release).
#
# Pulls the current release from S3 (open-swe-<env>-assets), builds the venv
# natively on the box, and restarts the service. This is the SINGLE source of the
# app-deploy procedure; it runs in two places:
#
# 1. first boot — user-data.sh decodes this script to /opt/open-swe/bin and
# calls it ONCE (non-fatal: if no release is published yet,
# nginx is already up and the box waits for the first deploy).
# 2. every release — the `open-swe-<env>-deploy` SSM document (CI fires it after
# uploading app.tar.gz / spa.tar.gz) runs this same script.
#
# It deploys CODE + STATIC ASSETS only. Secrets/config are NOT fetched here: the
# systemd unit's ExecStartPre=fetch-config.sh materializes the tmpfs .env on every
# (re)start, fail-fast — so `systemctl restart` below is what reloads config too.
#
# Contract:
# app.tar.gz = the Python source tree (pyproject.toml + uv.lock + agent/ +
# deploy/ + langgraph.json + README.md, NO ui/, NO .venv). The venv
# is built HERE with `uv sync` so it is native ARM64 and lives at
# the real runtime path (no cross-built / non-relocatable venv).
# spa.tar.gz = the built dashboard SPA (vite output: _shell.html + assets),
# extracted to the nginx web root.
set -euo pipefail
exec > >(tee -a /var/log/open-swe/deploy.log) 2>&1
echo "==> open-swe deploy start $(date -u +%FT%TZ)"
# Non-secret pointers written by user-data.sh (env, region, bucket, artifact prefix).
# shellcheck disable=SC1091
. /etc/open-swe/boot.env
export AWS_DEFAULT_REGION="${AWS_REGION:?boot.env missing AWS_REGION}"
: "${ASSETS_BUCKET:?boot.env missing ASSETS_BUCKET}"
: "${ARTIFACT_PREFIX:?boot.env missing ARTIFACT_PREFIX}"
# Fixed layout — must match provision.sh + user-data.sh + the templates.
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
WWW_ROOT="/var/www/open-swe"
ENV_FILE="/run/open-swe/.env"
UV_BIN="/usr/local/bin/uv"
UV_PYTHON_INSTALL_DIR="/opt/uv/python" # where provision.sh pre-installed py3.12
SERVICE_HOME="/opt/open-swe"
# Benign-vs-failure distinction: on a brand-new env no release is published yet.
# Treat "app.tar.gz absent in S3" as a benign no-op (exit 0) so first boot is not a
# scary failure; ONCE a release exists, any later step failing is loud (set -e).
if ! aws s3 ls "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" >/dev/null 2>&1; then
echo "==> no release published at s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/ yet — nothing to deploy"
exit 0
fi
echo "==> pull release from s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/"
tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/app.tar.gz" "${tmp}/app.tar.gz"
aws s3 cp "s3://${ASSETS_BUCKET}/${ARTIFACT_PREFIX}/spa.tar.gz" "${tmp}/spa.tar.gz"
# Replace app source + SPA atomically-ish: clear the dirs (drops files removed in
# this release) then extract. The venv is rebuilt below, so wiping .venv too is
# fine — uv's cache (in the service home) makes the rebuild fast.
#
# Hardening: deploy.sh runs as root, so extract with --no-same-owner
# --no-same-permissions — files take root:root + umask perms (NOT the archive's
# uid/mode), so a tarball cannot land a setuid/setgid binary or a foreign-owned
# file; the chown -R below then hands the tree to the service user. (GNU tar also
# refuses `..`-escaping members by default.) Defense-in-depth: the only writer of
# this bucket is the CI OIDC app role, but the box never trusts the archive's
# ownership/mode regardless.
echo "==> install app source -> ${APP_DIR}"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$APP_DIR" "$WWW_ROOT"
find "$APP_DIR" -mindepth 1 -delete
find "$WWW_ROOT" -mindepth 1 -delete
tar --no-same-owner --no-same-permissions -xzf "${tmp}/app.tar.gz" -C "$APP_DIR"
tar --no-same-owner --no-same-permissions -xzf "${tmp}/spa.tar.gz" -C "$WWW_ROOT"
# langgraph reads ./.env from WorkingDirectory; point it at the tmpfs file the
# systemd ExecStartPre materializes.
ln -sfn "$ENV_FILE" "${APP_DIR}/.env"
chown -R "$SERVICE_USER":"$SERVICE_USER" "$APP_DIR" "$WWW_ROOT"
echo "==> build venv natively (uv sync --frozen --no-dev)"
# Run as the service user so the venv + uv cache are owned by it. Pin the
# pre-baked interpreter dir so uv never reaches out to download Python at deploy.
cd "$APP_DIR"
sudo -u "$SERVICE_USER" env \
HOME="$SERVICE_HOME" \
UV_PYTHON_INSTALL_DIR="$UV_PYTHON_INSTALL_DIR" \
UV_CACHE_DIR="${SERVICE_HOME}/.cache/uv" \
"$UV_BIN" sync --frozen --no-dev
echo "==> restart open-swe.service + reload nginx"
# ExecStartPre=fetch-config.sh fails-fast if secrets/config are missing, so a
# restart here surfaces a bad config as a failed unit (non-zero exit below).
systemctl restart open-swe.service
nginx -t && systemctl reload nginx
if systemctl is-active --quiet open-swe.service; then
echo "==> open-swe deploy OK $(date -u +%FT%TZ)"
else
echo "!! open-swe.service is not active after deploy (check fetch-config/secrets)"
exit 1
fi

View file

@ -1,143 +0,0 @@
# Open SWE base AMI — ARM64 (Graviton) Ubuntu 24.04 LTS.
#
# Builds the immutable base image for the single EC2 instance per env
# (open-swe-dev / open-swe-prod) in seahaven-vpc. The image bakes the runtime
# (uv + Python 3.12, nginx, awscli v2, CloudWatch agent) and the service-user /
# systemd / nginx TEMPLATES. It bakes NO secrets and NO env-specific values —
# those are materialized at first boot by user-data + deploy/seahaven/fetch-config.sh
# (Secrets Manager + SSM -> root-only tmpfs .env, fail-fast).
#
# Build: packer init . && packer build open-swe-base.pkr.hcl
# The resulting AMI id is pinned in infra/cdk.context.json (CDK cachedInContext:true);
# see README.md "AMI -> cdk.context.json pinning contract".
packer {
required_version = ">= 1.11.0, < 2.0.0"
required_plugins {
amazon = {
source = "github.com/hashicorp/amazon"
version = "1.3.6"
}
}
}
variable "aws_region" {
type = string
default = "us-east-1"
}
variable "instance_type" {
type = string
default = "t4g.medium" # ARM64 (Graviton) build host; runtime instances are ~t4g.large
}
variable "ami_name_prefix" {
type = string
default = "open-swe-base-arm64"
}
# Versions baked into the image. Pin and bump deliberately.
variable "python_version" {
type = string
default = "3.12"
}
variable "node_major" {
type = string
default = "24"
}
variable "uv_version" {
type = string
default = "0.11.24"
}
variable "cloudwatch_agent_deb_url" {
type = string
default = "https://amazoncloudwatch-agent.s3.amazonaws.com/ubuntu/arm64/latest/amazon-cloudwatch-agent.deb"
}
variable "awscli_zip_url" {
type = string
default = "https://awscli.amazonaws.com/awscli-exe-linux-aarch64.zip"
}
locals {
timestamp = formatdate("YYYYMMDD-hhmmss", timestamp())
}
# Latest Canonical Ubuntu 24.04 (Noble) arm64 server image.
source "amazon-ebs" "open-swe" {
region = var.aws_region
instance_type = var.instance_type
ssh_username = "ubuntu"
ami_name = "${var.ami_name_prefix}-${local.timestamp}"
# ASCII only — AWS rejects non-ASCII in the AMI Description attribute.
ami_description = "Open SWE base - Ubuntu 24.04 arm64 + uv/py3.12 + nginx + CW agent (templates only, no secrets)"
source_ami_filter {
filters = {
name = "ubuntu/images/hvm-ssd*/ubuntu-noble-24.04-arm64-server-*"
architecture = "arm64"
root-device-type = "ebs"
virtualization-type = "hvm"
}
owners = ["099720109477"] # Canonical
most_recent = true
}
# IMDSv2 required on the build host.
metadata_options {
http_endpoint = "enabled"
http_tokens = "required"
http_put_response_hop_limit = 1
}
# gp3 root, encrypted. Runtime root size is set by CDK; this is just the build host.
launch_block_device_mappings {
device_name = "/dev/sda1"
volume_size = 20
volume_type = "gp3"
encrypted = true
delete_on_termination = true
}
tags = {
Name = "open-swe-base-arm64"
Purpose = "open-swe-runtime-base"
ManagedBy = "packer"
}
}
build {
name = "open-swe-base"
sources = ["source.amazon-ebs.open-swe"]
# Stage the boot-time templates into the image. The destination dir must exist
# BEFORE a trailing-slash (contents-only) file upload — packer's file provisioner
# does not create it, and uploading the directory itself trips scp ("Is a
# directory"). So mkdir first, then upload the contents into it.
provisioner "shell" {
inline = ["mkdir -p /tmp/open-swe-templates"]
}
provisioner "file" {
source = "${path.root}/templates/"
destination = "/tmp/open-swe-templates"
}
provisioner "shell" {
environment_vars = [
"PYTHON_VERSION=${var.python_version}",
"NODE_MAJOR=${var.node_major}",
"UV_VERSION=${var.uv_version}",
"CLOUDWATCH_AGENT_DEB_URL=${var.cloudwatch_agent_deb_url}",
"AWSCLI_ZIP_URL=${var.awscli_zip_url}",
]
# {{ .Vars }} MUST be included or the environment_vars above never reach the
# script (provision.sh runs under `set -u` and fails on the first reference).
execute_command = "chmod +x {{ .Path }}; {{ .Vars }} sudo -E bash '{{ .Path }}'"
script = "${path.root}/scripts/provision.sh"
}
}

View file

@ -1,105 +0,0 @@
#!/usr/bin/env bash
# Packer provisioner for the Open SWE base AMI (ARM64 Ubuntu 24.04).
#
# Bakes the runtime + boot-time templates ONLY. No secrets, no env-specific
# values. Everything env-specific is materialized at first boot by user-data.sh
# + deploy/seahaven/fetch-config.sh.
set -euo pipefail
PYTHON_VERSION="${PYTHON_VERSION:-3.12}"
NODE_MAJOR="${NODE_MAJOR:-24}"
UV_VERSION="${UV_VERSION:-0.11.24}"
CLOUDWATCH_AGENT_DEB_URL="${CLOUDWATCH_AGENT_DEB_URL:?}"
AWSCLI_ZIP_URL="${AWSCLI_ZIP_URL:?}"
# Layout (must match user-data.sh and the templates).
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
SERVICE_HOME="/opt/open-swe"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
LOG_DIR="/var/log/open-swe"
UV_BIN="/usr/local/bin/uv"
export DEBIAN_FRONTEND=noninteractive
echo "==> apt base packages"
apt-get update -y
apt-get upgrade -y
apt-get install -y --no-install-recommends \
nginx jq curl unzip ca-certificates gnupg lsb-release \
build-essential pkg-config git acl
echo "==> awscli v2 (aarch64)"
tmp="$(mktemp -d)"
curl -fsSL "$AWSCLI_ZIP_URL" -o "$tmp/awscliv2.zip"
unzip -q "$tmp/awscliv2.zip" -d "$tmp"
"$tmp/aws/install" --update
rm -rf "$tmp"
aws --version
echo "==> CloudWatch agent (arm64)"
tmp="$(mktemp -d)"
curl -fsSL "$CLOUDWATCH_AGENT_DEB_URL" -o "$tmp/amazon-cloudwatch-agent.deb"
dpkg -i -E "$tmp/amazon-cloudwatch-agent.deb"
rm -rf "$tmp"
# Do NOT enable/start the agent during the build; user-data fetches its config
# (with env-specific log-group names + 30-day retention) and starts it at boot.
systemctl disable amazon-cloudwatch-agent.service || true
echo "==> uv ${UV_VERSION} + Python ${PYTHON_VERSION} (system-wide)"
export UV_INSTALL_DIR=/usr/local/bin
curl -fsSL "https://astral.sh/uv/${UV_VERSION}/install.sh" | env UV_NO_MODIFY_PATH=1 sh
"$UV_BIN" --version
# Pre-install the interpreter so the box never reaches out at boot to build a venv.
UV_PYTHON_INSTALL_DIR=/opt/uv/python "$UV_BIN" python install "$PYTHON_VERSION"
echo "==> node ${NODE_MAJOR} + bun (build-time UI tooling only; the SPA is built in CI)"
curl -fsSL "https://deb.nodesource.com/setup_${NODE_MAJOR}.x" | bash -
apt-get install -y --no-install-recommends nodejs
node --version
# bun installed system-wide; used only if any UI tooling must run on-box. The
# production SPA build runs in GitHub Actions -> S3 (no on-box build, no swapfile).
export BUN_INSTALL=/usr/local
curl -fsSL https://bun.sh/install | bash
/usr/local/bin/bun --version || true
echo "==> non-login service user '${SERVICE_USER}'"
if ! id "$SERVICE_USER" >/dev/null 2>&1; then
useradd --system --create-home --home-dir "$SERVICE_HOME" \
--shell /usr/sbin/nologin "$SERVICE_USER"
fi
echo "==> directories"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$SERVICE_HOME" "$APP_DIR"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0755 "$WWW_ROOT"
install -d -o "$SERVICE_USER" -g "$SERVICE_USER" -m 0750 "$LOG_DIR"
install -d -o root -g root -m 0755 "$TEMPLATE_DIR"
echo "==> stage boot-time templates into the image"
cp /tmp/open-swe-templates/* "$TEMPLATE_DIR/"
chown root:root "$TEMPLATE_DIR"/*
chmod 0644 "$TEMPLATE_DIR"/*
rm -rf /tmp/open-swe-templates
echo "==> tmpfs for the runtime .env (root/owner-only, noexec/nosuid/nodev)"
# /run is already tmpfs on Ubuntu; this is an explicit, deliberately-small mount
# scoped to the service user so the materialized .env never touches disk.
if ! grep -q '/run/open-swe' /etc/fstab; then
cat >>/etc/fstab <<EOF
tmpfs /run/open-swe tmpfs rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m 0 0
EOF
fi
echo "==> disable nginx default site (open-swe site is installed at boot)"
rm -f /etc/nginx/sites-enabled/default
systemctl enable nginx
echo "==> harden: no password auth, IMDSv2 already enforced by launch template"
# (sshd is not exposed publicly — instance is in a private subnet, SG inbound = ALB only.)
echo "==> clean apt caches"
apt-get clean
rm -rf /var/lib/apt/lists/*
echo "==> provision complete"

View file

@ -1,55 +0,0 @@
{
"agent": {
"metrics_collection_interval": 60,
"run_as_user": "root"
},
"metrics": {
"namespace": "open-swe/@@OPENSWE_ENV@@",
"append_dimensions": {
"InstanceId": "${aws:InstanceId}"
},
"metrics_collected": {
"mem": { "measurement": ["mem_used_percent"] },
"disk": {
"measurement": ["used_percent"],
"resources": ["/"]
}
}
},
"logs": {
"logs_collected": {
"files": {
"collect_list": [
{
"file_path": "/var/log/open-swe/app.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/app",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/open-swe-user-data.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/user-data",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/access.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-access",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
},
{
"file_path": "/var/log/nginx/error.log",
"log_group_name": "/open-swe/@@OPENSWE_ENV@@/nginx-error",
"log_stream_name": "{instance_id}",
"retention_in_days": 30,
"timezone": "UTC"
}
]
}
}
}
}

View file

@ -1,62 +0,0 @@
# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy.
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
#
# nginx is the SOLE ingress and security boundary (T5 OSWE-IAC-03): the backend
# binds 127.0.0.1:2024 and is NOT network-reachable. nginx proxies exactly two
# prefixes to it — /dashboard/api/* and /webhooks/* — and nothing else. The
# unauthenticated LangGraph agent API (/threads, /runs, /assistants, /store) is
# NEVER proxied; those paths return the SPA shell.
#
# Both ALB target groups (dashboard host + hooks host) point at this nginx :80,
# not at :2024 directly, so there is no path to the raw control plane even from
# inside the SG. Webhook signature verification still happens in the app (the raw
# body + GitHub/Slack/Linear signature headers are passed through unmodified).
server {
listen 80 default_server;
listen [::]:80 default_server;
server_name @@SERVER_NAME@@;
root @@WWW_ROOT@@;
index _shell.html;
# ALB target-group health check (dashboard TG).
location = /healthz { default_type text/plain; return 200 "ok\n"; }
# Dashboard API + OAuth callback -> backend webapp.
location /dashboard/api/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
client_max_body_size 10m;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 300s;
}
# Inbound webhooks (GitHub/Slack/Linear) -> backend webapp. Routed through
# nginx so :2024 stays loopback-only (T5 OSWE-IAC-03). The raw request body +
# signature headers pass through unmodified for in-app signature verification.
location /webhooks/ {
proxy_pass http://@@BACKEND_ADDR@@;
proxy_http_version 1.1;
# GitHub permits webhook payloads up to 25 MB; nginx's 1 MB default would
# 413 large push/PR events at the edge BEFORE in-app signature verification
# runs, silently dropping them (OSWE-T12-01). proxy_request_buffering off
# does not relax the size cap — set it explicitly.
client_max_body_size 25m;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_request_buffering off;
proxy_read_timeout 300s;
}
# Static assets + SPA shell fallback (client-side routing).
location / {
try_files $uri $uri/ /_shell.html;
}
}

View file

@ -1,54 +0,0 @@
# Open SWE — stock LangGraph dev server (3+ graphs + FastAPI webapp, :2024).
# TEMPLATE: tokens (@@...@@) are rendered at first boot by user-data.sh.
# In-memory runtime (--no-reload) + ExecStartPost reseed; no Aegra/Postgres.
#
# Boot contract:
# ExecStartPre = fetch-config.sh <env> -> runs as root (`+`) ONLY to materialize
# the SERVICE-USER-owned tmpfs .env (@@ENV_FILE@@) from Secrets
# Manager + SSM and chown it to @@SERVICE_USER@@, fail-fast (the
# unit does NOT start if config can't be fetched).
# ExecStart = langgraph dev (as @@SERVICE_USER@@, bound to 127.0.0.1 — nginx
# is the sole ingress; never binds 0.0.0.0).
# ExecStartPost = seed_store.sh <env> -> reseeds team_settings + user_mappings
# that the in-memory store loses on every restart.
[Unit]
Description=Open SWE stock LangGraph dev server (graphs + webapp, :@@PORT@@)
After=network-online.target
Wants=network-online.target
RequiresMountsFor=/run/open-swe
[Service]
Type=simple
User=@@SERVICE_USER@@
Group=@@SERVICE_USER@@
WorkingDirectory=@@APP_DIR@@
# No EnvironmentFile: the secret-bearing app .env (@@ENV_FILE@@) is loaded by
# langgraph/dotenv (so the multiline GitHub App PEM never hits systemd's env
# parser), and seed_store.sh reads the same .env directly (without sourcing it).
#
# ExecStartPre runs as root (`+`) so it can chown the tmpfs .env to the service
# user; the env arg (@@OPENSWE_ENV@@) selects the SSM/Secrets prefix (T5 BOOT-01).
ExecStartPre=+@@FETCH_CONFIG@@ @@OPENSWE_ENV@@
# Bind 127.0.0.1 only — nginx proxies dashboard + webhooks; :@@PORT@@ is never
# directly network-reachable (T5 OSWE-IAC-03).
ExecStart=@@VENV@@/bin/langgraph dev --host 127.0.0.1 --port @@PORT@@ --no-browser --no-reload
ExecStartPost=@@SEED_STORE@@ @@OPENSWE_ENV@@
# App logs to a file CloudWatch collects (30-day retention set in the CW config).
StandardOutput=append:/var/log/open-swe/app.log
StandardError=append:/var/log/open-swe/app.log
Restart=on-failure
RestartSec=5
TimeoutStartSec=180
# Hardening — the box holds no durable state of its own.
NoNewPrivileges=true
ProtectSystem=full
ProtectHome=true
PrivateTmp=true
ReadWritePaths=/var/log/open-swe /var/www/open-swe /run/open-swe @@APP_DIR@@
[Install]
WantedBy=multi-user.target

View file

@ -1,145 +0,0 @@
#!/usr/bin/env bash
# Open SWE EC2 user-data — PROVISIONING-ONLY (runs once, at first boot).
#
# This is the rationale for `userDataCausesReplacement: true` in CDK: user-data
# does FIRST-BOOT provisioning, never durable runtime config. Editing it is a
# deliberate instance replacement. Durable runtime config is fetched fresh on
# every service start by deploy/seahaven/fetch-config.sh (ExecStartPre).
#
# The box holds NO durable state of its own:
# - secrets/config -> Secrets Manager + SSM, materialized to a tmpfs .env at boot
# - app artifact -> pulled from S3 (open-swe-<env>-assets) via the instance role
# - store state -> reseeded by seed_store.sh (ExecStartPost) on every start
# => there is no RETAIN volume to protect; replacement is tolerated. The EBS
# discipline (snapshot root + wait state=completed BEFORE any replacing deploy)
# is the safety net, not durable on-box state. See README "EBS discipline".
#
# Tokens (@@...@@) are substituted by CDK when it renders this script into the
# launch template. region is read from IMDSv2 as a fallback.
set -euo pipefail
exec > >(tee -a /var/log/open-swe-user-data.log) 2>&1
echo "==> open-swe user-data start $(date -u +%FT%TZ)"
# --- CDK-rendered values -----------------------------------------------------
# NOTE: CDK-substituted tokens use %%...%% (rendered by app-service.ts), DISTINCT
# from the @@...@@ tokens this script seds into the baked systemd/nginx templates.
# The two MUST NOT share a delimiter: a shared @@OPENSWE_ENV@@ / @@SERVER_NAME@@
# let CDK clobber the sed PATTERN, leaving the unit's token unsubstituted.
OPENSWE_ENV="%%OPENSWE_ENV%%" # dev | prod
ASSETS_BUCKET="%%ASSETS_BUCKET%%" # open-swe-<env>-assets
SERVER_NAME="%%SERVER_NAME%%" # openswe[-dev].seahaven.com
ARTIFACT_PREFIX="%%ARTIFACT_PREFIX%%" # e.g. releases/latest
# --- fixed layout (must match provision.sh + templates) ----------------------
SERVICE_USER="openswe"
APP_DIR="/opt/open-swe/app"
VENV="${APP_DIR}/.venv"
WWW_ROOT="/var/www/open-swe"
TEMPLATE_DIR="/opt/open-swe/templates"
ENV_FILE="/run/open-swe/.env"
PORT="2024"
FETCH_CONFIG="${APP_DIR}/deploy/seahaven/fetch-config.sh"
SEED_STORE="${APP_DIR}/deploy/seahaven/seed_store.sh"
# region from IMDSv2
TOKEN="$(curl -fsS -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 300" || true)"
AWS_REGION="$(curl -fsS -H "X-aws-ec2-metadata-token: ${TOKEN}" \
http://169.254.169.254/latest/meta-data/placement/region || echo us-east-1)"
export AWS_DEFAULT_REGION="$AWS_REGION"
echo "env=${OPENSWE_ENV} region=${AWS_REGION} bucket=${ASSETS_BUCKET} host=${SERVER_NAME}"
# --- boot.env: non-secret pointers fetch-config.sh reads ---------------------
install -d -o root -g root -m 0755 /etc/open-swe
cat >/etc/open-swe/boot.env <<EOF
OPENSWE_ENV=${OPENSWE_ENV}
AWS_REGION=${AWS_REGION}
ASSETS_BUCKET=${ASSETS_BUCKET}
ARTIFACT_PREFIX=${ARTIFACT_PREFIX}
ENV_FILE=${ENV_FILE}
SECRETS_PREFIX=open-swe-${OPENSWE_ENV}
SSM_PREFIX=/open-swe-${OPENSWE_ENV}
EOF
chmod 0644 /etc/open-swe/boot.env
# --- ensure the tmpfs for the materialized .env is mounted -------------------
# (baked into /etc/fstab by the AMI; mount it now in case it isn't yet.)
install -d -o root -g root -m 0755 /run/open-swe || true
mountpoint -q /run/open-swe || mount /run/open-swe || mount -t tmpfs \
-o rw,nosuid,nodev,noexec,mode=0700,uid=${SERVICE_USER},gid=${SERVICE_USER},size=8m \
tmpfs /run/open-swe
# --- install the deploy script (single source of the app-deploy procedure) ---
# deploy.sh (deploy/ami/deploy.sh) pulls the release from S3, builds the venv with
# `uv sync`, and restarts the service. CDK base64-renders the file into the
# %%DEPLOY_SH_B64%% token below so it is a normal reviewable repo file, not an
# inline heredoc. The `open-swe-<env>-deploy` SSM document runs this same script
# for every subsequent release.
echo "==> install /opt/open-swe/bin/deploy.sh"
install -d -o root -g root -m 0755 /opt/open-swe/bin
base64 -d >/opt/open-swe/bin/deploy.sh <<'DEPLOY_SH_B64'
%%DEPLOY_SH_B64%%
DEPLOY_SH_B64
chmod 0755 /opt/open-swe/bin/deploy.sh
# --- render + install the systemd unit ---------------------------------------
echo "==> install systemd unit"
sed \
-e "s|@@SERVICE_USER@@|${SERVICE_USER}|g" \
-e "s|@@APP_DIR@@|${APP_DIR}|g" \
-e "s|@@VENV@@|${VENV}|g" \
-e "s|@@PORT@@|${PORT}|g" \
-e "s|@@ENV_FILE@@|${ENV_FILE}|g" \
-e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
-e "s|@@FETCH_CONFIG@@|${FETCH_CONFIG}|g" \
-e "s|@@SEED_STORE@@|${SEED_STORE}|g" \
"${TEMPLATE_DIR}/open-swe.service" >/etc/systemd/system/open-swe.service
systemctl daemon-reload
# --- render + install the nginx site -----------------------------------------
echo "==> install nginx site"
sed \
-e "s|@@SERVER_NAME@@|${SERVER_NAME}|g" \
-e "s|@@WWW_ROOT@@|${WWW_ROOT}|g" \
-e "s|@@BACKEND_ADDR@@|127.0.0.1:${PORT}|g" \
"${TEMPLATE_DIR}/open-swe.nginx.conf" >/etc/nginx/sites-available/open-swe
ln -sfn /etc/nginx/sites-available/open-swe /etc/nginx/sites-enabled/open-swe
rm -f /etc/nginx/sites-enabled/default
nginx -t
# --- CloudWatch agent: 30-day log retention ----------------------------------
echo "==> configure CloudWatch agent (30-day retention)"
sed -e "s|@@OPENSWE_ENV@@|${OPENSWE_ENV}|g" \
"${TEMPLATE_DIR}/amazon-cloudwatch-agent.json" \
>/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
/opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
-a fetch-config -m ec2 -s \
-c file:/opt/aws/amazon-cloudwatch-agent/etc/open-swe-cw.json
# --- start nginx FIRST (the security boundary + health surface) --------------
# NOTE: intentionally NO swapfile here. The 8 GB-swapfile OOM hack existed only
# for the on-box Nitro SPA build, which now runs in GitHub Actions -> S3.
# nginx is brought up BEFORE the app is deployed so the ALB target-group health
# check (static `/healthz` -> 200) passes and the box is a healthy target even on
# the very first boot, before any release is published. open-swe.service is
# enabled (boot persistence) but STARTED by deploy.sh once the app is on disk.
echo "==> start nginx"
systemctl enable --now nginx
systemctl reload nginx
systemctl enable open-swe.service
# --- deploy the app (NON-FATAL on first boot) --------------------------------
# deploy.sh pulls the release, builds the venv, and starts open-swe.service. On a
# brand-new env no release exists yet, so this is allowed to fail WITHOUT aborting
# user-data: nginx is already up (healthy target), and the first `build-artifacts`
# run + `open-swe-<env>-deploy` SSM command will bring the app up. A failure here
# is logged, not fatal.
echo "==> initial app deploy (non-fatal if no release is published yet)"
if /opt/open-swe/bin/deploy.sh; then
echo "==> initial app deploy succeeded"
else
echo "==> no release yet (or deploy failed): open-swe.service deferred to the next SSM deploy"
fi
echo "==> open-swe user-data done $(date -u +%FT%TZ)"

View file

@ -1,280 +0,0 @@
# Sea Haven — Open SWE deployment runbook
How this fork is deployed at Sea Haven. **PROD is LIVE as of 2026-06-29.** The
runtime is the **stock LangGraph dev server** (not the Aegra path — see
[Aegra](#aegra-deferred)), running on a self-hosted ARM64 EC2 box behind the
shared `seahaven-com` ALB.
Account `328440206208`, region `us-east-1`. Internal addresses, ARNs, snapshot
ids, and account-scoped values that are sensitive are shown as `<PLACEHOLDERS>`;
the real values live in the private IT docs (Confluence "AWS Architecture Map",
id 1540098) and in AWS — **do not commit them to this public fork.**
This is the canonical deploy runbook. The CDK details live in
[`infra/README.md`](../../infra/README.md); the rotation procedure in
[`ROTATION.md`](ROTATION.md).
## Live prod facts (2026-06-29)
| | |
|---|---|
| Dashboard | `https://openswe.seahaven.com` |
| Webhooks | `https://hooks.seahaven.com/webhooks/*` |
| Ingress | shared internet-facing ALB `app/seahaven-com` → target group `open-swe-prod-tg` → EC2 `i-08a729e50779c4b07` (`t4g.large`, ARM64) on nginx `:80` |
| Backend | `langgraph dev` bound to `127.0.0.1:2024` (loopback only); nginx is the sole ingress |
| CDK stacks | `open-swe-iam` (OIDC roles) · `open-swe-dev` · `open-swe-prod` |
Dev mirrors prod with `-dev` hosts (`openswe-dev.seahaven.com` /
`hooks-dev.seahaven.com`), a `t4g.medium` box, and no GitHub-App/Slack/webhook
integration (it is a deployment-validation env, not a live-triggered agent).
> The retired on-prem `*.seahavenind.com` ALB routing and DNS were removed on
> 2026-06-29; prod is now live exclusively on `*.seahaven.com`.
## Hosting model
```
GitHub / Slack ──▶ hooks.seahaven.com ──┐
│ (shared ALB :443, host+path rules)
Browser ─────────▶ openswe.seahaven.com ──┤
▼
ALB app/seahaven-com ──▶ open-swe-prod-tg ──▶ EC2 box :80 (nginx)
├─ nginx — SPA + scoped proxy
│ /dashboard/api/* and /webhooks/*
└─ langgraph dev 127.0.0.1:2024
└─▶ LangSmith cloud sandbox (build/git/PR)
```
- A **single** VPC and a **single** internet-facing ALB (`app/seahaven-com`) are
shared with the on-prem `seahaven-site` stack. open-swe **imports** the VPC,
ALB SG, `:443` listener, and `seahaven.com` zone — it never owns/mutates them;
it only adds its own instance SG, a standalone ALB-egress rule, two listener
rules, a target group, and Route53 aliases.
- The EC2 box is in a **private** subnet (us-east-1a, same AZ as the single NAT
for in-AZ egress). It is reachable **only** from the shared ALB SG on `:80`.
- **nginx is the security boundary.** It serves the static dashboard SPA and
proxies exactly two prefixes to `:2024` — `/dashboard/api/*` and `/webhooks/*`.
The unauthenticated LangGraph API (`/threads`, `/runs`, `/assistants`,
`/store`) is never proxied; those paths return the SPA shell. `:2024` is
loopback-only and never network-reachable, even inside the SG.
- **Webhooks** ride listener rules below the on-prem host-agnostic `/webhooks/*`
rule (priority 2 dev / 3 prod, host-scoped to the open-swe hosts) so they reach
the open-swe box and never steal an on-prem host's webhooks.
The box holds **no durable state of its own**: secrets/config are materialized to
a tmpfs `.env` at boot, the app artifact is pulled from S3, and the in-memory
LangGraph store is re-seeded on every start. Replacement is tolerated; there is no
RETAIN volume.
---
## Deploy pipeline (end to end)
Two independent CD lanes, both OIDC-only (no static keys), both with a manual
approval gate on prod via the GitHub **`prod` Environment** (required reviewer:
Adam). The `environment: prod` declaration both fires the approval gate and makes
the OIDC subject `…:environment:prod`, which is the only subject the prod deploy
roles trust — so a dev-branch token can never reach prod.
### (a) Infra CD — `cd-infra.yml`
Deploys the CDK stacks. Path-filtered to `infra/**`.
```
push to dev → Infra CI (tsc + jest + cdk synth) → cdk deploy OpenSweDevStack (AUTO, CI-green-gated)
push to main → Infra CI → cdk deploy OpenSweProdStack (manual approval: env "prod")
```
- Roles: `githubdeploy-open-swe-infra-{dev,prod}` (in the `open-swe-iam` stack;
set as repo variables `AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}`).
- It targets **one stack explicitly per env** (`cdk deploy OpenSweDevStack` /
`OpenSweProdStack`), not `cdk deploy --all`, so a single-env push can never
deploy the other env or the shared IAM stack.
- The shared `open-swe-iam` stack (owns both envs' OIDC deploy roles) is **not**
deployed by CD — it is a privileged, human-gated apply.
**Stack order on a clean account:** `open-swe-iam` first (creates the OIDC roles;
set the repo deploy-role variables and configure the `prod` Environment reviewer
from its outputs), then `open-swe-dev`, then `open-swe-prod`.
### (b) Seed the config store — `put-config.sh <env>`
Run **after** `cdk deploy open-swe-<env>` and **before** the box first boots. CDK
creates the value-less Secrets Manager shells (`open-swe-<env>/<VAR>`) and the
IaC-managed SSM params (`/open-swe-<env>/<VAR>`); `put-config.sh` populates the
secret values plus the out-of-band SSM params that cannot live in IaC.
```bash
deploy/seahaven/put-config.sh <dev|prod> # set each value inline, via OPENSWE_PUT_<VAR>, or from a vault
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
```
`put-config.sh` ships `<FILL>` placeholders only — **no real secret values are
committed**. It does not touch the IaC-managed SSM params (CDK owns those).
**13 prod boot-required vars** — `fetch-config.sh` fail-fasts (refuses to write a
partial `.env`, the unit does not start) if any are missing/empty:
- **9 secrets** (Secrets Manager `open-swe-prod/<VAR>`): `DASHBOARD_JWT_SECRET`,
`TOKEN_ENCRYPTION_KEY`, `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`LANGSMITH_API_KEY_PROD`, `GITHUB_APP_PRIVATE_KEY`, `GITHUB_APP_CLIENT_SECRET`,
`GITHUB_WEBHOOK_SECRET`, `SLACK_SIGNING_SECRET`.
- **4 SSM params** (`/open-swe-prod/<VAR>`): `DEFAULT_SANDBOX_SNAPSHOT_ID`,
`GITHUB_APP_ID`, `GITHUB_APP_INSTALLATION_ID`, `GITHUB_APP_CLIENT_ID`.
(`ANTHROPIC_API_KEY` + `OPENAI_API_KEY` are required because that is the seeded
cross-family pair; the active set follows `REQUIRED_PROVIDER_KEYS`. The
`LANGSMITH_API_KEY_PROD` + `DEFAULT_SANDBOX_SNAPSHOT_ID` pair is required because
`SANDBOX_TYPE=langsmith`.) Dev boots without the GitHub-App / Slack / webhook
secrets — it has no such integration.
`fetch-config.sh` runs as an `ExecStartPre=+` hook (root, only long enough to
write the `openswe`-owned `0600` tmpfs `.env`), reads all `/open-swe-<env>/*` SSM
params + all `open-swe-<env>/*` secrets via the instance role, and forces
`DEFAULT_REPO_OWNER` away from the upstream `langchain-ai` org.
### (c) App artifact deploy — `build-artifacts.yml`
Builds the release and rolls the box. Path-filtered to `agent/**`, `ui/**`,
`deploy/**`, `langgraph.json`, `pyproject.toml`, `uv.lock`.
```
push to dev → build SPA + package → open-swe-dev-assets/releases/ → SSM open-swe-dev-deploy (AUTO)
push to main → build SPA + package → open-swe-prod-assets/releases/ → SSM open-swe-prod-deploy (manual approval: env "prod")
```
1. The dashboard SPA is built **on the runner** (`bun run build` → vite →
`ui/.output/public`) — the box is small, so the memory-heavy build runs in CI.
2. `package-artifacts.sh` produces two tarballs: `spa.tar.gz` (built SPA) and
`app.tar.gz` (Python source tree — no `ui/`, no `.venv`).
3. Both are uploaded to S3 `open-swe-<env>-assets` under `releases/<sha>/`
(immutable, auditable) and mirrored to `releases/latest/` (what the box pulls).
4. CI fires the `open-swe-<env>-deploy` SSM document (tag-scoped to
`project=open-swe,env=<env>`), which runs `/opt/open-swe/bin/deploy.sh` on the
box: pull the release from S3, build a native-ARM64 venv with
`uv sync --frozen --no-dev`, extract the SPA to the nginx web root,
`systemctl restart open-swe.service`, reload nginx, then gate on
`systemctl is-active --quiet open-swe.service` (a non-active unit exits the
deploy non-zero).
Roles: `githubdeploy-open-swe-app-{dev,prod}` (repo variables
`AWS_DEPLOY_ROLE_APP_{DEV,PROD}`) — tag-scoped `ssm:SendCommand` on the deploy
document only (not the generic `AWS-RunShellScript`) + write to the env's S3
bucket.
Secrets/config are **not** fetched by `deploy.sh`; the `systemctl restart`'s
`ExecStartPre=fetch-config.sh` re-materializes the `.env` on every restart, so a
bad config surfaces as a failed unit.
### (d) dev → main promotion + rollback
**Promotion — `promote-dev-to-prod.yml`** (nightly cron `0 8 * * *` + manual
dispatch): mints a GitHub App installation token (a bypass actor on the `main`
ruleset), gates on **every check-run on the dev HEAD commit being completed and
passing**, then **fast-forward-only** pushes `dev` → `main`. A diverged `main`
fails loudly rather than force-updating. The push to `main` is what triggers the
prod lanes of `cd-infra.yml` / `build-artifacts.yml` (each still behind the `prod`
Environment approval). Re-gating via a PR on `main` would be redundant since the
commit already passed every check on dev.
**Rollback — `rollback.yml`** (manual dispatch, `env` + optional `sha`): re-points
`releases/latest/` at a prior release and re-fires the `open-swe-<env>-deploy` SSM
document — same fire/wait/gate path as a forward deploy, no rebuild.
```
env=dev, sha blank → restore open-swe-dev-assets/releases/last-good/ (AUTO)
env=prod, sha blank → restore open-swe-prod-assets/releases/last-good/ (manual approval: env "prod")
sha=<commit> → restore that exact releases/<sha>/ instead
```
It reuses the existing `githubdeploy-open-swe-app-<env>` role (no new IAM).
---
## On-box layout (reference)
| Path | What |
|---|---|
| `open-swe.service` (systemd) | `langgraph dev --host 127.0.0.1 --port 2024 --no-browser --no-reload` as the unprivileged `openswe` user. In-memory runtime. |
| `fetch-config.sh` | `ExecStartPre=+` — materializes the tmpfs `.env` from Secrets Manager + SSM, fail-fast. |
| `seed_store.sh` | `ExecStartPost` — re-seeds `team_settings/default` + `user_mappings` (the in-memory store loses them on every restart). |
| nginx | SPA from `/var/www/open-swe`, proxy `/dashboard/api/` + `/webhooks/` → `127.0.0.1:2024`, `/healthz` → 200. |
| `deploy.sh` | the release procedure run on first boot (non-fatal) and by every SSM deploy. |
| CloudWatch logs | `/open-swe/<env>/{app,user-data,nginx-access,nginx-error}` at 30-day retention. |
The live systemd unit + nginx site are the **AMI templates**
(`deploy/ami/templates/open-swe.service`, `open-swe.nginx.conf`), rendered at
first boot by `deploy/ami/user-data.sh`. The AMI is the baked
`open-swe-base-arm64` image (Ubuntu 24.04 + uv/py3.12 + nginx + CW agent), pinned
by exact id in `infra/lib/constructs/ami-cache.ts`. There is intentionally **no**
on-box swapfile — the OOM-prone SPA build now runs in CI, not on the box.
> `deploy/seahaven/{nginx/openswe.conf,systemd/open-swe.service}` are the
> **retired on-prem VM** variants (run as `adam` from a home dir, bound `0.0.0.0`,
> Postgres-backed). They are kept only for on-prem-contrast reference and are not
> used by the AWS deployment.
### Models
Model selection is **store-driven**, not env. The `team_settings/default` store
doc wins (then per-user profile, then per-thread); `LLM_MODEL_ID` is only a
seed-time fallback. Defaults seeded by `seed_store.sh`:
- builder: `bedrock_converse:us.anthropic.claude-opus-4-8` (effort `high`)
- reviewer: `bedrock_converse:us.anthropic.claude-opus-4-8` (effort `high`) — set
`SEED_REVIEWER_MODEL` (or change it in the UI) to a Fireworks model if you want a
cross-family reviewer. Only ids present in `SUPPORTED_MODELS`
(`agent/dashboard/options.py`) are valid; OpenAI/Google models were removed in the
Bedrock/Fireworks migration.
- the `analyzer` graph is hardcoded to the code default and ignores team settings.
## Triggering
Mention **`@openswe`** (or `@open-swe` / `@seahaven-openswe`) in a GitHub issue or
PR comment, a Linear comment, or a Slack thread. The commenter must have a
`user_mappings` entry (seeded by `seed_store.sh` from `CONFIGURED_ADMINS` /
`SEED_USER_MAPPINGS`) or the run is skipped.
Live integration endpoints (set in each provider's app config):
| Integration | URL |
|---|---|
| GitHub webhook | `https://hooks.seahaven.com/webhooks/github` |
| Slack events | `https://hooks.seahaven.com/webhooks/slack` (+ `/webhooks/slack/interactivity`) |
| Linear webhook | `https://hooks.seahaven.com/webhooks/linear` |
| GitHub OAuth callback | `https://openswe.seahaven.com/dashboard/api/auth/callback` |
---
## Troubleshooting
**RETAIN secret-shell orphan on stack re-create.** The Secrets Manager shells use
`DeletionPolicy: Retain` + a fixed `open-swe-<env>/<VAR>` name. If a stack's first
create rolls back (or on a teardown/rebuild, a secret logical-id refactor, or
standing up a new env), the empty shells survive and keep their global names, so
every later create fails `AlreadyExists` — and a plain `delete-secret` does not
free the name (it stays reserved for the 7–30 day recovery window). Before
re-creating the stack, **force-delete the empty orphans** (only shells with no
value version — never a populated secret). Hit on prod 2026-06-29 (PR #51 deploy
failure). Full recovery command + rationale:
[`infra/README.md`](../../infra/README.md) (PR #52).
**`langgraph dev` won't start after a deploy.** `fetch-config.sh` fail-fasts on a
missing/empty required var and prints the offending variable **names** (never
values) to the unit journal. Confirm the 13 prod boot-required vars are populated
(`put-config.sh prod`), then `systemctl restart open-swe.service`.
**ALB target unhealthy.** The TG health check is `GET /healthz` on nginx `:80`
(static 200). nginx starts before the app on first boot, so an unhealthy target
usually means the box can't reach the ALB SG on `:80` (the standalone ALB-egress
rule) rather than an app fault.
---
## Aegra (deferred)
`aegra/aegra.json` + `aegra/aegra_entry.py` are the self-hosted-runtime
alternative (Apache-2.0, avoids the LangGraph-Platform Elastic license). Not
active on the stock deployment. To use: place both at the repo root, run
`aegra serve` (:2026), and point `LANGGRAPH_URL` at `:2026`. Aegra gives a
Postgres-backed durable store/checkpointer, which removes the need for
`seed_store.sh` and survives restarts (paused HITL interrupts persist).
</content>

View file

@ -1,92 +0,0 @@
# Sea Haven — Open SWE secret & config rotation
How secrets and config reach the running app, and how to rotate either one.
## How values flow at boot
```
AWS Secrets Manager open-swe-<env>/* ─┐
AWS SSM Param Store /open-swe-<env>/* ─┤── fetch-config.sh ──▶ tmpfs /run/open-swe/.env (root:root 0600)
│ (systemd ExecStartPre=+, EC2 role) │
▼ ▼
FAIL-FAST if a app symlink <APP_DIR>/.env
required var is empty python-dotenv reads at import
```
The app reads `.env` **once, at import**. There is no hot-reload of secrets.
Therefore the rotation contract is always the same two steps:
> **Rotation = (1) update the value in Secrets Manager / SSM, then (2) restart the
> service** so `fetch-config.sh` re-materializes the `.env`.
```bash
# after updating a secret/param in AWS:
sudo systemctl restart open-swe.service
# ExecStartPre=+ -> fetch-config.sh re-pulls + rewrites the tmpfs .env (fail-fast)
# ExecStartPost -> seed_store.sh re-seeds the in-memory store (team_settings + user_mappings)
```
There is **no zero-downtime path for most secrets** on the stock in-memory
runtime — a restart is required and it also wipes the in-memory store (re-seeded
by `seed_store.sh` automatically). The one secret built for zero-downtime overlap
is `TOKEN_ENCRYPTION_KEY` (see below), but even it needs the restart to load the
new key list.
## Rotating a secret (Secrets Manager)
```bash
ENV=prod # or dev
NAME=DASHBOARD_JWT_SECRET
aws secretsmanager put-secret-value \
--secret-id "open-swe-${ENV}/${NAME}" \
--secret-string 'NEW_VALUE' \
--region us-east-1
sudo systemctl restart open-swe.service # on the box
```
(`update-secret`/`put-secret-value` both create a new version; the boot hook
always reads `AWSCURRENT`.)
## Rotating a config param (SSM)
```bash
aws ssm put-parameter --overwrite \
--name "/open-swe-${ENV}/DASHBOARD_BASE_URL" \
--type String --value 'https://openswe.seahaven.com' \
--region us-east-1
sudo systemctl restart open-swe.service
```
## Per-secret rotation notes
| Secret | Rotation notes |
|---|---|
| **TOKEN_ENCRYPTION_KEY** | Fernet key(s). Supports a **comma/newline-separated list** (`agent/encryption.py`) for zero-downtime key rotation: prepend the NEW key, keep the OLD key(s) in the list. New data is encrypted with the first key; old data still decrypts with the trailing keys. After all encrypted-at-rest tokens (per-user GitHub OAuth tokens in thread metadata) have been re-encrypted/expired, drop the old key. Store the list as one secret value; `fetch-config.sh` writes it verbatim. **Never** rotate to a single new key in one step or every existing encrypted token becomes undecryptable. |
| **GITHUB_APP_PRIVATE_KEY** | Multiline PEM. Generate a new private key in the GitHub App settings (you may have **two active keys** during overlap), put the new PEM into the secret, restart, verify install-token minting + a webhook delivery, then delete the old key in GitHub. `fetch-config.sh` writes the PEM as a double-quoted multiline value (python-dotenv-safe); paste the full `-----BEGIN…-----END-----` block including newlines. |
| **GITHUB_WEBHOOK_SECRET** | Webhook HMAC. GitHub allows only **one** webhook secret per App, so this is a brief-break rotation: update the secret in AWS **and** the GitHub App webhook config, restart. Deliveries signed with the old secret during the gap will 401 (GitHub auto-redelivers). Required in **prod** (fail-fast). |
| **SLACK_SIGNING_SECRET** | Slack request-signature secret. Rotate in the Slack app config and AWS together, restart. Required in **prod** (fail-fast). A stale value silently 401s `url_verification`/events until restart (known gotcha). |
| **LINEAR_WEBHOOK_SECRET** | Linear webhook signature. Required in prod only when the Linear integration is wired (`LINEAR_API_KEY` present). Rotate in Linear + AWS together, restart. |
| **SLACK_CLIENT_SECRET / GITHUB_APP_CLIENT_SECRET** | OAuth client secrets (dashboard login / Slack OAuth). Rotate in the provider console + AWS, restart. Existing dashboard sessions are JWT-signed by `DASHBOARD_JWT_SECRET`, not these, so they survive. |
| **DASHBOARD_JWT_SECRET** | Signs dashboard session cookies. Rotating **invalidates all active sessions** (users re-login). Hard-required (RuntimeError if empty). No overlap list — single value. |
| **Model provider keys** (`FIREWORKS_API_KEY`; legacy `ANTHROPIC_API_KEY` / `OPENAI_API_KEY` / `GOOGLE_API_KEY` / `GROQ_API_KEY`) | Standard API-key rotation: issue new key, update AWS, restart, revoke old. Bedrock (the seeded Claude builder + reviewer) authenticates via the **host IAM role — no API key**; the only fail-fast-required provider key is now `FIREWORKS_API_KEY` (fallback / subagents / any non-Claude model). Keep `REQUIRED_PROVIDER_KEYS` in sync if you change the seeded models. |
| **LANGSMITH_API_KEY_PROD** | LangSmith key powering the `langsmith` sandbox (the only provider with working in-sandbox git/gh auth). Required when `SANDBOX_TYPE=langsmith` (fail-fast). Rotate in LangSmith + AWS, restart; existing sandboxes keep their already-injected proxy token until recycled. |
| **SLACK_BOT_TOKEN / LINEAR_API_KEY / GITHUB_PAT / EXA_API_KEY / DAYTONA_API_KEY / RUNLOOP_API_KEY / CORRIDOR_* / USER_ID_API_KEY_MAP / X_SERVICE_AUTH_JWT_SECRET** | Plain API-token rotation: update AWS, restart, revoke old at the provider. None support an overlap list. |
## Fail-fast safety
`fetch-config.sh` refuses to write the `.env` (exit 1) if any required var is
empty after a rotation — so a botched rotation (e.g. an empty `put-secret-value`)
stops the service at `ExecStartPre` instead of starting it with a partial `.env`.
The missing variable **names** are printed to the journal (values never are):
```bash
sudo journalctl -u open-swe.service -b | grep fetch-config
```
Required set enforced: `DASHBOARD_JWT_SECRET`, `TOKEN_ENCRYPTION_KEY`,
`GITHUB_APP_ID`, `GITHUB_APP_PRIVATE_KEY`, `GITHUB_APP_INSTALLATION_ID`,
`GITHUB_APP_CLIENT_ID`, `GITHUB_APP_CLIENT_SECRET`, the active provider keys
(`REQUIRED_PROVIDER_KEYS`, default `ANTHROPIC_API_KEY,OPENAI_API_KEY`), the
sandbox key for `SANDBOX_TYPE` (`langsmith` ⇒ `LANGSMITH_API_KEY_PROD` +
`DEFAULT_SANDBOX_SNAPSHOT_ID`), and — in **prod** — `GITHUB_WEBHOOK_SECRET`,
`SLACK_SIGNING_SECRET` (plus `LINEAR_WEBHOOK_SECRET` when Linear is wired).

View file

@ -1,11 +0,0 @@
{
"dependencies": ["."],
"graphs": {
"agent": "./aegra_entry.py:agent_graph",
"reviewer": "./aegra_entry.py:reviewer_graph",
"analyzer": "./aegra_entry.py:analyzer_graph"
},
"http": {
"app": "./aegra_entry.py:webapp_app"
}
}

View file

@ -1,19 +0,0 @@
"""Aegra entrypoint for Open SWE graphs.
Aegra loads graph files standalone via importlib.spec_from_file_location, which
gives them a synthetic module name and breaks agent/*.py's package-relative
imports (e.g. `from .dashboard.admin import ...`). Re-exporting the graphs here
through the installed ``agent`` package (absolute imports) restores correct
``__package__`` resolution, so the relative imports inside the agent modules work.
"""
from agent.analyzer import traced_analyzer as analyzer_graph
from agent.reviewer import traced_reviewer_agent as reviewer_graph
from agent.server import traced_agent as agent_graph
# Open SWE's FastAPI webapp (GitHub/Slack webhooks + dashboard API), mounted by
# Aegra via the "http" key in aegra.json. Absolute import for the same reason as
# the graphs above (relative imports break under Aegra's standalone file loader).
from agent.webapp import app as webapp_app
__all__ = ["agent_graph", "reviewer_graph", "analyzer_graph", "webapp_app"]

View file

@ -1,363 +0,0 @@
#!/usr/bin/env bash
# fetch-config.sh — AWS-sourced boot hook that materializes the app's .env.
#
# The stock `langgraph dev` runtime + the Open SWE app read a plain `.env` from
# the app working directory (python-dotenv). On the AWS lift-and-shift we do NOT
# commit a .env; instead every non-sensitive value lives in SSM Parameter Store
# (`/open-swe-<env>/*`) and every secret lives in AWS Secrets Manager
# (`open-swe-<env>/*`). This hook is run by systemd BEFORE the service starts; it
# pulls both sources via the EC2 instance role (no static keys), assembles a
# single .env on a tmpfs, and writes it owned by the unprivileged service user
# `chmod 600` (T5 SC-01: the privileged pre-hook materializes the secret; the app
# itself then runs as that NON-root service user, not root).
#
# It is intentionally FAIL-FAST: if any required secret/param is missing or empty
# it prints the offending variable NAMES (never values) and exits 1, so the
# service never starts with a partial .env.
#
# ---------------------------------------------------------------------------
# Naming contract (source of truth: T9 env/secret/config inventory)
# SSM /open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# Secrets open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# i.e. the last path segment IS the literal environment-variable name. This is a
# deliberate (documented) deviation from the handbook's kebab-case value-name
# example (`my-stack/slack-signing`): a .env materializer needs a lossless,
# unambiguous round-trip from store key -> env var, and the env var name is the
# only key that guarantees that. The `open-swe-<env>` stack prefix still follows
# kebab-case per naming-conventions.md.
# ---------------------------------------------------------------------------
#
# Wiring into systemd (AWS EC2 variant):
# The unit runs as the unprivileged service user (User=openswe). ONLY the
# ExecStartPre pre-hook runs as root (the `+` prefix) so it can pull from AWS,
# write the tmpfs .env, and chown it to the service user. The app (ExecStart)
# and the seeder (ExecStartPost) then run as openswe and read the openswe-owned
# 0600 .env — the agent never runs as root (T5 SC-01). Pass the env as the
# positional arg (T5 BOOT-01):
#
# [Service]
# User=openswe
# Group=openswe
# Environment=ENV_DIR=/run/open-swe SERVICE_USER=openswe
# # ExecStartPre runs as root (+) so it can chown the .env to the service user.
# ExecStartPre=+/opt/open-swe/deploy/seahaven/fetch-config.sh prod
# ExecStart=/opt/open-swe/.venv/bin/langgraph dev --host 127.0.0.1 --port 2024 \
# --no-browser --no-reload
# ExecStartPost=/opt/open-swe/deploy/seahaven/seed_store.sh prod
#
# tmpfs: /run is already a tmpfs on systemd hosts, so ENV_DIR=/run/open-swe is
# tmpfs-backed by default (the .env never touches disk). Set RUN_DEDICATED_TMPFS=1
# to mount a private tmpfs at ENV_DIR instead. The app's CWD `.env` is a symlink
# into ENV_DIR (created idempotently below), so python-dotenv finds it unchanged.
#
# Idempotent, re-runnable on every (re)start. No secret is ever echoed.
set -euo pipefail
umask 077
# --- Inputs ------------------------------------------------------------------
ENV="${1:-${OPENSWE_ENV:-}}"
case "$ENV" in
dev | prod) ;;
*)
echo "fetch-config: ENV must be 'dev' or 'prod' (got '${ENV:-<empty>}')" >&2
echo "usage: fetch-config.sh <dev|prod> (or set OPENSWE_ENV)" >&2
exit 2
;;
esac
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
SSM_PREFIX="/open-swe-${ENV}/"
SECRET_PREFIX="open-swe-${ENV}/"
ENV_DIR="${ENV_DIR:-/run/open-swe}" # tmpfs-backed (/run) by default
ENV_FILE="${ENV_DIR}/.env"
APP_DIR="${APP_DIR:-/opt/open-swe}" # where the app + its CWD .env live
APP_ENV_LINK="${APP_DIR}/.env" # symlink -> ENV_FILE
# The unprivileged service user that runs the app and OWNS the .env (T5 SC-01).
# fetch-config runs as root (ExecStartPre=+) only to chown the secret to it.
SERVICE_USER="${SERVICE_USER:-openswe}"
SERVICE_GROUP="${SERVICE_GROUP:-${SERVICE_USER}}"
# Sea Haven owner guard inputs (applied after the store is read, below). The owner
# normally comes from SSM /open-swe-<env>/DEFAULT_REPO_OWNER; OPENSWE_REPO_OWNER is an
# explicit operator override that wins over the store. FORBIDDEN = the upstream org
# the fork must never target; SAFE = the fallback when the resolved owner is blank or
# forbidden. (Upper/lower + whitespace are normalized before the guard check.)
OPENSWE_REPO_OWNER="${OPENSWE_REPO_OWNER:-}"
FORBIDDEN_REPO_OWNER="langchain-ai"
# Fallback org when the resolved owner is blank/forbidden — PER-ENV (mirrors the
# iacManagedSsm owner) so a dev box can NEVER fall back into the real Sea Haven org;
# it stays isolated in its own dev org. Defends the blank/upstream cases in-env.
case "$ENV" in
dev) SAFE_REPO_OWNER="seahaven-open-swe-dev" ;;
*) SAFE_REPO_OWNER="Sea-Haven-Industries" ;;
esac
for bin in aws jq; do
command -v "$bin" >/dev/null 2>&1 || { echo "fetch-config: '$bin' not found on PATH" >&2; exit 3; }
done
log() { echo "fetch-config[$ENV]: $*"; } # NAMES/counts only — never values
b64d() { base64 --decode; } # GNU coreutils on the EC2 host
# Accept a store key into VARS iff it is a valid env-var identifier and not a
# duplicate. Rejects non-identifier names (T5 SH-INJ-002 / set -e DoS hardening)
# and flat-namespace collisions (T5 SSM-05). $3 = source label for logs.
accept_var() {
local key="$1" value="$2" src="$3"
if ! [[ "$key" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
log "WARNING: skipping ${src} key with non-identifier name (rejected)"
return 0
fi
if [ -n "${VARS[$key]+set}" ]; then
echo "fetch-config[$ENV]: FAIL-FAST — duplicate key '${key}' from ${src} (flat-namespace collision)" >&2
exit 1
fi
VARS["$key"]="$value"
}
# --- tmpfs ------------------------------------------------------------------
mkdir -p "$ENV_DIR"
# Owned by the service user so the unprivileged app can traverse it (T5 SC-01).
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_DIR" 2>/dev/null || true
chmod 700 "$ENV_DIR"
if [ "${RUN_DEDICATED_TMPFS:-0}" = "1" ] && ! mountpoint -q "$ENV_DIR"; then
mount -t tmpfs -o nosuid,nodev,noexec,mode=0700,size=4m tmpfs "$ENV_DIR"
log "mounted dedicated tmpfs at $ENV_DIR"
fi
# --- Collect values into an associative array --------------------------------
declare -A VARS=()
# 1) SSM Parameter Store (non-sensitive config). NOT --recursive: the contract is
# a FLAT namespace /open-swe-<env>/<VAR>, so a non-recursive list returns exactly
# those keys and cannot collapse two nested paths onto one name (T5 SSM-05). aws
# CLI v2 auto-paginates NextToken.
log "reading SSM params under ${SSM_PREFIX} ..."
ssm_json="$(
aws ssm get-parameters-by-path \
--path "$SSM_PREFIX" \
--with-decryption \
--region "$REGION" \
--no-cli-pager \
--output json
)"
# Records are base64-encoded (name<TAB>value) so values with spaces/newlines/tabs
# survive the line-based read intact.
ssm_count=0
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}" # strip /open-swe-<env>/ prefix
[ -n "$key" ] || continue
accept_var "$key" "$value" "SSM"
ssm_count=$((ssm_count + 1))
done < <(jq -r '.Parameters[] | (.Name|@base64) + "\t" + (.Value|@base64)' <<<"$ssm_json")
log "loaded ${ssm_count} config param(s) from SSM"
# 2) Secrets Manager (sensitive values). We request secrets by EXPLICIT id
# (`batch-get-secret-value --secret-id-list ...`) rather than a name-prefix
# `--filters` collection scan (OSWE-IAC-SECRETS-LIST-01). Two wins:
# (a) Least privilege — an explicit id list lets the instance role scope
# BatchGetSecretValue to the per-secret ARN prefix and DROP the account-wide
# `secretsmanager:ListSecrets` grant that a filtered scan unavoidably forces
# (ListSecrets has no resource-level scoping). A filtered batch call also only
# authorizes against `*`; an id-list call authorizes per-secret ARN.
# (b) Deterministic set — the value set is the fixed SECRET_VARS shells created by
# the CDK ConfigStore, so we no longer depend on a list scan returning every
# page. Value-less shells and ids with no current value come back in the
# response `.Errors[]` (ResourceNotFound), never in `.SecretValues[]`, so a
# genuinely-missing REQUIRED secret is still caught by the FAIL-FAST check
# below — an absent optional secret is simply skipped.
# `--secret-id-list` is capped at 20 ids per call, so we chunk it. `--no-cli-pager`
# disables only the OUTPUT pager. Each record is base64(name<TAB>value).
#
# SOURCE OF TRUTH for this list: infra/lib/constructs/config-store.ts `SECRET_VARS`.
# Keep the two in lockstep — a new secret shell created there must be added here or it
# will never be fetched into the .env.
SECRET_VARS=(
ANTHROPIC_API_KEY CORRIDOR_API_TOKEN CORRIDOR_MCP_TOKEN CORRIDOR_TOKEN
DASHBOARD_JWT_SECRET DAYTONA_API_KEY EXA_API_KEY FIREWORKS_API_KEY
GITHUB_APP_CLIENT_SECRET GITHUB_APP_PRIVATE_KEY GITHUB_PAT GITHUB_WEBHOOK_SECRET
JUDGE_ANTHROPIC_API_KEY LANGSMITH_API_KEY
LANGSMITH_API_KEY_PROD LANGCHAIN_API_KEY LINEAR_API_KEY LINEAR_WEBHOOK_SECRET
RUNLOOP_API_KEY SLACK_BOT_TOKEN SLACK_CLIENT_SECRET
SLACK_SIGNING_SECRET TOKEN_ENCRYPTION_KEY USER_ID_API_KEY_MAP X_SERVICE_AUTH_JWT_SECRET
)
batch_get_secrets_tsv() {
local -a ids=()
local v
for v in "${SECRET_VARS[@]}"; do ids+=("${SECRET_PREFIX}${v}"); done
local i page
local -a chunk
for ((i = 0; i < ${#ids[@]}; i += 20)); do
chunk=("${ids[@]:i:20}")
# Capture the response into a variable FIRST so a non-zero `aws` exit (throttle,
# AccessDenied, KMS DecryptionFailure) aborts under set -e instead of being
# silently swallowed — then we'd FAIL-FAST below as "missing secret" with a wrong
# root cause. (Value-less / absent shells come back in .Errors[], not .SecretValues[].)
page="$(aws secretsmanager batch-get-secret-value \
--secret-id-list "${chunk[@]}" \
--region "$REGION" --no-cli-pager --output json)"
printf '%s' "$page" \
| jq -r '.SecretValues[] | select(.SecretString != null) | (.Name|@base64) + "\t" + (.SecretString|@base64)'
done
}
log "reading secrets under ${SECRET_PREFIX} ..."
secret_count=0
# Capture into a variable (NOT `done < <(...)` process substitution) so a non-zero
# exit from batch_get_secrets_tsv propagates under set -e — process substitution hides
# the producer's exit status from the parent shell, which would let a failed AWS call
# fall through to a misleading "missing required var" FAIL-FAST. Mirrors the SSM read.
secrets_tsv="$(batch_get_secrets_tsv)"
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
case "$name" in
"${SECRET_PREFIX}"*) ;; # defensive: exact-prefix only
*) continue ;;
esac
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}"
[ -n "$key" ] || continue
accept_var "$key" "$value" "Secrets"
secret_count=$((secret_count + 1))
done <<<"$secrets_tsv"
log "loaded ${secret_count} secret(s) from Secrets Manager"
# --- Sea Haven DEFAULT_REPO_OWNER guard --------------------------------------
# OSWE-OWNER-04 (revised for multi-org): HONOR the configured owner — the
# OPENSWE_REPO_OWNER env override if set, else the store value — so per-env orgs
# work (dev = seahaven-open-swe-dev, prod = Sea-Haven-Industries). But GUARD the two
# values that must NEVER reach the agent: blank, and the upstream 'langchain-ai' org
# (the fork's origin). Either falls back to the Sea Haven org so a stale/blank/mis-set
# value can never point the agent upstream. Comparison is case- and whitespace-
# insensitive. The POSITIVE org allowlist is enforced by the app (ALLOWED_GITHUB_ORGS).
resolved_owner="${OPENSWE_REPO_OWNER:-${VARS[DEFAULT_REPO_OWNER]:-}}"
# Normalize for the guard CHECK ONLY (the original value is what gets stored when
# allowed): lowercase, strip whitespace, take the FIRST path segment so a value like
# 'langchain-ai/open-swe' still trips the guard, and drop dots (GitHub owners contain
# none) so 'langchain-ai.' can't slip past. Homoglyph/unicode variants are out of scope
# here — the owner comes from admin-written SSM/IaC, not attacker-controlled input.
norm_owner="$(printf '%s' "$resolved_owner" | tr '[:upper:]' '[:lower:]' | tr -d '[:space:]')"
norm_owner="${norm_owner%%/*}"
norm_owner="${norm_owner//./}"
case "$norm_owner" in
"" | "$FORBIDDEN_REPO_OWNER")
log "WARNING: DEFAULT_REPO_OWNER ('${resolved_owner:-<blank>}') is blank or the upstream org -> forcing '${SAFE_REPO_OWNER}'"
resolved_owner="$SAFE_REPO_OWNER"
;;
esac
VARS[DEFAULT_REPO_OWNER]="$resolved_owner"
# --- FAIL-FAST: required vars -------------------------------------------------
# Hard-required regardless of mode:
required=(
DASHBOARD_JWT_SECRET # RuntimeError on startup if missing (oauth.py)
TOKEN_ENCRYPTION_KEY # Fernet key(s); decrypts per-user GitHub tokens
)
# NOTE: the GitHub App is NOT created/duplicated for dev — only prod owns the
# (single, shared) GitHub App + Slack app. So the GitHub App quintet + Slack +
# webhook-signing secrets are required for PROD only (see the prod block below).
# Dev boots without them: it has no GitHub-App/Slack/webhook integration — it is a
# deployment-validation env (boot/health/boundary), not a live-triggered agent.
# Active model-provider key(s): model selection is store-driven (team_settings),
# so fetch-config cannot infer it from .env. Bedrock (Claude builder + reviewer)
# authenticates via the host IAM role — no API key; Fireworks (fallback / subagents /
# any non-Claude model) needs its key. Override with a comma list if the active models change.
IFS=',' read -r -a provider_keys <<<"${REQUIRED_PROVIDER_KEYS:-FIREWORKS_API_KEY}"
for k in "${provider_keys[@]}"; do
k="${k//[[:space:]]/}"
[ -n "$k" ] && required+=("$k")
done
# Sandbox provider key(s) — depends on SANDBOX_TYPE (default langsmith).
sandbox_type="${VARS[SANDBOX_TYPE]:-langsmith}"
case "$sandbox_type" in
langsmith) required+=(LANGSMITH_API_KEY_PROD DEFAULT_SANDBOX_SNAPSHOT_ID) ;;
daytona) required+=(DAYTONA_API_KEY) ;;
runloop) required+=(RUNLOOP_API_KEY) ;;
modal | local) ;; # no key required
*) log "WARNING: unknown SANDBOX_TYPE='${sandbox_type}' — not enforcing a sandbox key" ;;
esac
# Prod-only: the GitHub App (installation-token minting + dashboard OAuth) and the
# webhook-signing secrets. Dev has no GitHub/Slack app, so none of these are
# required there; prod owns the single shared app and must have all of them.
if [ "$ENV" = "prod" ]; then
required+=(
GITHUB_APP_ID # GitHub App trio (installation-token minting) ...
GITHUB_APP_PRIVATE_KEY # ... multiline PEM ...
GITHUB_APP_INSTALLATION_ID # ... used by utils/github_app.py
GITHUB_APP_CLIENT_ID # dashboard OAuth login
GITHUB_APP_CLIENT_SECRET # dashboard OAuth login
GITHUB_WEBHOOK_SECRET # webhook signature verification
SLACK_SIGNING_SECRET # Slack webhook signature verification
)
if [ -n "${VARS[LINEAR_API_KEY]:-}" ] && [ "${OPENSWE_REQUIRE_LINEAR:-1}" = "1" ]; then
required+=(LINEAR_WEBHOOK_SECRET)
fi
fi
missing=()
for k in "${required[@]}"; do
[ -n "${VARS[$k]:-}" ] || missing+=("$k")
done
# de-dup the names for a clean report
if [ "${#missing[@]}" -gt 0 ]; then
mapfile -t missing < <(printf '%s\n' "${missing[@]}" | sort -u)
echo "fetch-config[$ENV]: FAIL-FAST — ${#missing[@]} required var(s) missing/empty:" >&2
printf ' - %s\n' "${missing[@]}" >&2
echo "fetch-config[$ENV]: refusing to write a partial .env; service will not start." >&2
exit 1
fi
# --- Write the .env atomically (root-only on tmpfs) --------------------------
# python-dotenv reads double-quoted values (incl. multiline PEMs). Its decoder
# unescapes ONLY backslash and double-quote (\\ -> \, \" -> "); it does NOT honor
# \$ or \` escapes, so escaping those would leave a spurious backslash. Escape
# exactly backslash then double-quote — real newlines stay literal (multiline OK).
# (Caveat: python-dotenv interpolates a literal `${VAR}` substring; the secret
# domain here — base64/hex/PEM keys — never contains one, so no extra guard.)
emit_var() {
local name="$1" value="$2" esc
esc="${value//\\/\\\\}"
esc="${esc//\"/\\\"}"
printf '%s="%s"\n' "$name" "$esc"
}
tmp="$(mktemp "${ENV_DIR}/.env.XXXXXX")"
chmod 600 "$tmp"
{
printf '# Generated by fetch-config.sh for env=%s at %s — DO NOT EDIT.\n' \
"$ENV" "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
printf '# Source: SSM /open-swe-%s/* + Secrets Manager open-swe-%s/*\n\n' "$ENV" "$ENV"
for k in $(printf '%s\n' "${!VARS[@]}" | sort); do
emit_var "$k" "${VARS[$k]}"
done
} >"$tmp"
mv -f "$tmp" "$ENV_FILE"
# Owned by the unprivileged service user (T5 SC-01) so the app reads it without
# running as root. fetch-config itself runs as root (ExecStartPre=+) to chown.
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_FILE"
chmod 600 "$ENV_FILE"
# Point the app's CWD .env at the tmpfs file (idempotent).
if [ "$APP_ENV_LINK" != "$ENV_FILE" ]; then
if [ -L "$APP_ENV_LINK" ] || [ ! -e "$APP_ENV_LINK" ]; then
ln -sfn "$ENV_FILE" "$APP_ENV_LINK"
elif [ "$(readlink -f "$APP_ENV_LINK" 2>/dev/null || true)" != "$(readlink -f "$ENV_FILE")" ]; then
log "WARNING: ${APP_ENV_LINK} exists and is not a symlink to ${ENV_FILE} — leaving it untouched"
fi
fi
total=$((ssm_count + secret_count))
log "wrote ${ENV_FILE} (${total} vars, sandbox=${sandbox_type}) — ${SERVICE_USER}:${SERVICE_GROUP} 0600"

View file

@ -1,36 +0,0 @@
# Open SWE dashboard frontend (TanStack Start SPA) + scoped API proxy.
# RETIRED on-prem VM variant — kept for on-prem-contrast reference only. The LIVE
# AWS nginx site is the AMI template deploy/ami/templates/open-swe.nginx.conf
# (rendered from an @@SERVER_NAME@@ token at first boot). See DEPLOYMENT.md.
#
# nginx is the security boundary: ONLY /dashboard/api/* reaches the backend;
# the unauthenticated LangGraph agent API (/threads,/runs,/assistants,/store) is NOT proxied.
server {
listen 80 default_server;
listen [::]:80 default_server;
server_name openswe.seahaven.com;
root /var/www/openswe;
index _shell.html;
# ALB health check
location = /healthz { default_type text/plain; return 200 "ok\n"; }
# Dashboard API + OAuth callback -> backend webapp on :2024 (the ONLY proxied path)
location /dashboard/api/ {
proxy_pass http://127.0.0.1:2024;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto https;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_read_timeout 300s;
}
# Static assets + SPA shell fallback (client-side routing)
location / {
try_files $uri $uri/ /_shell.html;
}
}

View file

@ -1,152 +0,0 @@
#!/usr/bin/env bash
# put-config.sh — out-of-band populator for the open-swe config store.
#
# T11 (CDK) creates the RESOURCE SHELLS:
# - 28 Secrets Manager secrets open-swe-<env>/<VAR> (value-LESS shells)
# - the IaC-managed SSM params /open-swe-<env>/<VAR> (real values, owned in CDK)
# This script sets the values that CANNOT live in IaC — every secret value, plus
# the out-of-band SSM params (operationally-variable / env-specific-unknown). Run
# it AFTER `cdk deploy open-swe-<env>` and BEFORE the EC2/T12 box first boots, so
# fetch-config.sh finds every REQUIRED var populated and never writes a partial .env.
#
# The secret list below mirrors config-store.ts SECRET_VARS (28 names — the
# inventory's "29" double-counted JUDGE_ANTHROPIC_BASE_URL, which is config
# (SSM), not a secret). Keep the two lists in lockstep.
#
# SAFETY:
# - NO real secret values live in this file — every secret is a <FILL> placeholder.
# Replace <FILL...> inline at run time, pipe from a vault, or export the
# matching OPENSWE_PUT_<VAR> env var; NEVER commit real values.
# - It does NOT touch the IaC-managed SSM params (SANDBOX_TYPE, DEFAULT_REPO_OWNER,
# ALLOWED_GITHUB_ORGS, DEFAULT_REPO_NAME, DASHBOARD_*_URL/ORIGINS, LLM_MODEL_ID)
# — CDK owns those; setting them here would cause drift.
# - Secrets go to Secrets Manager; the box's instance role grants read on
# open-swe-<env>/* and /open-swe-<env>/* (no kms:Decrypt — AWS-managed keys).
#
# Idempotent: put-secret-value adds a new AWSCURRENT version; put-parameter
# --overwrite updates in place.
set -euo pipefail
ENV="${1:-}"
case "$ENV" in
dev | prod) ;;
*)
echo "usage: put-config.sh <dev|prod>" >&2
exit 2
;;
esac
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
SECRET_PREFIX="open-swe-${ENV}/"
SSM_PREFIX="/open-swe-${ENV}/"
command -v aws >/dev/null 2>&1 || { echo "put-config: 'aws' not found on PATH" >&2; exit 3; }
# --- helpers -----------------------------------------------------------------
# put_secret VAR : set the value of an EXISTING secret shell open-swe-<env>/VAR.
# Resolution order for the value: $OPENSWE_PUT_<VAR> env var, else the literal
# <FILL> placeholder (which aborts so an unset secret is never silently shipped).
put_secret() {
local var="$1"
local override_name="OPENSWE_PUT_${var}"
local value="${!override_name:-<FILL>}"
if [ "$value" = "<FILL>" ]; then
echo "put-config[$ENV]: SKIP secret ${var} (no value; set ${override_name} or edit inline)" >&2
return 0
fi
aws secretsmanager put-secret-value \
--secret-id "${SECRET_PREFIX}${var}" \
--secret-string "$value" \
--region "$REGION" \
--no-cli-pager >/dev/null
echo "put-config[$ENV]: set secret ${var}"
}
# put_param VAR [TYPE] : create/update an out-of-band SSM param /open-swe-<env>/VAR.
# TYPE defaults to String; pass SecureString for anything sensitive-but-not-a-secret.
put_param() {
local var="$1" type="${2:-String}"
local override_name="OPENSWE_PUT_${var}"
local value="${!override_name:-<FILL>}"
if [ "$value" = "<FILL>" ]; then
echo "put-config[$ENV]: SKIP param ${var} (no value; set ${override_name} or edit inline)" >&2
return 0
fi
aws ssm put-parameter \
--name "${SSM_PREFIX}${var}" \
--value "$value" \
--type "$type" \
--overwrite \
--region "$REGION" \
--no-cli-pager >/dev/null
echo "put-config[$ENV]: set param ${var} (${type})"
}
# --- 1) Secrets (open-swe-<env>/<VAR>) — 28 shells from config-store.ts -------
# REQUIRED at boot (fetch-config fail-fast): DASHBOARD_JWT_SECRET,
# TOKEN_ENCRYPTION_KEY, GITHUB_APP_PRIVATE_KEY/CLIENT_SECRET, the active provider
# key(s) (ANTHROPIC_API_KEY + OPENAI_API_KEY by default), LANGSMITH_API_KEY_PROD,
# and (prod) GITHUB_WEBHOOK_SECRET + SLACK_SIGNING_SECRET.
put_secret ANTHROPIC_API_KEY # optional — eval judge only (JUDGE_ANTHROPIC_API_KEY fallback); Bedrock builder/reviewer use the host IAM role
put_secret DASHBOARD_JWT_SECRET # REQUIRED — dashboard session JWT signing
put_secret TOKEN_ENCRYPTION_KEY # REQUIRED — Fernet key(s) for GH-token crypto
put_secret GITHUB_APP_PRIVATE_KEY # REQUIRED — GitHub App PEM (multiline; quote it)
put_secret GITHUB_APP_CLIENT_SECRET # REQUIRED — dashboard OAuth login
put_secret LANGSMITH_API_KEY_PROD # REQUIRED (langsmith sandbox) — prod key
put_secret GITHUB_WEBHOOK_SECRET # prod-REQUIRED — GitHub webhook signature
put_secret SLACK_SIGNING_SECRET # prod-REQUIRED — Slack webhook signature
put_secret LINEAR_WEBHOOK_SECRET # required only when Linear is wired
# Optional / conditional secrets — set the ones this deployment actually uses.
put_secret CORRIDOR_API_TOKEN # optional — Corridor MCP
put_secret CORRIDOR_MCP_TOKEN # optional — Corridor MCP (alt name)
put_secret CORRIDOR_TOKEN # optional — Corridor MCP (alt name)
put_secret DAYTONA_API_KEY # only if SANDBOX_TYPE=daytona
put_secret EXA_API_KEY # optional — Exa web search
put_secret FIREWORKS_API_KEY # active non-Claude key (fallback/subagents); Bedrock uses the host IAM role
put_secret GITHUB_PAT # optional — PAT fallback
put_secret JUDGE_ANTHROPIC_API_KEY # optional — eval judge (falls back to ANTHROPIC)
put_secret LANGSMITH_API_KEY # optional — LangSmith (dev)
put_secret LANGCHAIN_API_KEY # optional — LangSmith alt name
put_secret LINEAR_API_KEY # optional — Linear API
put_secret RUNLOOP_API_KEY # only if SANDBOX_TYPE=runloop
put_secret SLACK_BOT_TOKEN # optional — Slack bot token
put_secret SLACK_CLIENT_SECRET # optional — Slack OAuth
put_secret USER_ID_API_KEY_MAP # optional — JSON map user id -> API key (per-user auth)
put_secret X_SERVICE_AUTH_JWT_SECRET # optional — service-auth JWT
# --- 2) Out-of-band SSM params (/open-swe-<env>/<VAR>) ------------------------
# These are NOT created by CDK (operationally-variable / env-specific-unknown).
# put_param creates them on first run.
put_param DEFAULT_SANDBOX_SNAPSHOT_ID # REQUIRED (langsmith) — changes every rebuild
put_param GITHUB_APP_ID # REQUIRED — GitHub App numeric id
put_param GITHUB_APP_INSTALLATION_ID # REQUIRED — GitHub App installation id
put_param GITHUB_APP_CLIENT_ID # REQUIRED — dashboard OAuth client id
put_param GITHUB_OAUTH_PROVIDER_ID # GitHub OAuth provider id
put_param LANGSMITH_TENANT_ID_PROD # LangSmith prod tenant id
put_param LANGSMITH_URL_PROD # LangSmith prod URL
put_param LANGSMITH_ENDPOINT # LangSmith API endpoint
put_param LANGSMITH_ENDPOINT_PROD # LangSmith prod API endpoint
put_param LANGSMITH_HOST_API_URL # LangSmith host API URL
put_param LANGGRAPH_URL # LangGraph server URL
put_param LANGGRAPH_URL_PROD # LangGraph server URL (prod)
put_param LANGCHAIN_REVISION_ID # LangChain revision id
put_param SLACK_CLIENT_ID # Slack OAuth client id
put_param SLACK_TEAM_ID # Slack workspace/team id
put_param SLACK_BOT_USER_ID # Slack bot user id
put_param SLACK_BOT_USERNAME # Slack bot username
put_param SLACK_REPO_OWNER # Slack default repo owner
put_param SLACK_REPO_NAME # Slack default repo name
put_param CONFIGURED_ADMINS # dashboard admin GitHub logins (comma list)
put_param OBSERVABILITY_AUTHORIZED_EMAILS # observability allowlist (comma list)
put_param PUBLIC_REPO_ORG_GATE # public-repo trigger gate (org name; empty=off)
put_param ALLOWED_GITHUB_REPOS # extra repo allowlist (comma list)
put_param LLM_FALLBACK_MODEL_ID # optional — model fallback id
put_param DATADOG_MCP_TOOLSETS # optional — Datadog MCP toolsets
put_param NOTION_MCP_CLIENT_NAME # optional — Notion MCP client name
put_param API_STANDARDS_SKILL_HANDLE # optional — API standards skill handle
put_param REPO_SNAPSHOT_BASE_IMAGE # optional — repo snapshot base image
put_param REPO_SNAPSHOT_BUILD_TIMEOUT_SECONDS # optional — snapshot build timeout
put_param REPO_SNAPSHOT_STALE_BUILD_SECONDS # optional — snapshot stale threshold
echo "put-config[$ENV]: done. Verify with fetch-config.sh ${ENV} before first boot."

View file

@ -1,184 +0,0 @@
#!/usr/bin/env bash
# Seed the LangGraph store after a (re)start.
#
# The stock `langgraph dev` server uses an IN-MEMORY store, so anything written
# to it (team model settings, user mappings) is lost on every restart. This
# script idempotently re-PUTs that state and is wired as a systemd
# ExecStartPost on the open-swe.service unit so it runs after each start.
# IT MUST RE-RUN ON EVERY RESTART — the in-memory store starts empty each boot.
#
# Replace this with Postgres-backed durability (Aegra / `langgraph up`) to make
# the store survive restarts and drop this script.
#
# AWS-env-aware: pass the env as $1 (dev|prod). Seed values (default repo, model
# ids, user mappings) come from the fetch-config-materialized .env.
#
# SECURITY (T5 /sh-security-review):
# - SH-INJ-001: this script NEVER `source`s the .env. python-dotenv and bash
# have incompatible escaping, and a config value like `$(cmd)` would execute
# when sourced. We extract the few single-line seed keys with a non-eval
# reader (read_env) instead.
# - SH-INJ-003: the store PUT bodies are built with `jq --arg`, so values are
# always JSON-encoded (no string interpolation into a JSON heredoc).
# - SH-INJ-004: BASE is pinned to loopback — never derived from store/SSM
# config (LANGGRAPH_URL) — so a tampered value can't redirect the PUTs.
# - SC-02: not sourcing the .env means secrets are never exported into this
# script's (or curl's) environment.
#
# Seed values read from the materialized .env (set in SSM /open-swe-<env>/*):
# DEFAULT_REPO_OWNER / DEFAULT_REPO_NAME -> team_settings default_repo
# LLM_MODEL_ID -> default builder model (fallback)
# SEED_AGENT_MODEL / SEED_AGENT_EFFORT -> builder model + effort (optional)
# SEED_REVIEWER_MODEL / SEED_REVIEWER_EFFORT -> reviewer model + effort (optional)
# SEED_USER_MAPPINGS -> "login:email,login:email" (optional)
# CONFIGURED_ADMINS -> "login,email" fallback for the mapping
# Legacy OPENSWE_* process-env overrides are still honored (highest precedence).
set -euo pipefail
ENV="${1:-${OPENSWE_ENV:-}}"
case "$ENV" in
dev | prod | "") ;; # empty allowed: pure-env / on-prem backward-compat mode
*)
echo "seed_store: ENV must be 'dev' or 'prod' (got '$ENV')" >&2
exit 2
;;
esac
ENV_DIR="${ENV_DIR:-/run/open-swe}"
ENV_FILE="${ENV_FILE:-${ENV_DIR}/.env}"
# read_env KEY -> prints the value of a SINGLE-LINE `KEY="..."` entry from the
# materialized .env WITHOUT shell evaluation (SH-INJ-001 fix). Seed keys are
# simple single-line values; multiline secrets (e.g. the PEM) are never read
# here. Returns empty if the key is absent/unreadable.
read_env() {
local key="$1" line
[ -r "$ENV_FILE" ] || return 0
line="$(grep -m1 -- "^${key}=" "$ENV_FILE" 2>/dev/null || true)"
[ -n "$line" ] || return 0
line="${line#*=}"
# strip one layer of surrounding double quotes (python-dotenv double-quoted form)
if [ "${line#\"}" != "$line" ]; then line="${line%\"}"; line="${line#\"}"; fi
# reverse python-dotenv double-quote escaping (only \" and \\ are escaped)
line="${line//\\\"/\"}"; line="${line//\\\\/\\}"
printf '%s' "$line"
}
# Process-env override (legacy/on-prem) -> .env value -> default.
pick() { # pick DEFAULT OVERRIDE_VALUE FILE_KEY...
local def="$1" override="$2"; shift 2
if [ -n "$override" ]; then printf '%s' "$override"; return; fi
local k v
for k in "$@"; do v="$(read_env "$k")"; [ -n "$v" ] && { printf '%s' "$v"; return; }; done
printf '%s' "$def"
}
# BASE is loopback-pinned (SH-INJ-004): this on-box seeder only talks to the
# local server; OPENSWE_PORT may override the port but never the host.
BASE="http://127.0.0.1:${OPENSWE_PORT:-2024}"
AGENT_MODEL="$(pick 'bedrock_converse:us.anthropic.claude-opus-4-8' "${OPENSWE_AGENT_MODEL:-}" SEED_AGENT_MODEL LLM_MODEL_ID)"
AGENT_EFFORT="$(pick 'high' "${OPENSWE_AGENT_EFFORT:-}" SEED_AGENT_EFFORT)"
REVIEWER_MODEL="$(pick 'bedrock_converse:us.anthropic.claude-opus-4-8' "${OPENSWE_REVIEWER_MODEL:-}" SEED_REVIEWER_MODEL)"
REVIEWER_EFFORT="$(pick 'high' "${OPENSWE_REVIEWER_EFFORT:-}" SEED_REVIEWER_EFFORT)"
# default_repo = owner/name from AWS config (DEFAULT_REPO_OWNER is hard-pinned
# away from upstream by fetch-config.sh).
REPO_OWNER="$(pick '' '' DEFAULT_REPO_OWNER)"
REPO_NAME="$(pick '' '' DEFAULT_REPO_NAME)"
if [ -n "${OPENSWE_DEFAULT_REPO:-}" ]; then
DEFAULT_REPO="$OPENSWE_DEFAULT_REPO"
elif [ -n "$REPO_OWNER" ] && [ -n "$REPO_NAME" ]; then
DEFAULT_REPO="${REPO_OWNER}/${REPO_NAME}"
else
echo "seed_store: set OPENSWE_DEFAULT_REPO=owner/repo (or DEFAULT_REPO_OWNER + DEFAULT_REPO_NAME in .env)" >&2
exit 1
fi
# user_mappings: explicit SEED_USER_MAPPINGS ("login:email,..."), then legacy
# OPENSWE_OWNER_LOGIN/EMAIL, then parse CONFIGURED_ADMINS ("login,email").
SEED_MAP="$(pick '' "${SEED_USER_MAPPINGS:-}" SEED_USER_MAPPINGS)"
ADMINS="$(pick '' "${CONFIGURED_ADMINS:-}" CONFIGURED_ADMINS)"
declare -a MAPPINGS=()
if [ -n "$SEED_MAP" ]; then
IFS=',' read -r -a _pairs <<<"$SEED_MAP"
for p in "${_pairs[@]}"; do
p="${p//[[:space:]]/}"
[ -n "$p" ] && MAPPINGS+=("$p")
done
elif [ -n "${OPENSWE_OWNER_LOGIN:-}" ] && [ -n "${OPENSWE_OWNER_EMAIL:-}" ]; then
MAPPINGS+=("${OPENSWE_OWNER_LOGIN}:${OPENSWE_OWNER_EMAIL}")
elif [ -n "$ADMINS" ]; then
_login="" _email=""
IFS=',' read -r -a _toks <<<"$ADMINS"
for t in "${_toks[@]}"; do
t="${t//[[:space:]]/}"
[ -z "$t" ] && continue
case "$t" in
*@*) [ -z "$_email" ] && _email="$t" ;;
*) [ -z "$_login" ] && _login="$t" ;;
esac
done
[ -n "$_login" ] && [ -n "$_email" ] && MAPPINGS+=("${_login}:${_email}")
fi
# No user mapping is NON-FATAL (OSWE-SEED-03 precedent: never fail the unit into a
# restart loop over a seeding gap — same as the server-not-ready path below). The
# server itself is healthy; an unseeded user_mappings table only means the @openswe
# trigger won't resolve a commenter, which a deployment-validation env (e.g. dev)
# does not need. team_settings is still seeded. Set SEED_USER_MAPPINGS (or
# CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL) to seed the mapping when wanted.
if [ "${#MAPPINGS[@]}" -eq 0 ]; then
echo "seed_store: no user mapping resolved — skipping user_mappings seed (set SEED_USER_MAPPINGS or OPENSWE_OWNER_LOGIN/EMAIL to enable)" >&2
fi
NOW="$(date -u +%Y-%m-%dT%H:%M:%S+00:00)"
# Wait for the server to accept requests (up to ~60s). Authoritative (OSWE-SEED-03):
# if it never comes up, log and exit 0 — do NOT fail the unit into a restart loop.
READY=0
for _ in $(seq 1 30); do
if [ "$(curl -s -o /dev/null -w '%{http_code}' "$BASE/ok" || true)" = "200" ]; then READY=1; break; fi
sleep 2
done
if [ "$READY" -ne 1 ]; then
echo "seed_store: server not ready at $BASE after ~60s; skipping seed (will reseed on next restart)" >&2
exit 0
fi
# 1) team_settings/default — JSON built with jq --arg (SH-INJ-003 fix).
team_body="$(jq -n \
--arg am "$AGENT_MODEL" --arg ae "$AGENT_EFFORT" \
--arg rm "$REVIEWER_MODEL" --arg re "$REVIEWER_EFFORT" \
--arg repo "$DEFAULT_REPO" --arg now "$NOW" \
'{namespace:["team_settings"],key:"default",value:{
review_draft_prs:false, pr_summaries:true, review_trace_links:true,
org_guidelines:null,
default_agent_model:$am, default_agent_reasoning_effort:$ae,
default_agent_subagent_model:$am, default_agent_subagent_reasoning_effort:$ae,
default_repo:$repo,
default_reviewer_model:$rm, default_reviewer_reasoning_effort:$re,
default_reviewer_subagent_model:$rm, default_reviewer_subagent_reasoning_effort:$re,
default_grouping_model:null, default_grouping_reasoning_effort:null,
default_chat_model:null, default_chat_reasoning_effort:null,
updated_at:$now}}')"
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$team_body" >/dev/null \
|| echo "seed_store: WARN team_settings PUT failed (will reseed next restart)" >&2
# 2) user_mappings/<login> — required, or the @openswe trigger ignores the commenter.
for pair in "${MAPPINGS[@]}"; do
login="${pair%%:*}"
email="${pair#*:}"
if [ -z "$login" ] || [ -z "$email" ] || [ "$login" = "$pair" ]; then
echo "seed_store: skipping malformed mapping '$pair' (want login:email)" >&2
continue
fi
map_body="$(jq -n --arg login "$login" --arg email "$email" --arg now "$NOW" \
'{namespace:["user_mappings"],key:$login,value:{
github_login:$login, work_email:$email, slack_user_id:null,
source:"slack_oauth", status:"active", created_at:$now, updated_at:$now}}')"
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$map_body" >/dev/null \
|| echo "seed_store: WARN user_mapping PUT failed for $login" >&2
done
echo "seed_store: done at $NOW (env=${ENV:-none}, repo=$DEFAULT_REPO, mappings=${#MAPPINGS[@]})"

View file

@ -1,17 +0,0 @@
[Unit]
Description=Open SWE stock LangGraph dev server (graphs + webapp, :2024)
After=network-online.target postgresql.service
Wants=network-online.target
[Service]
Type=simple
User=adam
WorkingDirectory=/home/adam/open-swe
ExecStart=/home/adam/open-swe/.venv/bin/langgraph dev --host 0.0.0.0 --port 2024 --no-browser --no-reload
ExecStartPost=/home/adam/open-swe/seed_store.sh
Restart=on-failure
RestartSec=5
TimeoutStartSec=120
[Install]
WantedBy=multi-user.target

21
infra/.gitignore vendored
View file

@ -1,21 +0,0 @@
# CDK / build output
cdk.out/
*.js
*.d.ts
*.js.map
# ...but jest.config.js is hand-authored config, not build output — keep it.
!jest.config.js
# deps
node_modules/
# env
.env
# coverage
coverage/
# NOTE: cdk.context.json IS committed on purpose (pins the AMI / lookups so
# deploys are reproducible and don't implicitly pick up a newer AMI — see
# lib/constructs/ami-cache.ts and feedback_inline_ebs_volumes).
!cdk.context.json

View file

@ -1,348 +0,0 @@
# open-swe infra (CDK TypeScript)
AWS infrastructure for the Open SWE → AWS migration. **Synth-only at this stage —
nothing here is deployed yet.** All IAM is applied only after the Phase-1 security
gate (T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review`) clears (T6).
## Layout
```
infra/
├── bin/
│ └── app.ts # CDK app entry — instantiates the 3 stacks, applies the naming Aspect
├── lib/
│ ├── config.ts # account/region/org constants, env type, OIDC trust subjects
│ ├── open-swe-iam-stack.ts # account-level: shared OIDC deploy roles
│ ├── open-swe-stack.ts # per-env stack (instance role + config store + AppService)
│ ├── aspects/
│ │ └── kebab-naming-aspect.ts # fails synth on any non-kebab-case explicit name
│ └── constructs/
│ ├── github-deploy-roles.ts # githubdeploy-open-swe-infra + githubdeploy-open-swe-app
│ ├── instance-role.ts # open-swe-<env>-instance-role (least-privilege)
│ ├── config-store.ts # Secrets Manager + SSM Parameter Store shells (T11)
│ ├── app-service.ts # EC2 box + imported-ALB ingress + Route53 + logs (T12)
│ └── ami-cache.ts # baked open-swe AMI pin (by id) + EBS/replacement docs
├── test/
│ └── kebab-naming-aspect.test.ts # jest: Aspect passes conforming names, flags bad ones
├── cdk.json
├── cdk.context.json # COMMITTED — {} (AMI is a static id pin; no lookups)
├── package.json # aws-cdk-lib pinned EXACT (2.260.0)
├── tsconfig.json
├── jest.config.js
└── .gitignore
```
## Stacks
| Stack name (kebab) | Construct | Contents |
|---|---|---|
| `open-swe-iam` | `OpenSweIamStack` | Account-level shared GitHub OIDC deploy roles (singletons). |
| `open-swe-dev` | `OpenSweStack` (`envName: dev`) | `open-swe-dev-instance-role`, config store (T11), and the EC2 box + ALB ingress (T12, `AppService`). |
| `open-swe-prod` | `OpenSweStack` (`envName: prod`) | `open-swe-prod-instance-role`, config store, and the EC2 box + ALB ingress. |
Account `328440206208`, region `us-east-1`. Stack names are set explicitly so CDK
never defaults to PascalCase; resource names follow `open-swe-<env>-*`.
> The two env stacks (`open-swe-dev` / `open-swe-prod`) are the required pair. The
> shared OIDC deploy roles are account-wide singletons (one `RoleName` each), so
> they live in their own dedicated `open-swe-iam` stack rather than being
> duplicated across the env stacks — and that stack deploys first (see ordering).
## IAM roles defined (unapplied)
- **`githubdeploy-open-swe-infra`** — GitHub OIDC role for CDK/CFN infra deploys.
Trust scoped to `repo:Sea-Haven-Industries/open-swe` on the `main`/`dev`
branches only. Permission is the org-standard CDK pattern: `sts:AssumeRole` on
the CDK bootstrap roles (`cdk-hnb659fds-*`) — the real CFN/IAM/resource scope
lives in the bootstrap `cfn-exec-role`, not in this role.
- **`githubdeploy-open-swe-app`** — GitHub OIDC role for app deploys. Tag-scoped
`ssm:SendCommand` (instances tagged `project=open-swe` + `env in {dev,prod}`) +
read-only access to the `open-swe-<env>-assets` S3 artifact buckets.
- **`open-swe-<env>-instance-role`** — EC2 instance role, least-privilege: read
`open-swe-<env>-assets` (S3), read `/open-swe-<env>/*` (SSM), read
`open-swe-<env>/*` (Secrets Manager), put `/open-swe/<env>/*` CloudWatch Logs,
plus `AmazonSSMManagedInstanceCore` for SSM agent registration. No admin.
The GitHub OIDC provider already exists account-wide (created for seahaven-site);
it is referenced by ARN, never re-created.
## CDK bootstrap qualifiers — per-env deploy isolation (B-1 / OSWE-IAC-01)
Each env's infra deploy role may assume **only its own bootstrap qualifier's**
roles, so a dev-branch token can never assume the bootstrap roles whose admin
`cfn-exec-role` deploys prod (closing the cross-env escalation that bypassed
prod's Environment approval gate). Mapping lives in `config.ts:bootstrapQualifier`:
| Env | Qualifier | Toolkit stack | Infra role assumes |
|---|---|---|---|
| dev | `oswedev` | `CDKToolkit-oswedev` | `cdk-oswedev-*` |
| prod | `hnb659fds` (default) | `CDKToolkit` | `cdk-hnb659fds-*` |
The dev stack synthesizes with `DefaultStackSynthesizer({ qualifier: "oswedev" })`
(`bin/app.ts`); prod uses the default. Bootstrap a new env qualifier with:
```bash
npx cdk bootstrap --qualifier <qual> --toolkit-stack-name CDKToolkit-<qual> \
--cloudformation-execution-policies arn:aws:iam::aws:policy/AdministratorAccess \
aws://328440206208/us-east-1
```
**Deploy order matters** when changing an env's qualifier: bootstrap the new
qualifier and deploy the env stack onto it **before** re-scoping that env's infra
role in `open-swe-iam` — otherwise a pipeline deploy with the re-scoped role would
fail to assume the not-yet-targeted bootstrap roles.
## Kebab-case naming Aspect
`KebabNamingAspect` (applied app-wide in `bin/app.ts`) fails synth via
`Annotations.addError` when a stack name or an explicit physical resource name
(`RoleName`, `BucketName`, …) is not kebab-case. Path-style names (Secrets
Manager `a/b`, SSM `/a/b`, log groups `/aws/.../x`) are validated per `/`-segment.
CDK logical construct ids are intentionally NOT validated (they are conventionally
PascalCase). Covered by `test/kebab-naming-aspect.test.ts`.
## Config store (Secrets Manager + SSM shells — T11)
`ConfigStore` (`lib/constructs/config-store.ts`, one per env from `OpenSweStack`)
renders the resource shells the boot hook `deploy/seahaven/fetch-config.sh` reads.
The naming contract (T9 inventory + the fetch-config header) is LITERAL env-var
names as the last path segment — `open-swe-<env>/<VAR>` for secrets,
`/open-swe-<env>/<VAR>` (FLAT) for config — because fetch-config strips the prefix
and exports that segment verbatim.
Three buckets:
1. **Secret shells (Secrets Manager) — 27 secrets.** Created value-LESS (an L1
`CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which
CloudFormation creates as an empty secret). The real value is set **out-of-band**
(`put-config.sh`) — CDK never owns it, so a later `cdk deploy` can never clobber
it. `UpdateReplacePolicy/DeletionPolicy: Retain` so a teardown can't destroy
operator-set secret material. AWS-managed key (no CMK — matches the instance
role, which omits `kms:Decrypt`).
> The T9 header says "29 secrets" but its table enumerates **27** distinct VAR
> names (the CORRIDOR row holds 3). We create 27 — we don't invent two to hit 29.
> **Confirm** the 27-vs-29 count (code-only candidates not in the table:
> `USER_ID_API_KEY_MAP`, `JUDGE_ANTHROPIC_BASE_URL`).
> ⚠️ **`Retain` + fixed name orphans these shells on a failed FIRST create.**
> If the stack's initial create fails and rolls back, `Retain` keeps the shells
> instead of deleting them. The stack is then gone, but the secrets survive,
> still holding the global `open-swe-<env>/<VAR>` names — so every later create
> fails with `AlreadyExists`. A plain `delete-secret` does **not** clear it
> (the name stays reserved for the 7–30 day recovery window). This bites on a
> **teardown/rebuild, a secret logical-id change/refactor, or standing up a new
> env** — never on routine updates of an already-created stack. **Recovery —**
> before re-creating the stack, force-delete the *empty* orphans so the names
> free immediately:
> ```bash
> aws secretsmanager list-secrets --region us-east-1 \
> --filters Key=name,Values=open-swe-<env>/ \
> --query 'SecretList[].Name' --output text | tr '\t' '\n' | while read -r n; do
> aws secretsmanager delete-secret --secret-id "$n" \
> --region us-east-1 --force-delete-without-recovery
> done
> ```
> Force-delete only shells with **no value version** — a populated secret holds
> real operator material. (Hit on prod 2026-06-29; see PR #51's deploy failure.)
2. **IaC-managed SSM config — 8 params, real values owned in code:**
| Param | dev | prod |
|---|---|---|
| `SANDBOX_TYPE` | `langsmith` | `langsmith` |
| `DEFAULT_REPO_OWNER` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `ALLOWED_GITHUB_ORGS` | `Sea-Haven-Industries` | `Sea-Haven-Industries` |
| `DEFAULT_REPO_NAME` | `open-swe-pilot` *(confirm)* | `open-swe-pilot` *(confirm)* |
| `DASHBOARD_BASE_URL` | `https://openswe-dev.seahaven.com` *(confirm host)* | `https://openswe.seahaven.com` *(confirm host)* |
| `DASHBOARD_API_BASE_URL` | same as base | same as base |
| `DASHBOARD_ALLOWED_ORIGINS` | same as base | same as base |
| `LLM_MODEL_ID` | `bedrock_converse:us.anthropic.claude-opus-4-8` | `bedrock_converse:us.anthropic.claude-opus-4-8` |
3. **Out-of-band SSM config — NOT created by CDK.** Operationally-variable or
env-specific-unknown values listed in `OUT_OF_BAND_SSM` and populated by
`put-config.sh`. The keystone is `DEFAULT_SANDBOX_SNAPSHOT_ID` (changes on every
snapshot rebuild → must NOT be CDK-managed or a deploy clobbers it); also the
GitHub App ids, Slack ids, and LangSmith tenant/urls.
### Kebab-Aspect deviation
`KebabNamingAspect` exempts `AWS::SecretsManager::Secret` and `AWS::SSM::Parameter`
from the kebab check (see the `KEBAB_EXEMPT_RESOURCE_TYPES` set) — the UPPER_SNAKE
env-var segment is a required, documented deviation for a lossless store→env
round-trip. Every other explicitly-named resource is still validated. Covered by a
dedicated case in `test/kebab-naming-aspect.test.ts`.
### Deploy ordering (values BEFORE the box boots)
The shells are synth-able now (T11). Population is out-of-band and happens **after**
`cdk deploy open-swe-<env>` but **before** the EC2/T12 box first boots:
```bash
cdk deploy open-swe-<env> # creates the 27 secret shells + 8 IaC params
deploy/seahaven/put-config.sh <dev|prod> # sets the 27 secret values + out-of-band SSM
deploy/seahaven/fetch-config.sh <dev|prod> # (on the box) fail-fast verify before first start
```
`put-config.sh` ships `<FILL>` placeholders only (no real secret values committed);
provide each value inline, via `OPENSWE_PUT_<VAR>` env vars, or from a vault. It does
NOT touch the IaC-managed params (CDK owns those — editing them here would drift).
## Compute + ingress (`AppService` — T12)
`AppService` (`lib/constructs/app-service.ts`, one per env from `OpenSweStack`)
builds the box and its path to the internet. A **single** internet-facing ALB
(`app/seahaven-com`) and a **single** VPC are shared with the on-prem
`seahaven-site` stack, so open-swe **imports** the VPC, the ALB security group
(`sg-0b0301deed193258a`), the `:443` listener, and the public `seahaven.com`
zone — and never owns/mutates them. It **adds**:
- **One ARM64 EC2 box** (`open-swe-<env>-box`, `t4g.medium` dev / `t4g.large`
prod) in **private1 (us-east-1a)** — same AZ as the single NAT for in-AZ egress.
`requireImdsv2`, gp3 **encrypted** root, `deleteOnTermination` (no RETAIN
volume — see below). `userDataCausesReplacement: true`; user-data is rendered
from `deploy/ami/user-data.sh`.
- **A standalone instance SG** reachable **only** from the shared ALB SG on `:80`
(nginx). Egress open (NAT). The ALB SG is opened to the box via a **standalone
`CfnSecurityGroupEgress`** so the imported (on-prem-owned) SG is never mutated.
- **A target group → instance `:80`** (nginx is the sole ingress; the LangGraph
control plane stays on loopback `:2024`). Health check `GET /healthz`.
- **Two listener rules** on the imported `:443` listener, both → the same TG:
- **Webhooks** (priority **2** dev / **3** prod): `host ∈ {openswe-<env>, hooks-<env>}.seahaven.com` **AND** path `/webhooks/*`.
- **Site** (priority **10** dev / **11** prod): `host = openswe-<env>.seahaven.com` (dashboard SPA + `/dashboard/api/`).
- **Route53 alias records** `openswe[-dev]` + `hooks[-dev]` → the shared ALB.
- **Four CloudWatch log groups** (`/open-swe/<env>/{app,user-data,nginx-access,nginx-error}`) at **30-day** retention (IaC-owned; mirrors the CW-agent config).
### Listener-rule ordering (load-bearing)
The shared listener already has a **host-agnostic** `/webhooks/*` PATH rule at
**priority 5** (on-prem). ALB rules are first-match by ascending priority, so the
open-swe webhook rule **must** sit below 5 or every `…/webhooks/*` request (any
host) is forwarded to the on-prem target first. Hence priority 2/3. The rule ANDs
a host condition, so it does **not** steal the on-prem hosts' webhooks. The
dashboard "site" rule carries no path that collides with rule 5, so it sits at
10/11.
**Cross-stack coordination (T13 review).** The `seahaven-site` (on-prem) and
`open-swe` stacks both add resources to the *same imported* listener and ALB SG.
This is safe: each stack owns only the resources it declares (its own logical
ids), so an on-prem deploy can't delete open-swe's rules/egress and vice-versa,
and the standalone `CfnSecurityGroupEgress` never mutates the shared SG's own
definition (the pattern on-prem itself uses). The one shared namespace that needs
care is **listener-rule priority** (globally unique per listener; a collision is
a fail-*safe* deploy error, not silent drift). Ownership — keep disjoint:
`seahaven-site` = **4-7 + default**; `open-swe` = **2, 3, 10, 11**. open-swe's
webhook rules are host-scoped to its own `*.seahaven.com` hosts, so they never
match an on-prem `seahavenind.com` host.
### Security review (T5/T12 `/sh-security-review`)
The T12 surface was run through the detector-fan-out + proof-or-kill verifier.
One **confirmed medium** (OSWE-T12-01: nginx's 1 MB default `client_max_body_size`
would 413 large GitHub webhooks before in-app signature verification) is fixed in
`open-swe.nginx.conf` (`25m` on `/webhooks/`, `10m` on `/dashboard/api/`). The
hooks hostname is scoped to `/webhooks/*` only (OSWE-T12-02 hygiene). An
X-Forwarded-For spoof candidate was **killed** — no code trusts the leftmost XFF.
No confirmed critical/high; no block.
## Baked AMI + EBS-replacement discipline
`bakedOpenSweArm64()` (in `lib/constructs/ami-cache.ts`) pins the custom
**open-swe-base-arm64** image by EXACT id (`BAKED_OPEN_SWE_AMI_ID`) via
`MachineImage.genericLinux({ "us-east-1": "<ami-id>" })` — no SSM lookup, so synth
and deploy are fully offline/deterministic. The image is built by
`deploy/ami/open-swe-base.pkr.hcl` (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW
agent + boot templates, **no secrets**); the box's `user-data.sh` assumes that
baked layout (`/opt/open-swe`, `openswe` user, nginx, CW agent).
Pinning by exact id (vs a `most_recent` name filter) is what prevents a routine
deploy from silently swapping the AMI → **EC2 instance replacement** (the
file-share data-loss root cause — memory `feedback_inline_ebs_volumes`).
- `userDataCausesReplacement: true` is **deliberate** — user-data is
provisioning-only and carries no durable state.
- **No durable state on the box → no RETAIN volume.** The in-memory langgraph
store is rebuilt on every boot from S3 + Secrets Manager / SSM, so there is
intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The goal is
replacement-*tolerance*, not avoidance.
- **Snapshot-before-replace** still applies operationally: before any replacing
deploy snapshot the root volume and wait `state=completed`, and re-verify "no
local-only durable state" first.
Refresh the AMI deliberately:
```bash
cd deploy/ami && packer build open-swe-base.pkr.hcl # prints the new ami-… id
# update BAKED_OPEN_SWE_AMI_ID in infra/lib/constructs/ami-cache.ts
cd infra && npx cdk diff OpenSweDevStack # WILL show "requires replacement"
```
> `cdk.context.json` is `{}` — nothing is resolved via context anymore (the AMI is
> a static id pin), so synth makes no live AWS call.
## Commands
```bash
npm install
npx cdk synth open-swe-iam
npx cdk synth open-swe-dev
npx cdk synth open-swe-prod
npm test # jest — naming Aspect
```
## CI/CD (T18 — `.github/workflows/ci-infra.yml` + `cd-infra.yml`)
Path-filtered, OIDC-only (no static keys). The Python agent keeps its own
`ci.yml` ("CI"); these two add the `/infra` half.
| Workflow | Trigger | Does |
|---|---|---|
| `ci-infra.yml` | PR touching `infra/**` | `tsc` + `jest` + `cdk synth` (reusable `ci-typescript-cdk.yaml`). |
| `cd-infra.yml` | push to `dev`/`main` touching `infra/**`, or dispatch | CI (pre-deploy) → per-env `cdk deploy`. |
`cd-infra.yml` flow:
- **push to `dev`** → CI green → **auto** `cdk deploy OpenSweDevStack` (assumes
`githubdeploy-open-swe-infra-dev`; the job declares **no** `environment:`, so the
OIDC subject is `…:ref:refs/heads/dev` — matching that role's trust).
- **push to `main`** → CI green → `cdk deploy OpenSweProdStack` behind the
**`prod` GitHub Environment** (required reviewer = Adam). The `environment: prod`
declaration both fires the manual-approval gate and makes the OIDC subject
`…:environment:prod` — matching `githubdeploy-open-swe-infra-prod`'s trust.
**Why not the reusable `cd-cdk.yaml`:** it runs `cdk deploy --all`, which from a
single-env push would deploy the *other* env + the shared IAM stack — breaking the
per-env boundary. So CD targets one stack explicitly per env. The shared
`open-swe-iam` stack is **not** deployed by CD (privileged, human-gated — T6).
**Gating note:** infra CI is enforced at the *deploy* boundary (`cd-infra`'s
`deploy-*` jobs `needs: ci`), not as a branch-protection required check —
path-filtering a *required* check would deadlock app-only PRs (a skipped required
check never satisfies). Making `Infra CI` a required check later needs a
skip-aware shim or dropping its path filter.
**Prerequisites (set post-T6, when the roles exist):**
- repo **variables** `AWS_DEPLOY_ROLE_INFRA_DEV` / `AWS_DEPLOY_ROLE_INFRA_PROD`
= the `githubdeploy-open-swe-infra-<env>` role ARNs (`open-swe-iam` outputs).
- a GitHub **Environment** named `prod` with Adam as a required reviewer.
> App-side CD (CI → S3 artifact → SSM deploy via `githubdeploy-open-swe-app-<env>`)
> is **T19**, not here.
## Deploy ordering (when the gate clears — NOT yet)
1. **`open-swe-iam` first** — apply the IAM stack (T6, human-gated), then set the
repo `AWS_DEPLOY_ROLE_INFRA_{DEV,PROD}` variables from its role-ARN outputs and
configure the `prod` Environment reviewer (BLOCK#3).
2. **Security gate** — T4 GPT-4.1 IAM cross-review + T5 `/sh-security-review` on
the synth; resolve every confirmed critical/high.
3. **IAM applied** (T6) — only after the gate.
4. Env stacks: first `open-swe-dev` (T14, manual validate), then CD auto-deploys
dev on push; `open-swe-prod` (T21) behind the `prod` Environment approval.
## Version policy
`aws-cdk-lib` is pinned EXACT (`2.260.0`) — no `^`/`~`. Dependabot keeps it
current; CI (`npm ci` + `cdk synth`) + dependency review gate each bump. See
`aws-infrastructure.md` "CDK Version Policy" and memory
`feedback_cdk_lib_bundled_deps`.

View file

@ -1,43 +0,0 @@
#!/usr/bin/env node
import "source-map-support/register";
import * as cdk from "aws-cdk-lib";
import { ACCOUNT, REGION, bootstrapQualifier } from "../lib/config";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
import { KebabNamingAspect } from "../lib/aspects/kebab-naming-aspect";
const app = new cdk.App();
const env = { account: ACCOUNT, region: REGION };
// Account-level shared OIDC deploy roles (singletons). Deployed FIRST.
new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env,
});
// The two env stacks — explicit kebab-case stackName (never let CDK default to
// PascalCase), env-parameterised so resources are `open-swe-<env>-*`.
//
// B-1 / OSWE-IAC-01: dev synthesizes against its OWN bootstrap qualifier
// (`oswedev`), so it deploys via the cdk-oswedev-* roles the dev infra role is
// scoped to — and NOT the default hnb659fds bootstrap roles that deploy prod.
// Prod stays on the default qualifier (no synthesizer override).
new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env,
envName: "dev",
synthesizer: new cdk.DefaultStackSynthesizer({
qualifier: bootstrapQualifier("dev"),
}),
});
new OpenSweStack(app, "OpenSweProdStack", {
stackName: "open-swe-prod",
env,
envName: "prod",
});
// Fail synth on any non-kebab-case explicit resource/stack name.
cdk.Aspects.of(app).add(new KebabNamingAspect());
app.synth();

View file

@ -1 +0,0 @@
{}

View file

@ -1,21 +0,0 @@
{
"app": "npx ts-node --prefer-ts-exts bin/app.ts",
"watch": {
"include": ["**"],
"exclude": [
"README.md",
"cdk*.json",
"**/*.d.ts",
"**/*.js",
"tsconfig.json",
"package*.json",
"node_modules",
"cdk.out"
]
},
"context": {
"@aws-cdk/aws-lambda:recognizeLayerVersion": true,
"@aws-cdk/core:checkSecretUsage": true,
"@aws-cdk/core:target-partitions": ["aws"]
}
}

View file

@ -1,9 +0,0 @@
module.exports = {
testEnvironment: "node",
roots: ["<rootDir>/test"],
testMatch: ["**/*.test.ts"],
preset: "ts-jest",
transform: {
"^.+\\.tsx?$": ["ts-jest", { tsconfig: "tsconfig.json" }],
},
};

View file

@ -1,119 +0,0 @@
import { Annotations, CfnResource, IAspect, Stack, Token } from "aws-cdk-lib";
import { IConstruct } from "constructs";
/**
* One "/"-delimited segment must be lower kebab-case: `a-b-c`, digits allowed.
*/
const KEBAB_SEGMENT = /^[a-z0-9]+(-[a-z0-9]+)*$/;
/**
* CloudFormation property keys that carry an *explicit physical name*. The
* codegen'd L1 stores these either camelCased (`roleName`) or CFN-cased
* (`RoleName`) depending on the construct, so the aspect matches keys
* case-insensitively.
*
* We deliberately validate physical NAMES + the stack name only — not CDK
* logical construct ids (those are conventionally PascalCase, e.g.
* `InfraDeployRole`, and validating them would be wrong).
*/
const NAME_PROPERTY_KEYS = [
"RoleName",
"BucketName",
"FunctionName",
"TableName",
"LogGroupName",
"QueueName",
"TopicName",
"SecretName",
"StreamName",
"RepositoryName",
"DBInstanceIdentifier",
"DBClusterIdentifier",
"StateMachineName",
"RuleName",
"UserPoolName",
];
// NOTE: `PolicyName` is intentionally NOT checked — CDK auto-generates inline
// `DefaultPolicy` names (e.g. "InstanceRoleDefaultPolicyF15F...") from the
// logical id; those are not explicit, user-controlled physical names and are
// outside the naming convention's scope.
const NAME_KEYS_LC = new Set(NAME_PROPERTY_KEYS.map((k) => k.toLowerCase()));
/**
* Resource types whose physical NAME is a REQUIRED deviation from kebab-case:
* the open-swe config store names secrets `open-swe-<env>/<ENV_VAR_NAME>` and SSM
* params `/open-swe-<env>/<ENV_VAR_NAME>`, where the last segment is the LITERAL
* UPPER_SNAKE environment-variable name. The boot hook
* (deploy/seahaven/fetch-config.sh) strips the prefix and exports that segment
* verbatim, so a lossless store→env round-trip needs the exact env-var name —
* it cannot be kebab-cased. These two resource types are therefore exempt; every
* OTHER explicitly-named resource is still validated. (The `open-swe-<env>`
* prefix is code-generated from `prefix(env)` and is always kebab-case.)
*/
const KEBAB_EXEMPT_RESOURCE_TYPES = new Set([
"AWS::SecretsManager::Secret",
"AWS::SSM::Parameter",
]);
/**
* `true` when every non-empty "/"-delimited segment is kebab-case.
*
* Path-style names are tolerated so the same check works for Secrets Manager
* (`open-swe-dev/foo`), SSM params (`/open-swe-dev/foo`) and log groups
* (`/open-swe/dev/agent`): each segment is validated independently, and a
* leading slash (empty first segment) is ignored.
*/
export function isKebabCase(value: string): boolean {
return value
.split("/")
.filter((seg) => seg.length > 0)
.every((seg) => KEBAB_SEGMENT.test(seg));
}
/**
* Aspect that FAILS synth (`Annotations.addError`) when an explicitly-named
* resource — or a stack name — is not kebab-case. Enforces the org naming
* convention (naming-conventions.md) deterministically at synth time so a
* non-conforming name can never reach a deploy. Wired in bin/app.ts via
* `Aspects.of(app).add(new KebabNamingAspect())`.
*/
export class KebabNamingAspect implements IAspect {
public visit(node: IConstruct): void {
if (node instanceof Stack) {
const name = node.stackName;
if (!Token.isUnresolved(name) && !isKebabCase(name)) {
Annotations.of(node).addError(
`Stack name "${name}" is not kebab-case (open-swe naming convention).`,
);
}
return;
}
if (node instanceof CfnResource) {
// The config store's Secret/Parameter names carry the literal UPPER_SNAKE
// env-var name per the fetch-config naming contract — a required deviation.
if (KEBAB_EXEMPT_RESOURCE_TYPES.has(node.cfnResourceType)) {
return;
}
// `_cfnProperties` is the props as set on the L1; resolve to collapse any
// intrinsic tokens (refs/getatt) so only literal strings are checked.
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const raw = (node as any)._cfnProperties ?? {};
const resolved = Stack.of(node).resolve(raw) ?? {};
for (const [key, value] of Object.entries(resolved)) {
if (
NAME_KEYS_LC.has(key.toLowerCase()) &&
typeof value === "string" &&
!Token.isUnresolved(value) &&
!isKebabCase(value)
) {
Annotations.of(node).addError(
`Resource "${node.node.path}" property ${key}="${value}" is not kebab-case ` +
`(open-swe naming convention).`,
);
}
}
}
}
}

View file

@ -1,59 +0,0 @@
/**
* Shared, non-sensitive constants for the open-swe infra app.
* Account / region are locked per the migration spec (TODO.md "Architecture (locked)").
*/
export const ACCOUNT = "328440206208";
export const REGION = "us-east-1";
export const GITHUB_ORG = "Sea-Haven-Industries";
export const GITHUB_REPO = "open-swe";
export type EnvName = "dev" | "prod";
/** `open-swe-dev` / `open-swe-prod` — kebab-case stack + resource prefix. */
export const prefix = (env: EnvName): string => `open-swe-${env}`;
/**
* Per-ENV GitHub OIDC trust subject for the deploy roles (T5 OSWE-IAC-01/02 fix:
* the dev/prod boundary is enforced in the IAM trust, not by convention).
*
* - `dev` → the `dev` integration branch ref (auto-deploy on push to dev).
* - `prod` → the **GitHub `prod` Environment** subject. A workflow can only mint
* a token with sub `…:environment:prod` by declaring `environment: prod`,
* which triggers the Environment's manual-approval gate (Adam, T18). So the
* prod approval is now expressed at the IAM layer: a dev-branch token can
* never assume a prod deploy role.
*
* Each env gets its OWN infra + app role (githubdeploy-open-swe-{infra,app}-<env>)
* so a dev token cannot reach prod. Exact subject → StringEquals (no `*`).
*
* Cross-env deploy isolation is enforced at the bootstrap layer too — see
* `bootstrapQualifier`: dev runs on its own qualifier so the dev infra role
* cannot assume the bootstrap roles that deploy prod.
*/
export const oidcSubject = (env: EnvName): string =>
env === "prod"
? `repo:${GITHUB_ORG}/${GITHUB_REPO}:environment:prod`
: `repo:${GITHUB_ORG}/${GITHUB_REPO}:ref:refs/heads/dev`;
/**
* Per-env CDK bootstrap qualifier (B-1 / OSWE-IAC-01 fix). Dev runs on its OWN
* qualifier `oswedev` (bootstrapped into the `CDKToolkit-oswedev` stack), so the
* dev infra deploy role only assumes `cdk-oswedev-*` and can NO LONGER assume the
* default `cdk-hnb659fds-*` set whose admin `cfn-exec-role` deploys prod. Prod
* stays on the default qualifier. This closes the cross-env escalation where a
* dev-branch token could `cdk deploy open-swe-prod` via the shared bootstrap
* roles, bypassing prod's Environment approval gate.
*/
export const DEFAULT_BOOTSTRAP_QUALIFIER = "hnb659fds";
export const bootstrapQualifier = (env: EnvName): string =>
env === "dev" ? "oswedev" : DEFAULT_BOOTSTRAP_QUALIFIER;
/**
* The GitHub Actions OIDC provider already exists account-wide (created for
* seahaven-site; see .github/oidc-deploy-roles.yaml `CreateOIDCProvider=false`).
* Reference it by ARN — never create a duplicate `AWS::IAM::OIDCProvider`
* (CloudFormation rejects a second provider for the same URL).
*/
export const GITHUB_OIDC_PROVIDER_ARN = `arn:aws:iam::${ACCOUNT}:oidc-provider/token.actions.githubusercontent.com`;

View file

@ -1,35 +0,0 @@
import * as ec2 from "aws-cdk-lib/aws-ec2";
import { REGION } from "../config";
/**
* The baked open-swe base AMI (ARM64 Ubuntu 24.04 + uv/py3.12 + nginx + CW agent
* + boot templates — NO secrets), produced by `deploy/ami/open-swe-base.pkr.hcl`.
* Pinned by EXACT id (not a name filter) so synth/deploy is fully offline and
* deterministic.
*
* Built 2026-06-26 from open-swe-base-arm64-20260626-203433.
*
* ── EBS / AMI replacement discipline (memory feedback_inline_ebs_volumes) ──
*
* Refresh DELIBERATELY: `cd deploy/ami && packer build open-swe-base.pkr.hcl`,
* then update this id. A new id → EC2 instance REPLACEMENT. Pinning by exact id
* (vs a `most_recent` name filter) is what prevents a routine deploy from silently
* swapping the AMI — the root cause of the file-share data-loss incidents
* (5/15, 5/27, 6/5).
*
* `userDataCausesReplacement: true` (AppService) is likewise DELIBERATE: user-data
* is provisioning-only and the box holds NO durable state (the langgraph store is
* in-memory, rebuilt every boot from S3 + Secrets Manager / SSM), so there is
* intentionally no standalone `ec2.Volume` + `removalPolicy.RETAIN`. The design
* goal is replacement-TOLERANCE, not avoidance.
*
* Operational guard before ANY replacing deploy (AMI / userData / instance-type):
* snapshot the root volume AND wait `state=completed`, re-verify "no local-only
* durable state", and review the `cdk diff` replacement at PR time.
*/
export const BAKED_OPEN_SWE_AMI_ID = "ami-00080084502093021";
/** The baked open-swe base image, pinned by id (offline, deterministic). */
export function bakedOpenSweArm64(): ec2.IMachineImage {
return ec2.MachineImage.genericLinux({ [REGION]: BAKED_OPEN_SWE_AMI_ID });
}

View file

@ -1,368 +0,0 @@
import * as fs from "fs";
import * as path from "path";
import * as cdk from "aws-cdk-lib";
import * as ec2 from "aws-cdk-lib/aws-ec2";
import * as elbv2 from "aws-cdk-lib/aws-elasticloadbalancingv2";
import * as elbTargets from "aws-cdk-lib/aws-elasticloadbalancingv2-targets";
import * as logs from "aws-cdk-lib/aws-logs";
import * as route53 from "aws-cdk-lib/aws-route53";
import * as iam from "aws-cdk-lib/aws-iam";
import * as ssm from "aws-cdk-lib/aws-ssm";
import { Construct, IConstruct } from "constructs";
import { EnvName, prefix } from "../config";
import { bakedOpenSweArm64 } from "./ami-cache";
/**
* Shared seahaven-vpc + internet-facing ALB facts (read-only recon 2026-06-26;
* scratchpad/T12-infra-facts.md). A SINGLE VPC and a SINGLE shared ALB front
* both the on-prem `seahaven-site` stack and open-swe. We IMPORT every one of
* these and NEVER own them - open-swe only ADDS its own instance SG, a standalone
* ALB-egress rule, listener rules, a target group, and DNS records.
*
* ── Cross-stack coordination on the SHARED listener + ALB SG (T13 review) ──
* Two CDK stacks (seahaven-site, open-swe) add resources to the same imported
* `:443` listener and ALB SG. This is safe because each stack owns ONLY the
* resources it declares (its own logical ids): an on-prem `cdk deploy` computes a
* changeset over its own template and cannot delete rules/egress it never
* declared. The standalone-egress pattern is what on-prem itself uses
* (sgr-0c57812752a3bca13), so it does not mutate the shared SG's own definition.
*
* The ONE shared namespace that REQUIRES coordination is listener-rule PRIORITY
* (globally unique per listener; a collision is a fail-SAFE deploy error, not
* silent drift). Ownership map - keep these disjoint when editing either stack:
* - seahaven-site (on-prem): priorities 4-7 + default.
* - open-swe: priorities 2, 3 (webhooks) and 10, 11 (site).
* open-swe's webhook rules are HOST-scoped to its own *.seahaven.com hosts, so
* they never match (let alone "steal") any seahavenind.com / on-prem host.
*/
const SHARED = {
vpcId: "vpc-0d3d4b67bd0cf8a68",
availabilityZones: ["us-east-1a", "us-east-1b"],
// Private subnets host the EC2 box. The single NAT gateway lives in 1a, so the
// box is pinned to private1 (1a) for in-AZ NAT egress (no cross-AZ data $).
privateSubnetIds: ["subnet-04e38c507e96f1926", "subnet-0a0b4fc6f296dfba5"],
instanceSubnetId: "subnet-04e38c507e96f1926",
instanceAz: "us-east-1a",
albDnsName: "seahaven-com-1856441924.us-east-1.elb.amazonaws.com",
albCanonicalHostedZoneId: "Z35SXDOTRQ7X7K",
albSecurityGroupId: "sg-0b0301deed193258a",
httpsListenerArn:
"arn:aws:elasticloadbalancing:us-east-1:328440206208:listener/app/seahaven-com/222c3257354ab559/bab8bcf0da0e2927",
publicZoneId: "Z06652411XKH89KTZD3XA",
publicZoneName: "seahaven.com",
} as const;
/**
* Per-env public hostnames, listener-rule priorities, and instance size.
*
* ── Listener-rule ordering hazard (load-bearing) ──
* The shared listener already has a HOST-AGNOSTIC `/webhooks/*` PATH rule at
* priority 5 (the on-prem seahaven-site stack owns it). ALB rules are first-match
* by ASCENDING priority, so a `…/webhooks/*` request to our host would match
* rule 5 (priority 5) and be forwarded to the on-prem target BEFORE any host rule
* at 10+. Therefore our webhook rule MUST sit below priority 5. The catch-all
* "site" rule (dashboard SPA + /dashboard/api/) carries no path that collides
* with rule 5, so it can sit at any free higher number (10/11). Free priorities
* confirmed by recon: 1-3 and 8+ (4=forgejo, 5=/webhooks/*, 6/7=seahavenind).
*/
const ENV_NET: Record<
EnvName,
{
dashboardHost: string;
hooksHost: string;
webhookPriority: number;
sitePriority: number;
instanceType: string;
}
> = {
dev: {
dashboardHost: "openswe-dev.seahaven.com",
hooksHost: "hooks-dev.seahaven.com",
webhookPriority: 2,
sitePriority: 10,
instanceType: "t4g.medium",
},
prod: {
dashboardHost: "openswe.seahaven.com",
hooksHost: "hooks.seahaven.com",
webhookPriority: 3,
sitePriority: 11,
instanceType: "t4g.large",
},
};
export interface AppServiceProps {
readonly envName: EnvName;
/** Least-privilege EC2 instance role (per-env; from InstanceRole). */
readonly instanceRole: iam.IRole;
/** S3 artifact key prefix the box pulls app.tar.gz / spa.tar.gz from. */
readonly artifactPrefix?: string;
}
/**
* The open-swe compute + ingress wiring for one env (T12):
* - one ARM64 EC2 box in private1 (1a), replacement-tolerant (no RETAIN volume),
* - a standalone instance SG reachable ONLY from the shared ALB SG on :80,
* - a target group -> instance:80 (nginx is the sole ingress; :2024 stays loopback),
* - two listener rules on the imported :443 listener (webhooks below the on-prem
* path rule; site catch-all above it), both -> the same TG,
* - Route53 alias records for both hostnames -> the shared ALB,
* - IaC-owned CloudWatch log groups at 30-day retention.
*
* Everything ALB/VPC/zone-side is IMPORTED. Synth is offline: the AMI is the
* cdk.context.json-pinned AL2023 ARM64 placeholder until the baked
* open-swe-base-arm64 id is pinned before the first real deploy.
*/
export class AppService extends Construct {
public readonly instance: ec2.Instance;
public readonly targetGroup: elbv2.ApplicationTargetGroup;
/** Name of the SSM document CI fires to roll the box to the latest release. */
public readonly deployDocumentName: string;
constructor(scope: Construct, id: string, props: AppServiceProps) {
super(scope, id);
const env = props.envName;
const p = prefix(env);
const net = ENV_NET[env];
const artifactPrefix = props.artifactPrefix ?? "releases/latest";
// Import the shared VPC with explicit attributes (no fromLookup -> offline synth).
const vpc = ec2.Vpc.fromVpcAttributes(this, "Vpc", {
vpcId: SHARED.vpcId,
availabilityZones: [...SHARED.availabilityZones],
privateSubnetIds: [...SHARED.privateSubnetIds],
});
// Standalone instance SG. Egress open (NAT path); ingress only from the ALB SG.
const instanceSg = new ec2.SecurityGroup(this, "InstanceSg", {
vpc,
securityGroupName: `${p}-instance-sg`,
description: `${p} instance SG - ingress only from the shared ALB SG on :80; egress via NAT.`,
allowAllOutbound: true,
});
instanceSg.addIngressRule(
ec2.Peer.securityGroupId(SHARED.albSecurityGroupId),
ec2.Port.tcp(80),
`${p}: shared ALB SG to nginx :80`,
);
// Open the IMPORTED ALB SG to our instance via a STANDALONE egress rule, so we
// never mutate the ALB SG's own (on-prem-owned) definition.
new ec2.CfnSecurityGroupEgress(this, "AlbToInstanceEgress", {
groupId: SHARED.albSecurityGroupId,
ipProtocol: "tcp",
fromPort: 80,
toPort: 80,
destinationSecurityGroupId: instanceSg.securityGroupId,
description: `${p}: ALB to instance nginx :80`,
});
// The app-deploy procedure (deploy/ami/deploy.sh) is a normal reviewable repo
// file; CDK base64-encodes it (single line - no `$`/regex-special chars in the
// base64 alphabet) and renders it into user-data's @@DEPLOY_SH_B64@@ token, so
// user-data writes it verbatim to /opt/open-swe/bin/deploy.sh at first boot.
// The same file is run by the open-swe-<env>-deploy SSM document on every
// release - a single source of truth for "pull release, build venv, restart".
const deployShPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "deploy.sh");
// Minify before embedding: strip full-line comments + blank lines (keep the
// shebang) so the base64 fits EC2's 25.6 KB user-data limit. The repo file
// keeps its comments; only the on-box copy is minified. deploy.sh becomes
// opaque base64 here, so this never affects user-data's heredoc parsing.
const deployShMin = fs
.readFileSync(deployShPath, "utf8")
.split("\n")
.filter((line, i) => i === 0 || (!/^\s*#/.test(line) && line.trim() !== ""))
.join("\n");
const deployShB64 = Buffer.from(deployShMin, "utf8").toString("base64");
// Render the provisioning script's @@tokens@@ into the instance user-data.
// userDataCausesReplacement makes a bootstrap change roll a fresh box (the box
// holds no durable state - see ami-cache.ts / user-data.sh). Editing deploy.sh
// therefore also rolls the box (its base64 is embedded here) - acceptable: the
// box is replacement-tolerant, and ongoing releases never touch user-data.
const userDataPath = path.join(__dirname, "..", "..", "..", "deploy", "ami", "user-data.sh");
const userData = ec2.UserData.custom(
fs
.readFileSync(userDataPath, "utf8")
// %%...%% tokens are CDK-substituted here; they are DELIBERATELY a
// different delimiter from the @@...@@ tokens user-data.sh seds into the
// baked systemd/nginx templates, so CDK can never clobber a sed pattern
// (a shared @@OPENSWE_ENV@@/@@SERVER_NAME@@ left the unit unsubstituted).
.replace(/%%OPENSWE_ENV%%/g, env)
.replace(/%%ASSETS_BUCKET%%/g, `${p}-assets`)
.replace(/%%SERVER_NAME%%/g, net.dashboardHost)
.replace(/%%ARTIFACT_PREFIX%%/g, artifactPrefix)
.replace(/%%DEPLOY_SH_B64%%/g, deployShB64),
);
this.instance = new ec2.Instance(this, "Instance", {
vpc,
vpcSubnets: {
subnets: [
ec2.Subnet.fromSubnetAttributes(this, "InstanceSubnet", {
subnetId: SHARED.instanceSubnetId,
availabilityZone: SHARED.instanceAz,
}),
],
},
instanceType: new ec2.InstanceType(net.instanceType),
// The baked open-swe base AMI (deploy/ami packer build) - ARM64 Ubuntu 24.04
// with the /opt/open-swe layout, openswe user, nginx, and CW agent that
// user-data.sh assumes. Pinned by exact id (see ami-cache.ts); refresh by
// rebuilding and updating BAKED_OPEN_SWE_AMI_ID.
machineImage: bakedOpenSweArm64(),
role: props.instanceRole,
securityGroup: instanceSg,
userData,
userDataCausesReplacement: true,
requireImdsv2: true,
instanceName: `${p}-box`,
blockDevices: [
{
deviceName: "/dev/xvda",
// gp3 encrypted root; deleteOnTermination (no durable on-box state ->
// intentionally NO standalone RETAIN volume; see ami-cache.ts).
volume: ec2.BlockDeviceVolume.ebs(30, {
volumeType: ec2.EbsDeviceVolumeType.GP3,
encrypted: true,
deleteOnTermination: true,
}),
},
],
});
// `requireImdsv2: true` makes CDK auto-create a launch template, and it names
// that LT from the construct id ("Instance" -> "InstanceLaunchTemplate") with NO
// env qualifier — so OpenSweDevStack and OpenSweProdStack both want the identical
// LT name and the second env to deploy fails with
// InvalidLaunchTemplateName.AlreadyExistsException (prod rollback, 2026-06-29).
// Force a per-env LT name. Done via an aspect because the LT is created at synth
// time by the requireImdsv2 handling, not in this constructor.
cdk.Aspects.of(this.instance).add({
visit(node: IConstruct) {
if (node instanceof ec2.CfnLaunchTemplate) {
node.launchTemplateName = `${p}-lt`;
}
// The instance references the LT BY NAME, so the reference must be renamed in
// lockstep (preserve the GetAtt version) or CFN can't find the template.
if (node instanceof ec2.CfnInstance && node.launchTemplate) {
const spec = node.launchTemplate as ec2.CfnInstance.LaunchTemplateSpecificationProperty;
node.launchTemplate = { ...spec, launchTemplateName: `${p}-lt` };
}
},
});
// SSM deploy document (open-swe-<env>-deploy): runs the baked
// /opt/open-swe/bin/deploy.sh to pull the latest release + restart. CI fires it
// (tag-scoped to project=open-swe,env=<env>) after uploading a release, so the
// app deploy role needs SendCommand ONLY on this document - NOT on the generic
// AWS-RunShellScript (closes the T4 BLOCK#3 arbitrary-shell timebox).
this.deployDocumentName = `${p}-deploy`;
new ssm.CfnDocument(this, "DeployDoc", {
name: this.deployDocumentName,
documentType: "Command",
documentFormat: "YAML",
updateMethod: "NewVersion",
content: {
schemaVersion: "2.2",
description: `Roll the ${p} box to the latest published release (runs /opt/open-swe/bin/deploy.sh).`,
mainSteps: [
{
action: "aws:runShellScript",
name: "deploy",
inputs: {
// Fixed command - no parameters, so nothing untrusted is interpolated
// into the shell. The script itself reads /etc/open-swe/boot.env.
runCommand: ["bash /opt/open-swe/bin/deploy.sh"],
},
},
],
},
});
// Target group -> instance:80 (nginx). Health check hits nginx's /healthz
// (returns 200; the dashboard TG health path defined in open-swe.nginx.conf).
this.targetGroup = new elbv2.ApplicationTargetGroup(this, "Tg", {
vpc,
targetGroupName: `${p}-tg`,
port: 80,
protocol: elbv2.ApplicationProtocol.HTTP,
targetType: elbv2.TargetType.INSTANCE,
targets: [new elbTargets.InstanceTarget(this.instance)],
deregistrationDelay: cdk.Duration.seconds(15),
healthCheck: {
path: "/healthz",
healthyHttpCodes: "200",
interval: cdk.Duration.seconds(30),
timeout: cdk.Duration.seconds(5),
healthyThresholdCount: 2,
unhealthyThresholdCount: 3,
},
});
// Import the shared :443 listener (with its ALB SG) and ADD our two rules.
const albSg = ec2.SecurityGroup.fromSecurityGroupId(this, "AlbSg", SHARED.albSecurityGroupId, {
mutable: false,
});
const listener = elbv2.ApplicationListener.fromApplicationListenerAttributes(this, "HttpsListener", {
listenerArn: SHARED.httpsListenerArn,
securityGroup: albSg,
});
// (1) Webhooks - accepted on EITHER host (integrations may target either), and
// MUST be below the on-prem path-only rule 5 (see ENV_NET note).
new elbv2.ApplicationListenerRule(this, "WebhooksRule", {
listener,
priority: net.webhookPriority,
conditions: [
elbv2.ListenerCondition.hostHeaders([net.dashboardHost, net.hooksHost]),
elbv2.ListenerCondition.pathPatterns(["/webhooks/*"]),
],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
// (2) Dashboard SPA + /dashboard/api/ (OAuth) - DASHBOARD host ONLY. The hooks
// host intentionally serves nothing but /webhooks/* (rule 1), so the OAuth /
// dashboard surface stays single-origin (OSWE-T12-02). Non-webhook paths on the
// hooks host fall through to the on-prem default.
new elbv2.ApplicationListenerRule(this, "SiteRule", {
listener,
priority: net.sitePriority,
conditions: [elbv2.ListenerCondition.hostHeaders([net.dashboardHost])],
action: elbv2.ListenerAction.forward([this.targetGroup]),
});
// Route53 ALIAS records -> the shared ALB, for both hostnames.
const zone = route53.HostedZone.fromHostedZoneAttributes(this, "PublicZone", {
hostedZoneId: SHARED.publicZoneId,
zoneName: SHARED.publicZoneName,
});
const albAlias: route53.IAliasRecordTarget = {
bind: () => ({
dnsName: SHARED.albDnsName,
hostedZoneId: SHARED.albCanonicalHostedZoneId,
}),
};
for (const [label, host] of [
["Dashboard", net.dashboardHost],
["Hooks", net.hooksHost],
] as const) {
new route53.ARecord(this, `${label}Alias`, {
zone,
recordName: host,
target: route53.RecordTarget.fromAlias(albAlias),
comment: `${p} ${label.toLowerCase()} -> shared seahaven-com ALB`,
});
}
// IaC-owned CloudWatch log groups at 30-day retention. Names mirror the
// CloudWatch-agent config (deploy/ami/templates/amazon-cloudwatch-agent.json);
// owning them here makes retention declarative rather than agent-set. Logs are
// not durable state -> DESTROY on stack delete.
for (const suffix of ["app", "user-data", "nginx-access", "nginx-error"]) {
new logs.LogGroup(this, `Log-${suffix}`, {
logGroupName: `/open-swe/${env}/${suffix}`,
retention: logs.RetentionDays.ONE_MONTH,
removalPolicy: cdk.RemovalPolicy.DESTROY,
});
}
}
}

View file

@ -1,58 +0,0 @@
import * as cdk from "aws-cdk-lib";
import * as s3 from "aws-cdk-lib/aws-s3";
import { Construct } from "constructs";
import { EnvName, prefix } from "../config";
/**
* The per-env S3 artifact bucket (`open-swe-<env>-assets`) the box pulls its
* release from (T7). CI builds the SPA + packages the app source and uploads
* `app.tar.gz` / `spa.tar.gz` under `releases/<sha>/` + `releases/latest/`
* (`build-artifacts.yml`, via the `githubdeploy-open-swe-app-<env>` OIDC role);
* the box pulls `releases/latest/*` at boot / on deploy via its instance role.
*
* The bucket holds ONLY build artifacts — no secrets (those live in Secrets
* Manager + SSM), no durable runtime state (the langgraph store is in-memory and
* rebuilt every boot). It is therefore safe to treat as reproducible-from-CI, but
* we RETAIN it on stack delete so an accidental `cdk destroy` cannot strand the
* box with no artifact to pull on its next replacement.
*
* Security posture (locked, reviewed in T7):
* - `BLOCK_ALL` public access (this is an internal artifact store; ALB/nginx is
* the only public surface — never S3 directly).
* - SSE-S3 encryption at rest + `enforceSSL` (deny any non-TLS request).
* - versioned, so a bad release can be rolled back to the previous object
* version (the last-good-artifact story in T19); a lifecycle rule expires
* NONcurrent versions after 30 days so history does not grow unbounded.
* - aborts incomplete multipart uploads after 7 days (cost hygiene).
*
* The name is the load-bearing contract: `instance-role.ts` (read), the app
* deploy role in `github-deploy-roles.ts` (write), and `user-data.sh` /
* `deploy.sh` (`@@ASSETS_BUCKET@@`) all reference `open-swe-<env>-assets` by
* literal name, so it is set explicitly here rather than auto-generated.
*/
export class AssetsBucket extends Construct {
public readonly bucket: s3.Bucket;
constructor(scope: Construct, id: string, envName: EnvName) {
super(scope, id);
const p = prefix(envName);
this.bucket = new s3.Bucket(this, "Bucket", {
bucketName: `${p}-assets`,
blockPublicAccess: s3.BlockPublicAccess.BLOCK_ALL,
encryption: s3.BucketEncryption.S3_MANAGED,
enforceSSL: true,
versioned: true,
// Artifacts are reproducible from CI, but RETAIN protects against an
// accidental stack delete leaving the box with nothing to pull (see above).
removalPolicy: cdk.RemovalPolicy.RETAIN,
lifecycleRules: [
{
id: "expire-noncurrent-artifact-versions",
noncurrentVersionExpiration: cdk.Duration.days(30),
abortIncompleteMultipartUploadAfter: cdk.Duration.days(7),
},
],
});
}
}

View file

@ -1,265 +0,0 @@
import * as cdk from "aws-cdk-lib";
import * as secretsmanager from "aws-cdk-lib/aws-secretsmanager";
import * as ssm from "aws-cdk-lib/aws-ssm";
import { Construct } from "constructs";
import { EnvName } from "../config";
/**
* Config / secret "shells" for the boot hook (`deploy/seahaven/fetch-config.sh`).
*
* The naming contract (source of truth: the T9 env/secret/config inventory + the
* fetch-config header) is LITERAL env-var names as the last path segment:
*
* Secrets open-swe-<env>/<ENV_VAR_NAME> (AWS Secrets Manager)
* Config /open-swe-<env>/<ENV_VAR_NAME> (AWS SSM Parameter Store, FLAT)
*
* fetch-config reads secrets with `batch-get-secret-value --filters
* Key=name,Values=open-swe-<env>/` and config with `get-parameters-by-path
* --path /open-swe-<env>/` (NON-recursive), then strips the prefix so the last
* segment IS the exported variable name. So these resources MUST carry the
* UPPER_SNAKE env-var name verbatim — which is why both resource types are
* exempted from the kebab-naming Aspect (see aspects/kebab-naming-aspect.ts).
*
* Three buckets:
*
* 1. SECRETS_SHELLS — the 29 secrets. Created as value-LESS shells (an L1
* `CfnSecret` with NEITHER `secretString` NOR `generateSecretString`, which
* CloudFormation creates as an empty secret with no version). The real value
* is set out-of-band via `deploy/seahaven/put-config.sh` (put-secret-value)
* BEFORE the box boots. Because CDK never owns the value, a later
* `cdk deploy` can never clobber the operator-set value. AWS-managed key
* (alias/aws/secretsmanager) — no CMK, matching the instance role which
* deliberately omits kms:Decrypt.
*
* 2. IAC_MANAGED_SSM — stable / derivable config. Real values are owned here in
* IaC (one StringParameter each) so they are reproducible and reviewed.
*
* 3. Out-of-band SSM (NOT created here) — operationally-variable or
* env-specific-unknown config (e.g. DEFAULT_SANDBOX_SNAPSHOT_ID, which
* changes on every snapshot rebuild and would be clobbered by a deploy if it
* were CDK-managed; GitHub App ids; Slack ids; LangSmith tenant/urls). These
* are listed in OUT_OF_BAND_SSM purely for documentation and are set by
* `put-config.sh`, never by CDK.
*/
/** The 29 Secrets Manager secret VAR names (T9 inventory SECRETS table). */
export const SECRET_VARS: readonly string[] = [
"ANTHROPIC_API_KEY",
"CORRIDOR_API_TOKEN",
"CORRIDOR_MCP_TOKEN",
"CORRIDOR_TOKEN",
"DASHBOARD_JWT_SECRET",
"DAYTONA_API_KEY",
"EXA_API_KEY",
"FIREWORKS_API_KEY",
"GITHUB_APP_CLIENT_SECRET",
"GITHUB_APP_PRIVATE_KEY",
"GITHUB_PAT",
"GITHUB_WEBHOOK_SECRET",
"JUDGE_ANTHROPIC_API_KEY",
"LANGSMITH_API_KEY",
"LANGSMITH_API_KEY_PROD",
"LANGCHAIN_API_KEY",
"LINEAR_API_KEY",
"LINEAR_WEBHOOK_SECRET",
"RUNLOOP_API_KEY",
"SLACK_BOT_TOKEN",
"SLACK_CLIENT_SECRET",
"SLACK_SIGNING_SECRET",
"TOKEN_ENCRYPTION_KEY",
"USER_ID_API_KEY_MAP",
"X_SERVICE_AUTH_JWT_SECRET",
] as const;
// 25 secret shells (OPENAI_API_KEY / GOOGLE_API_KEY / GROQ_API_KEY removed in the
// Bedrock/Fireworks migration — those providers are dropped from SUPPORTED_MODELS and
// their keys revoked + secret objects deleted). The T9 inventory header said "29" vs
// 27 enumerated; reconciled
// (Adam confirm 2026-06-26): USER_ID_API_KEY_MAP (maps user ids -> API keys; flagged
// sensitive by the T5 security review) is a SECRET and is included here.
// JUDGE_ANTHROPIC_BASE_URL is a URL (non-sensitive config, eval-only) -> SSM/default,
// NOT a secret. So the inventory's "29" was effectively a miscount.
/** Short, value-free descriptions for the secret shells (no secret material). */
const SECRET_DESCRIPTIONS: Record<string, string> = {
ANTHROPIC_API_KEY: "Claude LLM API key (primary builder provider).",
CORRIDOR_API_TOKEN: "Corridor MCP token (optional).",
CORRIDOR_MCP_TOKEN: "Corridor MCP token alt name (optional).",
CORRIDOR_TOKEN: "Corridor MCP token alt name (optional).",
DASHBOARD_JWT_SECRET: "JWT signing secret for dashboard session cookies (REQUIRED).",
DAYTONA_API_KEY: "Daytona sandbox key (only if SANDBOX_TYPE=daytona).",
EXA_API_KEY: "Exa web-search key (optional).",
FIREWORKS_API_KEY: "Fireworks LLM key (only if a fireworks: model is used).",
GITHUB_APP_CLIENT_SECRET: "GitHub App OAuth client secret (dashboard login).",
GITHUB_APP_PRIVATE_KEY: "GitHub App private key PEM (installation-token minting).",
GITHUB_PAT: "GitHub PAT fallback (optional).",
GITHUB_WEBHOOK_SECRET: "GitHub webhook signature secret (prod-required).",
JUDGE_ANTHROPIC_API_KEY: "Eval judge key (optional; falls back to ANTHROPIC_API_KEY).",
LANGSMITH_API_KEY: "LangSmith key (dev).",
LANGSMITH_API_KEY_PROD: "LangSmith key (prod / deployed sandbox).",
LANGCHAIN_API_KEY: "LangSmith key alt name (fallback).",
LINEAR_API_KEY: "Linear API key (optional).",
LINEAR_WEBHOOK_SECRET: "Linear webhook signature secret (required when Linear is wired).",
RUNLOOP_API_KEY: "Runloop sandbox key (only if SANDBOX_TYPE=runloop).",
SLACK_BOT_TOKEN: "Slack bot token (optional).",
SLACK_CLIENT_SECRET: "Slack OAuth client secret.",
SLACK_SIGNING_SECRET: "Slack webhook signing secret (prod-required).",
TOKEN_ENCRYPTION_KEY: "Fernet key(s) for per-user GitHub-token encryption (REQUIRED).",
USER_ID_API_KEY_MAP: "JSON map of user id -> API key for per-user auth (optional, sensitive).",
X_SERVICE_AUTH_JWT_SECRET: "Service-auth JWT secret (optional).",
};
/**
* IaC-managed SSM config: stable / derivable values owned in code, per env.
* Values are functions of envName so dev/prod render correct hosts.
*
* Anything operationally-variable or env-specific-unknown is deliberately NOT
* here — see OUT_OF_BAND_SSM.
*/
export function iacManagedSsm(env: EnvName): Record<string, string> {
// Public host = seahaven.com (the migration's new AWS public face; confirmed by
// recon: seahaven.com Route53 zone + *.seahaven.com ACM cert are live on the ALB
// — distinct from the on-prem seahavenind.com). dev = openswe-dev, prod = openswe.
const host = `https://openswe${env === "dev" ? "-dev" : ""}.seahaven.com`;
// Repo targeting is PER-ENV: dev drives the disposable sandbox repo in the
// dedicated seahaven-open-swe-dev org (isolates dev-agent activity from the real
// Sea Haven org); prod stays on the Sea Haven org pilot repo. fetch-config's owner
// GUARD honors this value (rejecting only blank / the upstream langchain-ai org).
const repo =
env === "dev"
? { owner: "seahaven-open-swe-dev", name: "openswe-dev-sandbox" }
: { owner: "Sea-Haven-Industries", name: "open-swe-pilot" };
const managed: Record<string, string> = {
// Sandbox provider — plan keeps stock langsmith (T9). Stable.
SANDBOX_TYPE: "langsmith",
DEFAULT_REPO_OWNER: repo.owner,
// Repo-owner allowlist (comma list) — the env's org. The app gates triggers on this.
ALLOWED_GITHUB_ORGS: repo.owner,
DEFAULT_REPO_NAME: repo.name,
// Dashboard URLs — derived from the public host.
DASHBOARD_BASE_URL: host,
DASHBOARD_API_BASE_URL: host,
DASHBOARD_ALLOWED_ORIGINS: host,
// Primary builder model. seed_store.sh's `pick` precedence is
// OPENSWE_AGENT_MODEL > SEED_AGENT_MODEL > LLM_MODEL_ID > script default, so this
// SSM value overrides the seed-script default — it MUST be a supported id. Post
// Bedrock/Fireworks migration the only Bedrock-Claude id is the inference profile;
// `anthropic:claude-opus-4-8` was removed from SUPPORTED_MODELS.
LLM_MODEL_ID: "bedrock_converse:us.anthropic.claude-opus-4-8",
};
// Dev e2e smoke: seed the owner's user_mapping so an @openswe comment from the
// triggering GitHub login resolves (an unmapped commenter is silently skipped).
// Dev-only — prod seeds its mappings via its own operator config. login:email.
if (env === "dev") {
managed.SEED_USER_MAPPINGS = "amoussa1229:adam@seahavenind.com";
}
return managed;
}
/**
* Out-of-band SSM config: NOT created by CDK. Listed for documentation and for
* `put-config.sh` to populate before the box boots. Each MUST stay out of IaC
* because its value is operationally-variable or env-specific and unknown at
* synth time — making it CDK-managed would either clobber the operator value on
* the next deploy (e.g. DEFAULT_SANDBOX_SNAPSHOT_ID) or hardcode a secret-ish id.
*/
export const OUT_OF_BAND_SSM: readonly string[] = [
// Sandbox snapshot id — changes on EVERY snapshot rebuild. MUST NOT be
// CDK-managed or a deploy clobbers it. Required for SANDBOX_TYPE=langsmith.
"DEFAULT_SANDBOX_SNAPSHOT_ID",
// GitHub App identifiers — set when the per-env GitHub App is created.
"GITHUB_APP_ID",
"GITHUB_APP_CLIENT_ID",
"GITHUB_APP_INSTALLATION_ID",
"GITHUB_OAUTH_PROVIDER_ID",
// LangSmith deployment coordinates (prod tenant/urls/endpoints).
"LANGSMITH_TENANT_ID_PROD",
"LANGSMITH_URL_PROD",
"LANGSMITH_ENDPOINT",
"LANGSMITH_ENDPOINT_PROD",
"LANGSMITH_HOST_API_URL",
"LANGGRAPH_URL",
"LANGGRAPH_URL_PROD",
"LANGCHAIN_REVISION_ID",
// Slack workspace ids — set after the Slack app is installed.
"SLACK_CLIENT_ID",
"SLACK_TEAM_ID",
"SLACK_BOT_USER_ID",
"SLACK_BOT_USERNAME",
"SLACK_REPO_OWNER",
"SLACK_REPO_NAME",
// Access / observability allowlists — operator-curated.
"CONFIGURED_ADMINS",
"OBSERVABILITY_AUTHORIZED_EMAILS",
"PUBLIC_REPO_ORG_GATE",
"ALLOWED_GITHUB_REPOS",
// Optional integrations + tuning knobs (left to code defaults unless set).
"LLM_FALLBACK_MODEL_ID",
"DATADOG_MCP_TOOLSETS",
"NOTION_MCP_CLIENT_NAME",
"API_STANDARDS_SKILL_HANDLE",
"REPO_SNAPSHOT_BASE_IMAGE",
"REPO_SNAPSHOT_BUILD_TIMEOUT_SECONDS",
"REPO_SNAPSHOT_STALE_BUILD_SECONDS",
] as const;
export interface ConfigStoreProps {
readonly envName: EnvName;
}
/**
* Per-env Secrets Manager + SSM Parameter Store shells the boot hook reads.
* Instantiated from OpenSweStack. Synth-able now (T11); values populated
* out-of-band BEFORE the EC2/T12 deploy. See infra/README.md "Config store".
*/
export class ConfigStore extends Construct {
public readonly secrets: secretsmanager.CfnSecret[] = [];
public readonly params: ssm.StringParameter[] = [];
constructor(scope: Construct, id: string, props: ConfigStoreProps) {
super(scope, id);
const env = props.envName;
// --- 1) Secret shells (value-LESS; populated out-of-band) ----------------
for (const varName of SECRET_VARS) {
const secret = new secretsmanager.CfnSecret(this, `Secret-${varName}`, {
name: `open-swe-${env}/${varName}`,
description: SECRET_DESCRIPTIONS[varName] ?? `open-swe ${varName}`,
// Deliberately NO secretString / generateSecretString: CloudFormation
// creates an empty secret, so the out-of-band value is never clobbered.
});
// RETAIN: a stack teardown must not destroy operator-set secret material.
//
// GOTCHA — RETAIN + fixed name orphans these shells on a FAILED FIRST
// CREATE. If the stack's initial create fails and rolls back, RETAIN keeps
// the shells instead of deleting them; the stack is then gone but the
// secrets survive, still holding the global `open-swe-<env>/<VAR>` names.
// Every later create then fails with `AlreadyExists` (and a normal
// delete-secret keeps the name reserved for the 7–30 day recovery window,
// so it does NOT clear the deadlock). Recovery: before re-creating the
// stack, force-delete the orphans so the names free immediately, e.g.
// aws secretsmanager list-secrets --filters Key=name,Values=open-swe-<env>/ \
// --query 'SecretList[].Name' --output text | tr '\t' '\n' | while read n; do
// aws secretsmanager delete-secret --secret-id "$n" \
// --force-delete-without-recovery; done
// Only force-delete shells that are EMPTY (no value version) — a populated
// secret holds real operator material. This bites on teardown/rebuild, a
// secret logical-id change/refactor, or standing up a new env — NOT on
// routine updates of an already-created stack. (Hit on prod 2026-06-29.)
secret.applyRemovalPolicy(cdk.RemovalPolicy.RETAIN);
this.secrets.push(secret);
}
// --- 2) IaC-managed SSM config (real, derivable values) ------------------
const managed = iacManagedSsm(env);
for (const [varName, value] of Object.entries(managed)) {
this.params.push(
new ssm.StringParameter(this, `Param-${varName}`, {
parameterName: `/open-swe-${env}/${varName}`,
stringValue: value,
description: `IaC-managed open-swe ${varName} (${env}).`,
tier: ssm.ParameterTier.STANDARD,
}),
);
}
}
}

View file

@ -1,173 +0,0 @@
import * as iam from "aws-cdk-lib/aws-iam";
import { Construct } from "constructs";
import {
ACCOUNT,
EnvName,
GITHUB_OIDC_PROVIDER_ARN,
REGION,
bootstrapQualifier,
oidcSubject,
} from "../config";
/**
* Per-ENV GitHub Actions OIDC deploy roles. Created ONCE per env in the
* dedicated `open-swe-iam` stack. Two roles per env, per the locked architecture's
* "dual OIDC roles":
*
* - githubdeploy-open-swe-infra-<env> → CFN/IAM (CDK) deploys of that env's stack
* - githubdeploy-open-swe-app-<env> → app deploys (env-tag-scoped SSM + S3 read)
*
* T5 OSWE-IAC-01/02 fix: roles are split per env and the trust subject is
* env-scoped (dev = dev branch ref; prod = the GitHub `prod` Environment subject,
* so the manual-approval gate is IAM-enforced). A dev-branch token therefore
* cannot SendCommand to the prod box nor assume a prod deploy role.
*
* Reviewed at T4 (GPT-4.1 IAM cross-review) + T5 (/sh-security-review) and
* deployed FIRST (BLOCK#3 "OIDC-role-first" ordering) before any other infra or
* secrets CI step.
*/
export class GithubDeployRoles extends Construct {
public readonly infraRole: iam.Role;
public readonly appRole: iam.Role;
constructor(scope: Construct, id: string, envName: EnvName) {
super(scope, id);
// The provider already exists account-wide — reference, never re-create.
const provider = iam.OpenIdConnectProvider.fromOpenIdConnectProviderArn(
this,
"GithubOidcProvider",
GITHUB_OIDC_PROVIDER_ARN,
);
// T4 BLOCK#2 + T5 IAC-01/02: exact env-scoped subject via StringEquals (no
// StringLike, no `*`). prod = environment:prod (manual-approval gate),
// dev = the dev branch ref.
const trust = new iam.WebIdentityPrincipal(provider.openIdConnectProviderArn, {
StringEquals: {
"token.actions.githubusercontent.com:aud": "sts.amazonaws.com",
"token.actions.githubusercontent.com:sub": oidcSubject(envName),
},
});
// ---- githubdeploy-open-swe-infra-<env> --------------------------------
this.infraRole = new iam.Role(this, "InfraDeployRole", {
roleName: `githubdeploy-open-swe-infra-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for CDK deploys of the open-swe-${envName} infra stack (assumes CDK bootstrap roles).`,
});
// Org-standard CDK deploy pattern (mirrors githubdeploy-seahaven-account-
// baseline / -forgejo / -apm-wo-analysis): the deploy role only needs to
// assume the CDK bootstrap roles. The actual CloudFormation + IAM + resource
// permissions are exercised by the bootstrap `cfn-exec-role`, whose scope is
// owned by the CDKToolkit stack — NOT granted directly here.
//
// B-1 / OSWE-IAC-01 fix: scope the assume to THIS env's bootstrap qualifier.
// Dev uses `oswedev` (its own CDKToolkit-oswedev bootstrap), prod uses the
// default `hnb659fds`. The dev infra role can therefore no longer assume the
// bootstrap roles whose admin cfn-exec-role deploys prod — closing the prior
// cross-env escalation (a dev-branch token could `cdk deploy open-swe-prod`
// via the shared account-wide bootstrap roles, bypassing prod's Environment
// approval gate). The qualifier wildcard still matches only the handful of
// roles `cdk bootstrap` creates for that qualifier.
this.infraRole.addToPolicy(
new iam.PolicyStatement({
sid: "AssumeCdkBootstrapRoles",
actions: ["sts:AssumeRole"],
resources: [`arn:aws:iam::${ACCOUNT}:role/cdk-${bootstrapQualifier(envName)}-*`],
}),
);
// ---- githubdeploy-open-swe-app-<env> ----------------------------------
this.appRole = new iam.Role(this, "AppDeployRole", {
roleName: `githubdeploy-open-swe-app-${envName}`,
assumedBy: trust,
description: `GitHub OIDC role for open-swe-${envName} app deploys: env-tag-scoped ssm:SendCommand + read of the ${envName} S3 artifact bucket.`,
});
// T5 OSWE-IAC-01 fix: SendCommand only to instances tagged project=open-swe
// AND env=<this env> (a SINGLE value, not {dev,prod}). The dev app role can
// never command the prod box and vice versa — env isolation in IAM.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandTagScoped",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ec2:${REGION}:${ACCOUNT}:instance/*`],
conditions: {
StringEquals: {
"ssm:resourceTag/project": "open-swe",
"ssm:resourceTag/env": envName,
},
},
}),
);
// SendCommand also has to reference the command document. Scope to this env's
// open-swe deploy document ONLY.
// T4 BLOCK#3 (CLOSED at T19): GPT-4.1 flagged AWS-RunShellScript as an
// arbitrary-shell escalation path. The dedicated `open-swe-${envName}-deploy`
// SSM document (app-service.ts) now runs the fixed, parameter-less command
// `bash /opt/open-swe/bin/deploy.sh`, so AWS-RunShellScript is dropped here:
// this role can run ONLY that one document, and only on its own env's box
// (tag-scoped by the SsmSendCommandTagScoped statement above).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmSendCommandDocuments",
actions: ["ssm:SendCommand"],
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:document/open-swe-${envName}-deploy`],
}),
);
// Poll command results. These read actions do not support resource-level
// scoping, so `*` is required by the API (T4 FIX: API limitation, documented).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "SsmReadCommandStatus",
actions: [
"ssm:GetCommandInvocation",
"ssm:ListCommands",
"ssm:ListCommandInvocations",
],
resources: ["*"],
}),
);
// Read+WRITE access to THIS env's artifact bucket only (T19): the
// build-artifacts workflow uploads app.tar.gz / spa.tar.gz under releases/*,
// then fires the deploy document so the box pulls them via its instance role.
// Object actions are scoped to releases/* (the only prefix CI writes), and to
// THIS env's bucket — a dev token can never write the prod bucket. No
// bucket-level mutation (no PutBucket*/Delete bucket) — that stays with CDK.
// GetObject + PutObject (S3-to-S3 copy = Get source + Put dest) is all the
// publish/rollback path uses; s3:DeleteObject is deliberately NOT granted so a
// CI token cannot erase an immutable release or the releases/last-good rollback
// fallback (lifecycle expiry handles old-version cleanup, not CI).
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ReadWriteArtifactObjects",
actions: ["s3:GetObject", "s3:PutObject"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets/releases/*`],
}),
);
// ListBucket is constrained to the releases/ prefix (F-1/IAC-04): the
// publish/rollback scripts only ever list under releases/, so a leaked CI
// token cannot enumerate anything else in the bucket. GetBucketLocation
// carries no s3:prefix, so it stays a separate, unconditioned statement.
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "ListArtifactBucket",
actions: ["s3:ListBucket"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
conditions: { StringLike: { "s3:prefix": ["releases/*"] } },
}),
);
this.appRole.addToPolicy(
new iam.PolicyStatement({
sid: "GetArtifactBucketLocation",
actions: ["s3:GetBucketLocation"],
resources: [`arn:aws:s3:::open-swe-${envName}-assets`],
}),
);
}
}

View file

@ -1,160 +0,0 @@
import * as iam from "aws-cdk-lib/aws-iam";
import { Construct } from "constructs";
import { ACCOUNT, EnvName, REGION, prefix } from "../config";
/**
* Least-privilege EC2 instance role for the open-swe box (one per env).
*
* Grants exactly what the boot/runtime flow needs and NOTHING ELSE — no admin,
* no `*` resources except where the AWS action genuinely has no resource-level
* scoping. Per-env so the dev box can never read prod secrets/config and vice
* versa. Reviewed at T4 (GPT-4.1 IAM cross-review) / T5 (/sh-security-review)
* before it is ever deployed (T6).
*/
export class InstanceRole extends Construct {
public readonly role: iam.Role;
constructor(scope: Construct, id: string, env: EnvName) {
super(scope, id);
const p = prefix(env);
this.role = new iam.Role(this, "Role", {
roleName: `${p}-instance-role`,
assumedBy: new iam.ServicePrincipal("ec2.amazonaws.com"),
description: `EC2 instance role for the ${p} open-swe box (least-privilege).`,
});
// AWS-managed: lets the SSM agent register the instance and RECEIVE the
// app-deploy `ssm:SendCommand` from githubdeploy-open-swe-app. This is the
// standard Session-Manager / RunCommand grant and is the only managed
// policy on the role. DELIBERATE — flag for T4 confirmation.
this.role.addManagedPolicy(
iam.ManagedPolicy.fromAwsManagedPolicyName("AmazonSSMManagedInstanceCore"),
);
// Read the build artifact from the env's S3 asset bucket (deploy = pull).
// Scoped to releases/* — the only prefix CI writes and the box pulls — so a
// compromised box (or stolen IMDS creds) cannot read anything else that might
// ever land in the bucket (least-privilege; mirrors the app role's write scope).
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadArtifactObjects",
actions: ["s3:GetObject"],
resources: [`arn:aws:s3:::${p}-assets/releases/*`],
}),
);
// ListBucket is constrained to the releases/ prefix (F-1/IAC-04) — the box
// only ever lists release artifacts, so a compromised box cannot enumerate
// any other object that might land in the bucket. GetBucketLocation has no
// s3:prefix in its request context, so it stays a separate, unconditioned
// statement (the condition would otherwise AccessDeny it).
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ListArtifactBucket",
actions: ["s3:ListBucket"],
resources: [`arn:aws:s3:::${p}-assets`],
conditions: { StringLike: { "s3:prefix": ["releases/*"] } },
}),
);
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "GetArtifactBucketLocation",
actions: ["s3:GetBucketLocation"],
resources: [`arn:aws:s3:::${p}-assets`],
}),
);
// Read non-sensitive config from SSM Parameter Store under /open-swe-<env>/*.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadSsmConfig",
actions: ["ssm:GetParameter", "ssm:GetParameters", "ssm:GetParametersByPath"],
resources: [`arn:aws:ssm:${REGION}:${ACCOUNT}:parameter/${p}/*`],
}),
);
// VALUE access — Secrets Manager under open-swe-<env>/*. Secret ARNs carry a
// random 6-char suffix, hence the trailing `*`. This is the statement that
// actually gates which secret VALUES the box can read: prefix-scoped, so the
// dev box can never read prod secret values (and vice versa). GetSecretValue is
// checked per-secret even when the value is returned via the batch call below.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "ReadSecretValues",
actions: ["secretsmanager:GetSecretValue", "secretsmanager:DescribeSecret"],
resources: [`arn:aws:secretsmanager:${REGION}:${ACCOUNT}:secret:${p}/*`],
}),
);
// BatchGetSecretValue MUST be granted on `*` — it is a collection action that
// AWS authorizes against the account, NOT the per-secret ARN, REGARDLESS of
// whether the caller uses `--filters` or `--secret-id-list`. A prefix-scoped
// BatchGetSecretValue AccessDenies the whole call ("no identity-based policy
// allows the secretsmanager:BatchGetSecretValue action") — VERIFIED on the live
// dev box 2026-06-29 (the OSWE-IAC-SECRETS-LIST-01 attempt to prefix-scope it
// crash-looped the box once the prior broad grant's eventual-consistency lapsed).
// This `*` does NOT widen VALUE access: a secret value is only returned when the
// prefix-scoped GetSecretValue above also allows it, so cross-env value isolation
// holds. The win that DID survive: fetch-config uses `--secret-id-list` (explicit
// names, no name filter), so `secretsmanager:ListSecrets` is NOT needed and is
// intentionally omitted — the box cannot enumerate secret names account-wide.
// F-2 (accepted residual): because the grant is `*`, a caller naming a secret
// in ANOTHER env's prefix learns whether that name EXISTS (an existence oracle
// via the per-secret AccessDenied-vs-not signal) even though the VALUE stays
// gated by the prefix-scoped GetSecretValue above. Accepted within Sea Haven's
// single-tenant account 328440206208 — cross-env VALUE isolation is preserved.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "BatchGetSecretValues",
actions: ["secretsmanager:BatchGetSecretValue"],
resources: ["*"],
}),
);
// NOTE (T11): SSM SecureString + Secrets Manager here are assumed to use the
// AWS-managed keys (alias/aws/ssm, alias/aws/secretsmanager) for which the
// service grants Decrypt implicitly — so NO kms:Decrypt is granted. If T11
// moves these to a customer CMK, add a scoped `kms:Decrypt` on that key ARN
// ONLY (not `*`).
// Invoke the Bedrock Claude model. DEFAULT_MODEL_ID is
// `bedrock_converse:us.anthropic.claude-opus-4-8`, and the model runs in the
// LangGraph server PROCESS on this box (not in the sandbox), so the EC2
// instance role is the calling principal. The `us.` cross-region inference
// profile fans out to us-east-1 / us-east-2 / us-west-2, and Bedrock authorizes
// InvokeModel against BOTH the inference-profile ARN AND the underlying
// foundation-model ARN in each routed region — all four resources are required
// or the call AccessDenies. Scoped to opus-4-8 ONLY (least-privilege): adding a
// new Bedrock model to SUPPORTED_MODELS means extending this resource list.
// IAM change — flag for T4 (GPT-4.1 IAM cross-review) / T5 (/sh-security-review).
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "InvokeBedrockClaude",
actions: ["bedrock:InvokeModel", "bedrock:InvokeModelWithResponseStream"],
resources: [
`arn:aws:bedrock:${REGION}:${ACCOUNT}:inference-profile/us.anthropic.claude-opus-4-8`,
"arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-opus-4-8",
"arn:aws:bedrock:us-east-2::foundation-model/anthropic.claude-opus-4-8",
"arn:aws:bedrock:us-west-2::foundation-model/anthropic.claude-opus-4-8",
],
}),
);
// Ship application logs to CloudWatch Logs under /open-swe/<env>/*.
this.role.addToPolicy(
new iam.PolicyStatement({
sid: "PutAppLogs",
actions: [
"logs:CreateLogGroup",
"logs:CreateLogStream",
"logs:PutLogEvents",
"logs:DescribeLogStreams",
],
resources: [
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*`,
`arn:aws:logs:${REGION}:${ACCOUNT}:log-group:/open-swe/${env}/*:*`,
],
}),
);
}
}

View file

@ -1,38 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { GithubDeployRoles } from "./constructs/github-deploy-roles";
/**
* Account-level IAM stack: the per-ENV GitHub OIDC deploy roles
* (githubdeploy-open-swe-{infra,app}-{dev,prod} — four roles).
*
* T5 OSWE-IAC-01/02 fix: roles are split per env with env-scoped OIDC trust, so
* a dev-branch token cannot reach prod (prod roles require the GitHub
* `prod` Environment manual-approval gate). They live in this dedicated stack
* rather than the env stacks because IAM roles are global and this stack ships
* FIRST (TODO.md BLOCK#3): the infra OIDC roles + the repo deploy-role-ARN
* secrets must exist before any infra/secrets CI step. Synth-only until the
* Phase-1 security gate (T4 + T5) clears (T6).
*/
export class OpenSweIamStack extends cdk.Stack {
constructor(scope: Construct, id: string, props?: cdk.StackProps) {
super(scope, id, props);
const dev = new GithubDeployRoles(this, "DeployRolesDev", "dev");
const prod = new GithubDeployRoles(this, "DeployRolesProd", "prod");
cdk.Tags.of(this).add("project", "open-swe");
cdk.Tags.of(this).add("ManagedBy", "cdk");
const out = (id: string, role: { roleName?: string }, env: string, kind: string) =>
new cdk.CfnOutput(this, id, {
value: `arn:aws:iam::${this.account}:role/${role.roleName}`,
description: `OIDC role ARN for ${env} ${kind} deploys — set as the ${env} deploy-role secret.`,
});
out("InfraDeployRoleDevArn", dev.infraRole, "dev", "infra (CDK)");
out("AppDeployRoleDevArn", dev.appRole, "dev", "app (tag-scoped SSM + S3)");
out("InfraDeployRoleProdArn", prod.infraRole, "prod", "infra (CDK)");
out("AppDeployRoleProdArn", prod.appRole, "prod", "app (tag-scoped SSM + S3)");
}
}

View file

@ -1,90 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Construct } from "constructs";
import { EnvName, prefix } from "./config";
import { AppService } from "./constructs/app-service";
import { AssetsBucket } from "./constructs/assets-bucket";
import { ConfigStore } from "./constructs/config-store";
import { InstanceRole } from "./constructs/instance-role";
import { BAKED_OPEN_SWE_AMI_ID } from "./constructs/ami-cache";
export interface OpenSweStackProps extends cdk.StackProps {
/** open-swe environment — drives the `open-swe-<env>-*` resource naming. */
readonly envName: EnvName;
}
/**
* Per-env open-swe stack (`open-swe-dev` / `open-swe-prod`). Resource names are
* prefixed `open-swe-<env>-*`.
*
* Composes: the per-env least-privilege instance role (T6), the Secrets/SSM
* config store (T11), and the compute + ingress wiring (T12, AppService — EC2
* box, instance SG, target group, imported-listener rules, Route53 aliases,
* 30-day log groups). The shared VPC and ALB are imported, never owned. Synth is
* offline (AMI is the cdk.context.json-pinned placeholder until T12-deploy).
*/
export class OpenSweStack extends cdk.Stack {
public readonly instanceRole: InstanceRole;
public readonly configStore: ConfigStore;
public readonly assetsBucket: AssetsBucket;
public readonly appService: AppService;
constructor(scope: Construct, id: string, props: OpenSweStackProps) {
super(scope, id, props);
const envName = props.envName;
const p = prefix(envName);
cdk.Tags.of(this).add("project", "open-swe");
cdk.Tags.of(this).add("env", envName);
cdk.Tags.of(this).add("ManagedBy", "cdk");
// Per-env least-privilege EC2 instance role (open-swe-<env>-instance-role).
this.instanceRole = new InstanceRole(this, "Instance", envName);
// Secrets Manager + SSM Parameter Store shells the boot hook reads
// (deploy/seahaven/fetch-config.sh). Secret shells are value-less and
// populated out-of-band; IaC-managed SSM params carry real derivable values.
// The instance role already grants read on open-swe-<env>/* + /open-swe-<env>/*.
this.configStore = new ConfigStore(this, "Config", { envName });
// T7: the S3 artifact bucket (open-swe-<env>-assets) CI uploads releases to
// and the box pulls app.tar.gz / spa.tar.gz from. The instance role already
// grants read on it by name; the app deploy role grants write.
this.assetsBucket = new AssetsBucket(this, "Assets", envName);
// Surface the baked open-swe base AMI id the box runs on (pinned by id in
// ami-cache.ts; refreshed by a deliberate packer rebuild → replacement).
new cdk.CfnOutput(this, "BakedAmiId", {
value: BAKED_OPEN_SWE_AMI_ID,
description: "Baked open-swe-base-arm64 AMI id consumed by the EC2 instance.",
});
// T12: compute + ingress. Imports the shared seahaven-vpc + ALB and adds the
// env's EC2 box, instance SG, target group, listener rules, DNS, log groups.
this.appService = new AppService(this, "App", {
envName,
instanceRole: this.instanceRole.role,
});
new cdk.CfnOutput(this, "InstanceRoleArn", {
value: this.instanceRole.role.roleArn,
description: `${p} EC2 instance role ARN.`,
});
new cdk.CfnOutput(this, "InstanceId", {
value: this.appService.instance.instanceId,
description: `${p} EC2 instance id.`,
});
new cdk.CfnOutput(this, "TargetGroupArn", {
value: this.appService.targetGroup.targetGroupArn,
description: `${p} ALB target group ARN (→ instance:80 nginx).`,
});
new cdk.CfnOutput(this, "AssetsBucketName", {
value: this.assetsBucket.bucket.bucketName,
description: `${p} S3 artifact bucket (CI uploads releases; box pulls).`,
});
new cdk.CfnOutput(this, "DeployDocumentName", {
value: this.appService.deployDocumentName,
description: `${p} SSM document that rolls the box to the latest release.`,
});
}
}

4487
infra/package-lock.json generated

File diff suppressed because it is too large Load diff

View file

@ -1,31 +0,0 @@
{
"name": "open-swe-infra",
"version": "1.0.0",
"description": "Open SWE AWS infrastructure (CDK TypeScript) — open-swe-dev / open-swe-prod stacks + shared OIDC deploy roles.",
"private": true,
"bin": {
"open-swe-infra": "bin/app.js"
},
"scripts": {
"build": "tsc",
"cdk": "cdk",
"synth": "cdk synth",
"diff": "cdk diff",
"test": "jest"
},
"devDependencies": {
"@types/jest": "^29.5.14",
"@types/node": "^24.0.0",
"@types/source-map-support": "^0.5.10",
"aws-cdk": "^2.1029.0",
"jest": "^29.7.0",
"source-map-support": "^0.5.21",
"ts-jest": "^29.2.5",
"ts-node": "^10.9.2",
"typescript": "~5.6.3"
},
"dependencies": {
"aws-cdk-lib": "2.260.0",
"constructs": "^10.0.0"
}
}

View file

@ -1,87 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Template } from "aws-cdk-lib/assertions";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
/**
* Several AWS APIs reject non-ASCII in fields that `cdk synth` happily emits and
* `tsc` happily compiles — so a stray em-dash/arrow only blows up at DEPLOY time
* (e.g. EC2 SecurityGroup GroupDescription: "Character sets beyond ASCII are not
* supported"). This has bitten us twice (the AMI Description, then the instance-SG
* description). This test fails the build at synth time instead.
*
* Scope: the EC2 fields with a documented ASCII/restricted-charset constraint —
* SecurityGroup GroupDescription and ingress/egress rule descriptions. (CloudFormation
* Output descriptions + Route53 comments accept UTF-8, so they are not asserted.)
*
* EC2 rule descriptions are stricter than ASCII: the allowed set is
* `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*` — note it EXCLUDES `<` and `>`, which is why a
* naive em-dash -> "->" replacement still fails at deploy. We assert that exact set.
*/
// Characters NOT in the EC2 description allowed set.
const DISALLOWED = /[^a-zA-Z0-9. _:/()#,@[\]+=&;{}!$*-]/;
function synthDev(): Record<string, { Type: string; Properties?: Record<string, unknown> }> {
const app = new cdk.App();
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
return Template.fromStack(dev).toJSON().Resources;
}
describe("ASCII-only EC2 description fields", () => {
const resources = synthDev();
it("SecurityGroup GroupDescription is ASCII", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::SecurityGroup") continue;
const desc = (r.Properties?.GroupDescription as string) ?? "";
expect(DISALLOWED.test(desc) ? `${id}: ${desc}` : "ascii").toBe("ascii");
}
});
it("SecurityGroup ingress/egress rule descriptions are ASCII", () => {
for (const [id, r] of Object.entries(resources)) {
const props = r.Properties ?? {};
const groups: Array<{ Description?: string }> = [];
if (Array.isArray(props.SecurityGroupIngress)) groups.push(...props.SecurityGroupIngress);
if (Array.isArray(props.SecurityGroupEgress)) groups.push(...props.SecurityGroupEgress);
// Standalone AWS::EC2::SecurityGroupEgress / ...Ingress resources.
if (r.Type === "AWS::EC2::SecurityGroupEgress" || r.Type === "AWS::EC2::SecurityGroupIngress") {
groups.push(props as { Description?: string });
}
for (const rule of groups) {
const desc = rule.Description ?? "";
expect(DISALLOWED.test(desc) ? `${id}: ${desc}` : "ascii").toBe("ascii");
}
}
});
// EC2 caps base64-encoded user-data at 25600 bytes; CDK + tsc don't check it, so
// an oversized boot script (e.g. an embedded deploy.sh) only fails at deploy.
it("EC2 user-data fits the 25600-byte encoded limit", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::Instance") continue;
const ud = (r.Properties?.UserData as { "Fn::Base64"?: string }) ?? {};
const script = typeof ud["Fn::Base64"] === "string" ? ud["Fn::Base64"] : "";
const encoded = Buffer.from(script, "utf8").toString("base64").length;
expect(`${id}: ${encoded} bytes`).toBe(encoded < 25600 ? `${id}: ${encoded} bytes` : "OVER 25600");
}
});
// CDK substitutes %%...%% tokens in user-data at synth. Any %%TOKEN%% left in the
// rendered script means a token wasn't wired in app-service.ts (the @@...@@ tokens
// are intentional — user-data seds those into the baked templates at boot).
it("user-data has no unresolved %%CDK%% tokens", () => {
for (const [id, r] of Object.entries(resources)) {
if (r.Type !== "AWS::EC2::Instance") continue;
const ud = (r.Properties?.UserData as { "Fn::Base64"?: string }) ?? {};
const script = typeof ud["Fn::Base64"] === "string" ? ud["Fn::Base64"] : "";
const leftover = script.match(/%%[A-Z0-9_]+%%/g) ?? [];
expect(`${id}: ${leftover.join(",")}`).toBe(`${id}: `);
}
});
});

View file

@ -1,55 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Match, Template } from "aws-cdk-lib/assertions";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { bootstrapQualifier } from "../lib/config";
const ENV = { account: "328440206208", region: "us-east-1" };
// B-1 / OSWE-IAC-01: each env's infra deploy role may assume ONLY its own
// bootstrap qualifier's roles. Dev runs on `oswedev`, so a dev-branch token can
// no longer assume the default `hnb659fds` bootstrap roles whose admin
// cfn-exec-role deploys prod. Prod stays on the default qualifier.
describe("Per-env CDK bootstrap qualifier isolation (B-1/OSWE-IAC-01)", () => {
it("maps dev -> oswedev and prod -> hnb659fds", () => {
expect(bootstrapQualifier("dev")).toBe("oswedev");
expect(bootstrapQualifier("prod")).toBe("hnb659fds");
});
it("dev infra deploy role assumes only cdk-oswedev-* bootstrap roles", () => {
const app = new cdk.App();
const stack = new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env: ENV,
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "AssumeCdkBootstrapRoles",
Action: "sts:AssumeRole",
Resource: "arn:aws:iam::328440206208:role/cdk-oswedev-*",
}),
]),
}),
});
});
it("prod infra deploy role stays on the default cdk-hnb659fds-* bootstrap roles", () => {
const app = new cdk.App();
const stack = new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env: ENV,
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "AssumeCdkBootstrapRoles",
Action: "sts:AssumeRole",
Resource: "arn:aws:iam::328440206208:role/cdk-hnb659fds-*",
}),
]),
}),
});
});
});

View file

@ -1,109 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Annotations, Match, Template } from "aws-cdk-lib/assertions";
import * as iam from "aws-cdk-lib/aws-iam";
import { KebabNamingAspect, isKebabCase } from "../lib/aspects/kebab-naming-aspect";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
describe("isKebabCase", () => {
it.each([
"open-swe-dev",
"open-swe-prod-instance-role",
"githubdeploy-open-swe-infra",
"open-swe-dev/slack-signing", // Secrets Manager path
"/open-swe-dev/feature-flag", // SSM param path
"/open-swe/dev/agent", // log group path
"abc123",
])("accepts conforming name %s", (name) => {
expect(isKebabCase(name)).toBe(true);
});
it.each([
"OpenSweDev",
"open_swe_dev",
"openSweDev",
"Open-Swe-Dev",
"open-swe-dev/SlackSigning",
])("rejects non-conforming name %s", (name) => {
expect(isKebabCase(name)).toBe(false);
});
});
describe("KebabNamingAspect", () => {
it("passes the real app stacks (no errors)", () => {
const app = new cdk.App();
cdk.Aspects.of(app).add(new KebabNamingAspect());
new OpenSweIamStack(app, "OpenSweIamStack", { stackName: "open-swe-iam", env: ENV });
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
const prod = new OpenSweStack(app, "OpenSweProdStack", {
stackName: "open-swe-prod",
env: ENV,
envName: "prod",
});
for (const s of [dev, prod]) {
Annotations.fromStack(s).hasNoError("*", Match.anyValue());
}
});
it("exempts Secrets Manager + SSM names that carry the literal env-var segment", () => {
// The config store names a secret open-swe-dev/ANTHROPIC_API_KEY and a param
// /open-swe-dev/SANDBOX_TYPE — the UPPER_SNAKE last segment is a REQUIRED
// deviation from kebab (fetch-config naming contract). Must NOT be flagged.
const app = new cdk.App();
cdk.Aspects.of(app).add(new KebabNamingAspect());
const dev = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
const tpl = Template.fromStack(dev);
// Sanity: the shells actually render with the literal env-var names.
tpl.hasResourceProperties("AWS::SecretsManager::Secret", {
Name: "open-swe-dev/ANTHROPIC_API_KEY",
});
tpl.hasResourceProperties("AWS::SSM::Parameter", {
Name: "/open-swe-dev/SANDBOX_TYPE",
Value: "langsmith",
});
Annotations.fromStack(dev).hasNoError("*", Match.anyValue());
});
it("flags a deliberately non-kebab-case resource name", () => {
const app = new cdk.App();
const stack = new cdk.Stack(app, "ConformingStackId", { stackName: "open-swe-test", env: ENV });
cdk.Aspects.of(stack).add(new KebabNamingAspect());
// Deliberately bad physical name — must be flagged.
new iam.Role(stack, "BadlyNamedRole", {
roleName: "OpenSweBadRole",
assumedBy: new iam.ServicePrincipal("ec2.amazonaws.com"),
});
Annotations.fromStack(stack).hasError(
"*",
Match.stringLikeRegexp("not kebab-case"),
);
});
it("flags a deliberately non-kebab-case stack name", () => {
const app = new cdk.App();
// PascalCase stackName — the convention CDK defaults to and that we forbid.
const stack = new cdk.Stack(app, "BadStack", { stackName: "OpenSweBadStack", env: ENV });
cdk.Aspects.of(stack).add(new KebabNamingAspect());
Annotations.fromStack(stack).hasError(
"*",
Match.stringLikeRegexp("Stack name .* is not kebab-case"),
);
});
});

View file

@ -1,71 +0,0 @@
import * as cdk from "aws-cdk-lib";
import { Match, Template } from "aws-cdk-lib/assertions";
import { OpenSweIamStack } from "../lib/open-swe-iam-stack";
import { OpenSweStack } from "../lib/open-swe-stack";
const ENV = { account: "328440206208", region: "us-east-1" };
// F-1 / IAC-04: s3:ListBucket must be constrained to the releases/ prefix so a
// compromised box / leaked CI token cannot enumerate the rest of the bucket.
const RELEASES_PREFIX_CONDITION = { StringLike: { "s3:prefix": ["releases/*"] } };
describe("S3 ListBucket prefix scoping (F-1/IAC-04)", () => {
it("instance role ListBucket is constrained to releases/*", () => {
const app = new cdk.App();
const stack = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "ListArtifactBucket",
Action: "s3:ListBucket",
Condition: RELEASES_PREFIX_CONDITION,
}),
]),
}),
});
});
it("github deploy app role ListBucket is constrained to releases/*", () => {
const app = new cdk.App();
const stack = new OpenSweIamStack(app, "OpenSweIamStack", {
stackName: "open-swe-iam",
env: ENV,
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "ListArtifactBucket",
Action: "s3:ListBucket",
Condition: RELEASES_PREFIX_CONDITION,
}),
]),
}),
});
});
it("GetBucketLocation stays a separate, unconditioned statement", () => {
const app = new cdk.App();
const stack = new OpenSweStack(app, "OpenSweDevStack", {
stackName: "open-swe-dev",
env: ENV,
envName: "dev",
});
Template.fromStack(stack).hasResourceProperties("AWS::IAM::Policy", {
PolicyDocument: Match.objectLike({
Statement: Match.arrayWith([
Match.objectLike({
Sid: "GetArtifactBucketLocation",
Action: "s3:GetBucketLocation",
Condition: Match.absent(),
}),
]),
}),
});
});
});

View file

@ -1,24 +0,0 @@
{
"compilerOptions": {
"target": "ES2022",
"module": "commonjs",
"lib": ["ES2022"],
"types": ["node", "jest"],
"declaration": true,
"strict": true,
"noImplicitAny": true,
"strictNullChecks": true,
"noImplicitReturns": true,
"noFallthroughCasesInSwitch": true,
"inlineSourceMap": true,
"inlineSources": true,
"strictPropertyInitialization": false,
"outDir": "./cdk.out",
"rootDir": ".",
"skipLibCheck": true,
"forceConsistentCasingInFileNames": true,
"resolveJsonModule": true,
"esModuleInterop": true
},
"exclude": ["node_modules", "cdk.out"]
}

View file

@ -732,6 +732,15 @@ export const api = {
request<UserMappingsPage>(
`/admin/user-mappings?page=${page}&page_size=${pageSize}`
),
adminUpsertUserMapping: (input: {
github_login: string
work_email: string
slack_user_id?: string | null
}) =>
request<UserMapping>("/admin/user-mappings", {
method: "POST",
body: JSON.stringify(input),
}),
adminDeleteUserMapping: (github_login: string) =>
request<{ deleted: boolean }>(
`/admin/user-mappings/${encodeURIComponent(github_login)}`,

View file

@ -171,6 +171,8 @@ const PAGE_SIZE = 20
function UserMappingsSection({ enabled }: { enabled: boolean }) {
const [error, setError] = useState<string | null>(null)
const [page, setPage] = useState(1)
const [newLogin, setNewLogin] = useState("")
const [newEmail, setNewEmail] = useState("")
const mappings = useQuery({
queryKey: ["adminUserMappings", page],
@ -193,16 +195,63 @@ function UserMappingsSection({ enabled }: { enabled: boolean }) {
onError: (e: Error) => setError(e.message),
})
const upsert = useMutation({
mutationFn: () =>
api.adminUpsertUserMapping({
github_login: newLogin.trim(),
work_email: newEmail.trim(),
}),
onSuccess: () => {
setNewLogin("")
setNewEmail("")
setError(null)
void mappings.refetch()
},
onError: (e: Error) => setError(e.message),
})
const canSubmit =
newLogin.trim().length > 0 &&
newEmail.trim().length > 0 &&
!upsert.isPending
const items = mappings.data?.items ?? []
return (
<SettingsSection
title="User mappings"
description="Mappings are created when users connect Slack from settings. Admins can remove stale mappings here."
description="Mappings link a GitHub login to a work email. They are created automatically when users connect Slack from settings; admins can also add or remove them here."
>
<div className="flex flex-col gap-3 p-4">
{error && <span className="text-xs text-destructive">{error}</span>}
<form
className="flex flex-col gap-2 sm:flex-row sm:items-center"
onSubmit={(e) => {
e.preventDefault()
if (canSubmit) upsert.mutate()
}}
>
<Input
value={newLogin}
onChange={(e) => setNewLogin(e.target.value)}
placeholder="github-login"
aria-label="GitHub login"
className="h-8 text-xs"
/>
<Input
value={newEmail}
onChange={(e) => setNewEmail(e.target.value)}
placeholder="work@example.com"
aria-label="Work email"
type="email"
className="h-8 text-xs"
/>
<Button type="submit" size="sm" disabled={!canSubmit}>
{upsert.isPending ? "Saving…" : "Add / update"}
</Button>
</form>
<div className="flex flex-col gap-0.5">
{mappings.isLoading ? (
<Skeleton className="h-32" />