mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 08:03:15 +00:00
* feat(infra): build + pin the baked open-swe-base-arm64 AMI (T12 AMI / item 3)
Packer-build the custom base image and repoint AppService off the AL2023
placeholder onto it.
deploy/ami/open-swe-base.pkr.hcl — fix two bugs that blocked the first real
`packer build` (the config had only ever been `packer validate`'d at T8):
- the file provisioner failed uploading the templates dir ('scp: …: Is a
directory') — a trailing-slash contents-upload needs the dest dir to exist;
added a 'mkdir -p /tmp/open-swe-templates' shell provisioner + dropped the
dest trailing slash.
- the shell provisioner's custom execute_command omitted {{ .Vars }}, so the
environment_vars never reached provision.sh (which runs under set -u and
aborted on CLOUDWATCH_AGENT_DEB_URL). Added {{ .Vars }}.
infra:
- ami-cache.ts: BAKED_OPEN_SWE_AMI_ID = ami-0545363bb147229ff (built 2026-06-26
from open-swe-base-arm64-20260626-201929) + bakedOpenSweArm64() pinning it by
exact id via MachineImage.genericLinux (offline, deterministic). Dropped the
now-dead AL2023 cachedInContext helper + context key; kept the EBS/replacement
discipline docs.
- app-service.ts: machineImage → bakedOpenSweArm64().
- open-swe-stack.ts: output BakedAmiId (was the AL2023 PinnedAmiId guard).
- cdk.context.json → {} (AMI is a static id pin; no context lookups remain).
- README: Baked AMI + EBS-replacement-discipline section.
tsc + cdk synth(dev+prod) + jest(16) clean; template ImageId = the baked AMI.
NOTE: held — do NOT merge until the open-swe-dev secret values are populated
(put-config.sh). The infra CD is live, so merging this to dev auto-deploys
OpenSweDevStack; without secrets the box boots but fetch-config fail-fasts →
unhealthy ALB target on the shared prod ALB. Merge once secrets are set (T14).
* fix(ami): ASCII-only AMI description + re-pin to ami-00080084502093021
Third packer bug: ami_description had an em-dash (non-ASCII); AWS rejects
non-ASCII in the AMI Description attribute, so packer registered then
DEREGISTERED the first AMI (ami-0545…) on the ModifyImageAttribute error.
Replaced with an ASCII '-'. Rebuilt clean → ami-00080084502093021 (available).
Re-pinned BAKED_OPEN_SWE_AMI_ID.
* fix(deploy): GitHub App + Slack required for prod only, not dev
Per the migration decision: do NOT create/duplicate a separate dev GitHub App or
Slack app — only prod owns the single shared app. So fetch-config.sh no longer
hard-requires the GitHub App quintet (ID/PRIVATE_KEY/INSTALLATION_ID/CLIENT_ID/
CLIENT_SECRET) + Slack/webhook secrets for dev; they move into the prod-only
block alongside the existing GITHUB_WEBHOOK_SECRET/SLACK_SIGNING_SECRET.
Dev now boots with just DASHBOARD_JWT_SECRET + TOKEN_ENCRYPTION_KEY + the active
provider key(s) + the langsmith sandbox keys. Dev is a deployment-validation env
(boot/health/boundary) with no GitHub/Slack/webhook integration; prod parity is
unchanged (prod still requires everything).
* feat: stand up dev properly — S3 assets bucket + artifact CD + on-box uv sync (T7+T19)
Make the dev/prod box deployable end-to-end: a real artifact pipeline and a
re-runnable on-box deploy, so OpenSweDevStack can come up genuinely healthy.
Infra (T7):
- assets-bucket.ts: open-swe-<env>-assets S3 bucket — BLOCK_ALL public access,
SSE-S3, enforceSSL (deny non-TLS), versioned, lifecycle (expire noncurrent +
abort MPU), RETAIN. Wired into OpenSweStack + CfnOutput.
- app-service.ts: open-swe-<env>-deploy SSM document that runs the baked
/opt/open-swe/bin/deploy.sh (tag-scoped roll-the-box). machineImage is the
baked open-swe-base-arm64 AMI (folds in the held #16).
IAM (app deploy role — cross-review gated):
- github-deploy-roles.ts: app role gains s3:PutObject/DeleteObject scoped to
open-swe-<env>-assets/releases/* (CI uploads releases). Drops the generic
AWS-RunShellScript grant now that the dedicated open-swe-<env>-deploy document
is the only SendCommand path — closes the T4 BLOCK#3 arbitrary-shell timebox.
Boot/deploy (T19):
- deploy/ami/deploy.sh: single, re-runnable app-deploy procedure — pull
app.tar.gz/spa.tar.gz from S3, `uv sync --frozen --no-dev` (native ARM64 venv
at the real path, py3.12 pre-baked), restart open-swe.service + reload nginx.
- user-data.sh: nginx starts BEFORE the app deploy (static /healthz -> the ALB
target is healthy even before the first release); deploy.sh is base64-rendered
by CDK into user-data (a normal reviewable repo file, not a heredoc) and the
first-boot deploy is NON-FATAL (no release yet -> wait for the first SSM deploy).
CI (T7+T19):
- build-artifacts.yml (+ .github/scripts): build the SPA with bun (vite ->
ui/.output/public -> spa.tar.gz), package the Python source via git archive
(app.tar.gz, no ui/ no .venv), upload to releases/<sha>/ + releases/latest/ via
the githubdeploy-open-swe-app-<env> OIDC role, then fire open-swe-<env>-deploy.
push dev -> dev (auto); push main -> prod (env "prod" approval gate).
Local: ruff/shellcheck clean, tsc clean, jest 16/16, cdk synth offline OK,
deploy.sh base64 round-trips exact.
* harden(sec-review): tar extraction, deploy gating, least-privilege, secret guard
Address the /sh-security-review fan-out + proof-or-kill verifier pass. Only one
confirmed-high surfaced and it is PRE-EXISTING and out-of-diff (OSWE-IAC-AUDIT-01,
the account-wide CDK cfn-exec residual already documented in config.ts; recorded in
.security-review/suppressions.json with justification + flagged for the per-env
bootstrap-qualifier follow-up). The rest were verifier-downgraded to unverified;
these are the cheap defense-in-depth fixes worth taking regardless:
- deploy.sh: extract tarballs with --no-same-owner --no-same-permissions (root
never honors an archive's uid/mode → no setuid/foreign-owned file can land); and
treat "no release in S3 yet" as a benign exit 0, distinct from a real deploy
failure (set -e stays loud once a release exists).
- publish-and-deploy.sh: gate on the AGGREGATE SSM Command.Status (+ TargetCount),
not CommandInvocations[0], so a partial failure across the brief 2-instance
replacement window can't be reported as success.
- instance-role.ts: scope the box's s3:GetObject to releases/* (mirrors the app
role's write scope) instead of the whole bucket.
- package-artifacts.sh: fail-closed secret-shaped-file guard on app.tar.gz
(defense in depth over .gitignore; scoped to data extensions so *_credentials.py
source is not a false positive — verified against the real tree).
Deferred as documented follow-ups (verifier: unverified, supply-chain-gated to the
CI OIDC writer; bucket is BLOCK_ALL + enforceSSL + versioned): SHA-pinned immutable
releases/<sha>/ pulls + signed checksum (vs mutable latest/), single-tarball release
to remove the torn-read window, and app-aware ALB health (vs static nginx /healthz).
shellcheck/tsc/jest(16) clean; both stacks synth offline.
* fix(infra): ASCII-only EC2 SecurityGroup descriptions + synth-time guard
The instance-SG GroupDescription + ingress/egress rule descriptions carried an
em-dash / arrow (—, →). `tsc` and `cdk synth` accept them, but the EC2 API rejects
non-ASCII in GroupDescription ("Character sets beyond ASCII are not supported"),
so OpenSweDevStack's first deploy failed at the SG and rolled back. (Pre-existing
from #14; same class as the AMI-description ASCII bug.)
- app-service.ts: replace —/→ with ASCII (- / ->) in the SG GroupDescription, the
ingress/egress rule descriptions, and the Route53 comment.
- test/ascii-aws-fields.test.ts: synth-time guard asserting EC2 SecurityGroup
GroupDescription + rule descriptions are pure ASCII, so this fails the build
instead of a deploy next time.
jest 18/18; tsc clean.
* fix(infra): SG rule descriptions use ASCII-charset-safe text (no `>`)
The first ASCII fix replaced the arrow with `->`, but EC2 SecurityGroup *rule*
descriptions allow a stricter set than ASCII — `a-zA-Z0-9. _-:/()#,@[]+=&;{}!$*`,
which EXCLUDES `<`/`>`. So OpenSweDevStack's second deploy still failed at the
ingress rule. Use "to" instead of "->", and tighten the guard test from "ASCII
only" to the exact EC2 allowed charset so it catches `>` (and `<`) too.
jest 18/18; tsc clean.
* fix(infra): minify embedded deploy.sh so user-data fits EC2's 25.6 KB limit
The base64 deploy.sh embedded in user-data pushed the encoded boot script to
27184 bytes, over EC2's 25600-byte cap, so OpenSweDevStack's instance failed with
"Encoded User data is limited to 25600 bytes". Strip full-line comments + blank
lines from deploy.sh before base64-embedding it (repo file keeps comments; only
the on-box copy is minified; the script is opaque base64 so user-data heredocs are
unaffected) -> rendered user-data drops to 16424 bytes (9 KB margin). Add a
synth-time guard test asserting EC2 user-data stays under 25600 bytes encoded.
jest 19/19; minified deploy.sh passes bash -n + shellcheck.
291 lines
13 KiB
Bash
Executable file
291 lines
13 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# fetch-config.sh — AWS-sourced boot hook that materializes the app's .env.
|
|
#
|
|
# The stock `langgraph dev` runtime + the Open SWE app read a plain `.env` from
|
|
# the app working directory (python-dotenv). On the AWS lift-and-shift we do NOT
|
|
# commit a .env; instead every non-sensitive value lives in SSM Parameter Store
|
|
# (`/open-swe-<env>/*`) and every secret lives in AWS Secrets Manager
|
|
# (`open-swe-<env>/*`). This hook is run by systemd BEFORE the service starts; it
|
|
# pulls both sources via the EC2 instance role (no static keys), assembles a
|
|
# single .env on a tmpfs, and writes it owned by the unprivileged service user
|
|
# `chmod 600` (T5 SC-01: the privileged pre-hook materializes the secret; the app
|
|
# itself then runs as that NON-root service user, not root).
|
|
#
|
|
# It is intentionally FAIL-FAST: if any required secret/param is missing or empty
|
|
# it prints the offending variable NAMES (never values) and exits 1, so the
|
|
# service never starts with a partial .env.
|
|
#
|
|
# ---------------------------------------------------------------------------
|
|
# Naming contract (source of truth: T9 env/secret/config inventory)
|
|
# SSM /open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
|
|
# Secrets open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
|
|
# i.e. the last path segment IS the literal environment-variable name. This is a
|
|
# deliberate (documented) deviation from the handbook's kebab-case value-name
|
|
# example (`my-stack/slack-signing`): a .env materializer needs a lossless,
|
|
# unambiguous round-trip from store key -> env var, and the env var name is the
|
|
# only key that guarantees that. The `open-swe-<env>` stack prefix still follows
|
|
# kebab-case per naming-conventions.md.
|
|
# ---------------------------------------------------------------------------
|
|
#
|
|
# Wiring into systemd (AWS EC2 variant):
|
|
# The unit runs as the unprivileged service user (User=openswe). ONLY the
|
|
# ExecStartPre pre-hook runs as root (the `+` prefix) so it can pull from AWS,
|
|
# write the tmpfs .env, and chown it to the service user. The app (ExecStart)
|
|
# and the seeder (ExecStartPost) then run as openswe and read the openswe-owned
|
|
# 0600 .env — the agent never runs as root (T5 SC-01). Pass the env as the
|
|
# positional arg (T5 BOOT-01):
|
|
#
|
|
# [Service]
|
|
# User=openswe
|
|
# Group=openswe
|
|
# Environment=ENV_DIR=/run/open-swe SERVICE_USER=openswe
|
|
# # ExecStartPre runs as root (+) so it can chown the .env to the service user.
|
|
# ExecStartPre=+/opt/open-swe/deploy/seahaven/fetch-config.sh prod
|
|
# ExecStart=/opt/open-swe/.venv/bin/langgraph dev --host 127.0.0.1 --port 2024 \
|
|
# --no-browser --no-reload
|
|
# ExecStartPost=/opt/open-swe/deploy/seahaven/seed_store.sh prod
|
|
#
|
|
# tmpfs: /run is already a tmpfs on systemd hosts, so ENV_DIR=/run/open-swe is
|
|
# tmpfs-backed by default (the .env never touches disk). Set RUN_DEDICATED_TMPFS=1
|
|
# to mount a private tmpfs at ENV_DIR instead. The app's CWD `.env` is a symlink
|
|
# into ENV_DIR (created idempotently below), so python-dotenv finds it unchanged.
|
|
#
|
|
# Idempotent, re-runnable on every (re)start. No secret is ever echoed.
|
|
|
|
set -euo pipefail
|
|
umask 077
|
|
|
|
# --- Inputs ------------------------------------------------------------------
|
|
ENV="${1:-${OPENSWE_ENV:-}}"
|
|
case "$ENV" in
|
|
dev | prod) ;;
|
|
*)
|
|
echo "fetch-config: ENV must be 'dev' or 'prod' (got '${ENV:-<empty>}')" >&2
|
|
echo "usage: fetch-config.sh <dev|prod> (or set OPENSWE_ENV)" >&2
|
|
exit 2
|
|
;;
|
|
esac
|
|
|
|
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
|
|
SSM_PREFIX="/open-swe-${ENV}/"
|
|
SECRET_PREFIX="open-swe-${ENV}/"
|
|
|
|
ENV_DIR="${ENV_DIR:-/run/open-swe}" # tmpfs-backed (/run) by default
|
|
ENV_FILE="${ENV_DIR}/.env"
|
|
APP_DIR="${APP_DIR:-/opt/open-swe}" # where the app + its CWD .env live
|
|
APP_ENV_LINK="${APP_DIR}/.env" # symlink -> ENV_FILE
|
|
|
|
# The unprivileged service user that runs the app and OWNS the .env (T5 SC-01).
|
|
# fetch-config runs as root (ExecStartPre=+) only to chown the secret to it.
|
|
SERVICE_USER="${SERVICE_USER:-openswe}"
|
|
SERVICE_GROUP="${SERVICE_GROUP:-${SERVICE_USER}}"
|
|
|
|
# Sea Haven override: upstream defaults DEFAULT_REPO_OWNER to "langchain-ai".
|
|
# For the Sea Haven deployment it MUST be the org. fetch-config forces this so a
|
|
# stale/blank SSM value can never point the agent at the upstream org.
|
|
SH_REPO_OWNER="${OPENSWE_REPO_OWNER:-Sea-Haven-Industries}"
|
|
|
|
for bin in aws jq; do
|
|
command -v "$bin" >/dev/null 2>&1 || { echo "fetch-config: '$bin' not found on PATH" >&2; exit 3; }
|
|
done
|
|
|
|
log() { echo "fetch-config[$ENV]: $*"; } # NAMES/counts only — never values
|
|
b64d() { base64 --decode; } # GNU coreutils on the EC2 host
|
|
|
|
# Accept a store key into VARS iff it is a valid env-var identifier and not a
|
|
# duplicate. Rejects non-identifier names (T5 SH-INJ-002 / set -e DoS hardening)
|
|
# and flat-namespace collisions (T5 SSM-05). $3 = source label for logs.
|
|
accept_var() {
|
|
local key="$1" value="$2" src="$3"
|
|
if ! [[ "$key" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
|
|
log "WARNING: skipping ${src} key with non-identifier name (rejected)"
|
|
return 0
|
|
fi
|
|
if [ -n "${VARS[$key]+set}" ]; then
|
|
echo "fetch-config[$ENV]: FAIL-FAST — duplicate key '${key}' from ${src} (flat-namespace collision)" >&2
|
|
exit 1
|
|
fi
|
|
VARS["$key"]="$value"
|
|
}
|
|
|
|
# --- tmpfs ------------------------------------------------------------------
|
|
mkdir -p "$ENV_DIR"
|
|
# Owned by the service user so the unprivileged app can traverse it (T5 SC-01).
|
|
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_DIR" 2>/dev/null || true
|
|
chmod 700 "$ENV_DIR"
|
|
if [ "${RUN_DEDICATED_TMPFS:-0}" = "1" ] && ! mountpoint -q "$ENV_DIR"; then
|
|
mount -t tmpfs -o nosuid,nodev,noexec,mode=0700,size=4m tmpfs "$ENV_DIR"
|
|
log "mounted dedicated tmpfs at $ENV_DIR"
|
|
fi
|
|
|
|
# --- Collect values into an associative array --------------------------------
|
|
declare -A VARS=()
|
|
|
|
# 1) SSM Parameter Store (non-sensitive config). NOT --recursive: the contract is
|
|
# a FLAT namespace /open-swe-<env>/<VAR>, so a non-recursive list returns exactly
|
|
# those keys and cannot collapse two nested paths onto one name (T5 SSM-05). aws
|
|
# CLI v2 auto-paginates NextToken.
|
|
log "reading SSM params under ${SSM_PREFIX} ..."
|
|
ssm_json="$(
|
|
aws ssm get-parameters-by-path \
|
|
--path "$SSM_PREFIX" \
|
|
--with-decryption \
|
|
--region "$REGION" \
|
|
--no-cli-pager \
|
|
--output json
|
|
)"
|
|
# Records are base64-encoded (name<TAB>value) so values with spaces/newlines/tabs
|
|
# survive the line-based read intact.
|
|
ssm_count=0
|
|
while IFS=$'\t' read -r nb vb; do
|
|
[ -n "$nb" ] || continue
|
|
name="$(printf '%s' "$nb" | b64d)"
|
|
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
|
|
key="${name##*/}" # strip /open-swe-<env>/ prefix
|
|
[ -n "$key" ] || continue
|
|
accept_var "$key" "$value" "SSM"
|
|
ssm_count=$((ssm_count + 1))
|
|
done < <(jq -r '.Parameters[] | (.Name|@base64) + "\t" + (.Value|@base64)' <<<"$ssm_json")
|
|
log "loaded ${ssm_count} config param(s) from SSM"
|
|
|
|
# 2) Secrets Manager (sensitive values). batch-get-secret-value filters by name
|
|
# prefix and auto-paginates; one secret per env var, SecretString = the value.
|
|
log "reading secrets under ${SECRET_PREFIX} ..."
|
|
secret_count=0
|
|
while IFS=$'\t' read -r nb vb; do
|
|
[ -n "$nb" ] || continue
|
|
name="$(printf '%s' "$nb" | b64d)"
|
|
case "$name" in
|
|
"${SECRET_PREFIX}"*) ;; # defensive: exact-prefix only
|
|
*) continue ;;
|
|
esac
|
|
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
|
|
key="${name##*/}"
|
|
[ -n "$key" ] || continue
|
|
accept_var "$key" "$value" "Secrets"
|
|
secret_count=$((secret_count + 1))
|
|
done < <(
|
|
aws secretsmanager batch-get-secret-value \
|
|
--filters "Key=name,Values=${SECRET_PREFIX}" \
|
|
--region "$REGION" \
|
|
--no-cli-pager \
|
|
--output json \
|
|
| jq -r '.SecretValues[] | select(.SecretString != null) | (.Name|@base64) + "\t" + (.SecretString|@base64)'
|
|
)
|
|
log "loaded ${secret_count} secret(s) from Secrets Manager"
|
|
|
|
# --- Sea Haven DEFAULT_REPO_OWNER hard pin -----------------------------------
|
|
# T5 OSWE-OWNER-04: a HARD pin, not a deny-list. Whatever the store holds (blank,
|
|
# upstream 'langchain-ai', a case/space variant, or any other org), the owner is
|
|
# unconditionally forced to the Sea Haven org so the agent can never target the
|
|
# wrong owner.
|
|
cur_owner="${VARS[DEFAULT_REPO_OWNER]:-}"
|
|
if [ -n "$cur_owner" ] && [ "$cur_owner" != "$SH_REPO_OWNER" ]; then
|
|
log "WARNING: SSM DEFAULT_REPO_OWNER differs from the pinned org -> hard-pinning to '${SH_REPO_OWNER}'"
|
|
fi
|
|
VARS[DEFAULT_REPO_OWNER]="$SH_REPO_OWNER"
|
|
|
|
# --- FAIL-FAST: required vars -------------------------------------------------
|
|
# Hard-required regardless of mode:
|
|
required=(
|
|
DASHBOARD_JWT_SECRET # RuntimeError on startup if missing (oauth.py)
|
|
TOKEN_ENCRYPTION_KEY # Fernet key(s); decrypts per-user GitHub tokens
|
|
)
|
|
# NOTE: the GitHub App is NOT created/duplicated for dev — only prod owns the
|
|
# (single, shared) GitHub App + Slack app. So the GitHub App quintet + Slack +
|
|
# webhook-signing secrets are required for PROD only (see the prod block below).
|
|
# Dev boots without them: it has no GitHub-App/Slack/webhook integration — it is a
|
|
# deployment-validation env (boot/health/boundary), not a live-triggered agent.
|
|
|
|
# Active model-provider key(s): model selection is store-driven (team_settings),
|
|
# so fetch-config cannot infer it from .env. Default to the seeded cross-family
|
|
# pair (anthropic builder + openai reviewer). Override with a comma list.
|
|
IFS=',' read -r -a provider_keys <<<"${REQUIRED_PROVIDER_KEYS:-ANTHROPIC_API_KEY,OPENAI_API_KEY}"
|
|
for k in "${provider_keys[@]}"; do
|
|
k="${k//[[:space:]]/}"
|
|
[ -n "$k" ] && required+=("$k")
|
|
done
|
|
|
|
# Sandbox provider key(s) — depends on SANDBOX_TYPE (default langsmith).
|
|
sandbox_type="${VARS[SANDBOX_TYPE]:-langsmith}"
|
|
case "$sandbox_type" in
|
|
langsmith) required+=(LANGSMITH_API_KEY_PROD DEFAULT_SANDBOX_SNAPSHOT_ID) ;;
|
|
daytona) required+=(DAYTONA_API_KEY) ;;
|
|
runloop) required+=(RUNLOOP_API_KEY) ;;
|
|
modal | local) ;; # no key required
|
|
*) log "WARNING: unknown SANDBOX_TYPE='${sandbox_type}' — not enforcing a sandbox key" ;;
|
|
esac
|
|
|
|
# Prod-only: the GitHub App (installation-token minting + dashboard OAuth) and the
|
|
# webhook-signing secrets. Dev has no GitHub/Slack app, so none of these are
|
|
# required there; prod owns the single shared app and must have all of them.
|
|
if [ "$ENV" = "prod" ]; then
|
|
required+=(
|
|
GITHUB_APP_ID # GitHub App trio (installation-token minting) ...
|
|
GITHUB_APP_PRIVATE_KEY # ... multiline PEM ...
|
|
GITHUB_APP_INSTALLATION_ID # ... used by utils/github_app.py
|
|
GITHUB_APP_CLIENT_ID # dashboard OAuth login
|
|
GITHUB_APP_CLIENT_SECRET # dashboard OAuth login
|
|
GITHUB_WEBHOOK_SECRET # webhook signature verification
|
|
SLACK_SIGNING_SECRET # Slack webhook signature verification
|
|
)
|
|
if [ -n "${VARS[LINEAR_API_KEY]:-}" ] && [ "${OPENSWE_REQUIRE_LINEAR:-1}" = "1" ]; then
|
|
required+=(LINEAR_WEBHOOK_SECRET)
|
|
fi
|
|
fi
|
|
|
|
missing=()
|
|
for k in "${required[@]}"; do
|
|
[ -n "${VARS[$k]:-}" ] || missing+=("$k")
|
|
done
|
|
# de-dup the names for a clean report
|
|
if [ "${#missing[@]}" -gt 0 ]; then
|
|
mapfile -t missing < <(printf '%s\n' "${missing[@]}" | sort -u)
|
|
echo "fetch-config[$ENV]: FAIL-FAST — ${#missing[@]} required var(s) missing/empty:" >&2
|
|
printf ' - %s\n' "${missing[@]}" >&2
|
|
echo "fetch-config[$ENV]: refusing to write a partial .env; service will not start." >&2
|
|
exit 1
|
|
fi
|
|
|
|
# --- Write the .env atomically (root-only on tmpfs) --------------------------
|
|
# python-dotenv reads double-quoted values (incl. multiline PEMs). Its decoder
|
|
# unescapes ONLY backslash and double-quote (\\ -> \, \" -> "); it does NOT honor
|
|
# \$ or \` escapes, so escaping those would leave a spurious backslash. Escape
|
|
# exactly backslash then double-quote — real newlines stay literal (multiline OK).
|
|
# (Caveat: python-dotenv interpolates a literal `${VAR}` substring; the secret
|
|
# domain here — base64/hex/PEM keys — never contains one, so no extra guard.)
|
|
emit_var() {
|
|
local name="$1" value="$2" esc
|
|
esc="${value//\\/\\\\}"
|
|
esc="${esc//\"/\\\"}"
|
|
printf '%s="%s"\n' "$name" "$esc"
|
|
}
|
|
|
|
tmp="$(mktemp "${ENV_DIR}/.env.XXXXXX")"
|
|
chmod 600 "$tmp"
|
|
{
|
|
printf '# Generated by fetch-config.sh for env=%s at %s — DO NOT EDIT.\n' \
|
|
"$ENV" "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
|
|
printf '# Source: SSM /open-swe-%s/* + Secrets Manager open-swe-%s/*\n\n' "$ENV" "$ENV"
|
|
for k in $(printf '%s\n' "${!VARS[@]}" | sort); do
|
|
emit_var "$k" "${VARS[$k]}"
|
|
done
|
|
} >"$tmp"
|
|
|
|
mv -f "$tmp" "$ENV_FILE"
|
|
# Owned by the unprivileged service user (T5 SC-01) so the app reads it without
|
|
# running as root. fetch-config itself runs as root (ExecStartPre=+) to chown.
|
|
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_FILE"
|
|
chmod 600 "$ENV_FILE"
|
|
|
|
# Point the app's CWD .env at the tmpfs file (idempotent).
|
|
if [ "$APP_ENV_LINK" != "$ENV_FILE" ]; then
|
|
if [ -L "$APP_ENV_LINK" ] || [ ! -e "$APP_ENV_LINK" ]; then
|
|
ln -sfn "$ENV_FILE" "$APP_ENV_LINK"
|
|
elif [ "$(readlink -f "$APP_ENV_LINK" 2>/dev/null || true)" != "$(readlink -f "$ENV_FILE")" ]; then
|
|
log "WARNING: ${APP_ENV_LINK} exists and is not a symlink to ${ENV_FILE} — leaving it untouched"
|
|
fi
|
|
fi
|
|
|
|
total=$((ssm_count + secret_count))
|
|
log "wrote ${ENV_FILE} (${total} vars, sandbox=${sandbox_type}) — ${SERVICE_USER}:${SERVICE_GROUP} 0600"
|