open-swe/deploy/seahaven/fetch-config.sh
Adam Moussa a4ed19ba61
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)

Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.

- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
  region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
  FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module

* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8

The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.

Map profile effort to additional_model_request_fields:
  {thinking: {type: adaptive, display: summarized},
   output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.

* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids

Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
  (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
  FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
  otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.

Surfaced by the cross-family review + verified against deploy/.

* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip

From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
  validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
  error code only, so the role ARN + account id in the raw botocore message never
  reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
  (Converse emits reasoning_content, not thinking) so the middleware is not a no-op
  on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)

* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids

Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
  us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
  each routed region (us-east-1/2, us-west-2). The model runs in the server process
  on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
  for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
  mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
  bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
  seed_store.sh's default via pick precedence, so the seed-script fix alone was
  insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
  ids to the Bedrock id (config.toml's model_id was an active, now-broken value).

AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.

* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)

Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
  (eval judge only — Bedrock builder/reviewer auth via the host IAM role).

REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00

363 lines
18 KiB
Bash
Executable file

#!/usr/bin/env bash
# fetch-config.sh — AWS-sourced boot hook that materializes the app's .env.
#
# The stock `langgraph dev` runtime + the Open SWE app read a plain `.env` from
# the app working directory (python-dotenv). On the AWS lift-and-shift we do NOT
# commit a .env; instead every non-sensitive value lives in SSM Parameter Store
# (`/open-swe-<env>/*`) and every secret lives in AWS Secrets Manager
# (`open-swe-<env>/*`). This hook is run by systemd BEFORE the service starts; it
# pulls both sources via the EC2 instance role (no static keys), assembles a
# single .env on a tmpfs, and writes it owned by the unprivileged service user
# `chmod 600` (T5 SC-01: the privileged pre-hook materializes the secret; the app
# itself then runs as that NON-root service user, not root).
#
# It is intentionally FAIL-FAST: if any required secret/param is missing or empty
# it prints the offending variable NAMES (never values) and exits 1, so the
# service never starts with a partial .env.
#
# ---------------------------------------------------------------------------
# Naming contract (source of truth: T9 env/secret/config inventory)
# SSM /open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# Secrets open-swe-<env>/<ENV_VAR_NAME> -> exported as ENV_VAR_NAME
# i.e. the last path segment IS the literal environment-variable name. This is a
# deliberate (documented) deviation from the handbook's kebab-case value-name
# example (`my-stack/slack-signing`): a .env materializer needs a lossless,
# unambiguous round-trip from store key -> env var, and the env var name is the
# only key that guarantees that. The `open-swe-<env>` stack prefix still follows
# kebab-case per naming-conventions.md.
# ---------------------------------------------------------------------------
#
# Wiring into systemd (AWS EC2 variant):
# The unit runs as the unprivileged service user (User=openswe). ONLY the
# ExecStartPre pre-hook runs as root (the `+` prefix) so it can pull from AWS,
# write the tmpfs .env, and chown it to the service user. The app (ExecStart)
# and the seeder (ExecStartPost) then run as openswe and read the openswe-owned
# 0600 .env — the agent never runs as root (T5 SC-01). Pass the env as the
# positional arg (T5 BOOT-01):
#
# [Service]
# User=openswe
# Group=openswe
# Environment=ENV_DIR=/run/open-swe SERVICE_USER=openswe
# # ExecStartPre runs as root (+) so it can chown the .env to the service user.
# ExecStartPre=+/opt/open-swe/deploy/seahaven/fetch-config.sh prod
# ExecStart=/opt/open-swe/.venv/bin/langgraph dev --host 127.0.0.1 --port 2024 \
# --no-browser --no-reload
# ExecStartPost=/opt/open-swe/deploy/seahaven/seed_store.sh prod
#
# tmpfs: /run is already a tmpfs on systemd hosts, so ENV_DIR=/run/open-swe is
# tmpfs-backed by default (the .env never touches disk). Set RUN_DEDICATED_TMPFS=1
# to mount a private tmpfs at ENV_DIR instead. The app's CWD `.env` is a symlink
# into ENV_DIR (created idempotently below), so python-dotenv finds it unchanged.
#
# Idempotent, re-runnable on every (re)start. No secret is ever echoed.
set -euo pipefail
umask 077
# --- Inputs ------------------------------------------------------------------
ENV="${1:-${OPENSWE_ENV:-}}"
case "$ENV" in
dev | prod) ;;
*)
echo "fetch-config: ENV must be 'dev' or 'prod' (got '${ENV:-<empty>}')" >&2
echo "usage: fetch-config.sh <dev|prod> (or set OPENSWE_ENV)" >&2
exit 2
;;
esac
REGION="${AWS_REGION:-${AWS_DEFAULT_REGION:-us-east-1}}"
SSM_PREFIX="/open-swe-${ENV}/"
SECRET_PREFIX="open-swe-${ENV}/"
ENV_DIR="${ENV_DIR:-/run/open-swe}" # tmpfs-backed (/run) by default
ENV_FILE="${ENV_DIR}/.env"
APP_DIR="${APP_DIR:-/opt/open-swe}" # where the app + its CWD .env live
APP_ENV_LINK="${APP_DIR}/.env" # symlink -> ENV_FILE
# The unprivileged service user that runs the app and OWNS the .env (T5 SC-01).
# fetch-config runs as root (ExecStartPre=+) only to chown the secret to it.
SERVICE_USER="${SERVICE_USER:-openswe}"
SERVICE_GROUP="${SERVICE_GROUP:-${SERVICE_USER}}"
# Sea Haven owner guard inputs (applied after the store is read, below). The owner
# normally comes from SSM /open-swe-<env>/DEFAULT_REPO_OWNER; OPENSWE_REPO_OWNER is an
# explicit operator override that wins over the store. FORBIDDEN = the upstream org
# the fork must never target; SAFE = the fallback when the resolved owner is blank or
# forbidden. (Upper/lower + whitespace are normalized before the guard check.)
OPENSWE_REPO_OWNER="${OPENSWE_REPO_OWNER:-}"
FORBIDDEN_REPO_OWNER="langchain-ai"
# Fallback org when the resolved owner is blank/forbidden — PER-ENV (mirrors the
# iacManagedSsm owner) so a dev box can NEVER fall back into the real Sea Haven org;
# it stays isolated in its own dev org. Defends the blank/upstream cases in-env.
case "$ENV" in
dev) SAFE_REPO_OWNER="seahaven-open-swe-dev" ;;
*) SAFE_REPO_OWNER="Sea-Haven-Industries" ;;
esac
for bin in aws jq; do
command -v "$bin" >/dev/null 2>&1 || { echo "fetch-config: '$bin' not found on PATH" >&2; exit 3; }
done
log() { echo "fetch-config[$ENV]: $*"; } # NAMES/counts only — never values
b64d() { base64 --decode; } # GNU coreutils on the EC2 host
# Accept a store key into VARS iff it is a valid env-var identifier and not a
# duplicate. Rejects non-identifier names (T5 SH-INJ-002 / set -e DoS hardening)
# and flat-namespace collisions (T5 SSM-05). $3 = source label for logs.
accept_var() {
local key="$1" value="$2" src="$3"
if ! [[ "$key" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]]; then
log "WARNING: skipping ${src} key with non-identifier name (rejected)"
return 0
fi
if [ -n "${VARS[$key]+set}" ]; then
echo "fetch-config[$ENV]: FAIL-FAST — duplicate key '${key}' from ${src} (flat-namespace collision)" >&2
exit 1
fi
VARS["$key"]="$value"
}
# --- tmpfs ------------------------------------------------------------------
mkdir -p "$ENV_DIR"
# Owned by the service user so the unprivileged app can traverse it (T5 SC-01).
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_DIR" 2>/dev/null || true
chmod 700 "$ENV_DIR"
if [ "${RUN_DEDICATED_TMPFS:-0}" = "1" ] && ! mountpoint -q "$ENV_DIR"; then
mount -t tmpfs -o nosuid,nodev,noexec,mode=0700,size=4m tmpfs "$ENV_DIR"
log "mounted dedicated tmpfs at $ENV_DIR"
fi
# --- Collect values into an associative array --------------------------------
declare -A VARS=()
# 1) SSM Parameter Store (non-sensitive config). NOT --recursive: the contract is
# a FLAT namespace /open-swe-<env>/<VAR>, so a non-recursive list returns exactly
# those keys and cannot collapse two nested paths onto one name (T5 SSM-05). aws
# CLI v2 auto-paginates NextToken.
log "reading SSM params under ${SSM_PREFIX} ..."
ssm_json="$(
aws ssm get-parameters-by-path \
--path "$SSM_PREFIX" \
--with-decryption \
--region "$REGION" \
--no-cli-pager \
--output json
)"
# Records are base64-encoded (name<TAB>value) so values with spaces/newlines/tabs
# survive the line-based read intact.
ssm_count=0
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}" # strip /open-swe-<env>/ prefix
[ -n "$key" ] || continue
accept_var "$key" "$value" "SSM"
ssm_count=$((ssm_count + 1))
done < <(jq -r '.Parameters[] | (.Name|@base64) + "\t" + (.Value|@base64)' <<<"$ssm_json")
log "loaded ${ssm_count} config param(s) from SSM"
# 2) Secrets Manager (sensitive values). We request secrets by EXPLICIT id
# (`batch-get-secret-value --secret-id-list ...`) rather than a name-prefix
# `--filters` collection scan (OSWE-IAC-SECRETS-LIST-01). Two wins:
# (a) Least privilege — an explicit id list lets the instance role scope
# BatchGetSecretValue to the per-secret ARN prefix and DROP the account-wide
# `secretsmanager:ListSecrets` grant that a filtered scan unavoidably forces
# (ListSecrets has no resource-level scoping). A filtered batch call also only
# authorizes against `*`; an id-list call authorizes per-secret ARN.
# (b) Deterministic set — the value set is the fixed SECRET_VARS shells created by
# the CDK ConfigStore, so we no longer depend on a list scan returning every
# page. Value-less shells and ids with no current value come back in the
# response `.Errors[]` (ResourceNotFound), never in `.SecretValues[]`, so a
# genuinely-missing REQUIRED secret is still caught by the FAIL-FAST check
# below — an absent optional secret is simply skipped.
# `--secret-id-list` is capped at 20 ids per call, so we chunk it. `--no-cli-pager`
# disables only the OUTPUT pager. Each record is base64(name<TAB>value).
#
# SOURCE OF TRUTH for this list: infra/lib/constructs/config-store.ts `SECRET_VARS`.
# Keep the two in lockstep — a new secret shell created there must be added here or it
# will never be fetched into the .env.
SECRET_VARS=(
ANTHROPIC_API_KEY CORRIDOR_API_TOKEN CORRIDOR_MCP_TOKEN CORRIDOR_TOKEN
DASHBOARD_JWT_SECRET DAYTONA_API_KEY EXA_API_KEY FIREWORKS_API_KEY
GITHUB_APP_CLIENT_SECRET GITHUB_APP_PRIVATE_KEY GITHUB_PAT GITHUB_WEBHOOK_SECRET
JUDGE_ANTHROPIC_API_KEY LANGSMITH_API_KEY
LANGSMITH_API_KEY_PROD LANGCHAIN_API_KEY LINEAR_API_KEY LINEAR_WEBHOOK_SECRET
RUNLOOP_API_KEY SLACK_BOT_TOKEN SLACK_CLIENT_SECRET
SLACK_SIGNING_SECRET TOKEN_ENCRYPTION_KEY USER_ID_API_KEY_MAP X_SERVICE_AUTH_JWT_SECRET
)
batch_get_secrets_tsv() {
local -a ids=()
local v
for v in "${SECRET_VARS[@]}"; do ids+=("${SECRET_PREFIX}${v}"); done
local i page
local -a chunk
for ((i = 0; i < ${#ids[@]}; i += 20)); do
chunk=("${ids[@]:i:20}")
# Capture the response into a variable FIRST so a non-zero `aws` exit (throttle,
# AccessDenied, KMS DecryptionFailure) aborts under set -e instead of being
# silently swallowed — then we'd FAIL-FAST below as "missing secret" with a wrong
# root cause. (Value-less / absent shells come back in .Errors[], not .SecretValues[].)
page="$(aws secretsmanager batch-get-secret-value \
--secret-id-list "${chunk[@]}" \
--region "$REGION" --no-cli-pager --output json)"
printf '%s' "$page" \
| jq -r '.SecretValues[] | select(.SecretString != null) | (.Name|@base64) + "\t" + (.SecretString|@base64)'
done
}
log "reading secrets under ${SECRET_PREFIX} ..."
secret_count=0
# Capture into a variable (NOT `done < <(...)` process substitution) so a non-zero
# exit from batch_get_secrets_tsv propagates under set -e — process substitution hides
# the producer's exit status from the parent shell, which would let a failed AWS call
# fall through to a misleading "missing required var" FAIL-FAST. Mirrors the SSM read.
secrets_tsv="$(batch_get_secrets_tsv)"
while IFS=$'\t' read -r nb vb; do
[ -n "$nb" ] || continue
name="$(printf '%s' "$nb" | b64d)"
case "$name" in
"${SECRET_PREFIX}"*) ;; # defensive: exact-prefix only
*) continue ;;
esac
value="$(printf '%s' "$vb" | b64d; printf 'x')"; value="${value%x}"
key="${name##*/}"
[ -n "$key" ] || continue
accept_var "$key" "$value" "Secrets"
secret_count=$((secret_count + 1))
done <<<"$secrets_tsv"
log "loaded ${secret_count} secret(s) from Secrets Manager"
# --- Sea Haven DEFAULT_REPO_OWNER guard --------------------------------------
# OSWE-OWNER-04 (revised for multi-org): HONOR the configured owner — the
# OPENSWE_REPO_OWNER env override if set, else the store value — so per-env orgs
# work (dev = seahaven-open-swe-dev, prod = Sea-Haven-Industries). But GUARD the two
# values that must NEVER reach the agent: blank, and the upstream 'langchain-ai' org
# (the fork's origin). Either falls back to the Sea Haven org so a stale/blank/mis-set
# value can never point the agent upstream. Comparison is case- and whitespace-
# insensitive. The POSITIVE org allowlist is enforced by the app (ALLOWED_GITHUB_ORGS).
resolved_owner="${OPENSWE_REPO_OWNER:-${VARS[DEFAULT_REPO_OWNER]:-}}"
# Normalize for the guard CHECK ONLY (the original value is what gets stored when
# allowed): lowercase, strip whitespace, take the FIRST path segment so a value like
# 'langchain-ai/open-swe' still trips the guard, and drop dots (GitHub owners contain
# none) so 'langchain-ai.' can't slip past. Homoglyph/unicode variants are out of scope
# here — the owner comes from admin-written SSM/IaC, not attacker-controlled input.
norm_owner="$(printf '%s' "$resolved_owner" | tr '[:upper:]' '[:lower:]' | tr -d '[:space:]')"
norm_owner="${norm_owner%%/*}"
norm_owner="${norm_owner//./}"
case "$norm_owner" in
"" | "$FORBIDDEN_REPO_OWNER")
log "WARNING: DEFAULT_REPO_OWNER ('${resolved_owner:-<blank>}') is blank or the upstream org -> forcing '${SAFE_REPO_OWNER}'"
resolved_owner="$SAFE_REPO_OWNER"
;;
esac
VARS[DEFAULT_REPO_OWNER]="$resolved_owner"
# --- FAIL-FAST: required vars -------------------------------------------------
# Hard-required regardless of mode:
required=(
DASHBOARD_JWT_SECRET # RuntimeError on startup if missing (oauth.py)
TOKEN_ENCRYPTION_KEY # Fernet key(s); decrypts per-user GitHub tokens
)
# NOTE: the GitHub App is NOT created/duplicated for dev — only prod owns the
# (single, shared) GitHub App + Slack app. So the GitHub App quintet + Slack +
# webhook-signing secrets are required for PROD only (see the prod block below).
# Dev boots without them: it has no GitHub-App/Slack/webhook integration — it is a
# deployment-validation env (boot/health/boundary), not a live-triggered agent.
# Active model-provider key(s): model selection is store-driven (team_settings),
# so fetch-config cannot infer it from .env. Bedrock (Claude builder + reviewer)
# authenticates via the host IAM role — no API key; Fireworks (fallback / subagents /
# any non-Claude model) needs its key. Override with a comma list if the active models change.
IFS=',' read -r -a provider_keys <<<"${REQUIRED_PROVIDER_KEYS:-FIREWORKS_API_KEY}"
for k in "${provider_keys[@]}"; do
k="${k//[[:space:]]/}"
[ -n "$k" ] && required+=("$k")
done
# Sandbox provider key(s) — depends on SANDBOX_TYPE (default langsmith).
sandbox_type="${VARS[SANDBOX_TYPE]:-langsmith}"
case "$sandbox_type" in
langsmith) required+=(LANGSMITH_API_KEY_PROD DEFAULT_SANDBOX_SNAPSHOT_ID) ;;
daytona) required+=(DAYTONA_API_KEY) ;;
runloop) required+=(RUNLOOP_API_KEY) ;;
modal | local) ;; # no key required
*) log "WARNING: unknown SANDBOX_TYPE='${sandbox_type}' — not enforcing a sandbox key" ;;
esac
# Prod-only: the GitHub App (installation-token minting + dashboard OAuth) and the
# webhook-signing secrets. Dev has no GitHub/Slack app, so none of these are
# required there; prod owns the single shared app and must have all of them.
if [ "$ENV" = "prod" ]; then
required+=(
GITHUB_APP_ID # GitHub App trio (installation-token minting) ...
GITHUB_APP_PRIVATE_KEY # ... multiline PEM ...
GITHUB_APP_INSTALLATION_ID # ... used by utils/github_app.py
GITHUB_APP_CLIENT_ID # dashboard OAuth login
GITHUB_APP_CLIENT_SECRET # dashboard OAuth login
GITHUB_WEBHOOK_SECRET # webhook signature verification
SLACK_SIGNING_SECRET # Slack webhook signature verification
)
if [ -n "${VARS[LINEAR_API_KEY]:-}" ] && [ "${OPENSWE_REQUIRE_LINEAR:-1}" = "1" ]; then
required+=(LINEAR_WEBHOOK_SECRET)
fi
fi
missing=()
for k in "${required[@]}"; do
[ -n "${VARS[$k]:-}" ] || missing+=("$k")
done
# de-dup the names for a clean report
if [ "${#missing[@]}" -gt 0 ]; then
mapfile -t missing < <(printf '%s\n' "${missing[@]}" | sort -u)
echo "fetch-config[$ENV]: FAIL-FAST — ${#missing[@]} required var(s) missing/empty:" >&2
printf ' - %s\n' "${missing[@]}" >&2
echo "fetch-config[$ENV]: refusing to write a partial .env; service will not start." >&2
exit 1
fi
# --- Write the .env atomically (root-only on tmpfs) --------------------------
# python-dotenv reads double-quoted values (incl. multiline PEMs). Its decoder
# unescapes ONLY backslash and double-quote (\\ -> \, \" -> "); it does NOT honor
# \$ or \` escapes, so escaping those would leave a spurious backslash. Escape
# exactly backslash then double-quote — real newlines stay literal (multiline OK).
# (Caveat: python-dotenv interpolates a literal `${VAR}` substring; the secret
# domain here — base64/hex/PEM keys — never contains one, so no extra guard.)
emit_var() {
local name="$1" value="$2" esc
esc="${value//\\/\\\\}"
esc="${esc//\"/\\\"}"
printf '%s="%s"\n' "$name" "$esc"
}
tmp="$(mktemp "${ENV_DIR}/.env.XXXXXX")"
chmod 600 "$tmp"
{
printf '# Generated by fetch-config.sh for env=%s at %s — DO NOT EDIT.\n' \
"$ENV" "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
printf '# Source: SSM /open-swe-%s/* + Secrets Manager open-swe-%s/*\n\n' "$ENV" "$ENV"
for k in $(printf '%s\n' "${!VARS[@]}" | sort); do
emit_var "$k" "${VARS[$k]}"
done
} >"$tmp"
mv -f "$tmp" "$ENV_FILE"
# Owned by the unprivileged service user (T5 SC-01) so the app reads it without
# running as root. fetch-config itself runs as root (ExecStartPre=+) to chown.
chown "${SERVICE_USER}:${SERVICE_GROUP}" "$ENV_FILE"
chmod 600 "$ENV_FILE"
# Point the app's CWD .env at the tmpfs file (idempotent).
if [ "$APP_ENV_LINK" != "$ENV_FILE" ]; then
if [ -L "$APP_ENV_LINK" ] || [ ! -e "$APP_ENV_LINK" ]; then
ln -sfn "$ENV_FILE" "$APP_ENV_LINK"
elif [ "$(readlink -f "$APP_ENV_LINK" 2>/dev/null || true)" != "$(readlink -f "$ENV_FILE")" ]; then
log "WARNING: ${APP_ENV_LINK} exists and is not a symlink to ${ENV_FILE} — leaving it untouched"
fi
fi
total=$((ssm_count + secret_count))
log "wrote ${ENV_FILE} (${total} vars, sandbox=${sandbox_type}) — ${SERVICE_USER}:${SERVICE_GROUP} 0600"