mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 09:13:14 +00:00
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
184 lines
8.7 KiB
Bash
Executable file
184 lines
8.7 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Seed the LangGraph store after a (re)start.
|
|
#
|
|
# The stock `langgraph dev` server uses an IN-MEMORY store, so anything written
|
|
# to it (team model settings, user mappings) is lost on every restart. This
|
|
# script idempotently re-PUTs that state and is wired as a systemd
|
|
# ExecStartPost on the open-swe.service unit so it runs after each start.
|
|
# IT MUST RE-RUN ON EVERY RESTART — the in-memory store starts empty each boot.
|
|
#
|
|
# Replace this with Postgres-backed durability (Aegra / `langgraph up`) to make
|
|
# the store survive restarts and drop this script.
|
|
#
|
|
# AWS-env-aware: pass the env as $1 (dev|prod). Seed values (default repo, model
|
|
# ids, user mappings) come from the fetch-config-materialized .env.
|
|
#
|
|
# SECURITY (T5 /sh-security-review):
|
|
# - SH-INJ-001: this script NEVER `source`s the .env. python-dotenv and bash
|
|
# have incompatible escaping, and a config value like `$(cmd)` would execute
|
|
# when sourced. We extract the few single-line seed keys with a non-eval
|
|
# reader (read_env) instead.
|
|
# - SH-INJ-003: the store PUT bodies are built with `jq --arg`, so values are
|
|
# always JSON-encoded (no string interpolation into a JSON heredoc).
|
|
# - SH-INJ-004: BASE is pinned to loopback — never derived from store/SSM
|
|
# config (LANGGRAPH_URL) — so a tampered value can't redirect the PUTs.
|
|
# - SC-02: not sourcing the .env means secrets are never exported into this
|
|
# script's (or curl's) environment.
|
|
#
|
|
# Seed values read from the materialized .env (set in SSM /open-swe-<env>/*):
|
|
# DEFAULT_REPO_OWNER / DEFAULT_REPO_NAME -> team_settings default_repo
|
|
# LLM_MODEL_ID -> default builder model (fallback)
|
|
# SEED_AGENT_MODEL / SEED_AGENT_EFFORT -> builder model + effort (optional)
|
|
# SEED_REVIEWER_MODEL / SEED_REVIEWER_EFFORT -> reviewer model + effort (optional)
|
|
# SEED_USER_MAPPINGS -> "login:email,login:email" (optional)
|
|
# CONFIGURED_ADMINS -> "login,email" fallback for the mapping
|
|
# Legacy OPENSWE_* process-env overrides are still honored (highest precedence).
|
|
set -euo pipefail
|
|
|
|
ENV="${1:-${OPENSWE_ENV:-}}"
|
|
case "$ENV" in
|
|
dev | prod | "") ;; # empty allowed: pure-env / on-prem backward-compat mode
|
|
*)
|
|
echo "seed_store: ENV must be 'dev' or 'prod' (got '$ENV')" >&2
|
|
exit 2
|
|
;;
|
|
esac
|
|
|
|
ENV_DIR="${ENV_DIR:-/run/open-swe}"
|
|
ENV_FILE="${ENV_FILE:-${ENV_DIR}/.env}"
|
|
|
|
# read_env KEY -> prints the value of a SINGLE-LINE `KEY="..."` entry from the
|
|
# materialized .env WITHOUT shell evaluation (SH-INJ-001 fix). Seed keys are
|
|
# simple single-line values; multiline secrets (e.g. the PEM) are never read
|
|
# here. Returns empty if the key is absent/unreadable.
|
|
read_env() {
|
|
local key="$1" line
|
|
[ -r "$ENV_FILE" ] || return 0
|
|
line="$(grep -m1 -- "^${key}=" "$ENV_FILE" 2>/dev/null || true)"
|
|
[ -n "$line" ] || return 0
|
|
line="${line#*=}"
|
|
# strip one layer of surrounding double quotes (python-dotenv double-quoted form)
|
|
if [ "${line#\"}" != "$line" ]; then line="${line%\"}"; line="${line#\"}"; fi
|
|
# reverse python-dotenv double-quote escaping (only \" and \\ are escaped)
|
|
line="${line//\\\"/\"}"; line="${line//\\\\/\\}"
|
|
printf '%s' "$line"
|
|
}
|
|
|
|
# Process-env override (legacy/on-prem) -> .env value -> default.
|
|
pick() { # pick DEFAULT OVERRIDE_VALUE FILE_KEY...
|
|
local def="$1" override="$2"; shift 2
|
|
if [ -n "$override" ]; then printf '%s' "$override"; return; fi
|
|
local k v
|
|
for k in "$@"; do v="$(read_env "$k")"; [ -n "$v" ] && { printf '%s' "$v"; return; }; done
|
|
printf '%s' "$def"
|
|
}
|
|
|
|
# BASE is loopback-pinned (SH-INJ-004): this on-box seeder only talks to the
|
|
# local server; OPENSWE_PORT may override the port but never the host.
|
|
BASE="http://127.0.0.1:${OPENSWE_PORT:-2024}"
|
|
|
|
AGENT_MODEL="$(pick 'bedrock_converse:us.anthropic.claude-opus-4-8' "${OPENSWE_AGENT_MODEL:-}" SEED_AGENT_MODEL LLM_MODEL_ID)"
|
|
AGENT_EFFORT="$(pick 'high' "${OPENSWE_AGENT_EFFORT:-}" SEED_AGENT_EFFORT)"
|
|
REVIEWER_MODEL="$(pick 'bedrock_converse:us.anthropic.claude-opus-4-8' "${OPENSWE_REVIEWER_MODEL:-}" SEED_REVIEWER_MODEL)"
|
|
REVIEWER_EFFORT="$(pick 'high' "${OPENSWE_REVIEWER_EFFORT:-}" SEED_REVIEWER_EFFORT)"
|
|
|
|
# default_repo = owner/name from AWS config (DEFAULT_REPO_OWNER is hard-pinned
|
|
# away from upstream by fetch-config.sh).
|
|
REPO_OWNER="$(pick '' '' DEFAULT_REPO_OWNER)"
|
|
REPO_NAME="$(pick '' '' DEFAULT_REPO_NAME)"
|
|
if [ -n "${OPENSWE_DEFAULT_REPO:-}" ]; then
|
|
DEFAULT_REPO="$OPENSWE_DEFAULT_REPO"
|
|
elif [ -n "$REPO_OWNER" ] && [ -n "$REPO_NAME" ]; then
|
|
DEFAULT_REPO="${REPO_OWNER}/${REPO_NAME}"
|
|
else
|
|
echo "seed_store: set OPENSWE_DEFAULT_REPO=owner/repo (or DEFAULT_REPO_OWNER + DEFAULT_REPO_NAME in .env)" >&2
|
|
exit 1
|
|
fi
|
|
|
|
# user_mappings: explicit SEED_USER_MAPPINGS ("login:email,..."), then legacy
|
|
# OPENSWE_OWNER_LOGIN/EMAIL, then parse CONFIGURED_ADMINS ("login,email").
|
|
SEED_MAP="$(pick '' "${SEED_USER_MAPPINGS:-}" SEED_USER_MAPPINGS)"
|
|
ADMINS="$(pick '' "${CONFIGURED_ADMINS:-}" CONFIGURED_ADMINS)"
|
|
declare -a MAPPINGS=()
|
|
if [ -n "$SEED_MAP" ]; then
|
|
IFS=',' read -r -a _pairs <<<"$SEED_MAP"
|
|
for p in "${_pairs[@]}"; do
|
|
p="${p//[[:space:]]/}"
|
|
[ -n "$p" ] && MAPPINGS+=("$p")
|
|
done
|
|
elif [ -n "${OPENSWE_OWNER_LOGIN:-}" ] && [ -n "${OPENSWE_OWNER_EMAIL:-}" ]; then
|
|
MAPPINGS+=("${OPENSWE_OWNER_LOGIN}:${OPENSWE_OWNER_EMAIL}")
|
|
elif [ -n "$ADMINS" ]; then
|
|
_login="" _email=""
|
|
IFS=',' read -r -a _toks <<<"$ADMINS"
|
|
for t in "${_toks[@]}"; do
|
|
t="${t//[[:space:]]/}"
|
|
[ -z "$t" ] && continue
|
|
case "$t" in
|
|
*@*) [ -z "$_email" ] && _email="$t" ;;
|
|
*) [ -z "$_login" ] && _login="$t" ;;
|
|
esac
|
|
done
|
|
[ -n "$_login" ] && [ -n "$_email" ] && MAPPINGS+=("${_login}:${_email}")
|
|
fi
|
|
|
|
# No user mapping is NON-FATAL (OSWE-SEED-03 precedent: never fail the unit into a
|
|
# restart loop over a seeding gap — same as the server-not-ready path below). The
|
|
# server itself is healthy; an unseeded user_mappings table only means the @openswe
|
|
# trigger won't resolve a commenter, which a deployment-validation env (e.g. dev)
|
|
# does not need. team_settings is still seeded. Set SEED_USER_MAPPINGS (or
|
|
# CONFIGURED_ADMINS / OPENSWE_OWNER_LOGIN+EMAIL) to seed the mapping when wanted.
|
|
if [ "${#MAPPINGS[@]}" -eq 0 ]; then
|
|
echo "seed_store: no user mapping resolved — skipping user_mappings seed (set SEED_USER_MAPPINGS or OPENSWE_OWNER_LOGIN/EMAIL to enable)" >&2
|
|
fi
|
|
|
|
NOW="$(date -u +%Y-%m-%dT%H:%M:%S+00:00)"
|
|
|
|
# Wait for the server to accept requests (up to ~60s). Authoritative (OSWE-SEED-03):
|
|
# if it never comes up, log and exit 0 — do NOT fail the unit into a restart loop.
|
|
READY=0
|
|
for _ in $(seq 1 30); do
|
|
if [ "$(curl -s -o /dev/null -w '%{http_code}' "$BASE/ok" || true)" = "200" ]; then READY=1; break; fi
|
|
sleep 2
|
|
done
|
|
if [ "$READY" -ne 1 ]; then
|
|
echo "seed_store: server not ready at $BASE after ~60s; skipping seed (will reseed on next restart)" >&2
|
|
exit 0
|
|
fi
|
|
|
|
# 1) team_settings/default — JSON built with jq --arg (SH-INJ-003 fix).
|
|
team_body="$(jq -n \
|
|
--arg am "$AGENT_MODEL" --arg ae "$AGENT_EFFORT" \
|
|
--arg rm "$REVIEWER_MODEL" --arg re "$REVIEWER_EFFORT" \
|
|
--arg repo "$DEFAULT_REPO" --arg now "$NOW" \
|
|
'{namespace:["team_settings"],key:"default",value:{
|
|
review_draft_prs:false, pr_summaries:true, review_trace_links:true,
|
|
org_guidelines:null,
|
|
default_agent_model:$am, default_agent_reasoning_effort:$ae,
|
|
default_agent_subagent_model:$am, default_agent_subagent_reasoning_effort:$ae,
|
|
default_repo:$repo,
|
|
default_reviewer_model:$rm, default_reviewer_reasoning_effort:$re,
|
|
default_reviewer_subagent_model:$rm, default_reviewer_subagent_reasoning_effort:$re,
|
|
default_grouping_model:null, default_grouping_reasoning_effort:null,
|
|
default_chat_model:null, default_chat_reasoning_effort:null,
|
|
updated_at:$now}}')"
|
|
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$team_body" >/dev/null \
|
|
|| echo "seed_store: WARN team_settings PUT failed (will reseed next restart)" >&2
|
|
|
|
# 2) user_mappings/<login> — required, or the @openswe trigger ignores the commenter.
|
|
for pair in "${MAPPINGS[@]}"; do
|
|
login="${pair%%:*}"
|
|
email="${pair#*:}"
|
|
if [ -z "$login" ] || [ -z "$email" ] || [ "$login" = "$pair" ]; then
|
|
echo "seed_store: skipping malformed mapping '$pair' (want login:email)" >&2
|
|
continue
|
|
fi
|
|
map_body="$(jq -n --arg login "$login" --arg email "$email" --arg now "$NOW" \
|
|
'{namespace:["user_mappings"],key:$login,value:{
|
|
github_login:$login, work_email:$email, slack_user_id:null,
|
|
source:"slack_oauth", status:"active", created_at:$now, updated_at:$now}}')"
|
|
curl -fsS -X PUT "$BASE/store/items" -H "Content-Type: application/json" -d "$map_body" >/dev/null \
|
|
|| echo "seed_store: WARN user_mapping PUT failed for $login" >&2
|
|
done
|
|
|
|
echo "seed_store: done at $NOW (env=${ENV:-none}, repo=$DEFAULT_REPO, mappings=${#MAPPINGS[@]})"
|