* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
6.5 KiB
Sea Haven — Open SWE secret & config rotation
How secrets and config reach the running app, and how to rotate either one.
How values flow at boot
AWS Secrets Manager open-swe-<env>/* ─┐
AWS SSM Param Store /open-swe-<env>/* ─┤── fetch-config.sh ──▶ tmpfs /run/open-swe/.env (root:root 0600)
│ (systemd ExecStartPre=+, EC2 role) │
▼ ▼
FAIL-FAST if a app symlink <APP_DIR>/.env
required var is empty python-dotenv reads at import
The app reads .env once, at import. There is no hot-reload of secrets.
Therefore the rotation contract is always the same two steps:
Rotation = (1) update the value in Secrets Manager / SSM, then (2) restart the service so
fetch-config.shre-materializes the.env.
# after updating a secret/param in AWS:
sudo systemctl restart open-swe.service
# ExecStartPre=+ -> fetch-config.sh re-pulls + rewrites the tmpfs .env (fail-fast)
# ExecStartPost -> seed_store.sh re-seeds the in-memory store (team_settings + user_mappings)
There is no zero-downtime path for most secrets on the stock in-memory
runtime — a restart is required and it also wipes the in-memory store (re-seeded
by seed_store.sh automatically). The one secret built for zero-downtime overlap
is TOKEN_ENCRYPTION_KEY (see below), but even it needs the restart to load the
new key list.
Rotating a secret (Secrets Manager)
ENV=prod # or dev
NAME=DASHBOARD_JWT_SECRET
aws secretsmanager put-secret-value \
--secret-id "open-swe-${ENV}/${NAME}" \
--secret-string 'NEW_VALUE' \
--region us-east-1
sudo systemctl restart open-swe.service # on the box
(update-secret/put-secret-value both create a new version; the boot hook
always reads AWSCURRENT.)
Rotating a config param (SSM)
aws ssm put-parameter --overwrite \
--name "/open-swe-${ENV}/DASHBOARD_BASE_URL" \
--type String --value 'https://openswe.seahaven.com' \
--region us-east-1
sudo systemctl restart open-swe.service
Per-secret rotation notes
| Secret | Rotation notes |
|---|---|
| TOKEN_ENCRYPTION_KEY | Fernet key(s). Supports a comma/newline-separated list (agent/encryption.py) for zero-downtime key rotation: prepend the NEW key, keep the OLD key(s) in the list. New data is encrypted with the first key; old data still decrypts with the trailing keys. After all encrypted-at-rest tokens (per-user GitHub OAuth tokens in thread metadata) have been re-encrypted/expired, drop the old key. Store the list as one secret value; fetch-config.sh writes it verbatim. Never rotate to a single new key in one step or every existing encrypted token becomes undecryptable. |
| GITHUB_APP_PRIVATE_KEY | Multiline PEM. Generate a new private key in the GitHub App settings (you may have two active keys during overlap), put the new PEM into the secret, restart, verify install-token minting + a webhook delivery, then delete the old key in GitHub. fetch-config.sh writes the PEM as a double-quoted multiline value (python-dotenv-safe); paste the full -----BEGIN…-----END----- block including newlines. |
| GITHUB_WEBHOOK_SECRET | Webhook HMAC. GitHub allows only one webhook secret per App, so this is a brief-break rotation: update the secret in AWS and the GitHub App webhook config, restart. Deliveries signed with the old secret during the gap will 401 (GitHub auto-redelivers). Required in prod (fail-fast). |
| SLACK_SIGNING_SECRET | Slack request-signature secret. Rotate in the Slack app config and AWS together, restart. Required in prod (fail-fast). A stale value silently 401s url_verification/events until restart (known gotcha). |
| LINEAR_WEBHOOK_SECRET | Linear webhook signature. Required in prod only when the Linear integration is wired (LINEAR_API_KEY present). Rotate in Linear + AWS together, restart. |
| SLACK_CLIENT_SECRET / GITHUB_APP_CLIENT_SECRET | OAuth client secrets (dashboard login / Slack OAuth). Rotate in the provider console + AWS, restart. Existing dashboard sessions are JWT-signed by DASHBOARD_JWT_SECRET, not these, so they survive. |
| DASHBOARD_JWT_SECRET | Signs dashboard session cookies. Rotating invalidates all active sessions (users re-login). Hard-required (RuntimeError if empty). No overlap list — single value. |
Model provider keys (FIREWORKS_API_KEY; legacy ANTHROPIC_API_KEY / OPENAI_API_KEY / GOOGLE_API_KEY / GROQ_API_KEY) |
Standard API-key rotation: issue new key, update AWS, restart, revoke old. Bedrock (the seeded Claude builder + reviewer) authenticates via the host IAM role — no API key; the only fail-fast-required provider key is now FIREWORKS_API_KEY (fallback / subagents / any non-Claude model). Keep REQUIRED_PROVIDER_KEYS in sync if you change the seeded models. |
| LANGSMITH_API_KEY_PROD | LangSmith key powering the langsmith sandbox (the only provider with working in-sandbox git/gh auth). Required when SANDBOX_TYPE=langsmith (fail-fast). Rotate in LangSmith + AWS, restart; existing sandboxes keep their already-injected proxy token until recycled. |
| SLACK_BOT_TOKEN / LINEAR_API_KEY / GITHUB_PAT / EXA_API_KEY / DAYTONA_API_KEY / RUNLOOP_API_KEY / CORRIDOR_ / USER_ID_API_KEY_MAP / X_SERVICE_AUTH_JWT_SECRET* | Plain API-token rotation: update AWS, restart, revoke old at the provider. None support an overlap list. |
Fail-fast safety
fetch-config.sh refuses to write the .env (exit 1) if any required var is
empty after a rotation — so a botched rotation (e.g. an empty put-secret-value)
stops the service at ExecStartPre instead of starting it with a partial .env.
The missing variable names are printed to the journal (values never are):
sudo journalctl -u open-swe.service -b | grep fetch-config
Required set enforced: DASHBOARD_JWT_SECRET, TOKEN_ENCRYPTION_KEY,
GITHUB_APP_ID, GITHUB_APP_PRIVATE_KEY, GITHUB_APP_INSTALLATION_ID,
GITHUB_APP_CLIENT_ID, GITHUB_APP_CLIENT_SECRET, the active provider keys
(REQUIRED_PROVIDER_KEYS, default ANTHROPIC_API_KEY,OPENAI_API_KEY), the
sandbox key for SANDBOX_TYPE (langsmith ⇒ LANGSMITH_API_KEY_PROD +
DEFAULT_SANDBOX_SNAPSHOT_ID), and — in prod — GITHUB_WEBHOOK_SECRET,
SLACK_SIGNING_SECRET (plus LINEAR_WEBHOOK_SECRET when Linear is wired).