mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 09:13:14 +00:00
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
234 lines
9.2 KiB
Python
234 lines
9.2 KiB
Python
import os
|
|
from typing import Literal, TypedDict, Unpack
|
|
|
|
from langchain.chat_models import init_chat_model
|
|
|
|
from ..dashboard.options import DEFAULT_MODEL_ID
|
|
|
|
OPENAI_RESPONSES_WS_BASE_URL = "wss://api.openai.com/v1"
|
|
|
|
# Anthropic SDK default is 2; a 529 burst can outlive that. Bump to give the
|
|
# primary provider a fair chance before the fallback middleware kicks in.
|
|
DEFAULT_MAX_RETRIES = 6
|
|
|
|
OpenAIReasoningEffort = Literal["none", "low", "medium", "high", "xhigh"]
|
|
# OpenAI's Responses API only returns human-readable reasoning text when a
|
|
# summary is requested; without it, reasoning happens silently (billed in
|
|
# output tokens) and the reasoning content block arrives empty.
|
|
OpenAIReasoningSummary = Literal["auto", "concise", "detailed"]
|
|
AnthropicThinkingType = Literal["adaptive"]
|
|
AnthropicThinkingDisplay = Literal["summarized", "omitted"]
|
|
AnthropicEffort = Literal["low", "medium", "high", "xhigh", "max"]
|
|
GoogleThinkingLevel = Literal["minimal", "low", "medium", "high"]
|
|
FireworksReasoningEffort = Literal["none", "low", "medium", "high", "xhigh", "max"]
|
|
|
|
|
|
class OpenAIReasoning(TypedDict, total=False):
|
|
effort: OpenAIReasoningEffort
|
|
summary: OpenAIReasoningSummary
|
|
|
|
|
|
DEFAULT_LLM_REASONING: "OpenAIReasoning" = {"effort": "medium", "summary": "auto"}
|
|
|
|
|
|
class AnthropicThinking(TypedDict, total=False):
|
|
type: AnthropicThinkingType
|
|
display: AnthropicThinkingDisplay
|
|
|
|
|
|
class ModelKwargs(TypedDict, total=False):
|
|
max_tokens: int | None
|
|
reasoning: OpenAIReasoning | None
|
|
thinking: AnthropicThinking | None
|
|
effort: AnthropicEffort | None
|
|
thinking_level: GoogleThinkingLevel | None
|
|
temperature: float | None
|
|
max_retries: int | None
|
|
model_kwargs: dict[str, object] | None
|
|
additional_model_request_fields: dict[str, object] | None
|
|
region_name: str | None
|
|
|
|
|
|
_ANTHROPIC_EFFORTS: set[AnthropicEffort] = {"low", "medium", "high", "xhigh", "max"}
|
|
|
|
|
|
def make_model(model_id: str, **kwargs: Unpack[ModelKwargs]):
|
|
model_kwargs: dict[str, object] = kwargs.copy()
|
|
model_kwargs.setdefault("max_retries", DEFAULT_MAX_RETRIES)
|
|
|
|
if model_id.startswith("openai:"):
|
|
model_kwargs["base_url"] = OPENAI_RESPONSES_WS_BASE_URL
|
|
model_kwargs["use_responses_api"] = True
|
|
elif model_id.startswith("bedrock_converse:"):
|
|
# Resolve region with the same precedence validate_local_dev_llm_config accepts
|
|
# (AWS_REGION or AWS_DEFAULT_REGION), so the validated value is the one actually used.
|
|
model_kwargs.setdefault(
|
|
"region_name",
|
|
os.environ.get("AWS_REGION") or os.environ.get("AWS_DEFAULT_REGION") or "us-east-1",
|
|
)
|
|
|
|
return init_chat_model(model=model_id, **model_kwargs)
|
|
|
|
|
|
def fallback_model_id_for(primary_model_id: str) -> str | None:
|
|
"""Return the cross-provider fallback model id for a given primary, if any.
|
|
|
|
Bedrock (Claude) primaries fall back to Fireworks and vice versa. Returns
|
|
``None`` when the provider has no configured cross-provider fallback (e.g.
|
|
local or self-hosted providers we don't want to silently route off-host).
|
|
"""
|
|
if primary_model_id.startswith("bedrock_converse:"):
|
|
return "fireworks:accounts/fireworks/models/deepseek-v4-pro"
|
|
if primary_model_id.startswith("fireworks:"):
|
|
return "bedrock_converse:us.anthropic.claude-opus-4-8"
|
|
return None
|
|
|
|
|
|
def is_gemini_3_family(model_id: str) -> bool:
|
|
model_name = model_id.split(":", 1)[-1]
|
|
return model_name.startswith("gemini-3")
|
|
|
|
|
|
def openai_reasoning_for(
|
|
profile_effort: str | None,
|
|
*,
|
|
default_effort: OpenAIReasoningEffort | None = None,
|
|
) -> OpenAIReasoning | None:
|
|
"""Return an OpenAI reasoning kwarg from a profile effort string.
|
|
|
|
Requests ``summary: "auto"`` for every reasoning effort so the Responses
|
|
API emits visible reasoning text. ``effort: "none"`` disables reasoning
|
|
entirely, so no summary is attached.
|
|
"""
|
|
effort = profile_effort or default_effort or DEFAULT_LLM_REASONING.get("effort")
|
|
if effort == "none":
|
|
return {"effort": "none"}
|
|
if effort == "low":
|
|
return {"effort": "low", "summary": "auto"}
|
|
if effort == "medium":
|
|
return {"effort": "medium", "summary": "auto"}
|
|
if effort == "high":
|
|
return {"effort": "high", "summary": "auto"}
|
|
if effort == "xhigh":
|
|
return {"effort": "xhigh", "summary": "auto"}
|
|
return None
|
|
|
|
|
|
def anthropic_thinking_for(profile_effort: str | None) -> AnthropicThinking | None:
|
|
if profile_effort in _ANTHROPIC_EFFORTS:
|
|
# `display: "summarized"` makes Opus 4.7+ return the (summarized) reasoning
|
|
# text in the response. The adaptive default is "omitted", which streams a
|
|
# reasoning block carrying only a signature and no visible thinking — so the
|
|
# dashboard never has any text to render.
|
|
return {"type": "adaptive", "display": "summarized"}
|
|
return None
|
|
|
|
|
|
def anthropic_effort_for(profile_effort: str | None) -> AnthropicEffort | None:
|
|
if profile_effort in _ANTHROPIC_EFFORTS:
|
|
return profile_effort
|
|
return None
|
|
|
|
|
|
def fireworks_reasoning_effort_for(profile_effort: str | None) -> FireworksReasoningEffort | None:
|
|
"""Map profile effort to a Fireworks ``reasoning_effort`` value.
|
|
|
|
Fireworks' OpenAI-compatible API accepts ``reasoning_effort`` on its reasoning
|
|
models. ``none`` disables reasoning; ``xhigh``/``max`` are only honored by models
|
|
that advertise them (e.g. DeepSeek V4 Pro). The per-model ``efforts`` lists in
|
|
``dashboard/options.py`` gate which values can actually reach this function.
|
|
"""
|
|
if profile_effort == "none":
|
|
return "none"
|
|
if profile_effort == "low":
|
|
return "low"
|
|
if profile_effort == "medium":
|
|
return "medium"
|
|
if profile_effort == "high":
|
|
return "high"
|
|
if profile_effort == "xhigh":
|
|
return "xhigh"
|
|
if profile_effort == "max":
|
|
return "max"
|
|
return None
|
|
|
|
|
|
def google_thinking_level_for(profile_effort: str | None) -> GoogleThinkingLevel | None:
|
|
"""Map profile effort to Gemini 3+ ``thinking_level``."""
|
|
if profile_effort in ("minimal", "none"):
|
|
return "minimal"
|
|
if profile_effort == "low":
|
|
return "low"
|
|
if profile_effort == "medium":
|
|
return "medium"
|
|
if profile_effort in ("high", "xhigh", "max"):
|
|
return "high"
|
|
return None
|
|
|
|
|
|
def provider_model_kwargs(
|
|
model_id: str,
|
|
profile_effort: str | None,
|
|
*,
|
|
max_tokens: int,
|
|
openai_reasoning_default: OpenAIReasoning | None = None,
|
|
) -> ModelKwargs:
|
|
"""Build provider-specific kwargs for ``make_model`` from a model id and effort."""
|
|
kwargs: ModelKwargs = {"max_tokens": max_tokens}
|
|
if model_id.startswith("openai:"):
|
|
reasoning = openai_reasoning_for(profile_effort)
|
|
if reasoning is not None:
|
|
kwargs["reasoning"] = reasoning
|
|
elif openai_reasoning_default is not None:
|
|
kwargs["reasoning"] = openai_reasoning_default
|
|
elif model_id.startswith("anthropic:"):
|
|
thinking = anthropic_thinking_for(profile_effort)
|
|
if thinking is not None:
|
|
kwargs["thinking"] = thinking
|
|
effort = anthropic_effort_for(profile_effort)
|
|
if effort is not None:
|
|
kwargs["effort"] = effort
|
|
elif model_id.startswith("bedrock_converse:"):
|
|
# Opus 4.7+ on Bedrock rejects thinking.type "enabled"/budget_tokens with a
|
|
# ValidationException; it requires adaptive thinking plus output_config.effort,
|
|
# passed through Converse's additional_model_request_fields.
|
|
fields: dict[str, object] = {}
|
|
thinking = anthropic_thinking_for(profile_effort)
|
|
if thinking is not None:
|
|
fields["thinking"] = thinking
|
|
effort = anthropic_effort_for(profile_effort)
|
|
if effort is not None:
|
|
fields["output_config"] = {"effort": effort}
|
|
if fields:
|
|
kwargs["additional_model_request_fields"] = fields
|
|
elif model_id.startswith("google_genai:") and is_gemini_3_family(model_id):
|
|
thinking_level = google_thinking_level_for(profile_effort)
|
|
if thinking_level is not None:
|
|
kwargs["thinking_level"] = thinking_level
|
|
elif model_id.startswith("fireworks:"):
|
|
effort = fireworks_reasoning_effort_for(profile_effort)
|
|
if effort is not None:
|
|
kwargs["model_kwargs"] = {"reasoning_effort": effort}
|
|
return kwargs
|
|
|
|
|
|
def validate_local_dev_llm_config() -> None:
|
|
"""Validate API keys for the locally configured default model.
|
|
|
|
This check only runs in localhost development environments and is
|
|
intended to catch missing credentials for the default model specified
|
|
via LLM_MODEL_ID/DEFAULT_MODEL_ID. Runtime model selection may come
|
|
from team, profile, or thread configuration and is not validated here.
|
|
"""
|
|
dashboard_url = os.environ.get("DASHBOARD_BASE_URL", "")
|
|
if not dashboard_url.startswith("http://localhost"):
|
|
return
|
|
|
|
model_id = os.environ.get("LLM_MODEL_ID", DEFAULT_MODEL_ID)
|
|
|
|
if model_id.startswith("bedrock_converse:") and not (
|
|
os.environ.get("AWS_REGION") or os.environ.get("AWS_DEFAULT_REGION")
|
|
):
|
|
raise ValueError(f"AWS_REGION is required for configured model {model_id}")
|
|
elif model_id.startswith("fireworks:") and not os.environ.get("FIREWORKS_API_KEY"):
|
|
raise ValueError(f"FIREWORKS_API_KEY is required for configured model {model_id}")
|