open-swe/tests/models/test_model_fallback_resolution.py

195 lines
7.1 KiB
Python
Raw Permalink Normal View History

from unittest.mock import AsyncMock, patch
import pytest
from agent.dashboard.agent_overrides import normalize_profile_overrides
from agent.dashboard.options import (
DEFAULT_MODEL_ID,
FABLE_MODEL_IDS,
default_model_pair,
fable_disabled_fallback,
gate_fable_model,
provider_fallback_pair,
)
from agent.dashboard.team_settings import (
TeamSettingsUpdate,
get_team_default_model,
normalize_team_settings_for_response,
)
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) * feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude) Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash) entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay freely selectable for the agent and reviewer graphs and via team/profile defaults. - pyproject: add langchain-aws (ChatBedrockConverse + boto3) - options.py: Bedrock Claude entry + default; remove openai/google entries - model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget), region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/ FIREWORKS_API_KEY local-dev validation - server.py: provider-aware fallback kwargs build - sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks - model_fallback: treat transient botocore ClientError codes as fallback-worthy - eval_jobs: repoint hardcoded eval model id to Bedrock Claude - tests: repoint dropped model ids; drop obsolete google test module * fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8 The handoff spec wired Bedrock Converse thinking as {type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a ValidationException: thinking.type "enabled" is not supported; it requires thinking.type "adaptive" plus output_config.effort. Verified by live invoke against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1): the enabled+budget shape 400s, adaptive+effort returns normally. Map profile effort to additional_model_request_fields: {thinking: {type: adaptive, display: summarized}, output_config: {effort: <low|medium|high|xhigh|max>}} reusing anthropic_thinking_for/anthropic_effort_for. Update the two subagent-model tests asserting the old shape. * fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids Model selection is store-driven, so seed_store.sh's team_settings/default seed is what runs in prod. It still seeded the removed providers, which would fail at runtime after the migration: - agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8 - reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8 (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer) - fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY -> FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would otherwise fail-fast at boot) - docs (DEPLOYMENT/ROTATION/put-config) updated to match. Surfaced by the cross-family review + verified against deploy/. * fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip From /sh-security-review (all confirmed-low): - model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches validate_local_dev_llm_config) so the validated region is the one actually used. - model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the error code only, so the role ARN + account id in the raw botocore message never reach logs or the user channel (CWE-209). - sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks (Converse emits reasoning_content, not thinking) so the middleware is not a no-op on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.) * deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids Deployment-readiness for the Bedrock migration (PR #62): - instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in each routed region (us-east-1/2, us-west-2). The model runs in the server process on the box, so the EC2 instance role is the principal. Simulator-verified (allowed for opus-4-8, implicitDeny for other models) and synth-verified. Passed the mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed). - config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides seed_store.sh's default via pick precedence, so the seed-script fix alone was insufficient — both sources now point at the supported Bedrock id. - infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai: ids to the Bedrock id (config.toml's model_id was an active, now-broken value). AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there. * chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed) Those three providers were dropped in the Bedrock/Fireworks migration and their keys revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY) were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not recreate the shells, and from fetch-config's mirror array so boot stops requesting them: - config-store.ts SECRET_VARS + descriptions (28 -> 25 shells) - fetch-config.sh SECRET_VARS array (kept in lockstep) - put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional (eval judge only — Bedrock builder/reviewer auth via the host IAM role). REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
STALE_ANTHROPIC = "bedrock_converse:us.anthropic.claude-opus-4-7"
SUPPORTED_ANTHROPIC = "bedrock_converse:us.anthropic.claude-opus-4-8"
STALE_SONNET = "bedrock_converse:us.anthropic.claude-sonnet-4-9"
SUPPORTED_SONNET = "bedrock_converse:us.anthropic.claude-sonnet-5"
RETIRED_DIRECT_ANTHROPIC = "anthropic:claude-opus-4-8"
RETIRED_FABLE = "anthropic:claude-fable-5"
def test_provider_fallback_preserves_provider_and_effort() -> None:
assert provider_fallback_pair(STALE_ANTHROPIC, "xhigh") == (SUPPORTED_ANTHROPIC, "xhigh")
def test_provider_fallback_keeps_bedrock_sonnet_in_family() -> None:
# A dropped Bedrock Sonnet must prefer the current Bedrock Sonnet, not cross to
# the Bedrock Opus that happens to sit first in the provider's list. Requires
# _claude_family_of to understand bedrock_converse ids, not just anthropic:.
assert provider_fallback_pair(STALE_SONNET, "high") == (SUPPORTED_SONNET, "high")
def test_provider_fallback_uses_default_effort_when_unsupported() -> None:
assert provider_fallback_pair(STALE_ANTHROPIC, "bogus") == (SUPPORTED_ANTHROPIC, "high")
assert provider_fallback_pair(STALE_ANTHROPIC, None) == (SUPPORTED_ANTHROPIC, "high")
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) * feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude) Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash) entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay freely selectable for the agent and reviewer graphs and via team/profile defaults. - pyproject: add langchain-aws (ChatBedrockConverse + boto3) - options.py: Bedrock Claude entry + default; remove openai/google entries - model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget), region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/ FIREWORKS_API_KEY local-dev validation - server.py: provider-aware fallback kwargs build - sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks - model_fallback: treat transient botocore ClientError codes as fallback-worthy - eval_jobs: repoint hardcoded eval model id to Bedrock Claude - tests: repoint dropped model ids; drop obsolete google test module * fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8 The handoff spec wired Bedrock Converse thinking as {type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a ValidationException: thinking.type "enabled" is not supported; it requires thinking.type "adaptive" plus output_config.effort. Verified by live invoke against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1): the enabled+budget shape 400s, adaptive+effort returns normally. Map profile effort to additional_model_request_fields: {thinking: {type: adaptive, display: summarized}, output_config: {effort: <low|medium|high|xhigh|max>}} reusing anthropic_thinking_for/anthropic_effort_for. Update the two subagent-model tests asserting the old shape. * fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids Model selection is store-driven, so seed_store.sh's team_settings/default seed is what runs in prod. It still seeded the removed providers, which would fail at runtime after the migration: - agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8 - reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8 (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer) - fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY -> FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would otherwise fail-fast at boot) - docs (DEPLOYMENT/ROTATION/put-config) updated to match. Surfaced by the cross-family review + verified against deploy/. * fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip From /sh-security-review (all confirmed-low): - model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches validate_local_dev_llm_config) so the validated region is the one actually used. - model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the error code only, so the role ARN + account id in the raw botocore message never reach logs or the user channel (CWE-209). - sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks (Converse emits reasoning_content, not thinking) so the middleware is not a no-op on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.) * deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids Deployment-readiness for the Bedrock migration (PR #62): - instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in each routed region (us-east-1/2, us-west-2). The model runs in the server process on the box, so the EC2 instance role is the principal. Simulator-verified (allowed for opus-4-8, implicitDeny for other models) and synth-verified. Passed the mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed). - config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides seed_store.sh's default via pick precedence, so the seed-script fix alone was insufficient — both sources now point at the supported Bedrock id. - infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai: ids to the Bedrock id (config.toml's model_id was an active, now-broken value). AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there. * chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed) Those three providers were dropped in the Bedrock/Fireworks migration and their keys revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY) were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not recreate the shells, and from fetch-config's mirror array so boot stops requesting them: - config-store.ts SECRET_VARS + descriptions (28 -> 25 shells) - fetch-config.sh SECRET_VARS array (kept in lockstep) - put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional (eval judge only — Bedrock builder/reviewer auth via the host IAM role). REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
def test_provider_fallback_resolves_fireworks_within_provider() -> None:
model, effort = provider_fallback_pair("fireworks:accounts/fireworks/models/legacy-old", "low")
assert model.startswith("fireworks:")
assert effort == "low"
@pytest.mark.parametrize("model_id", ["unknown:model", "no-colon", "", None, 123])
def test_provider_fallback_returns_none_without_provider_match(model_id: object) -> None:
assert provider_fallback_pair(model_id, "high") is None
@pytest.mark.asyncio
async def test_team_default_stale_anthropic_stays_on_provider() -> None:
settings = {
"default_agent_model": STALE_ANTHROPIC,
"default_agent_reasoning_effort": "xhigh",
}
with patch(
"agent.dashboard.team_settings.get_team_settings",
new_callable=AsyncMock,
return_value=settings,
):
assert await get_team_default_model("agent") == (SUPPORTED_ANTHROPIC, "xhigh")
@pytest.mark.asyncio
async def test_team_default_unknown_provider_falls_back_to_global() -> None:
settings = {
"default_reviewer_model": "mystery:model",
"default_reviewer_reasoning_effort": "high",
}
with patch(
"agent.dashboard.team_settings.get_team_settings",
new_callable=AsyncMock,
return_value=settings,
):
assert await get_team_default_model("reviewer") == default_model_pair()
def test_profile_stale_anthropic_upgrades_to_supported() -> None:
profile = {"default_model": STALE_ANTHROPIC, "reasoning_effort": "high"}
assert normalize_profile_overrides(profile) == (SUPPORTED_ANTHROPIC, "high")
def test_team_settings_update_normalizes_retired_models() -> None:
update = TeamSettingsUpdate(
default_agent_model=SUPPORTED_ANTHROPIC,
default_agent_reasoning_effort="medium",
default_agent_subagent_model=RETIRED_DIRECT_ANTHROPIC,
default_agent_subagent_reasoning_effort="medium",
default_reviewer_model=RETIRED_DIRECT_ANTHROPIC,
default_reviewer_reasoning_effort="medium",
default_reviewer_subagent_model=RETIRED_DIRECT_ANTHROPIC,
default_reviewer_subagent_reasoning_effort="low",
)
assert update.default_agent_subagent_model == SUPPORTED_ANTHROPIC
assert update.default_reviewer_model == SUPPORTED_ANTHROPIC
assert update.default_reviewer_subagent_model == SUPPORTED_ANTHROPIC
def test_team_settings_update_maps_retired_fable_to_opus_not_fable() -> None:
update = TeamSettingsUpdate(
fable_enabled=True,
default_agent_model=RETIRED_FABLE,
default_agent_reasoning_effort="high",
)
assert update.default_agent_model == SUPPORTED_ANTHROPIC
assert update.default_agent_model not in FABLE_MODEL_IDS
def test_team_settings_update_rejects_unknown_model() -> None:
with pytest.raises(ValueError, match="unsupported agent model"):
TeamSettingsUpdate(
default_agent_model="openai:gpt-6",
default_agent_reasoning_effort="medium",
)
def test_team_settings_update_rejects_invalid_effort_for_retired_model() -> None:
with pytest.raises(ValueError, match="effort 'bogus' not supported"):
TeamSettingsUpdate(
default_agent_model=RETIRED_DIRECT_ANTHROPIC,
default_agent_reasoning_effort="bogus",
)
def test_team_settings_response_normalizes_retired_models() -> None:
settings = normalize_team_settings_for_response(
{
"default_agent_subagent_model": RETIRED_DIRECT_ANTHROPIC,
"default_agent_subagent_reasoning_effort": "medium",
"default_reviewer_model": "openai:gpt-5.5",
"default_reviewer_reasoning_effort": "medium",
"default_reviewer_subagent_model": RETIRED_FABLE,
"default_reviewer_subagent_reasoning_effort": "low",
}
)
assert settings["default_agent_subagent_model"] == SUPPORTED_ANTHROPIC
assert settings["default_reviewer_model"] == SUPPORTED_ANTHROPIC
assert settings["default_reviewer_subagent_model"] == SUPPORTED_ANTHROPIC
def test_profile_without_model_defers_to_team_default() -> None:
assert normalize_profile_overrides({"reasoning_effort": "high"}) == (None, None)
def test_profile_unknown_provider_defers_to_team_default() -> None:
profile = {"default_model": "mystery:model", "reasoning_effort": "high"}
assert normalize_profile_overrides(profile) == (None, None)
def test_global_default_is_supported() -> None:
model, _ = default_model_pair()
assert model == DEFAULT_MODEL_ID
def test_gate_fable_passthrough_when_enabled() -> None:
assert gate_fable_model(
"bedrock_converse:us.anthropic.claude-fable-5", "high", fable_enabled=True
) == (
"bedrock_converse:us.anthropic.claude-fable-5",
"high",
)
def test_gate_fable_swaps_to_opus_when_disabled() -> None:
assert gate_fable_model(
"bedrock_converse:us.anthropic.claude-fable-5", "high", fable_enabled=False
) == (
"bedrock_converse:us.anthropic.claude-opus-4-8",
"high",
)
def test_gate_fable_leaves_non_fable_ids_alone() -> None:
assert gate_fable_model(
"fireworks:accounts/fireworks/models/deepseek-v4-pro", "high", fable_enabled=False
) == (
"fireworks:accounts/fireworks/models/deepseek-v4-pro",
"high",
)
def test_fable_disabled_fallback_is_non_fable_claude() -> None:
model, effort = fable_disabled_fallback("high")
assert model == "bedrock_converse:us.anthropic.claude-opus-4-8"
assert model not in FABLE_MODEL_IDS
assert effort == "high"