mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 09:13:14 +00:00
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
324 lines
11 KiB
Python
324 lines
11 KiB
Python
from __future__ import annotations
|
|
|
|
from typing import Any
|
|
|
|
import pytest
|
|
from fastapi import HTTPException
|
|
|
|
from agent.dashboard import thread_api
|
|
|
|
|
|
class _FakeThreads:
|
|
def __init__(self, metadata: dict[str, Any]) -> None:
|
|
self.metadata = metadata
|
|
self.updates: list[dict[str, Any]] = []
|
|
|
|
async def get(self, thread_id: str) -> dict[str, Any]:
|
|
return {"thread_id": thread_id, "metadata": self.metadata}
|
|
|
|
async def update(self, *, thread_id: str, metadata: dict[str, Any]) -> None:
|
|
self.updates.append(metadata)
|
|
self.metadata.update(metadata)
|
|
|
|
|
|
class _FakeRuns:
|
|
def __init__(self) -> None:
|
|
self.created: list[dict[str, Any]] = []
|
|
|
|
async def create(self, *args: Any, **kwargs: Any) -> dict[str, str]:
|
|
self.created.append({"args": args, "kwargs": kwargs})
|
|
return {"run_id": "run-1"}
|
|
|
|
|
|
class _FakeClient:
|
|
def __init__(self, metadata: dict[str, Any]) -> None:
|
|
self.threads = _FakeThreads(metadata)
|
|
self.runs = _FakeRuns()
|
|
|
|
|
|
async def _inactive_thread(thread_id: str) -> bool:
|
|
return False
|
|
|
|
|
|
async def _active_thread(thread_id: str) -> bool:
|
|
return True
|
|
|
|
|
|
async def _noop_token_check(login: str) -> None:
|
|
return None
|
|
|
|
|
|
async def _empty_profile(login: str) -> dict[str, Any]:
|
|
return {}
|
|
|
|
|
|
async def _run_email(login: str, profile: dict[str, Any]) -> str:
|
|
return "octocat@example.com"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_on_slack_thread_uses_dashboard_source(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "slack",
|
|
"github_login": "octocat",
|
|
"triggering_user_email": "octocat@example.com",
|
|
"repo_owner": "octo",
|
|
"repo_name": "repo",
|
|
"source_context": {
|
|
"slack_thread": {"channel_id": "C1", "thread_ts": "123.45"},
|
|
},
|
|
}
|
|
client = _FakeClient(metadata)
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _inactive_thread)
|
|
monkeypatch.setattr(thread_api, "_ensure_dashboard_github_token", _noop_token_check)
|
|
monkeypatch.setattr(thread_api, "get_profile", _empty_profile)
|
|
monkeypatch.setattr(thread_api, "_resolve_run_email", _run_email)
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(content="continue in web"),
|
|
email="octocat@example.com",
|
|
)
|
|
|
|
assert exc_info.value.status_code == 409
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_sends_image_content_blocks(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "dashboard",
|
|
"github_login": "octocat",
|
|
"repo_owner": "octo",
|
|
"repo_name": "repo",
|
|
}
|
|
client = _FakeClient(metadata)
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _inactive_thread)
|
|
monkeypatch.setattr(thread_api, "_ensure_dashboard_github_token", _noop_token_check)
|
|
monkeypatch.setattr(thread_api, "get_profile", _empty_profile)
|
|
monkeypatch.setattr(thread_api, "_resolve_run_email", _run_email)
|
|
monkeypatch.setattr(
|
|
thread_api,
|
|
"create_image_block",
|
|
lambda *, base64, mime_type: {"type": "image", "data": base64, "mime_type": mime_type},
|
|
)
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(
|
|
content="describe this",
|
|
images=[
|
|
thread_api.DashboardImageBody(
|
|
base64="aW1hZ2U=",
|
|
mimeType="image/png",
|
|
fileName="screenshot.png",
|
|
)
|
|
],
|
|
),
|
|
)
|
|
|
|
assert exc_info.value.status_code == 409
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_on_busy_thread_queues_dashboard_handoff(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "slack",
|
|
"github_login": "octocat",
|
|
"triggering_user_email": "octocat@example.com",
|
|
}
|
|
client = _FakeClient(metadata)
|
|
queued_messages: list[object] = []
|
|
|
|
async def fake_queue_message_for_thread(thread_id: str, message_content: object) -> bool:
|
|
queued_messages.append(message_content)
|
|
return True
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _active_thread)
|
|
monkeypatch.setattr(thread_api, "queue_message_for_thread", fake_queue_message_for_thread)
|
|
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(content="continue in web"),
|
|
email="octocat@example.com",
|
|
)
|
|
|
|
assert client.threads.updates[0]["source"] == "dashboard"
|
|
assert queued_messages == [{"text": "continue in web", "source": "dashboard"}]
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_on_busy_thread_queues_images(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "dashboard",
|
|
"github_login": "octocat",
|
|
"resolved_model": "bedrock_converse:us.anthropic.claude-opus-4-8",
|
|
}
|
|
client = _FakeClient(metadata)
|
|
queued_messages: list[object] = []
|
|
|
|
async def fake_queue_message_for_thread(thread_id: str, message_content: object) -> bool:
|
|
queued_messages.append(message_content)
|
|
return True
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _active_thread)
|
|
monkeypatch.setattr(thread_api, "queue_message_for_thread", fake_queue_message_for_thread)
|
|
monkeypatch.setattr(
|
|
thread_api,
|
|
"create_image_block",
|
|
lambda *, base64, mime_type: {"type": "image", "data": base64, "mime_type": mime_type},
|
|
)
|
|
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(
|
|
content="continue in web",
|
|
images=[thread_api.DashboardImageBody(base64="aW1hZ2U=", mimeType="image/png")],
|
|
),
|
|
)
|
|
|
|
assert queued_messages == [
|
|
{
|
|
"text": "continue in web",
|
|
"source": "dashboard",
|
|
"images": [{"type": "image", "data": "aW1hZ2U=", "mime_type": "image/png"}],
|
|
}
|
|
]
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_on_busy_text_only_thread_rejects_images(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "dashboard",
|
|
"github_login": "octocat",
|
|
"resolved_model": "fireworks:accounts/fireworks/models/deepseek-v4-pro",
|
|
}
|
|
client = _FakeClient(metadata)
|
|
queued_messages: list[object] = []
|
|
|
|
async def fake_queue_message_for_thread(thread_id: str, message_content: object) -> bool:
|
|
queued_messages.append(message_content)
|
|
return True
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _active_thread)
|
|
monkeypatch.setattr(thread_api, "queue_message_for_thread", fake_queue_message_for_thread)
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(
|
|
content="continue in web",
|
|
images=[thread_api.DashboardImageBody(base64="aW1hZ2U=", mimeType="image/png")],
|
|
model_id="fireworks:accounts/fireworks/models/deepseek-v4-pro",
|
|
effort="medium",
|
|
),
|
|
)
|
|
|
|
assert exc_info.value.status_code == 422
|
|
assert "does not support image input" in exc_info.value.detail
|
|
assert queued_messages == []
|
|
assert client.threads.updates == []
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_on_busy_unknown_model_rejects_images(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "dashboard",
|
|
"github_login": "octocat",
|
|
}
|
|
client = _FakeClient(metadata)
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _active_thread)
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(
|
|
content="continue in web",
|
|
images=[thread_api.DashboardImageBody(base64="aW1hZ2U=", mimeType="image/png")],
|
|
),
|
|
)
|
|
|
|
assert exc_info.value.status_code == 422
|
|
assert "does not support image input" in exc_info.value.detail
|
|
assert client.threads.updates == []
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_preserves_explicit_repo_less_thread(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "dashboard",
|
|
"github_login": "octocat",
|
|
"repo_explicitly_none": True,
|
|
}
|
|
client = _FakeClient(metadata)
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _inactive_thread)
|
|
monkeypatch.setattr(thread_api, "_ensure_dashboard_github_token", _noop_token_check)
|
|
monkeypatch.setattr(thread_api, "get_profile", _empty_profile)
|
|
monkeypatch.setattr(thread_api, "_resolve_run_email", _run_email)
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(content="continue in web"),
|
|
)
|
|
|
|
assert exc_info.value.status_code == 409
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_dashboard_followup_without_repo_metadata_allows_team_default(
|
|
monkeypatch: pytest.MonkeyPatch,
|
|
) -> None:
|
|
metadata = {
|
|
"source": "dashboard",
|
|
"github_login": "octocat",
|
|
}
|
|
client = _FakeClient(metadata)
|
|
|
|
monkeypatch.setattr(thread_api, "langgraph_client", lambda: client)
|
|
monkeypatch.setattr(thread_api, "get_thread_active_status", _inactive_thread)
|
|
monkeypatch.setattr(thread_api, "_ensure_dashboard_github_token", _noop_token_check)
|
|
monkeypatch.setattr(thread_api, "get_profile", _empty_profile)
|
|
monkeypatch.setattr(thread_api, "_resolve_run_email", _run_email)
|
|
|
|
with pytest.raises(HTTPException) as exc_info:
|
|
await thread_api.send_dashboard_message(
|
|
"thread-1",
|
|
"octocat",
|
|
thread_api.ThreadMessageBody(content="continue in web"),
|
|
)
|
|
|
|
assert exc_info.value.status_code == 409
|