open-swe/tests
Adam Moussa a4ed19ba61
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)

Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.

- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
  region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
  FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module

* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8

The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.

Map profile effort to additional_model_request_fields:
  {thinking: {type: adaptive, display: summarized},
   output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.

* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids

Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
  (set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
  FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
  otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.

Surfaced by the cross-family review + verified against deploy/.

* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip

From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
  validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
  error code only, so the role ARN + account id in the raw botocore message never
  reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
  (Converse emits reasoning_content, not thinking) so the middleware is not a no-op
  on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)

* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids

Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
  us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
  each routed region (us-east-1/2, us-west-2). The model runs in the server process
  on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
  for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
  mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
  bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
  seed_store.sh's default via pick precedence, so the seed-script fix alone was
  insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
  ids to the Bedrock id (config.toml's model_id was an active, now-broken value).

AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.

* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)

Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
  (eval judge only — Bedrock builder/reviewer auth via the host IAM role).

REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
2026-06-29 15:57:19 -04:00
..
e2e feat(open-swe): copy-plan-as-markdown button + fix premature "ready" banner (#1603) 2026-06-24 15:07:25 -04:00
middleware fix: keep sandbox backend stable across recovery (#1294)w 2026-05-11 16:03:38 -07:00
conftest.py Remove reviewer env allowlist (#1353) 2026-05-28 14:29:33 -07:00
test_account_link.py fix: use Slack OIDC mappings for Slack thread ownership (#1410) 2026-06-04 19:58:30 +00:00
test_agent_instructions.py feat: per-repo custom instructions for the coding agent (#1460) 2026-06-09 15:46:29 -07:00
test_agent_schedules.py feat: add scheduled web agents (#1422) 2026-06-05 02:20:24 +00:00
test_agent_subagent_models.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_agent_thread_pr_state.py feat: track PR lifecycle state per thread for sidebar (#1492) 2026-06-11 12:21:59 -07:00
test_agent_usage.py Revert out-of-process usage-snapshot builder (#1473) 2026-06-09 13:19:11 -07:00
test_agents_md.py feat: reviewer enforces AGENTS.md/CLAUDE.md repo rules as mandatory pass (#1569) 2026-06-18 11:13:58 -07:00
test_analyzer_cron.py feat: outcomes dataset + bootstrap/continual split via skills (#1365) 2026-06-01 13:25:12 -07:00
test_analyzer_skills.py feat: outcomes dataset + bootstrap/continual split via skills (#1365) 2026-06-01 13:25:12 -07:00
test_anthropic_effort.py feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475) 2026-06-11 09:54:35 -07:00
test_api_standards_skill.py feat: apply API standards skill in PR reviews for API changes (#1452) 2026-06-08 14:37:04 -07:00
test_app_bot_identity.py feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60) 2026-06-29 14:22:33 -04:00
test_auth_error_leak.py fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) 2026-06-29 12:21:19 -04:00
test_auth_sources.py feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60) 2026-06-29 14:22:33 -04:00
test_authorship.py Adopt Sea Haven agent conventions, no attribution (#30) 2026-06-27 22:08:45 -04:00
test_autofix_state.py feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561) 2026-06-17 14:12:04 -07:00
test_autofix_webhook.py feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561) 2026-06-17 14:12:04 -07:00
test_check_message_queue.py feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561) 2026-06-17 14:12:04 -07:00
test_ci_autofix.py feat: activate PR babysitting UI toggles for autofix and trigger mode (#1561) 2026-06-17 14:12:04 -07:00
test_corridor_mcp.py feat: Add Corridor MCP analyzePlan integration (#1572) 2026-06-18 14:01:25 -07:00
test_currents_tools.py feat: add user-scoped Currents.dev API key for e2e test investigation (#1566) 2026-06-17 13:59:55 -07:00
test_dashboard_admin.py fix: allow GitHub logins for dashboard admins (#1582) 2026-06-20 09:44:12 -07:00
test_dashboard_csrf.py fix: harden origin parsing and home prompt submit failure (#1499) 2026-06-11 11:54:25 -07:00
test_dashboard_links.py fix: point reviewer "Open in Web" link to the review page (#1519) 2026-06-12 14:15:58 -07:00
test_dashboard_org_login_gate.py fix: Lock dashboard login to GitHub org members (#1367) 2026-06-01 20:56:30 +00:00
test_dashboard_repo_optional.py feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475) 2026-06-11 09:54:35 -07:00
test_dashboard_repos.py fix: Reduce graph load and dashboard repo failures (#1412) 2026-06-04 13:54:06 -07:00
test_dashboard_reviews.py perf: speed up Reviews list (My PRs + All PRs) (#1518) 2026-06-12 14:14:50 -07:00
test_dashboard_run_email.py fix: use Slack OIDC mappings for Slack thread ownership (#1410) 2026-06-04 19:58:30 +00:00
test_dashboard_thread_api.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_dashboard_thread_api_activity.py fix: restore org-wide read access to agent threads (#1474) 2026-06-09 16:41:07 -07:00
test_dashboard_web_handoff.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_daytona_integration.py fix(daytona): make sandbox snapshot configurable (#1220) 2026-05-01 22:51:34 +00:00
test_encryption.py feat: support TOKEN_ENCRYPTION_KEY rotation via MultiFernet [closes AB-2323] (#1275) 2026-05-08 14:29:23 -07:00
test_eval_jobs.py feat: Run reviewer eval in a GitHub Action; dashboard becomes read-only (#1556) 2026-06-16 19:38:36 -07:00
test_eval_store_reporter.py feat: Run reviewer eval in a GitHub Action; dashboard becomes read-only (#1556) 2026-06-16 19:38:36 -07:00
test_fireworks_model.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_github_app.py perf: cut review-chat time-to-first-token (#1598) 2026-06-23 12:32:10 -07:00
test_github_checks.py feat: shared GitHub HTTP helper with retries, rate-limit handling [closes OPE-45] (#1565) 2026-06-17 14:45:57 -07:00
test_github_ci.py feat: shared GitHub HTTP helper with retries, rate-limit handling [closes OPE-45] (#1565) 2026-06-17 14:45:57 -07:00
test_github_comment_prompts.py Make PR title repo-aware and auto-link issues (#46) 2026-06-27 22:56:01 -04:00
test_github_feedback.py feat: reconcile reviewer comment lifecycle (#1332) 2026-05-26 16:24:34 -07:00
test_github_http.py feat: shared GitHub HTTP helper with retries, rate-limit handling [closes OPE-45] (#1565) 2026-06-17 14:45:57 -07:00
test_github_issue_webhook.py fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) 2026-06-29 12:21:19 -04:00
test_github_oauth_refresh.py fix: auto-recover from expired GitHub refresh tokens (#1491) 2026-06-11 12:05:27 -07:00
test_github_proxy_refresh.py fix: refresh sandbox GitHub proxy token before mid-run expiry (#1496) 2026-06-11 10:59:21 -07:00
test_github_token_ttl.py fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) 2026-06-29 12:21:19 -04:00
test_http_security.py fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321] (#1277) 2026-05-08 22:06:25 +00:00
test_langsmith_sandbox_config.py feat: add idle TTL and delete-after-stop sandbox lifecycle controls (#1265) 2026-05-08 00:30:14 -04:00
test_langsmith_sandbox_timeout.py fix: enforce client-side deadline on sandbox execute (#1451) 2026-06-08 11:16:08 -07:00
test_langsmith_trace_url.py feat: plan mode with model-driven entry and collaborative review (#1580) 2026-06-23 12:06:58 -07:00
test_linear_webhook_replay.py fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) 2026-06-29 12:21:19 -04:00
test_local_integration.py fix: auto-create local sandbox root dir (#1402) 2026-06-03 23:47:08 +00:00
test_model_fallback_middleware.py feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475) 2026-06-11 09:54:35 -07:00
test_model_fallback_resolution.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_multimodal.py feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560) 2026-06-17 09:12:52 -07:00
test_normalize_repo.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_notify_step_limit_middleware.py fix: notify users via Slack when agent hits model call step limit (#1204) 2026-05-01 14:24:25 -07:00
test_notion_oauth.py feat: add user-scoped Notion MCP OAuth (#1593) 2026-06-23 12:07:13 -07:00
test_observability_tools.py feat: add user-scoped Notion MCP OAuth (#1593) 2026-06-23 12:07:13 -07:00
test_open_pull_request.py feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60) 2026-06-29 14:22:33 -04:00
test_plan_mode.py feat: plan mode with model-driven entry and collaborative review (#1580) 2026-06-23 12:06:58 -07:00
test_plan_review.py refactor(open-swe): plain HTTP comments instead of Yjs/BlockNote collab (#1601) 2026-06-23 22:12:42 +00:00
test_pr_ready_auto_review.py fix: reviewer resolves GitHub App token at run start, not via cross-process cache (#1409) 2026-06-04 12:11:23 -07:00
test_prompt_default_repo.py fix: make default repository configurable (#1429) 2026-06-05 13:48:47 -07:00
test_proxy_auth.py fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) 2026-06-29 12:21:19 -04:00
test_public_repo_org_gate.py fix: remove manual review trigger surfaces (#1396) 2026-06-03 12:33:22 -07:00
test_recent_comments.py chore: Drop monorepo (#1029) 2026-03-06 16:10:34 -08:00
test_refresh_slack_status_middleware.py feat: add optional Slack Assistants API typing status indicator (#1269) 2026-05-08 10:21:55 -07:00
test_repair_orphaned_tool_calls.py fix: repair orphaned tool calls before model calls (#1604) 2026-06-24 12:58:10 -07:00
test_repo_binding_isolation.py fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) 2026-06-29 12:21:19 -04:00
test_repo_extraction.py feat: resolve Slack repo from channel topic/purpose (#1453) 2026-06-08 13:53:16 -07:00
test_repo_prep.py fix: reviewer can silently review a stale checkout on reused sandboxes (#1503) 2026-06-11 13:42:10 -07:00
test_repo_snapshots.py feat: repo-scoped dynamic sandbox snapshots (#1595) 2026-06-23 12:24:11 -07:00
test_review_api.py perf: cut review-chat time-to-first-token (#1598) 2026-06-23 12:32:10 -07:00
test_review_chat.py perf: cut review-chat time-to-first-token (#1598) 2026-06-23 12:32:10 -07:00
test_review_style_collector.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_review_style_sync.py fix: review style prompts UX, stale runs, and OAuth refresh (#1321) 2026-05-21 17:44:04 +00:00
test_review_styles_store.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
test_reviewer.py fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54) 2026-06-29 12:21:19 -04:00
test_reviewer_diff.py fix: fetch PR diff via GitHub API to re-enable add_finding validation (#1339) 2026-05-27 17:03:33 +00:00
test_reviewer_eval_run.py Reviewer eval admin: configurable runs + stacked form layout (#1540) 2026-06-16 10:12:35 -07:00
test_reviewer_eval_target.py feat: Run reviewer eval in a GitHub Action; dashboard becomes read-only (#1556) 2026-06-16 19:38:36 -07:00
test_reviewer_findings.py fix: honest publish_review reporting + structured thread-not-found errors (#1481) 2026-06-10 10:44:13 -07:00
test_reviewer_groups.py fix: Simplify review explanation: full-width, plain prose, no diff links (#1547) 2026-06-16 15:32:24 -07:00
test_reviewer_outcomes.py feat: outcomes dataset + bootstrap/continual split via skills (#1365) 2026-06-01 13:25:12 -07:00
test_reviewer_publish.py feat: surface sub-threshold findings in review summary with web app link (#1571) 2026-06-18 11:15:42 -07:00
test_reviewer_reconcile.py feat: let reviewer set comment titles (#1356) 2026-05-28 16:04:38 -07:00
test_reviewer_tools.py feat: disable out-of-diff reviewer findings (#1516) 2026-06-12 10:19:30 -07:00
test_reviewer_watch.py fix: keep review check on follow-up commits, drop neutral conclusion (#1486) 2026-06-10 13:33:34 -07:00
test_sandbox_paths.py fix: better custom backend support (#1071) 2026-03-17 11:55:36 -07:00
test_sanitize_thinking_blocks.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_sanitize_tool_inputs.py fix: coerce malformed integer strings in read_file offset/limit params (#1216) 2026-05-01 14:29:48 -07:00
test_schedule_thread_wakeup.py feat: add schedule_thread_wakeup tool for self-polling (#1592) 2026-06-23 11:06:01 -07:00
test_slack_assistants_status.py feat: add Slack Block Kit reply options (#1407) 2026-06-04 10:26:09 -07:00
test_slack_context.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_slack_feedback.py feat: add Slack reaction feedback to LangSmith (#1231) 2026-05-08 14:24:12 -07:00
test_slack_oauth.py feat: open Slack-triggered PRs as the triggering user (#1375) 2026-06-02 15:04:20 -07:00
test_slack_thread_reply_tool.py feat: add Slack Block Kit reply options (#1407) 2026-06-04 10:26:09 -07:00
test_stale_sandbox_creating.py feat: repo-scoped dynamic sandbox snapshots (#1595) 2026-06-23 12:24:11 -07:00
test_team_credentials.py feat: server-side Datadog/LangSmith observability tools + team creds [closes OPE-54] (#1476) 2026-06-10 11:07:42 -07:00
test_team_settings_grouping.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_team_settings_org_guidelines.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
test_tool_artifact_middleware.py feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475) 2026-06-11 09:54:35 -07:00
test_user_credentials.py feat: add user-scoped Notion MCP OAuth (#1593) 2026-06-23 12:07:13 -07:00
test_user_mappings.py fix: use Slack OIDC mappings for Slack thread ownership (#1410) 2026-06-04 19:58:30 +00:00