Applies the plan's C5 step: git mv every test per the domain-reorg
move-map (movemap-m50.txt) into tests/{agent,analyzer,auth,dashboard,
github,middleware,models,reviewer,sandbox,slack,tools,webhooks}/, plus
the 13 fork-only placements from the scoping report §2c (Atlassian
webhook tests -> tests/webhooks/, test_atlassian_connect.py and
test_auth_error_leak.py -> tests/auth/, jira/confluence util tests ->
tests/tools/, test_repo_binding_isolation.py -> tests/sandbox/,
bot-identity/autofix tests -> tests/github/).
Path-only move: the only content edits are parents[1] -> parents[2]
fixes in test_e2b_integration.py and test_daytona_integration.py,
required because their __file__-relative ROOT path gained one more
directory level in the move.
Monkeypatch retargets for these files were already completed in C4;
none remained outstanding here.
Drive dashboard-triggered reviewer eval runs with per-run model, effort,
score mode, severity threshold, cap, limit, and concurrency overrides, plus
per-example start/finish/error logging in the eval target.
Rework the admin eval form from the label-left/control-right SettingsRow
(which crushed the description column when packing 3-4 wide inputs) into
stacked field groups with captioned inputs in a responsive grid.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: trigger reviewer evals from the admin page
Add an admin-only "Reviewer eval" section + endpoints that launch the
reviewer benchmark as an isolated subprocess against the running
deployment, with live status and the LangSmith experiment link. Route
eval traces to a dedicated open-swe-evals project so they stay out of
the production tracing project.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: reconcile reviewer eval status via heartbeat, not local process
The persisted record is shared across workers but _PROCS is process-local.
The owning worker now refreshes a heartbeat while the subprocess runs, and
status is only reconciled to failed once the heartbeat is stale, so a poll on
a worker without the local handle no longer kills a live run (and a duplicate
start is rejected across workers).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: upgrade default agent + reviewer model to Opus 4.8
Replace Opus 4.7 with Opus 4.8 (claude-opus-4-8) as the supported
Anthropic model surfaced in the profile editor and used by the main
agent and reviewer graphs. Effort levels (low/medium/high/xhigh/max)
and the high default are unchanged, matching the official Opus 4.8
docs. Updates eval config comment and tests accordingly.
* fix: provider-aware fallback for stale stored model ids
Dropping claude-opus-4-7 from the supported set meant persisted
profile/team-settings still holding it failed the SUPPORTED_MODEL_IDS
check and fell through to default_model_pair() — a cross-provider jump
to the OpenAI global default.
Add provider_fallback_pair: when a stored id is no longer supported but
its provider still has a supported model, resolve to that provider's
newest supported model (anthropic:claude-opus-4-7 -> 4.8), preserving
effort when valid. Resolution order is now: valid stored pair ->
same-provider fallback -> global default_model_pair(). Profile overrides
keep deferring to the team default when no model is set or the provider
is unknown.
* feat: tighten reviewer eval workflow
Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.
* chore: move reviewer eval settings to config
Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.
* feat: allow reviewer eval model overrides
Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.
* fix: use adaptive thinking for Opus 4.7
Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.
* refactor: use latest Anthropic effort API
Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.
* revert prompting