mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 09:13:14 +00:00
* feat: tighten reviewer eval workflow Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews. * chore: move reviewer eval settings to config Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface. * feat: allow reviewer eval model overrides Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking. * fix: use adaptive thinking for Opus 4.7 Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API. * refactor: use latest Anthropic effort API Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort. * revert prompting |
||
|---|---|---|
| .. | ||
| auth.py | ||
| authorship.py | ||
| comments.py | ||
| github_app.py | ||
| github_comments.py | ||
| github_org_membership.py | ||
| github_token.py | ||
| github_user_email_map.py | ||
| langsmith.py | ||
| linear.py | ||
| linear_team_repo_map.py | ||
| messages.py | ||
| model.py | ||
| multimodal.py | ||
| repo.py | ||
| sandbox.py | ||
| sandbox_paths.py | ||
| sandbox_state.py | ||
| slack.py | ||
| slack_feedback.py | ||