mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 15:03:16 +00:00
* feat: tighten reviewer eval workflow Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews. * chore: move reviewer eval settings to config Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface. * feat: allow reviewer eval model overrides Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking. * fix: use adaptive thinking for Opus 4.7 Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API. * refactor: use latest Anthropic effort API Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort. * revert prompting |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| add_finding.py | ||
| fetch_url.py | ||
| http_request.py | ||
| linear_comment.py | ||
| linear_create_issue.py | ||
| linear_delete_issue.py | ||
| linear_get_issue.py | ||
| linear_get_issue_comments.py | ||
| linear_list_teams.py | ||
| linear_update_issue.py | ||
| list_findings.py | ||
| publish_review.py | ||
| request_pr_review.py | ||
| slack_read_thread_messages.py | ||
| slack_thread_reply.py | ||
| update_finding.py | ||
| web_search.py | ||