open-swe/agent
Johannes du Plessis 834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00
..
dashboard feat: open-swe dashboard for per-user profile config (#1302) 2026-05-15 11:23:53 -07:00
integrations feat: add idle TTL and delete-after-stop sandbox lifecycle controls (#1265) 2026-05-08 00:30:14 -04:00
middleware fix: keep sandbox backend stable across recovery (#1294)w 2026-05-11 16:03:38 -07:00
tools feat: Adds ability to run evals against deployment (#1311) 2026-05-18 15:47:13 -07:00
utils feat: Adds ability to run evals against deployment (#1311) 2026-05-18 15:47:13 -07:00
encryption.py feat: support TOKEN_ENCRYPTION_KEY rotation via MultiFernet [closes AB-2323] (#1275) 2026-05-08 14:29:23 -07:00
prompt.py slack: move feedback-reaction ask out of every reply, into the tip rotation (#1279) 2026-05-08 21:59:14 +00:00
reviewer.py feat: Adds ability to run evals against deployment (#1311) 2026-05-18 15:47:13 -07:00
reviewer_diff.py fix(open-swe): surface reviewer diff-prep failures instead of swallowing (#1260) 2026-05-07 18:20:27 -07:00
reviewer_findings.py feat(open-swe): always anchor findings to a single line (#1264) 2026-05-07 20:26:59 -07:00
reviewer_publish.py fix: publish_review tool returns generic "Failed to POST PR review" without GitHub API status/body, agent retries with no signal (#1299) 2026-05-12 13:55:39 -07:00
server.py feat: Adds ability to run evals against deployment (#1311) 2026-05-18 15:47:13 -07:00
webapp.py fix: stop inferring Slack repos from message text (#1306) 2026-05-15 22:36:44 +00:00