open-swe/tests/test_anthropic_effort.py
Johannes du Plessis 834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00

13 lines
453 B
Python

from __future__ import annotations
from agent.server import _anthropic_effort_for, _anthropic_thinking_for
def test_anthropic_uses_adaptive_thinking_and_effort() -> None:
assert _anthropic_thinking_for("high") == {"type": "adaptive"}
assert _anthropic_effort_for("high") == "high"
def test_anthropic_ignores_unknown_effort() -> None:
assert _anthropic_thinking_for("unknown") is None
assert _anthropic_effort_for("unknown") is None