Commit graph

3 commits

Author SHA1 Message Date
Johannes du Plessis
834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00
Johannes du Plessis
ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Johannes du Plessis
6cecf429c7
feat: add reviewer eval harness (#1239)
* feat: add reviewer eval harness

Offline LangSmith eval scaffolding for the upcoming Open SWE Reviewer graph.
Imports the 50 PRs from withmartian/code-review-benchmark goldens, resolves
base/head SHAs via gh, and runs a claude-opus-4-5 LLM judge using martian's
verbatim prompt so scores are directly comparable to their published Devin
Review numbers. Reviewer graph itself is not part of this change.

* fix: ruff lint and format on reviewer eval files
2026-05-05 13:04:23 -07:00