open-swe/evals/reviewer
Johannes du Plessis ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
..
golden_comments feat: add reviewer eval harness (#1239) 2026-05-05 13:04:23 -07:00
build_dataset.py feat: add reviewer eval harness (#1239) 2026-05-05 13:04:23 -07:00
judge.py feat: add reviewer graph + eval target wiring (#1241) 2026-05-06 10:15:58 -07:00
README.md feat: add reviewer eval harness (#1239) 2026-05-05 13:04:23 -07:00
run_eval.py feat: add reviewer graph + eval target wiring (#1241) 2026-05-06 10:15:58 -07:00
target.py feat: add reviewer graph + eval target wiring (#1241) 2026-05-06 10:15:58 -07:00

Reviewer Eval

Offline LangSmith eval for the Open SWE Reviewer graph against the 50 PRs from withmartian/code-review-benchmark. See REVIEWER_EVAL_PLAN.md at the repo root for the full design.

Layout

evals/reviewer/
├── golden_comments/      # 50 PRs × golden comments (copied from martian benchmark)
├── build_dataset.py      # martian JSON → LangSmith dataset (resolves SHAs via gh)
├── judge.py              # claude-opus-4-5 pairwise match evaluator + aggregate
├── target.py             # invokes the reviewer graph over langgraph_sdk
└── run_eval.py           # client.aevaluate entrypoint

Prerequisites

  • LANGSMITH_API_KEY set in your env.
  • gh authenticated (gh auth status) — needed for build_dataset.py.
  • ANTHROPIC_API_KEY set — judge runs claude-opus-4-5.
  • A running reviewer graph (local langgraph dev or deployed assistant id) with REVIEWER_ASSISTANT_ID env var pointing at it. Defaults to assistant reviewer on http://localhost:2024.

1. Build the dataset (once)

# Dry run — writes evals/reviewer/dataset_dryrun.json without uploading
uv run python -m evals.reviewer.build_dataset --dry-run

# Upload for real
uv run python -m evals.reviewer.build_dataset --dataset-name openswe-reviewer-v1

Each example carries: repo, pr_number, pr_url, base_sha, head_sha, base_ref, head_ref, pr_title. The dataset is frozen at upload time — upstream PR drift can't invalidate it.

2. Run the eval

The reviewer graph must be running and accept a pr input matching the example schema, and must emit a submit_review tool call (or set state["review"]["comments"]) with [{file, line, severity, body}, ...].

uv run python -m evals.reviewer.run_eval \
    --experiment-prefix openswe-reviewer-baseline \
    --max-concurrency 5

Smoke-test with 3 PRs first:

uv run python -m evals.reviewer.run_eval --limit 3

Comparing against Devin Review

Both tools are scored on the same 50 PRs with the same judge model (claude-opus-4-5) and the same judge prompt (verbatim from martian step3_judge_comments.py). Pull martian's published Devin numbers from their dashboard and compare against the LangSmith experiment's micro_* / macro_* summary metrics.

Notes

  • No GitHub forks needed — both upstream repos and martian's benchmark forks (ai-code-review-evaluation/*) are public.
  • judge_match charges judge LLM tokens proportional to n_candidates × n_goldens per example. For 50 PRs with ~3 goldens each and agents emitting ~10 candidates, expect ~1500 judge calls per experiment.