* feat: add reviewer graph + eval target wiring
- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
alongside the main `agent` graph. Reuses the same sandbox lifecycle,
GH proxy auth, and middleware primitives from `agent.server`, but with
a narrower tool set, a reviewer-specific system prompt, no
commit/push, and the `task` (subagent) tool stripped via
`_ToolExclusionMiddleware` so review stays in one context.
- New `github_comment` tool: agents call it once per issue with
`(file, line, body, severity)` and the eval scores those calls
against golden comments.
- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
*not* on the reviewer's stack — that middleware exists to enforce the
main agent's "always finalize via Slack/Linear/PR" contract, which
the reviewer doesn't have. The main agent's behavior is unchanged.
- `evals/reviewer/target.py`: send PR info as a user message, extract
every `github_comment` tool call (multiple expected per review) into
the run output.
- `evals/reviewer/judge.py`: per-example evaluator now returns a list
of metrics under `{"results": [...]}` so LangSmith averages each
numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
the UI. Dropped the broken `aggregate_pr` summary evaluator that
reached for an attribute that doesn't exist on `RunTree`.
- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
`client.list_examples(limit=N)` since `aevaluate` doesn't accept
`max_examples`.
- Makefile: `dev` and `run` targets now use `uv run` so they work
without an activated venv.
* resolve comments
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
|
||
|---|---|---|
| .. | ||
| golden_comments | ||
| build_dataset.py | ||
| judge.py | ||
| README.md | ||
| run_eval.py | ||
| target.py | ||
Reviewer Eval
Offline LangSmith eval for the Open SWE Reviewer graph against the 50 PRs from
withmartian/code-review-benchmark. See REVIEWER_EVAL_PLAN.md at the repo
root for the full design.
Layout
evals/reviewer/
├── golden_comments/ # 50 PRs × golden comments (copied from martian benchmark)
├── build_dataset.py # martian JSON → LangSmith dataset (resolves SHAs via gh)
├── judge.py # claude-opus-4-5 pairwise match evaluator + aggregate
├── target.py # invokes the reviewer graph over langgraph_sdk
└── run_eval.py # client.aevaluate entrypoint
Prerequisites
LANGSMITH_API_KEYset in your env.ghauthenticated (gh auth status) — needed forbuild_dataset.py.ANTHROPIC_API_KEYset — judge runsclaude-opus-4-5.- A running reviewer graph (local
langgraph devor deployed assistant id) withREVIEWER_ASSISTANT_IDenv var pointing at it. Defaults to assistantrevieweronhttp://localhost:2024.
1. Build the dataset (once)
# Dry run — writes evals/reviewer/dataset_dryrun.json without uploading
uv run python -m evals.reviewer.build_dataset --dry-run
# Upload for real
uv run python -m evals.reviewer.build_dataset --dataset-name openswe-reviewer-v1
Each example carries: repo, pr_number, pr_url, base_sha, head_sha,
base_ref, head_ref, pr_title. The dataset is frozen at upload time —
upstream PR drift can't invalidate it.
2. Run the eval
The reviewer graph must be running and accept a pr input matching the
example schema, and must emit a submit_review tool call (or set
state["review"]["comments"]) with [{file, line, severity, body}, ...].
uv run python -m evals.reviewer.run_eval \
--experiment-prefix openswe-reviewer-baseline \
--max-concurrency 5
Smoke-test with 3 PRs first:
uv run python -m evals.reviewer.run_eval --limit 3
Comparing against Devin Review
Both tools are scored on the same 50 PRs with the same judge model
(claude-opus-4-5) and the same judge prompt (verbatim from martian
step3_judge_comments.py). Pull martian's published Devin numbers from their
dashboard and compare against the LangSmith experiment's micro_* /
macro_* summary metrics.
Notes
- No GitHub forks needed — both upstream repos and martian's benchmark forks
(
ai-code-review-evaluation/*) are public. judge_matchcharges judge LLM tokens proportional ton_candidates × n_goldensper example. For 50 PRs with ~3 goldens each and agents emitting ~10 candidates, expect ~1500 judge calls per experiment.