* feat: implement reviewer findings, publish_review, and watch mode
Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:
- Findings as first-class state on the reviewer thread metadata
(`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
ranges, suggestion text for ```suggestion blocks, github_review_comment_id
for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
compute_diff_line_set for in-diff validation, extract_diff_hunk for
caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
ranges fail at creation, not at GitHub-publish), `update_finding`,
`list_findings`, `publish_review`. The reviewer agent's tool list is
swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
one POST /reviews call with body + inline comments + ```suggestion blocks,
per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
before the agent's first model call (warm- and cold-path symmetric);
computed diff and in-diff line set passed via runnable config; system
prompt rewritten for the single-evolving-findings model, severity ladder,
in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
added to supported events. New `process_github_push_event` resolves the
open PR for the pushed branch, gates on the reviewer thread's `watch`
flag, builds a re-review configurable, and triggers a run on the same
canonical thread. `process_github_pr_close` toggles watch on
closed/reopened. `set_reviewer_thread_metadata` is called on first
review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
legacy {file, line, body, severity} shape the judge expects) and passes
the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
publish rendering + GraphQL resolve, and watch-mode webhook handlers
(push triggers re-review only when watching, idempotent on unchanged
head SHA, PR close disables watch). Updated existing reviewer-webhook
tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
into REVIEWER_DESIGN.md.
* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL
Address PR #1253 review findings:
- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
(`option no-prefix takes no value` — every prep run was failing
silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
now uses three-dot `base...head` (the merge-base diff GitHub renders
on Files-changed) so we don't pick up changes that landed on the base
branch after the PR diverged. Re-review delta keeps two-dot
`last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
`github_review_comment_id`. Without this, watched re-reviews
re-posted every previously surfaced finding, and only the most-recent
duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
`/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
is the canonical endpoint; the old form 404s, so comment ids were
never stored and watch-mode resolution couldn't run.
Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.
* fix(reviewer): default publish cap from 15 to 4
A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
|
||
|---|---|---|
| .. | ||
| golden_comments | ||
| build_dataset.py | ||
| judge.py | ||
| README.md | ||
| run_eval.py | ||
| target.py | ||
Reviewer Eval
Offline LangSmith eval for the Open SWE Reviewer graph against the 50 PRs from
withmartian/code-review-benchmark. See REVIEWER_EVAL_PLAN.md at the repo
root for the full design.
Layout
evals/reviewer/
├── golden_comments/ # 50 PRs × golden comments (copied from martian benchmark)
├── build_dataset.py # martian JSON → LangSmith dataset (resolves SHAs via gh)
├── judge.py # claude-opus-4-5 pairwise match evaluator + aggregate
├── target.py # invokes the reviewer graph over langgraph_sdk
└── run_eval.py # client.aevaluate entrypoint
Prerequisites
LANGSMITH_API_KEYset in your env.ghauthenticated (gh auth status) — needed forbuild_dataset.py.ANTHROPIC_API_KEYset — judge runsclaude-opus-4-5.- A running reviewer graph (local
langgraph devor deployed assistant id) withREVIEWER_ASSISTANT_IDenv var pointing at it. Defaults to assistantrevieweronhttp://localhost:2024.
1. Build the dataset (once)
# Dry run — writes evals/reviewer/dataset_dryrun.json without uploading
uv run python -m evals.reviewer.build_dataset --dry-run
# Upload for real
uv run python -m evals.reviewer.build_dataset --dataset-name openswe-reviewer-v1
Each example carries: repo, pr_number, pr_url, base_sha, head_sha,
base_ref, head_ref, pr_title. The dataset is frozen at upload time —
upstream PR drift can't invalidate it.
2. Run the eval
The reviewer graph must be running and accept a pr input matching the
example schema, and must emit a submit_review tool call (or set
state["review"]["comments"]) with [{file, line, severity, body}, ...].
uv run python -m evals.reviewer.run_eval \
--experiment-prefix openswe-reviewer-baseline \
--max-concurrency 5
Smoke-test with 3 PRs first:
uv run python -m evals.reviewer.run_eval --limit 3
Comparing against Devin Review
Both tools are scored on the same 50 PRs with the same judge model
(claude-opus-4-5) and the same judge prompt (verbatim from martian
step3_judge_comments.py). Pull martian's published Devin numbers from their
dashboard and compare against the LangSmith experiment's micro_* /
macro_* summary metrics.
Notes
- No GitHub forks needed — both upstream repos and martian's benchmark forks
(
ai-code-review-evaluation/*) are public. judge_matchcharges judge LLM tokens proportional ton_candidates × n_goldensper example. For 50 PRs with ~3 goldens each and agents emitting ~10 candidates, expect ~1500 judge calls per experiment.