Part of the domain-reorg adoption (build plan step C2): fork content,
upstream layout. Nine 1:1 module moves (reviewer_diff/eval_store/
findings/groups/publish/reconcile/trace_context + review_style_
collector/guidance) into agent/review/, with internal relative
imports re-wired to the new package depth. agent/review/__init__.py
mirrors upstream's thin re-export shim (one of the 21 verified "A"
structural adds).
Rewrote the 38 grep hits across importer files (agent/{analyzer,
ci_autofix,reviewer,webapp}.py, agent/dashboard/*, agent/middleware/
settle_review_check.py, agent/tools/*, agent/utils/github_feedback.py,
agent/webhooks/github.py, evals/reviewer/*, and the reviewer test
suite) to point at agent.review.*; 4 of the 38 hits were name
collisions (list_reviewer_findings, reviewer_outcomes,
_reviewer_thread_id, reviewer_thread_id — not the moved modules) and
were left untouched. tests/test_github_checks.py's module-alias
import (`from agent import reviewer_publish`) follows upstream's own
`from agent.review import publish as reviewer_publish` pattern so
downstream `reviewer_publish.*` call sites needed no changes.
agent/reviewer.py and agent/webapp.py stay in place per the hard
rule (fork content, import-only rewire) and are not part of this
package.
Gates: ruff check + ruff format --check, pytest --co -q (1637
collected), full unit suite (1637 passed), and the reviewer/findings
suite in isolation (pytest -k "review or finding", 421 passed).
|
||
|---|---|---|
| .. | ||
| golden_comments | ||
| build_dataset.py | ||
| config.toml | ||
| judge.py | ||
| README.md | ||
| run_eval.py | ||
| store_reporter.py | ||
| target.py | ||
Reviewer Eval
Offline LangSmith eval for the Open SWE Reviewer graph against the 50 PRs and
136 reference findings from withmartian/code-review-benchmark. Examples have
1–6 references (mean 2.72).
Layout
evals/reviewer/
├── golden_comments/ # 50 PRs × golden comments (copied from martian benchmark)
├── build_dataset.py # martian JSON → LangSmith dataset (resolves SHAs via gh)
├── config.toml # default benchmark run config
├── judge.py # claude-opus-4-5 pairwise match evaluator + aggregate
├── target.py # invokes the reviewer graph over langgraph_sdk
├── store_reporter.py # publishes live progress to the dashboard store record
└── run_eval.py # client.aevaluate entrypoint
Prerequisites
LANGSMITH_API_KEYset in your env.ghauthenticated (gh auth status) — needed forbuild_dataset.py.ANTHROPIC_API_KEYset — judge runsclaude-opus-4-5.- A running reviewer graph (local
langgraph devor deployed assistant id) withREVIEWER_ASSISTANT_IDenv var pointing at it. Defaults to assistantrevieweronhttp://localhost:2024.
1. Build the dataset (once)
# Dry run — writes evals/reviewer/dataset_dryrun.json without uploading
uv run python -m evals.reviewer.build_dataset --dry-run
# Upload for real
uv run python -m evals.reviewer.build_dataset --dataset-name openswe-reviewer-v1
Each example carries: repo, pr_number, pr_url, base_sha, head_sha,
base_ref, head_ref, pr_title. The dataset is frozen at upload time —
upstream PR drift can't invalidate it.
2. Run the eval
The reviewer graph must be running and accept the benchmark message/config
input. Eval runs record findings with add_finding and finish with
publish_review, which persists the exact ordered publication snapshot scored
by the harness.
uv run python -m evals.reviewer.run_eval
Smoke-test with 3 PRs first:
uv run python -m evals.reviewer.run_eval --limit 3
From the GitHub Action (recommended for full runs)
Trigger the Reviewer eval workflow (.github/workflows/reviewer-eval.yml)
from the Actions UI or gh workflow run reviewer-eval.yml --ref prod -f limit=3.
Run it on the prod branch so the harness/judge match the deployed reviewer it
scores. Running it on a durable runner (instead of inside the serving deployment)
means a deploy or container recycle can't kill a long run.
The Action sets REVIEWER_EVAL_REPORT_STORE=1, so run_eval publishes live
status/progress/logs to the LangGraph store record the dashboard reads — watch it
at Admin → Reviewer eval (/admin/evals), which is now a read-only progress
view (status, completed / total, log tail, LangSmith experiment link, and a link
back to the GitHub run). If the Action is cancelled/killed, the heartbeat goes
stale and the dashboard flips the run to failed within ~60s.
Required repository config:
- secrets:
LANGSMITH_API_KEY,ANTHROPIC_API_KEY(the judge runs in-process; reviewer-model keys are not needed — the reviewer runs in the deployment). - secret or var:
LANGGRAPH_URL— the deployment URL the eval drives and reports to.
Tracing project
Eval traces are routed to the open-swe-evals LangSmith project (set via
langsmith_project in config.toml, default open-swe-evals) so they stay out
of the deployment's production tracing project. The admin-triggered run forces
the same project via the LANGSMITH_PROJECT env var; override the default with
EVAL_LANGSMITH_PROJECT.
The runner reads benchmark settings from evals/reviewer/config.toml. Set the
deployment URL there (or leave it blank to use LANGGRAPH_URL / local dev).
The target sets reviewer_eval for every run, so publish_review does not post
to GitHub.
Per-repo review style prompts
At runtime the reviewer loads a custom style guide from LangGraph Store when
configurable.repo is set (owner + name → store key owner/name). This
applies to eval runs too, as long as a completed style profile exists for
that repo.
The Martian benchmark uses these upstream repos (10 PRs each):
getsentry/sentrykeycloak/keycloakgrafana/grafanadiscourse/discoursecalcom/cal.com
Before scoring with repo-specific styles, run Review styles analysis in the
dashboard for each repo (or copy prompts into store). Re-run make dev so the
reviewer graph sees the same store.
By default the judge scores the exact final surfaced_findings snapshot,
including only renderable findings selected by publish_review. Set
score_mode = "all_findings" only to diagnose deduplicated add_finding
calls before publication.
model_id and reasoning_effort in the config are passed to the reviewer run,
so isolated benchmark deployments can test a specific model/effort without
changing deployment-wide defaults.
Notes
- No GitHub forks needed — both upstream repos and martian's benchmark forks
(
ai-code-review-evaluation/*) are public. judge_matchevaluates the full deduplicatedn_candidates × n_goldensmatrix so matching is order-independent and its reasoning remains auditable in LangSmith.