open-swe/evals/reviewer
Johannes du Plessis 4e6c445246
feat: upgrade default agent + reviewer model to Opus 4.8 (#1350)
* feat: upgrade default agent + reviewer model to Opus 4.8

Replace Opus 4.7 with Opus 4.8 (claude-opus-4-8) as the supported
Anthropic model surfaced in the profile editor and used by the main
agent and reviewer graphs. Effort levels (low/medium/high/xhigh/max)
and the high default are unchanged, matching the official Opus 4.8
docs. Updates eval config comment and tests accordingly.

* fix: provider-aware fallback for stale stored model ids

Dropping claude-opus-4-7 from the supported set meant persisted
profile/team-settings still holding it failed the SUPPORTED_MODEL_IDS
check and fell through to default_model_pair() — a cross-provider jump
to the OpenAI global default.

Add provider_fallback_pair: when a stored id is no longer supported but
its provider still has a supported model, resolve to that provider's
newest supported model (anthropic:claude-opus-4-7 -> 4.8), preserving
effort when valid. Resolution order is now: valid stored pair ->
same-provider fallback -> global default_model_pair(). Profile overrides
keep deferring to the team default when no model is set or the provider
is unknown.
2026-05-28 10:59:29 -07:00
..
golden_comments feat: add reviewer eval harness (#1239) 2026-05-05 13:04:23 -07:00
build_dataset.py feat: add reviewer eval harness (#1239) 2026-05-05 13:04:23 -07:00
config.toml feat: upgrade default agent + reviewer model to Opus 4.8 (#1350) 2026-05-28 10:59:29 -07:00
judge.py feat: add reviewer graph + eval target wiring (#1241) 2026-05-06 10:15:58 -07:00
README.md feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
run_eval.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00
target.py feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312) 2026-05-20 18:35:00 +00:00

Reviewer Eval

Offline LangSmith eval for the Open SWE Reviewer graph against the 50 PRs from withmartian/code-review-benchmark. See REVIEWER_EVAL_PLAN.md at the repo root for the full design.

Layout

evals/reviewer/
├── golden_comments/      # 50 PRs × golden comments (copied from martian benchmark)
├── build_dataset.py      # martian JSON → LangSmith dataset (resolves SHAs via gh)
├── config.toml           # default benchmark run config
├── judge.py              # claude-opus-4-5 pairwise match evaluator + aggregate
├── target.py             # invokes the reviewer graph over langgraph_sdk
└── run_eval.py           # client.aevaluate entrypoint

Prerequisites

  • LANGSMITH_API_KEY set in your env.
  • gh authenticated (gh auth status) — needed for build_dataset.py.
  • ANTHROPIC_API_KEY set — judge runs claude-opus-4-5.
  • A running reviewer graph (local langgraph dev or deployed assistant id) with REVIEWER_ASSISTANT_ID env var pointing at it. Defaults to assistant reviewer on http://localhost:2024.

1. Build the dataset (once)

# Dry run — writes evals/reviewer/dataset_dryrun.json without uploading
uv run python -m evals.reviewer.build_dataset --dry-run

# Upload for real
uv run python -m evals.reviewer.build_dataset --dataset-name openswe-reviewer-v1

Each example carries: repo, pr_number, pr_url, base_sha, head_sha, base_ref, head_ref, pr_title. The dataset is frozen at upload time — upstream PR drift can't invalidate it.

2. Run the eval

The reviewer graph must be running and accept a pr input matching the example schema, and must emit a submit_review tool call (or set state["review"]["comments"]) with [{file, line, severity, body}, ...].

uv run python -m evals.reviewer.run_eval

Smoke-test with 3 PRs first:

uv run python -m evals.reviewer.run_eval --limit 3

The runner reads benchmark settings from evals/reviewer/config.toml. Set the deployment URL there (or leave it blank to use LANGGRAPH_URL / local dev). The target sets reviewer_eval for every run, so publish_review does not post to GitHub.

Per-repo review style prompts

At runtime the reviewer loads a custom style guide from LangGraph Store when configurable.repo is set (owner + name → store key owner/name). This applies to eval runs too, as long as a completed style profile exists for that repo.

The Martian benchmark uses these upstream repos (10 PRs each):

  • getsentry/sentry
  • keycloak/keycloak
  • grafana/grafana
  • discourse/discourse
  • calcom/cal.com

Before scoring with repo-specific styles, run Review styles analysis in the dashboard for each repo (or copy prompts into store). Re-run make dev so the reviewer graph sees the same store.

By default the judge scores final add_finding calls. Set score_mode = "surfaced_findings" in the config to score only findings that would pass the production threshold/cap.

model_id and reasoning_effort in the config are passed to the reviewer run, so isolated benchmark deployments can test a specific model/effort without changing deployment-wide defaults.

Comparing against Devin Review

Both tools are scored on the same 50 PRs with the same judge model (claude-opus-4-5) and the same judge prompt (verbatim from martian step3_judge_comments.py). Pull martian's published Devin numbers from their dashboard and compare against the LangSmith experiment's micro_* / macro_* summary metrics.

Notes

  • No GitHub forks needed — both upstream repos and martian's benchmark forks (ai-code-review-evaluation/*) are public.
  • judge_match charges judge LLM tokens proportional to n_candidates × n_goldens per example. For 50 PRs with ~3 goldens each and agents emitting ~10 candidates, expect ~1500 judge calls per experiment.