* feat: upgrade default agent + reviewer model to Opus 4.8 Replace Opus 4.7 with Opus 4.8 (claude-opus-4-8) as the supported Anthropic model surfaced in the profile editor and used by the main agent and reviewer graphs. Effort levels (low/medium/high/xhigh/max) and the high default are unchanged, matching the official Opus 4.8 docs. Updates eval config comment and tests accordingly. * fix: provider-aware fallback for stale stored model ids Dropping claude-opus-4-7 from the supported set meant persisted profile/team-settings still holding it failed the SUPPORTED_MODEL_IDS check and fell through to default_model_pair() — a cross-provider jump to the OpenAI global default. Add provider_fallback_pair: when a stored id is no longer supported but its provider still has a supported model, resolve to that provider's newest supported model (anthropic:claude-opus-4-7 -> 4.8), preserving effort when valid. Resolution order is now: valid stored pair -> same-provider fallback -> global default_model_pair(). Profile overrides keep deferring to the team default when no model is set or the provider is unknown. |
||
|---|---|---|
| .. | ||
| golden_comments | ||
| build_dataset.py | ||
| config.toml | ||
| judge.py | ||
| README.md | ||
| run_eval.py | ||
| target.py | ||
Reviewer Eval
Offline LangSmith eval for the Open SWE Reviewer graph against the 50 PRs from
withmartian/code-review-benchmark. See REVIEWER_EVAL_PLAN.md at the repo
root for the full design.
Layout
evals/reviewer/
├── golden_comments/ # 50 PRs × golden comments (copied from martian benchmark)
├── build_dataset.py # martian JSON → LangSmith dataset (resolves SHAs via gh)
├── config.toml # default benchmark run config
├── judge.py # claude-opus-4-5 pairwise match evaluator + aggregate
├── target.py # invokes the reviewer graph over langgraph_sdk
└── run_eval.py # client.aevaluate entrypoint
Prerequisites
LANGSMITH_API_KEYset in your env.ghauthenticated (gh auth status) — needed forbuild_dataset.py.ANTHROPIC_API_KEYset — judge runsclaude-opus-4-5.- A running reviewer graph (local
langgraph devor deployed assistant id) withREVIEWER_ASSISTANT_IDenv var pointing at it. Defaults to assistantrevieweronhttp://localhost:2024.
1. Build the dataset (once)
# Dry run — writes evals/reviewer/dataset_dryrun.json without uploading
uv run python -m evals.reviewer.build_dataset --dry-run
# Upload for real
uv run python -m evals.reviewer.build_dataset --dataset-name openswe-reviewer-v1
Each example carries: repo, pr_number, pr_url, base_sha, head_sha,
base_ref, head_ref, pr_title. The dataset is frozen at upload time —
upstream PR drift can't invalidate it.
2. Run the eval
The reviewer graph must be running and accept a pr input matching the
example schema, and must emit a submit_review tool call (or set
state["review"]["comments"]) with [{file, line, severity, body}, ...].
uv run python -m evals.reviewer.run_eval
Smoke-test with 3 PRs first:
uv run python -m evals.reviewer.run_eval --limit 3
The runner reads benchmark settings from evals/reviewer/config.toml. Set the
deployment URL there (or leave it blank to use LANGGRAPH_URL / local dev).
The target sets reviewer_eval for every run, so publish_review does not post
to GitHub.
Per-repo review style prompts
At runtime the reviewer loads a custom style guide from LangGraph Store when
configurable.repo is set (owner + name → store key owner/name). This
applies to eval runs too, as long as a completed style profile exists for
that repo.
The Martian benchmark uses these upstream repos (10 PRs each):
getsentry/sentrykeycloak/keycloakgrafana/grafanadiscourse/discoursecalcom/cal.com
Before scoring with repo-specific styles, run Review styles analysis in the
dashboard for each repo (or copy prompts into store). Re-run make dev so the
reviewer graph sees the same store.
By default the judge scores final add_finding calls. Set
score_mode = "surfaced_findings" in the config to score only findings that
would pass the production threshold/cap.
model_id and reasoning_effort in the config are passed to the reviewer run,
so isolated benchmark deployments can test a specific model/effort without
changing deployment-wide defaults.
Comparing against Devin Review
Both tools are scored on the same 50 PRs with the same judge model
(claude-opus-4-5) and the same judge prompt (verbatim from martian
step3_judge_comments.py). Pull martian's published Devin numbers from their
dashboard and compare against the LangSmith experiment's micro_* /
macro_* summary metrics.
Notes
- No GitHub forks needed — both upstream repos and martian's benchmark forks
(
ai-code-review-evaluation/*) are public. judge_matchcharges judge LLM tokens proportional ton_candidates × n_goldensper example. For 50 PRs with ~3 goldens each and agents emitting ~10 candidates, expect ~1500 judge calls per experiment.