evals/reviewer/{target,run_eval,build_dataset} called load_dotenv() at
import time, so importing them in tests injected the real .env (live
LANGGRAPH_URL, tokens) into the whole pytest process. The slack-context
default-repo tests then reached the real LangGraph store via
get_team_default_repo() and picked up the developer's actual team
default repo, failing in full-suite runs while passing in isolation.
Move load_dotenv() into the CLI entrypoints (all env reads were already
lazy), and patch get_team_default_repo in the two affected tests so they
stay hermetic regardless of environment.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add reviewer eval harness
Offline LangSmith eval scaffolding for the upcoming Open SWE Reviewer graph.
Imports the 50 PRs from withmartian/code-review-benchmark goldens, resolves
base/head SHAs via gh, and runs a claude-opus-4-5 LLM judge using martian's
verbatim prompt so scores are directly comparable to their published Devin
Review numbers. Reviewer graph itself is not part of this change.
* fix: ruff lint and format on reviewer eval files