mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 16:13:15 +00:00
* feat: add reviewer graph + eval target wiring
- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
alongside the main `agent` graph. Reuses the same sandbox lifecycle,
GH proxy auth, and middleware primitives from `agent.server`, but with
a narrower tool set, a reviewer-specific system prompt, no
commit/push, and the `task` (subagent) tool stripped via
`_ToolExclusionMiddleware` so review stays in one context.
- New `github_comment` tool: agents call it once per issue with
`(file, line, body, severity)` and the eval scores those calls
against golden comments.
- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
*not* on the reviewer's stack — that middleware exists to enforce the
main agent's "always finalize via Slack/Linear/PR" contract, which
the reviewer doesn't have. The main agent's behavior is unchanged.
- `evals/reviewer/target.py`: send PR info as a user message, extract
every `github_comment` tool call (multiple expected per review) into
the run output.
- `evals/reviewer/judge.py`: per-example evaluator now returns a list
of metrics under `{"results": [...]}` so LangSmith averages each
numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
the UI. Dropped the broken `aggregate_pr` summary evaluator that
reached for an attribute that doesn't exist on `RunTree`.
- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
`client.list_examples(limit=N)` since `aevaluate` doesn't accept
`max_examples`.
- Makefile: `dev` and `run` targets now use `uv run` so they work
without an activated venv.
* resolve comments
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
53 lines
1.8 KiB
Python
53 lines
1.8 KiB
Python
from typing import Any, Literal
|
|
|
|
Severity = Literal["Low", "Medium", "High", "Critical"]
|
|
|
|
_VALID_SEVERITIES: frozenset[str] = frozenset({"Low", "Medium", "High", "Critical"})
|
|
|
|
|
|
def _normalize_severity(value: str) -> Severity:
|
|
"""Title-case `value` and validate it against the allowed set.
|
|
|
|
The model occasionally emits "low"/"HIGH" instead of the title-cased
|
|
canonical form. Normalize before recording so we don't burn an LLM
|
|
turn on a Pydantic ValidationError retry.
|
|
"""
|
|
titled = value.strip().title()
|
|
if titled not in _VALID_SEVERITIES:
|
|
valid = ", ".join(sorted(_VALID_SEVERITIES))
|
|
raise ValueError(f"severity must be one of {valid}; got {value!r}")
|
|
return titled # type: ignore[return-value]
|
|
|
|
|
|
def github_comment(
|
|
file: str,
|
|
line: int,
|
|
body: str,
|
|
severity: str,
|
|
) -> dict[str, Any]:
|
|
"""Record a single inline review comment on the PR under review.
|
|
|
|
Call this tool once per issue you find. Multiple calls are expected — one
|
|
per distinct concern. The eval harness records every github_comment call
|
|
you make and scores them against the PR's golden comments.
|
|
|
|
**Do not** use this tool to summarize the PR or make general remarks. Each
|
|
call must point at a specific file and line and describe one concrete
|
|
issue (bug, security concern, perf problem, correctness issue, etc.).
|
|
|
|
Args:
|
|
file: Repo-relative path to the file the comment applies to.
|
|
line: 1-based line number in the file.
|
|
body: The review comment text. Be specific about the issue.
|
|
severity: One of "Low", "Medium", "High", "Critical" (case-insensitive).
|
|
|
|
Returns:
|
|
{"recorded": True, "file", "line", "severity", "body"}.
|
|
"""
|
|
return {
|
|
"recorded": True,
|
|
"file": file,
|
|
"line": line,
|
|
"severity": _normalize_severity(severity),
|
|
"body": body,
|
|
}
|