2026-02-10 11:22:52 -08:00
|
|
|
from .check_message_queue import check_message_queue_before_model
|
2026-03-04 17:59:52 -08:00
|
|
|
from .ensure_no_empty_msg import ensure_no_empty_msg
|
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring
- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
alongside the main `agent` graph. Reuses the same sandbox lifecycle,
GH proxy auth, and middleware primitives from `agent.server`, but with
a narrower tool set, a reviewer-specific system prompt, no
commit/push, and the `task` (subagent) tool stripped via
`_ToolExclusionMiddleware` so review stays in one context.
- New `github_comment` tool: agents call it once per issue with
`(file, line, body, severity)` and the eval scores those calls
against golden comments.
- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
*not* on the reviewer's stack — that middleware exists to enforce the
main agent's "always finalize via Slack/Linear/PR" contract, which
the reviewer doesn't have. The main agent's behavior is unchanged.
- `evals/reviewer/target.py`: send PR info as a user message, extract
every `github_comment` tool call (multiple expected per review) into
the run output.
- `evals/reviewer/judge.py`: per-example evaluator now returns a list
of metrics under `{"results": [...]}` so LangSmith averages each
numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
the UI. Dropped the broken `aggregate_pr` summary evaluator that
reached for an attribute that doesn't exist on `RunTree`.
- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
`client.list_examples(limit=N)` since `aevaluate` doesn't accept
`max_examples`.
- Makefile: `dev` and `run` targets now use `uv run` so they work
without an activated venv.
* resolve comments
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
|
|
|
from .exclude_tools import ExcludeToolsMiddleware
|
2026-05-01 14:24:25 -07:00
|
|
|
from .notify_step_limit import notify_step_limit_reached
|
2026-05-01 14:29:48 -07:00
|
|
|
from .sanitize_tool_inputs import SanitizeToolInputsMiddleware
|
2026-02-09 12:53:34 -08:00
|
|
|
from .tool_error_handler import ToolErrorMiddleware
|
|
|
|
|
|
2026-02-09 15:34:18 -08:00
|
|
|
__all__ = [
|
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring
- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
alongside the main `agent` graph. Reuses the same sandbox lifecycle,
GH proxy auth, and middleware primitives from `agent.server`, but with
a narrower tool set, a reviewer-specific system prompt, no
commit/push, and the `task` (subagent) tool stripped via
`_ToolExclusionMiddleware` so review stays in one context.
- New `github_comment` tool: agents call it once per issue with
`(file, line, body, severity)` and the eval scores those calls
against golden comments.
- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
*not* on the reviewer's stack — that middleware exists to enforce the
main agent's "always finalize via Slack/Linear/PR" contract, which
the reviewer doesn't have. The main agent's behavior is unchanged.
- `evals/reviewer/target.py`: send PR info as a user message, extract
every `github_comment` tool call (multiple expected per review) into
the run output.
- `evals/reviewer/judge.py`: per-example evaluator now returns a list
of metrics under `{"results": [...]}` so LangSmith averages each
numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
the UI. Dropped the broken `aggregate_pr` summary evaluator that
reached for an attribute that doesn't exist on `RunTree`.
- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
`client.list_examples(limit=N)` since `aevaluate` doesn't accept
`max_examples`.
- Makefile: `dev` and `run` targets now use `uv run` so they work
without an activated venv.
* resolve comments
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
|
|
|
"ExcludeToolsMiddleware",
|
2026-05-01 14:29:48 -07:00
|
|
|
"SanitizeToolInputsMiddleware",
|
2026-02-09 15:34:18 -08:00
|
|
|
"ToolErrorMiddleware",
|
|
|
|
|
"check_message_queue_before_model",
|
2026-03-04 17:59:52 -08:00
|
|
|
"ensure_no_empty_msg",
|
2026-05-01 14:24:25 -07:00
|
|
|
"notify_step_limit_reached",
|
2026-02-09 15:34:18 -08:00
|
|
|
]
|