Commit graph

5 commits

Author SHA1 Message Date
Johannes du Plessis
ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
langsmith-forge[bot]
48965a82a6
fix: coerce malformed integer strings in read_file offset/limit params (#1216)
- Root cause: LLM occasionally generates strings like '1, 80' or '170, "limit": 60'
  for integer fields, causing a Pydantic ValidationError and wasting an LLM turn
- Change: add SanitizeToolInputsMiddleware in agent/middleware/sanitize_tool_inputs.py
  that extracts the leading integer from any string value in offset/limit before
  the call reaches Pydantic validation; registered before ToolErrorMiddleware in server.py
- Verified: 14 unit tests covering all three production trace patterns pass

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:29:48 -07:00
langsmith-forge[bot]
e5bc27a0ad
fix: notify users via Slack when agent hits model call step limit (#1204)
* fix: notify users via Slack when agent hits model call step limit

- Root cause: GraphRecursionError at 1000 steps bypassed all @after_agent
  middleware including open_pr_if_needed, leaving users with no notification
- Change: Added ModelCallLimitMiddleware(run_limit=60) to intercept gracefully
  before the hard recursion limit, and added notify_step_limit_reached
  @after_agent middleware to post a Slack thread reply when the limit fires
- Verified: 107 existing tests pass, no regressions

* fix: harden step-limit Slack notification

Ensure the step-limit notification runs after the PR safety net and cover the new middleware behavior with focused unit tests.

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:24:25 -07:00
Brace Sproul
bd52e5e09d
chore: Drop monorepo (#1029)
* chore: Drop monorepo

* cr
2026-03-06 16:10:34 -08:00
Renamed from apps/agent/agent/middleware/__init__.py (Browse further)