Commit graph

12 commits

Author SHA1 Message Date
open-swe[bot]
96f97710ad
feat: add optional Slack Assistants API typing status indicator (#1269)
* feat: add optional Slack Assistants API typing status indicator

Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.

* fix(slack): drop redundant clear, add status heartbeat across model calls

- Slack auto-clears the typing indicator on bot post; remove the explicit
  assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
  that refreshes it on every model tick so it stays visible across long
  agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
  configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
  way out per Slack docs); no scope or app-config change required.

* feat(slack): contextual status text + rotating loading_messages

- set_slack_assistant_status now accepts an optional loading_messages list
  (capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
  payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
  assistant message's tool calls (e.g. "searching the codebase…" after
  grep, "running commands…" after execute), falling back to the default
  "is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
  contextual status on each refresh.

* fix slack assistant status lifecycle

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 10:21:55 -07:00
Johannes du Plessis
bdea1fb6da
feat(open-swe): always anchor findings to a single line (#1264)
* feat(open-swe): always anchor findings to a single line

GitHub renders multi-line review comment ranges as walls of context above
the comment, which buries the point. Match Devin Review's behavior and
always collapse `end_line` to `start_line` so each finding is anchored to
the single most relevant line.

Removes the (now unused) `MAX_FINDING_RANGE_LINES` cap and
`clip_finding_range` helper.

* fix(open-swe): collapse end_line before diff-range check
2026-05-07 20:26:59 -07:00
Johannes du Plessis
089db06135
feat(open-swe): collapse oversized finding ranges to start line (#1263)
Anchoring a finding to a 25-line range (e.g. an entire function) makes
GitHub render the whole block as context above the comment, burying the
review text. Cap ranges at 10 lines and collapse to the start line when
exceeded; small ranges (≤10 lines) still render as a multi-line anchor.
Updated the reviewer prompt to anchor tightly rather than range-select
the whole function.
2026-05-07 19:59:25 -07:00
Johannes du Plessis
0212a60f10
feat(open-swe): cap reviewer suggestions at 4 lines (#1262)
* feat(open-swe): cap reviewer suggestions at 4 lines

Long suggestion blocks read as the reviewer rewriting the code rather
than flagging an issue, which clutters PR comments. Steer the reviewer
toward description-only findings for non-trivial fixes, and enforce the
cap in `add_finding` / `update_finding` so suggestions over 4 lines are
dropped (the finding itself still publishes).

* fix(open-swe): don't clobber prior suggestion on over-cap update

`update_finding` was setting `suggestion=None` whenever clip_suggestion
dropped the input, which silently wiped any existing suggestion on the
finding. Distinguish the three cases: empty string clears, valid value
sets, over-cap value is rejected without touching the stored field.
2026-05-07 19:27:16 -07:00
Johannes du Plessis
3c330cdd40
fix(open-swe): hotfix — reviewer fetches its own diff via gh (#1261)
The host-side prep (clone + git diff in the sandbox) was producing empty
diffs for some PRs. The agent then silently published a "no issues
found" review (e.g. langchainplus#24069, #24458). Fix is to rip out the
prep entirely and instruct the agent to fetch the diff itself with
\`gh pr diff <num> --repo <owner>/<repo>\` (or the compare API on
re-review). \`add_finding\` line-range validation is skipped when no
\`diff_line_set\` is provided in config — we trust the agent's anchors.

We can revisit a host-side diff path once we understand why the prep
was failing; this hotfix unblocks reviews on private PRs in the
meantime.
2026-05-07 18:40:12 -07:00
Johannes du Plessis
8d52e17614
fix(open-swe): surface reviewer diff-prep failures instead of swallowing (#1260)
When `_ensure_repo_checked_out` or `compute_diff_in_sandbox` fail, the
reviewer used to log+continue and hand the agent an empty diff. The
agent would then emit a misleading "no issues found" review and the
underlying error stayed buried in server logs.

Now the helpers raise on non-zero exit codes (with the sandbox output)
and the prep block re-raises, so the failure shows up in the LangSmith
run trace and we can debug the real issue.
2026-05-07 18:20:27 -07:00
Johannes du Plessis
88a55057ea
fix(open-swe): host-format the review summary, drop agent prose (#1257)
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.

Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
2026-05-07 15:41:40 -07:00
Johannes du Plessis
03d9f23645
fix(open-swe): always post a review summary, even with no findings (#1256)
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable

Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.

Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
  (with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
  already-reviewed head_sha)

* fix(reviewer): always post a summary review, even with no findings

The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.

- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
  findings list is empty, vs. `**No issues at or above \`<sev>\`
  severity.**` with hidden count when only sub-threshold findings
  exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
  body is meaningful; calls out specifically not to skip on a clean PR.
2026-05-07 22:16:20 +00:00
Johannes du Plessis
378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
Johannes du Plessis
6ded489751
fix: reuse reviewer thread github token (#1247) 2026-05-06 18:09:39 -07:00
Johannes du Plessis
65f6b4636b
feat(open-swe): remove github_comment and enable reviewer trigger (#1244)
* Use gh for reviewer inline comments

* Trigger reviewer on PR review requests

* Format reviewer webhook test

* Add GitHub repo allowlist

* Separate reviewer repo allowlist

* only implement allowlist for reviewer
2026-05-06 16:14:38 -07:00
Johannes du Plessis
ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00