Commit graph

25 commits

Author SHA1 Message Date
Johannes du Plessis
8461979b0d
fix: honest publish_review reporting + structured thread-not-found errors (#1481)
* fix: honest publish_review reporting + structured thread-not-found errors

- Document skipped_empty_re_review and dry_run in the publish_review
  docstring and add a closing-summary contract to the reviewer prompt so
  the agent never claims a review was published when review_id is null.
- Raise ReviewerThreadMissingError from replace_findings on SDK
  NotFoundError; add_finding/update_finding/publish_review return a
  structured do-not-retry result instead of raising, so the agent reports
  the blocker after one failure instead of retrying 10-30 times.

* fix: translate thread 404s across all reviewer tool boundaries

get_thread_metadata now raises ReviewerThreadMissingError instead of
swallowing a missing thread as {} (which produced misleading 'No finding
found' results), set_reviewer_thread_metadata translates the SDK 404 the
same way, and every reviewer tool entrypoint (add/update/list findings,
publish_review incl. eval dry-run, resolve/reply thread) returns the
structured do-not-retry result.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:44:13 -07:00
Johannes du Plessis
7cd882bb67
fix: make review publish idempotent (partial-failure recovery) (#1477)
Publish had partial-failure windows that double-posted summaries or
corrupted findings state. This hardens the recovery paths:

- open_swe_review_exists is now tri-state (True/False/None). On a
  pagination/API failure it returns None ("unknown") instead of False,
  and the empty-summary dedup keys off the durable last_reviewed_sha
  before consulting GitHub, so a transient failure never double-posts a
  "no issues found" summary.
- Comment-id backfill matches strictly on the embedded open-swe marker;
  the colliding (path, line, body) fallback is gone, so similar findings
  no longer share a comment id and break resolve-on-fix.
- Review-id and comment-id stamping collapse into one guarded
  read-modify-write (re-reads latest before writing), removing the
  half-stamped intermediate states the prior multi-write flow left open.
- New mutate_findings primitive centralizes read-modify-write so finding
  updates operate on the freshest persisted list and skip no-op writes.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:17:31 -07:00
Johannes du Plessis
1cd26dda53
fix: reword reviewer feedback prompt (#1449)
* fix: reword reviewer feedback prompt

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refine reviewer feedback prompt

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-08 10:02:02 -07:00
Ramon Nogueira
d2268871eb
feat: surface dashboard UI link on PR reviews [INF-0000] (#1440)
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]

Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.

The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.

* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
2026-06-07 05:17:09 +00:00
Johannes du Plessis
f64ab2bcd7
feat: surface out-of-diff findings in a collapsed dropdown (#1427)
* fix: stop reviewer retrying out-of-diff findings

add_finding rejects findings anchored outside the PR diff, but the agent
retried the same finding 2-3x with adjacent line ranges before giving up,
burning a model turn each. Add a reviewer-prompt recovery block telling the
agent the rejection is authoritative (drop or re-anchor to a + line, don't
retry adjacent), and enrich the rejection payload with nearby in-diff line
ranges for the file so a single re-anchor needs no guessing.

* feat: surface out-of-diff findings in a collapsed dropdown

Instead of rejecting findings anchored outside the PR diff, accept them
(marked in_diff=false) and surface them in a collapsed <details> section of
the review summary, Devin-style. Inline comments stay reserved for in-diff
findings; out-of-diff are severity-gated and capped the same way.

Re-review normally suppresses the empty summary, but now makes an exception
when there are new out-of-diff findings to surface. Surfaced out-of-diff
findings carry a github_review_id so they aren't reposted on later pushes.

Supersedes the earlier 'drop/re-anchor out-of-diff' prompt guidance.

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-05 10:41:48 -07:00
Johannes du Plessis
6a9fd85295
fix: guard reviewer GraphQL against null repository (#1426)
GitHub returns repository: null when the token can't read the repo (SAML,
expired token, private/deleted). dict.get(k, {}) doesn't coalesce explicit
null, so fetch_pr_review_threads crashed with AttributeError and publish_review
could never post a review. Guard with isinstance checks and return collected
threads on null repository; sweep the same pattern in resolve_review_thread.

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-05 09:37:42 -07:00
Johannes du Plessis
1d4f1aed33
fix: reviewer publishes against stale head_sha on mid-run re-review (#1393)
* fix: resolve reviewer head_sha from thread metadata, not frozen run config

A push that lands while a reviewer run is in flight is delivered as a
queued message into that run. The run's configurable is frozen at
creation, so its head_sha still names the commit the run was created for
— not the commit just pushed. publish_review then anchored the GitHub
review to the stale commit and regressed last_reviewed_sha to it, and
add_finding/update_finding stamped findings with the stale SHA.

Persist the current head in thread metadata at every reviewer dispatch
(both the ready-for-review and push paths, before they branch to create
a run or queue a message), and add resolve_review_head_sha() which
prefers the metadata head over the run config. Wire it into
publish_review (review commit_id + last_reviewed_sha), add_finding
(first_seen_sha) and update_finding (last_confirmed_sha). Falls back to
the run config when metadata carries no head (first review, eval, tests).

* fix: persist head_sha in manual review dispatch (trigger_pr_review_from_ref)

resolve_review_head_sha prefers metadata[head_sha] over the run config,
and the push/ready dispatchers write it — but trigger_pr_review_from_ref
(Slack/GitHub @open-swe review, request_pr_review tool) created a run
with a freshly-fetched config head while leaving metadata's head stale
from a prior dispatch. A manual re-review at a newer commit would then
resolve to the old head and publish/advance findings against it.

Persist head_sha in that dispatch's metadata write too, so every
run-creating reviewer dispatch keeps metadata in sync with the head its
run targets. Caught by the Open SWE reviewer on this PR.
2026-06-03 11:38:56 -07:00
Johannes du Plessis
18f8ca56fb
fix: dedup empty reviewer summary by PR state, not stale re_review flag (#1391)
A push that lands while a reviewer run is in flight is delivered as a
queued message into the still-running first-review run, whose
configurable still has re_review=False. The empty-review guard in
publish_review only skipped the 'No issues found' summary when
is_re_review was True, so the queued reconcile published a second,
duplicate top-level 'No issues found' review.

Key the empty-review skip off actual PR state instead: add
open_swe_review_exists(), which detects the marker render_review_body
embeds in every Open SWE review body, and skip the summary when a prior
Open SWE review already exists (regardless of the re_review flag). Fails
open on API error so a genuine first review is never suppressed.
2026-06-03 10:30:13 -07:00
open-swe[bot]
a361ee8f2e
feat: let reviewer set comment titles (#1356)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-28 16:04:38 -07:00
open-swe[bot]
9998119921
feat: generate reviewer finding titles (#1355)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-28 15:28:55 -07:00
open-swe[bot]
d8d3794649
fix: include reviewer trace links (#1351)
* fix: include reviewer trace links

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* feat: move reviewer trace-link toggle to dashboard

Replace the OPEN_SWE_REVIEW_TRACE_LINK_ENABLED env var with a team-level
'Trace Links' toggle in the Open SWE Review dashboard tab. The toggle is
read per-publish via get_team_review_trace_links_enabled(); the per-run
review_trace_link_enabled config override still forces it off.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-28 13:48:14 -07:00
Johannes du Plessis
fee209601b
feat: structured reviewer comments + auto resolution comments (#1352)
* feat: structured reviewer comments + auto resolution comments

Restructure inline review comment bodies (severity emoji, bold title from
the first line, line reference, feedback footer) without duplicating the
first description line, and post an automatic resolution/dismissal comment
to the GitHub thread when a finding is resolved or dismissed.

- Centralize render_resolution_comment in reviewer_publish; fix a crash when
  last_reconciliation_note is None and drop the misleading generic fallback.
- Post the resolution comment in resolve_finding_thread (the normal
  update_finding path), not only in publish_review, so it actually fires on
  re-review. Dedupe via github_posted_resolution_comment_ids.
- Add pytest coverage for rendering and the resolution-comment flow.

* fix: post resolution comment to every closed thread in resolve_finding_thread

Per-thread iteration (matching _resolve_threads_for_resolved_findings) so
duplicate threads after the first also receive the resolved/dismissed
explanation before being closed.
2026-05-28 12:44:00 -07:00
Johannes du Plessis
0d4d1c5a3b
fix: dedupe reviewer comments from PR state (#1341)
* fix: dedupe reviewer comments from PR state

Use GitHub review-thread markers to repair reviewer publication state before posting or resolving findings, so re-reviews do not duplicate comments and resolved findings close all matching PR threads.

* fix: require all duplicate reviewer threads resolved

Avoid treating a marker-backed finding as resolved when only one duplicate thread is outdated while another matching thread remains open.
2026-05-27 10:52:30 -07:00
Johannes du Plessis
fa147ce256
fix: fetch PR diff via GitHub API to re-enable add_finding validation (#1339)
* reviewer: fetch PR diff via GitHub API to re-enable add_finding validation

The previous hotfix in reviewer.py set diff_line_set=None because the
sandbox-based diff prep was sometimes producing empty diffs. That made
every bad anchor a publish-time 422 instead of a creation-time
rejection — the agent burned tokens producing unanchorable findings,
and we had to add a publish-time retry safety net (#1338) to clean up.

Fetch the PR's unified diff via the GitHub REST API at reviewer
startup and populate diff_text + diff_line_set so add_finding can
reject bad anchors immediately. The API path is reliable and is the
same diff GitHub validates against when posting inline review
comments. If the fetch fails, fall back to the previous behavior
(validation disabled, publish-time retry handles it).

Also extract the PR-diff fetch into reviewer_diff.fetch_pr_diff so
both reviewer.py and publish_review.py share one implementation
instead of two copies.

* reviewer: make diff_line_set validation side-aware

compute_diff_line_set previously returned only new-side line numbers,
so re-enabling add_finding's validation would wrongly reject findings
with side=LEFT (deleted-line bugs whose only anchor is an old-side
line). Return {file: {"RIGHT": {new_lines}, "LEFT": {old_lines}}}
instead, and have is_range_in_diff select the matching side from the
finding's recorded side. add_finding and publish_review's retry
filter both pass the finding's side through.
2026-05-27 17:03:33 +00:00
langsmith-engine[bot]
58b1d52fee
fix: publish_review HTTP 422 "Path/Line could not be resolved" — agent retries with identical args instead of dropping unresolvable findings (#1338)
* publish_review: drop unresolvable findings and retry once on GitHub 422

GitHub returns 422 with 'Path could not be resolved' or 'Line could not be
resolved' when an inline comment anchors to a file/line not in the PR diff.
Previously the agent retried publish_review with byte-identical args
multiple times before draining to skipped_empty_re_review=true, silently
losing findings.

- reviewer_publish.post_pull_request_review: parse 422 body and tag with
  _error_kind='unresolved_anchor' plus _raw_errors so callers can act.
- tools/publish_review._publish_review_async: when that signal fires,
  cross-check each finding's range against the run config's diff_line_set,
  drop the bad ones, and re-POST once with only the valid findings. Return
  unresolvable_findings + hint so the agent calls update_finding instead of
  retrying the same payload.
- reviewer.py: one-line prompt addendum telling the agent that
  unresolvable_findings means update_finding, not retry.
- tests: cover 422 tagging (path + line), the drop-and-retry success path,
  the retry-still-fails path, and the don't-blind-retry path when no
  diff_line_set is available.

* publish_review: fetch PR diff on demand for 422 retry filter

Reviewer runs clear configurable['diff_line_set'] before the agent
starts, so the unresolved-anchor retry path had no diff data to filter
against — in the reachable production case it dropped nothing and
returned success=False with empty unresolvable_findings, losing the
otherwise-valid comments.

Fall back to fetching the PR's unified diff via the GitHub REST API
and recomputing the line set on the fly when no cached set is
available. The cached set is still preferred when present.

---------

Co-authored-by: issues-agent <issues-agent@langchain.dev>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-26 18:37:56 -07:00
Johannes du Plessis
5702a9d452
feat: reconcile reviewer comment lifecycle (#1332)
* feat: reconcile reviewer comment lifecycle

Track GitHub review threads for reviewer findings so re-reviews can resolve or reply to existing comments, and collect thumbs feedback on new review comments in LangSmith.

* fix: clarify reviewer comment lifecycle

* chore: apply reviewer formatting

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-26 16:24:34 -07:00
Johannes du Plessis
40162a6d9e
fix: inject existing PR review threads into reviewer context (#1331)
* fix(reviewer): inject existing PR review threads into reviewer context

The reviewer agent was filing the same inline comment on every re-review
because it only saw findings recorded on its own thread metadata — not
the live PR review-thread state on GitHub. When a previous finding was
still open (code unchanged, or a human reply explained it), the agent
rediscovered the same defect on the next push and called `add_finding`
again, producing duplicate comments.

This change fetches the PR's review threads (across all reviewers, with
replies and isResolved status) via GraphQL and renders them into the
first-review and re-review contexts as a "Pre-existing PR review
threads" block. The system prompt now lists overlap with that block as
a hard "Do NOT file" rule, and treats threads addressed by a human
reply as resolved.

This also gives the reviewer comment-awareness on its very first run
on a PR, so it skips findings already raised by another reviewer or
bot.

* fix(reviewer): wrap PR review threads in untrusted-data XML block

Addresses the reviewer comment on this PR
(https://github.com/langchain-ai/open-swe/pull/1331#discussion_r3295497533):
PR review comment bodies are attacker-controlled (anyone who can comment
on the PR can put anything in them), and they were being concatenated
into the reviewer's system prompt with instruction-priority.

Switches the existing-threads section from a Markdown block to an XML
data block:

  <pr_review_threads>
    <thread location="path:line" status="open">
      <comment author="open-swe[bot]">
        <body>...</body>
      </comment>
      <comment author="romain-priour-lc">
        <body>We added defaults in the template</body>
      </comment>
    </thread>
  </pr_review_threads>

The system prompt now explicitly names the wrapper, tells the agent that
everything inside it is untrusted data from the PR (not instructions),
and that prompt-injection payloads inside bodies must be disregarded.
We keep the bodies so the agent can actually read engineer replies —
that's the whole point of comment-awareness — but they're delimited as
data, not concatenated as prose. Modern frontier models are well-trained
to honor this contract.

Additional defenses:
- Author logins are validated against the GitHub username grammar; any
  unexpected value is rendered as "unknown" so the `author` attribute
  can't smuggle freeform text.
- Literal closing tags (`</body>`, `</pr_review_threads>`, etc.) in
  bodies are neutered so a body can't break out of its wrapper.
- Body length is capped at 4000 chars per comment to bound the prompt.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-24 16:36:50 -07:00
Johannes du Plessis
b2a0ac3b79
feat: auto-review PRs on opened / ready-for-review (#1325)
* feat: auto-review PRs on opened / ready-for-review

Trigger Open SWE Review on `pull_request` actions `opened` and
`ready_for_review` against the canonical reviewer thread (no need to
request open-swe[bot] as a reviewer). `converted_to_draft` now also
flips watch=False on the existing reviewer thread.

Draft PRs are gated by a tri-state user setting on the profile:
inherit team default, always on, or always off. The team-wide
`review_draft_prs` setting is the org-wide default; each user can
override it in My Settings.

External contributors with no Open SWE profile fall back to the team
default.

* fix: PR review comments — auth source + draft-aware watch toggle

- `process_github_pr_ready` now dispatches with `source="github"` so the
  auth resolver finds the bot token persisted on the thread. The previous
  `source="github_auto"` fell through to the email-based path in non
  bot-token-only deployments and failed with a missing-user-email error.

- `converted_to_draft` no longer unconditionally clears `watch`. When the
  PR author's effective `review_draft_prs` setting is on, watch stays on
  so subsequent pushes still trigger re-reviews while the PR is in draft.

* feat(reviewer): skip "no issues found" comment on empty re-reviews

A re-review run with no new findings to surface no longer posts another
"Open SWE Review: No issues found" comment on the PR. The "no issues"
summary now only appears on the first review of a PR — matching Devin's
behavior, where subsequent reviews are silent unless there's something
new to flag.

Resolved-thread reconciliation and ``last_reviewed_sha`` persistence
still happen on the skipped path, so findings the user just fixed still
get their GitHub threads marked resolved, and the next push event sees
an up-to-date dedup SHA.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-22 13:36:14 -07:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
Johannes du Plessis
834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00
langsmith-forge[bot]
74f5df5df8
fix: publish_review tool returns generic "Failed to POST PR review" without GitHub API status/body, agent retries with no signal (#1299)
* fix(reviewer): surface HTTP status and body excerpt for non-dict GitHub PR review responses

When post_pull_request_review received a non-dict body, it returned
None and publish_review surfaced a generic 'Failed to POST PR review'
string with no signal for the agent to adapt — leading to blind
retries with permuted cap/severity_threshold args.

Now the non-dict-body path mirrors the existing HTTPStatusError /
HTTPError paths: it returns {'_error': 'HTTP <status>: non-dict
response body: <excerpt>'} so the user-facing tool can include the
underlying detail. The bare-None branch in publish_review.py is kept
as a defensive guard with a clearer message.

* ci: apply ruff format to reviewer_publish.py

Collapse the multi-line return dict into a single line so it matches the
output of `ruff format`, unblocking the Agent lint / format-check CI jobs.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: LangSmith Issues Agent <issues-agent@langsmith.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-12 13:55:39 -07:00
Johannes du Plessis
892347041f
feat(open-swe): post review summary to Slack on first review (#1258)
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.

The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
2026-05-07 16:11:36 -07:00
Johannes du Plessis
88a55057ea
fix(open-swe): host-format the review summary, drop agent prose (#1257)
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.

Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
2026-05-07 15:41:40 -07:00
Johannes du Plessis
03d9f23645
fix(open-swe): always post a review summary, even with no findings (#1256)
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable

Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.

Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
  (with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
  already-reviewed head_sha)

* fix(reviewer): always post a summary review, even with no findings

The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.

- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
  findings list is empty, vs. `**No issues at or above \`<sev>\`
  severity.**` with hidden count when only sub-threshold findings
  exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
  body is meaningful; calls out specifically not to skip on a clean PR.
2026-05-07 22:16:20 +00:00
Johannes du Plessis
378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00