Commit graph

18 commits

Author SHA1 Message Date
Johannes du Plessis
3e105f5027
fix: Reviews tab — anchored finding card, paginated list, file tree truncation (#1507)
* fix: Reviews tab — anchored finding card, paginated list, file tree truncation

- Finding card now tracks the diff anchor while scrolling instead of staying
  frozen in the viewport; auto-hides when its diff card collapses (including
  collapse via mark-as-viewed) and on click outside
- Checks section capped with max height + scroll
- /reviews paginated (page size 20) with has_more, filtered to the current
  user's PRs by default with an All toggle; PR author login now stored in
  reviewer thread metadata
- File tree truncation marker overlapped filenames because the sidebar bg
  was transparent; use the opaque sidebar color

* fix: finding card tracks anchor 1:1 while scrolling

Drop the vertical viewport clamp — it pinned the card at the clamp
boundary while the highlighted lines kept scrolling, breaking the
attachment.

* fix: anchor finding card with Base UI popover

Replace manual fixed-position tracking (laggy: setState per scroll
frame) with a Popover anchored to the finding's diff row. Floating UI
tracks the anchor outside React renders, so the card moves 1:1 with
the content and scrolls out of view with it. Unanchored findings keep
the fixed top-right card.

* fix: lock finding card to diff scroll

Replace the Base UI popover (async repositioning, paints a frame behind
native scroll) with a card absolutely positioned inside the scroll
container, so it scrolls with the diff in the same compositor frame.
Scroll moves to the ReviewBody root, side panel becomes sticky.
Position recomputes only on layout shifts via ResizeObserver.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:52:20 -07:00
Johannes du Plessis
8461979b0d
fix: honest publish_review reporting + structured thread-not-found errors (#1481)
* fix: honest publish_review reporting + structured thread-not-found errors

- Document skipped_empty_re_review and dry_run in the publish_review
  docstring and add a closing-summary contract to the reviewer prompt so
  the agent never claims a review was published when review_id is null.
- Raise ReviewerThreadMissingError from replace_findings on SDK
  NotFoundError; add_finding/update_finding/publish_review return a
  structured do-not-retry result instead of raising, so the agent reports
  the blocker after one failure instead of retrying 10-30 times.

* fix: translate thread 404s across all reviewer tool boundaries

get_thread_metadata now raises ReviewerThreadMissingError instead of
swallowing a missing thread as {} (which produced misleading 'No finding
found' results), set_reviewer_thread_metadata translates the SDK 404 the
same way, and every reviewer tool entrypoint (add/update/list findings,
publish_review incl. eval dry-run, resolve/reply thread) returns the
structured do-not-retry result.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:44:13 -07:00
Johannes du Plessis
7cd882bb67
fix: make review publish idempotent (partial-failure recovery) (#1477)
Publish had partial-failure windows that double-posted summaries or
corrupted findings state. This hardens the recovery paths:

- open_swe_review_exists is now tri-state (True/False/None). On a
  pagination/API failure it returns None ("unknown") instead of False,
  and the empty-summary dedup keys off the durable last_reviewed_sha
  before consulting GitHub, so a transient failure never double-posts a
  "no issues found" summary.
- Comment-id backfill matches strictly on the embedded open-swe marker;
  the colliding (path, line, body) fallback is gone, so similar findings
  no longer share a comment id and break resolve-on-fix.
- Review-id and comment-id stamping collapse into one guarded
  read-modify-write (re-reads latest before writing), removing the
  half-stamped intermediate states the prior multi-write flow left open.
- New mutate_findings primitive centralizes read-modify-write so finding
  updates operate on the freshest persisted list and skip no-op writes.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:17:31 -07:00
Johannes du Plessis
f64ab2bcd7
feat: surface out-of-diff findings in a collapsed dropdown (#1427)
* fix: stop reviewer retrying out-of-diff findings

add_finding rejects findings anchored outside the PR diff, but the agent
retried the same finding 2-3x with adjacent line ranges before giving up,
burning a model turn each. Add a reviewer-prompt recovery block telling the
agent the rejection is authoritative (drop or re-anchor to a + line, don't
retry adjacent), and enrich the rejection payload with nearby in-diff line
ranges for the file so a single re-anchor needs no guessing.

* feat: surface out-of-diff findings in a collapsed dropdown

Instead of rejecting findings anchored outside the PR diff, accept them
(marked in_diff=false) and surface them in a collapsed <details> section of
the review summary, Devin-style. Inline comments stay reserved for in-diff
findings; out-of-diff are severity-gated and capped the same way.

Re-review normally suppresses the empty summary, but now makes an exception
when there are new out-of-diff findings to surface. Surfaced out-of-diff
findings carry a github_review_id so they aren't reposted on later pushes.

Supersedes the earlier 'drop/re-anchor out-of-diff' prompt guidance.

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-05 10:41:48 -07:00
Brace Sproul
2070a770c2
fix: stop storing GitHub tokens in metadata (#1405)
* fix: stop storing GitHub tokens in metadata

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: bound in-process GitHub token cache with 24h TTL + sweep

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-04 09:33:51 -07:00
Johannes du Plessis
1d4f1aed33
fix: reviewer publishes against stale head_sha on mid-run re-review (#1393)
* fix: resolve reviewer head_sha from thread metadata, not frozen run config

A push that lands while a reviewer run is in flight is delivered as a
queued message into that run. The run's configurable is frozen at
creation, so its head_sha still names the commit the run was created for
— not the commit just pushed. publish_review then anchored the GitHub
review to the stale commit and regressed last_reviewed_sha to it, and
add_finding/update_finding stamped findings with the stale SHA.

Persist the current head in thread metadata at every reviewer dispatch
(both the ready-for-review and push paths, before they branch to create
a run or queue a message), and add resolve_review_head_sha() which
prefers the metadata head over the run config. Wire it into
publish_review (review commit_id + last_reviewed_sha), add_finding
(first_seen_sha) and update_finding (last_confirmed_sha). Falls back to
the run config when metadata carries no head (first review, eval, tests).

* fix: persist head_sha in manual review dispatch (trigger_pr_review_from_ref)

resolve_review_head_sha prefers metadata[head_sha] over the run config,
and the push/ready dispatchers write it — but trigger_pr_review_from_ref
(Slack/GitHub @open-swe review, request_pr_review tool) created a run
with a freshly-fetched config head while leaving metadata's head stale
from a prior dispatch. A manual re-review at a newer commit would then
resolve to the old head and publish/advance findings against it.

Persist head_sha in that dispatch's metadata write too, so every
run-creating reviewer dispatch keeps metadata in sync with the head its
run targets. Caught by the Open SWE reviewer on this PR.
2026-06-03 11:38:56 -07:00
open-swe[bot]
a361ee8f2e
feat: let reviewer set comment titles (#1356)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-28 16:04:38 -07:00
open-swe[bot]
9998119921
feat: generate reviewer finding titles (#1355)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-28 15:28:55 -07:00
Johannes du Plessis
fee209601b
feat: structured reviewer comments + auto resolution comments (#1352)
* feat: structured reviewer comments + auto resolution comments

Restructure inline review comment bodies (severity emoji, bold title from
the first line, line reference, feedback footer) without duplicating the
first description line, and post an automatic resolution/dismissal comment
to the GitHub thread when a finding is resolved or dismissed.

- Centralize render_resolution_comment in reviewer_publish; fix a crash when
  last_reconciliation_note is None and drop the misleading generic fallback.
- Post the resolution comment in resolve_finding_thread (the normal
  update_finding path), not only in publish_review, so it actually fires on
  re-review. Dedupe via github_posted_resolution_comment_ids.
- Add pytest coverage for rendering and the resolution-comment flow.

* fix: post resolution comment to every closed thread in resolve_finding_thread

Per-thread iteration (matching _resolve_threads_for_resolved_findings) so
duplicate threads after the first also receive the resolved/dismissed
explanation before being closed.
2026-05-28 12:44:00 -07:00
Johannes du Plessis
197d339df4
fix: Reconcile reviewer findings with PR threads (#1346)
* feat: reconcile reviewer findings with PR threads

* fix: harden reviewer finding reply handling

* fix: queue reviewer finding reply body

Ensure review-comment replies that arrive during an active reviewer run include the sanitized reply body in the queued reassessment prompt.

* fix: apply reviewer reply formatting

Apply the repository formatter so the reviewer reply handling fix passes CI format checks.
2026-05-27 17:26:08 -07:00
Johannes du Plessis
0d4d1c5a3b
fix: dedupe reviewer comments from PR state (#1341)
* fix: dedupe reviewer comments from PR state

Use GitHub review-thread markers to repair reviewer publication state before posting or resolving findings, so re-reviews do not duplicate comments and resolved findings close all matching PR threads.

* fix: require all duplicate reviewer threads resolved

Avoid treating a marker-backed finding as resolved when only one duplicate thread is outdated while another matching thread remains open.
2026-05-27 10:52:30 -07:00
Johannes du Plessis
5702a9d452
feat: reconcile reviewer comment lifecycle (#1332)
* feat: reconcile reviewer comment lifecycle

Track GitHub review threads for reviewer findings so re-reviews can resolve or reply to existing comments, and collect thumbs feedback on new review comments in LangSmith.

* fix: clarify reviewer comment lifecycle

* chore: apply reviewer formatting

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-26 16:24:34 -07:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
Johannes du Plessis
bdea1fb6da
feat(open-swe): always anchor findings to a single line (#1264)
* feat(open-swe): always anchor findings to a single line

GitHub renders multi-line review comment ranges as walls of context above
the comment, which buries the point. Match Devin Review's behavior and
always collapse `end_line` to `start_line` so each finding is anchored to
the single most relevant line.

Removes the (now unused) `MAX_FINDING_RANGE_LINES` cap and
`clip_finding_range` helper.

* fix(open-swe): collapse end_line before diff-range check
2026-05-07 20:26:59 -07:00
Johannes du Plessis
089db06135
feat(open-swe): collapse oversized finding ranges to start line (#1263)
Anchoring a finding to a 25-line range (e.g. an entire function) makes
GitHub render the whole block as context above the comment, burying the
review text. Cap ranges at 10 lines and collapse to the start line when
exceeded; small ranges (≤10 lines) still render as a multi-line anchor.
Updated the reviewer prompt to anchor tightly rather than range-select
the whole function.
2026-05-07 19:59:25 -07:00
Johannes du Plessis
0212a60f10
feat(open-swe): cap reviewer suggestions at 4 lines (#1262)
* feat(open-swe): cap reviewer suggestions at 4 lines

Long suggestion blocks read as the reviewer rewriting the code rather
than flagging an issue, which clutters PR comments. Steer the reviewer
toward description-only findings for non-trivial fixes, and enforce the
cap in `add_finding` / `update_finding` so suggestions over 4 lines are
dropped (the finding itself still publishes).

* fix(open-swe): don't clobber prior suggestion on over-cap update

`update_finding` was setting `suggestion=None` whenever clip_suggestion
dropped the input, which silently wiped any existing suggestion on the
finding. Distinguish the three cases: empty string clears, valid value
sets, over-cap value is rejected without touching the stored field.
2026-05-07 19:27:16 -07:00
Johannes du Plessis
892347041f
feat(open-swe): post review summary to Slack on first review (#1258)
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.

The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
2026-05-07 16:11:36 -07:00
Johannes du Plessis
378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00