* fix: include reviewer trace links
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: move reviewer trace-link toggle to dashboard
Replace the OPEN_SWE_REVIEW_TRACE_LINK_ENABLED env var with a team-level
'Trace Links' toggle in the Open SWE Review dashboard tab. The toggle is
read per-publish via get_team_review_trace_links_enabled(); the per-run
review_trace_link_enabled config override still forces it off.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: structured reviewer comments + auto resolution comments
Restructure inline review comment bodies (severity emoji, bold title from
the first line, line reference, feedback footer) without duplicating the
first description line, and post an automatic resolution/dismissal comment
to the GitHub thread when a finding is resolved or dismissed.
- Centralize render_resolution_comment in reviewer_publish; fix a crash when
last_reconciliation_note is None and drop the misleading generic fallback.
- Post the resolution comment in resolve_finding_thread (the normal
update_finding path), not only in publish_review, so it actually fires on
re-review. Dedupe via github_posted_resolution_comment_ids.
- Add pytest coverage for rendering and the resolution-comment flow.
* fix: post resolution comment to every closed thread in resolve_finding_thread
Per-thread iteration (matching _resolve_threads_for_resolved_findings) so
duplicate threads after the first also receive the resolved/dismissed
explanation before being closed.
* fix: dedupe reviewer comments from PR state
Use GitHub review-thread markers to repair reviewer publication state before posting or resolving findings, so re-reviews do not duplicate comments and resolved findings close all matching PR threads.
* fix: require all duplicate reviewer threads resolved
Avoid treating a marker-backed finding as resolved when only one duplicate thread is outdated while another matching thread remains open.
* publish_review: drop unresolvable findings and retry once on GitHub 422
GitHub returns 422 with 'Path could not be resolved' or 'Line could not be
resolved' when an inline comment anchors to a file/line not in the PR diff.
Previously the agent retried publish_review with byte-identical args
multiple times before draining to skipped_empty_re_review=true, silently
losing findings.
- reviewer_publish.post_pull_request_review: parse 422 body and tag with
_error_kind='unresolved_anchor' plus _raw_errors so callers can act.
- tools/publish_review._publish_review_async: when that signal fires,
cross-check each finding's range against the run config's diff_line_set,
drop the bad ones, and re-POST once with only the valid findings. Return
unresolvable_findings + hint so the agent calls update_finding instead of
retrying the same payload.
- reviewer.py: one-line prompt addendum telling the agent that
unresolvable_findings means update_finding, not retry.
- tests: cover 422 tagging (path + line), the drop-and-retry success path,
the retry-still-fails path, and the don't-blind-retry path when no
diff_line_set is available.
* publish_review: fetch PR diff on demand for 422 retry filter
Reviewer runs clear configurable['diff_line_set'] before the agent
starts, so the unresolved-anchor retry path had no diff data to filter
against — in the reachable production case it dropped nothing and
returned success=False with empty unresolvable_findings, losing the
otherwise-valid comments.
Fall back to fetching the PR's unified diff via the GitHub REST API
and recomputing the line set on the fly when no cached set is
available. The cached set is still preferred when present.
---------
Co-authored-by: issues-agent <issues-agent@langchain.dev>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix(reviewer): inject existing PR review threads into reviewer context
The reviewer agent was filing the same inline comment on every re-review
because it only saw findings recorded on its own thread metadata — not
the live PR review-thread state on GitHub. When a previous finding was
still open (code unchanged, or a human reply explained it), the agent
rediscovered the same defect on the next push and called `add_finding`
again, producing duplicate comments.
This change fetches the PR's review threads (across all reviewers, with
replies and isResolved status) via GraphQL and renders them into the
first-review and re-review contexts as a "Pre-existing PR review
threads" block. The system prompt now lists overlap with that block as
a hard "Do NOT file" rule, and treats threads addressed by a human
reply as resolved.
This also gives the reviewer comment-awareness on its very first run
on a PR, so it skips findings already raised by another reviewer or
bot.
* fix(reviewer): wrap PR review threads in untrusted-data XML block
Addresses the reviewer comment on this PR
(https://github.com/langchain-ai/open-swe/pull/1331#discussion_r3295497533):
PR review comment bodies are attacker-controlled (anyone who can comment
on the PR can put anything in them), and they were being concatenated
into the reviewer's system prompt with instruction-priority.
Switches the existing-threads section from a Markdown block to an XML
data block:
<pr_review_threads>
<thread location="path:line" status="open">
<comment author="open-swe[bot]">
<body>...</body>
</comment>
<comment author="romain-priour-lc">
<body>We added defaults in the template</body>
</comment>
</thread>
</pr_review_threads>
The system prompt now explicitly names the wrapper, tells the agent that
everything inside it is untrusted data from the PR (not instructions),
and that prompt-injection payloads inside bodies must be disregarded.
We keep the bodies so the agent can actually read engineer replies —
that's the whole point of comment-awareness — but they're delimited as
data, not concatenated as prose. Modern frontier models are well-trained
to honor this contract.
Additional defenses:
- Author logins are validated against the GitHub username grammar; any
unexpected value is rendered as "unknown" so the `author` attribute
can't smuggle freeform text.
- Literal closing tags (`</body>`, `</pr_review_threads>`, etc.) in
bodies are neutered so a body can't break out of its wrapper.
- Body length is capped at 4000 chars per comment to bound the prompt.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix(reviewer): surface HTTP status and body excerpt for non-dict GitHub PR review responses
When post_pull_request_review received a non-dict body, it returned
None and publish_review surfaced a generic 'Failed to POST PR review'
string with no signal for the agent to adapt — leading to blind
retries with permuted cap/severity_threshold args.
Now the non-dict-body path mirrors the existing HTTPStatusError /
HTTPError paths: it returns {'_error': 'HTTP <status>: non-dict
response body: <excerpt>'} so the user-facing tool can include the
underlying detail. The bare-None branch in publish_review.py is kept
as a defensive guard with a clearer message.
* ci: apply ruff format to reviewer_publish.py
Collapse the multi-line return dict into a single line so it matches the
output of `ruff format`, unblocking the Agent lint / format-check CI jobs.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: LangSmith Issues Agent <issues-agent@langsmith.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]
Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* webapp: forward installation-token expiry to reviewer cache writes
The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
post_pull_request_review used to swallow the HTTPError and return None,
so the tool result was just "Failed to POST PR review" with no status
or body — the actual cause (e.g. 422 invalid inline comment, 404 app
not installed) only lived in logs. Capture status + body and propagate
into the tool result so the agent (and traces) can see why.
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.
Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable
Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.
Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
(with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
already-reviewed head_sha)
* fix(reviewer): always post a summary review, even with no findings
The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.
- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
findings list is empty, vs. `**No issues at or above \`<sev>\`
severity.**` with hidden count when only sub-threshold findings
exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
body is meaningful; calls out specifically not to skip on a clean PR.
* feat: implement reviewer findings, publish_review, and watch mode
Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:
- Findings as first-class state on the reviewer thread metadata
(`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
ranges, suggestion text for ```suggestion blocks, github_review_comment_id
for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
compute_diff_line_set for in-diff validation, extract_diff_hunk for
caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
ranges fail at creation, not at GitHub-publish), `update_finding`,
`list_findings`, `publish_review`. The reviewer agent's tool list is
swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
one POST /reviews call with body + inline comments + ```suggestion blocks,
per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
before the agent's first model call (warm- and cold-path symmetric);
computed diff and in-diff line set passed via runnable config; system
prompt rewritten for the single-evolving-findings model, severity ladder,
in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
added to supported events. New `process_github_push_event` resolves the
open PR for the pushed branch, gates on the reviewer thread's `watch`
flag, builds a re-review configurable, and triggers a run on the same
canonical thread. `process_github_pr_close` toggles watch on
closed/reopened. `set_reviewer_thread_metadata` is called on first
review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
legacy {file, line, body, severity} shape the judge expects) and passes
the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
publish rendering + GraphQL resolve, and watch-mode webhook handlers
(push triggers re-review only when watching, idempotent on unchanged
head SHA, PR close disables watch). Updated existing reviewer-webhook
tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
into REVIEWER_DESIGN.md.
* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL
Address PR #1253 review findings:
- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
(`option no-prefix takes no value` — every prep run was failing
silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
now uses three-dot `base...head` (the merge-base diff GitHub renders
on Files-changed) so we don't pick up changes that landed on the base
branch after the PR diverged. Re-review delta keeps two-dot
`last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
`github_review_comment_id`. Without this, watched re-reviews
re-posted every previously surfaced finding, and only the most-recent
duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
`/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
is the canonical endpoint; the old form 404s, so comment ids were
never stored and watch-mode resolution couldn't run.
Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.
* fix(reviewer): default publish cap from 15 to 4
A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.