open-swe/agent/tools/publish_review.py
Johannes du Plessis 378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00

297 lines
10 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""Tool: ``publish_review``. Post the findings list to GitHub as a PR Review."""
from __future__ import annotations
import asyncio
from typing import Any
from langgraph.config import get_config
from ..reviewer_findings import (
Severity,
filter_findings_for_publish,
get_thread_id_from_runtime,
replace_findings,
set_reviewer_thread_metadata,
)
from ..reviewer_findings import (
list_findings as list_findings_async,
)
from ..reviewer_publish import (
fetch_review_comments,
fetch_review_thread_id_for_comment,
post_pull_request_review,
render_inline_comment_payload,
render_review_body,
resolve_review_thread,
)
from ..utils.github_token import get_github_token
def publish_review(
summary: str | None = None,
severity_threshold: str = "medium",
cap: int = 4,
) -> dict[str, Any]:
"""Post all current findings to the PR as a GitHub Review.
Call this once at the end of a review run, after you have finished adding
findings (and, on a re-review, after marking resolved findings via
``update_finding``). It will:
1. Read findings from the reviewer thread.
2. Filter to status=open and severity ≥ ``severity_threshold``, capped
at ``cap`` to avoid review spam.
3. POST a single GitHub PR Review with the eligible findings as inline
comments. ``finding.suggestion`` becomes a ```suggestion``` block
(the "Commit suggestion" UX).
4. Store the returned per-comment IDs back on each finding so a future
re-review can resolve those threads on GitHub when the issues are fixed.
5. For findings whose status moved ``open`` → ``resolved`` since the last
publish, resolve their existing GitHub review threads via the GraphQL
``resolveReviewThread`` mutation.
6. Update ``last_reviewed_sha`` on the thread to the current head SHA.
Args:
summary: Optional 1–2 sentence top-level take on the PR. Rendered as
the review body. Skip if you have nothing useful to say beyond
the per-finding comments.
severity_threshold: Lowest severity to surface to GitHub (default
``medium``). Lower-severity findings stay in state and surface in
the future UI but not on the PR.
cap: Maximum number of inline comments to publish (default 4).
Returns:
Dictionary with ``success``, ``review_id``, ``surfaced_count``,
``hidden_count``, ``resolved_thread_count``.
"""
if severity_threshold not in {"informational", "low", "medium", "high", "critical"}:
return {"success": False, "error": f"Invalid severity_threshold: {severity_threshold}"}
config = get_config()
configurable = config.get("configurable", {}) if isinstance(config, dict) else {}
repo_config = configurable.get("repo") if isinstance(configurable, dict) else None
pr_number = configurable.get("pr_number") if isinstance(configurable, dict) else None
head_sha = configurable.get("head_sha") if isinstance(configurable, dict) else None
if (
not isinstance(repo_config, dict)
or not repo_config.get("owner")
or not repo_config.get("name")
):
return {"success": False, "error": "Missing repo info in run config"}
if not isinstance(pr_number, int):
return {"success": False, "error": "Missing pr_number in run config"}
if not isinstance(head_sha, str) or not head_sha:
return {"success": False, "error": "Missing head_sha in run config"}
token = get_github_token()
if not token:
return {"success": False, "error": "No GitHub token available"}
return asyncio.run(
_publish_review_async(
owner=str(repo_config["owner"]),
repo=str(repo_config["name"]),
pr_number=pr_number,
head_sha=head_sha,
token=token,
summary=summary,
severity_threshold=_cast_severity(severity_threshold),
cap=cap,
)
)
def _cast_severity(value: str) -> Severity:
return value # type: ignore[return-value]
async def _publish_review_async(
*,
owner: str,
repo: str,
pr_number: int,
head_sha: str,
token: str,
summary: str | None,
severity_threshold: Severity,
cap: int,
) -> dict[str, Any]:
thread_id = get_thread_id_from_runtime()
findings = await list_findings_async(thread_id)
# Re-reviews only post NEW findings. Anything with a github_review_comment_id
# already lives on GitHub from a prior publish — reposting would create
# duplicate inline comments and break the resolve-on-fix flow (only
# whichever duplicate id we'd cache last would resolve later).
unpublished_findings = [
f for f in findings if not isinstance(f.get("github_review_comment_id"), int)
]
open_unpublished = [f for f in unpublished_findings if f.get("status", "open") == "open"]
eligible = filter_findings_for_publish(
unpublished_findings, severity_threshold=severity_threshold, cap=cap
)
inline_comments: list[dict[str, Any]] = []
eligible_with_payload: list[tuple[dict[str, Any], dict[str, Any]]] = []
for finding in eligible:
payload = render_inline_comment_payload(finding)
if payload is None:
continue
inline_comments.append(payload)
eligible_with_payload.append((dict(finding), payload))
review_body = render_review_body(
pr_number=pr_number,
surfaced_count=len(inline_comments),
total_open_count=len(open_unpublished),
severity_threshold=severity_threshold,
summary=summary,
)
review_id: int | None = None
if inline_comments or summary:
review_response = await post_pull_request_review(
owner=owner,
repo=repo,
pr_number=pr_number,
head_sha=head_sha,
body=review_body,
inline_comments=inline_comments,
token=token,
)
if review_response is None:
return {"success": False, "error": "Failed to POST PR review"}
review_id = review_response.get("id") if isinstance(review_response, dict) else None
if review_id is not None and inline_comments:
comment_records = await fetch_review_comments(
owner=owner,
repo=repo,
pr_number=pr_number,
review_id=review_id,
token=token,
)
await _store_comment_ids_on_findings(
thread_id=thread_id,
findings=findings,
eligible_with_payload=eligible_with_payload,
comment_records=comment_records,
)
resolved_thread_count = await _resolve_threads_for_resolved_findings(
owner=owner,
repo=repo,
pr_number=pr_number,
token=token,
findings=await list_findings_async(thread_id),
)
await set_reviewer_thread_metadata(thread_id, last_reviewed_sha=head_sha)
return {
"success": True,
"review_id": review_id,
"surfaced_count": len(inline_comments),
"hidden_count": max(len(open_unpublished) - len(inline_comments), 0),
"resolved_thread_count": resolved_thread_count,
}
async def _store_comment_ids_on_findings(
*,
thread_id: str,
findings: list[dict[str, Any]],
eligible_with_payload: list[tuple[dict[str, Any], dict[str, Any]]],
comment_records: list[dict[str, Any]],
) -> None:
"""Match returned GitHub comment ids back to the findings that produced them.
Match key is ``(path, line, body)`` since we don't have a server-side hint
pointing each REST comment to its source finding.
"""
by_key: dict[tuple[str, int, str], int] = {}
for record in comment_records:
path = record.get("path")
line = record.get("line") or record.get("original_line")
body = record.get("body", "")
comment_id = record.get("id")
if (
isinstance(path, str)
and isinstance(line, int)
and isinstance(body, str)
and isinstance(comment_id, int)
):
by_key[(path, line, body)] = comment_id
updated = False
findings_by_id = {f.get("id"): f for f in findings}
for finding_snapshot, payload in eligible_with_payload:
line_value = payload.get("line")
if not isinstance(line_value, int):
continue
key = (
str(payload.get("path", "")),
line_value,
str(payload.get("body", "")),
)
comment_id = by_key.get(key)
if comment_id is None:
continue
finding = findings_by_id.get(finding_snapshot.get("id"))
if finding is None:
continue
finding["github_review_comment_id"] = comment_id
updated = True
if updated:
await replace_findings(thread_id, list(findings_by_id.values()))
async def _resolve_threads_for_resolved_findings(
*,
owner: str,
repo: str,
pr_number: int,
token: str,
findings: list[dict[str, Any]],
) -> int:
"""Resolve GitHub review threads for findings that just transitioned to resolved.
A finding qualifies if:
- status == ``resolved``
- has a ``github_review_comment_id`` from a prior publish
- has not already been GitHub-resolved (tracked via
``github_thread_resolved`` flag we write back here)
"""
resolved_count = 0
mutated = False
for finding in findings:
if finding.get("status") != "resolved":
continue
comment_id = finding.get("github_review_comment_id")
if not isinstance(comment_id, int):
continue
if finding.get("github_thread_resolved"):
continue
thread_node_id = await fetch_review_thread_id_for_comment(
owner=owner,
repo=repo,
pr_number=pr_number,
review_comment_id=comment_id,
token=token,
)
if not thread_node_id:
continue
ok = await resolve_review_thread(thread_node_id=thread_node_id, token=token)
if ok:
finding["github_thread_resolved"] = True
mutated = True
resolved_count += 1
if mutated:
thread_id = get_thread_id_from_runtime()
await replace_findings(thread_id, findings)
return resolved_count