open-swe/agent/reviewer_diff.py

240 lines
8.3 KiB
Python
Raw Normal View History

feat: implement reviewer findings, publish_review, and watch mode (#1253) * feat: implement reviewer findings, publish_review, and watch mode Build out the reviewer agent end-to-end against the design in REVIEWER_DESIGN.md: - Findings as first-class state on the reviewer thread metadata (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line ranges, suggestion text for ```suggestion blocks, github_review_comment_id for cross-run reconciliation, diff_hunk for UI rendering. Thread-level metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a future frontend can list reviewer threads via the langgraph SDK. - Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff, compute_diff_line_set for in-diff validation, extract_diff_hunk for caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA diffs against the prepped repo. - Tools: `add_finding` (validates against the diff line set so out-of-diff ranges fail at creation, not at GitHub-publish), `update_finding`, `list_findings`, `publish_review`. The reviewer agent's tool list is swapped from `[]` (direct shell `gh api` calls) to these four. - Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`): one POST /reviews call with body + inline comments + ```suggestion blocks, per-comment IDs stored back on findings, GraphQL `resolveReviewThread` fired for findings transitioning open->resolved on a re-review. - Reviewer graph: deterministic clone-or-fetch + checkout in the factory before the agent's first model call (warm- and cold-path symmetric); computed diff and in-diff line set passed via runnable config; system prompt rewritten for the single-evolving-findings model, severity ladder, in-diff-only discipline, and watch-mode reconciliation flow. - Watch mode in webapp.py: `push` event + `pull_request` closed/reopened added to supported events. New `process_github_push_event` resolves the open PR for the pushed branch, gates on the reviewer thread's `watch` flag, builds a re-review configurable, and triggers a run on the same canonical thread. `process_github_pr_close` toggles watch on closed/reopened. `set_reviewer_thread_metadata` is called on first review to install `kind=reviewer` + PR identity + watch=True. - Eval harness: target.py now extracts `add_finding` calls (mapped to the legacy {file, line, body, severity} shape the judge expects) and passes the right configurable so the prep step has base/head SHAs. - Tests: new unit suites for findings helpers, diff parsing, finding tools, publish rendering + GraphQL resolve, and watch-mode webhook handlers (push triggers re-review only when watching, idempotent on unchanged head SHA, PR close disables watch). Updated existing reviewer-webhook tests to mock `set_reviewer_thread_metadata`. - REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context into REVIEWER_DESIGN.md. * fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL Address PR #1253 review findings: - compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag (`option no-prefix takes no value` — every prep run was failing silently and the agent saw an empty diff). - compute_diff_in_sandbox grew a `merge_base` flag. First-review path now uses three-dot `base...head` (the merge-base diff GitHub renders on Files-changed) so we don't pick up changes that landed on the base branch after the PR diverged. Re-review delta keeps two-dot `last_reviewed_sha..head` since that's exactly the new commits. - publish_review skips findings that already carry `github_review_comment_id`. Without this, watched re-reviews re-posted every previously surfaced finding, and only the most-recent duplicate's id would later resolve when the issue got addressed. - fetch_review_comments URL now includes `{pull_number}` — `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments` is the canonical endpoint; the old form 404s, so comment ids were never stored and watch-mode resolution couldn't run. Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix` flag in the executed command, and that publish_review does not re-post findings whose `github_review_comment_id` is set. * fix(reviewer): default publish cap from 15 to 4 A clean PR with one critical issue padded out by three lower-severity findings is fine; fifteen is review spam. The agent can override per call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
"""Diff utilities for the reviewer agent.
The reviewer needs three things from a PR diff:
1. The set of (file, line) tuples that are part of the diff, so ``add_finding``
can validate at creation time rather than at GitHub-publish time.
2. The hunk text relevant to a given (file, start_line, end_line) range, so we
can stash it on the Finding (``diff_hunk``) for rendering in the future UI
without re-fetching from GitHub or the (evictable) sandbox.
3. A way to compute the diff in the sandbox between two SHAs, used both on
first review (``base_sha..head_sha``) and on watched re-review
(``last_reviewed_sha..new_head_sha``).
"""
from __future__ import annotations
import asyncio
import logging
import re
from dataclasses import dataclass
from typing import TYPE_CHECKING
if TYPE_CHECKING:
from deepagents.backends.protocol import SandboxBackendProtocol
logger = logging.getLogger(__name__)
_DIFF_FILE_HEADER_RE = re.compile(r"^diff --git a/(?P<a>.+?) b/(?P<b>.+?)$")
_HUNK_HEADER_RE = re.compile(
r"^@@ -(?P<old_start>\d+)(?:,(?P<old_count>\d+))? "
r"\+(?P<new_start>\d+)(?:,(?P<new_count>\d+))? @@"
)
@dataclass(frozen=True)
class DiffHunk:
"""One hunk for one file in a unified diff.
``new_start``/``new_end`` are inclusive 1-based line numbers in the
post-PR (RIGHT side) file. ``body`` is the raw hunk text including the
``@@`` header — what gets stored on a Finding's ``diff_hunk``.
"""
file: str
new_start: int
new_end: int
old_start: int
old_end: int
body: str
@dataclass(frozen=True)
class FileDiff:
"""All hunks for one file in a unified diff."""
file: str
hunks: tuple[DiffHunk, ...]
def parse_unified_diff(diff_text: str) -> list[FileDiff]:
"""Parse a unified diff into per-file hunk records.
Skips ``--- ``/``+++ `` and binary file markers. Returns one ``FileDiff``
per file with at least one hunk; files with no hunks (e.g., pure renames)
are dropped.
"""
files: list[FileDiff] = []
lines = diff_text.splitlines()
i = 0
while i < len(lines):
header_match = _DIFF_FILE_HEADER_RE.match(lines[i])
if not header_match:
i += 1
continue
file_path = header_match.group("b")
i += 1
# Skip metadata lines until first hunk or next file header
hunks: list[DiffHunk] = []
current_hunk_lines: list[str] = []
current_meta: tuple[int, int, int, int] | None = None
while i < len(lines) and not _DIFF_FILE_HEADER_RE.match(lines[i]):
line = lines[i]
hunk_match = _HUNK_HEADER_RE.match(line)
if hunk_match:
if current_meta is not None and current_hunk_lines:
hunks.append(
DiffHunk(
file=file_path,
old_start=current_meta[0],
old_end=current_meta[1],
new_start=current_meta[2],
new_end=current_meta[3],
body="\n".join(current_hunk_lines),
)
)
old_start = int(hunk_match.group("old_start"))
old_count = int(hunk_match.group("old_count") or "1")
new_start = int(hunk_match.group("new_start"))
new_count = int(hunk_match.group("new_count") or "1")
# End line is inclusive; if count is 0 (deletion-only), end == start
old_end = old_start + max(old_count - 1, 0)
new_end = new_start + max(new_count - 1, 0)
current_meta = (old_start, old_end, new_start, new_end)
current_hunk_lines = [line]
elif current_meta is not None:
current_hunk_lines.append(line)
i += 1
if current_meta is not None and current_hunk_lines:
hunks.append(
DiffHunk(
file=file_path,
old_start=current_meta[0],
old_end=current_meta[1],
new_start=current_meta[2],
new_end=current_meta[3],
body="\n".join(current_hunk_lines),
)
)
if hunks:
files.append(FileDiff(file=file_path, hunks=tuple(hunks)))
return files
def compute_diff_line_set(diff_text: str) -> dict[str, set[int]]:
"""Return ``{file: {line, ...}}`` for the new-side lines covered by the diff.
A ``Finding`` whose ``(file, start_line..end_line)`` range falls outside
this set cannot be rendered as an inline GitHub review comment, so
``add_finding`` rejects it.
"""
out: dict[str, set[int]] = {}
for file_diff in parse_unified_diff(diff_text):
lines = out.setdefault(file_diff.file, set())
for hunk in file_diff.hunks:
for line in range(hunk.new_start, hunk.new_end + 1):
lines.add(line)
return out
def extract_diff_hunk(
diff_text: str,
file: str,
start_line: int | None,
end_line: int | None,
) -> str | None:
"""Extract the hunk body covering ``file:start_line..end_line``.
Returns ``None`` if no hunk overlaps. For file-level findings (both lines
None) returns the first hunk in the file as best-effort context.
"""
file_diffs = [fd for fd in parse_unified_diff(diff_text) if fd.file == file]
if not file_diffs:
return None
hunks = file_diffs[0].hunks
if not hunks:
return None
if start_line is None or end_line is None:
return hunks[0].body
for hunk in hunks:
if hunk.new_start <= end_line and start_line <= hunk.new_end:
return hunk.body
return None
def is_range_in_diff(
line_set: dict[str, set[int]],
file: str,
start_line: int | None,
end_line: int | None,
) -> bool:
"""Return True if every line in ``start_line..end_line`` is in the diff.
File-level findings (both None) are always allowed.
"""
if start_line is None and end_line is None:
return True
if start_line is None or end_line is None:
return False
file_lines = line_set.get(file)
if not file_lines:
return False
return all(line in file_lines for line in range(start_line, end_line + 1))
async def compute_diff_in_sandbox(
sandbox_backend: SandboxBackendProtocol,
work_dir: str,
base_ref: str,
head_ref: str,
*,
merge_base: bool = False,
) -> str:
"""Run ``git diff`` inside the sandbox and return its stdout.
Refs can be SHAs or branch names. Caller is responsible for ensuring both
refs exist locally (e.g., having fetched the PR head).
Args:
merge_base: When ``True``, use three-dot ``base...head`` (the merge-base
diff — what GitHub shows on the PR's "Files changed" tab). Use this
for first review so we don't pick up changes that landed on the
base branch after the PR diverged. When ``False``, use two-dot
``base..head`` — appropriate for re-review deltas where ``base`` is
the previously reviewed SHA and we want exactly the commits added
since.
"""
operator = "..." if merge_base else ".."
cmd = f"cd {work_dir} && git diff --no-color {base_ref}{operator}{head_ref}"
result = await asyncio.to_thread(sandbox_backend.execute, cmd)
exit_code = getattr(result, "exit_code", None)
if exit_code not in (0, None):
output = _stdout_from_result(result)
raise RuntimeError(
f"git diff failed (exit {exit_code}) for "
f"{base_ref}{operator}{head_ref} in {work_dir}. Output:\n{output}"
)
feat: implement reviewer findings, publish_review, and watch mode (#1253) * feat: implement reviewer findings, publish_review, and watch mode Build out the reviewer agent end-to-end against the design in REVIEWER_DESIGN.md: - Findings as first-class state on the reviewer thread metadata (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line ranges, suggestion text for ```suggestion blocks, github_review_comment_id for cross-run reconciliation, diff_hunk for UI rendering. Thread-level metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a future frontend can list reviewer threads via the langgraph SDK. - Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff, compute_diff_line_set for in-diff validation, extract_diff_hunk for caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA diffs against the prepped repo. - Tools: `add_finding` (validates against the diff line set so out-of-diff ranges fail at creation, not at GitHub-publish), `update_finding`, `list_findings`, `publish_review`. The reviewer agent's tool list is swapped from `[]` (direct shell `gh api` calls) to these four. - Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`): one POST /reviews call with body + inline comments + ```suggestion blocks, per-comment IDs stored back on findings, GraphQL `resolveReviewThread` fired for findings transitioning open->resolved on a re-review. - Reviewer graph: deterministic clone-or-fetch + checkout in the factory before the agent's first model call (warm- and cold-path symmetric); computed diff and in-diff line set passed via runnable config; system prompt rewritten for the single-evolving-findings model, severity ladder, in-diff-only discipline, and watch-mode reconciliation flow. - Watch mode in webapp.py: `push` event + `pull_request` closed/reopened added to supported events. New `process_github_push_event` resolves the open PR for the pushed branch, gates on the reviewer thread's `watch` flag, builds a re-review configurable, and triggers a run on the same canonical thread. `process_github_pr_close` toggles watch on closed/reopened. `set_reviewer_thread_metadata` is called on first review to install `kind=reviewer` + PR identity + watch=True. - Eval harness: target.py now extracts `add_finding` calls (mapped to the legacy {file, line, body, severity} shape the judge expects) and passes the right configurable so the prep step has base/head SHAs. - Tests: new unit suites for findings helpers, diff parsing, finding tools, publish rendering + GraphQL resolve, and watch-mode webhook handlers (push triggers re-review only when watching, idempotent on unchanged head SHA, PR close disables watch). Updated existing reviewer-webhook tests to mock `set_reviewer_thread_metadata`. - REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context into REVIEWER_DESIGN.md. * fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL Address PR #1253 review findings: - compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag (`option no-prefix takes no value` — every prep run was failing silently and the agent saw an empty diff). - compute_diff_in_sandbox grew a `merge_base` flag. First-review path now uses three-dot `base...head` (the merge-base diff GitHub renders on Files-changed) so we don't pick up changes that landed on the base branch after the PR diverged. Re-review delta keeps two-dot `last_reviewed_sha..head` since that's exactly the new commits. - publish_review skips findings that already carry `github_review_comment_id`. Without this, watched re-reviews re-posted every previously surfaced finding, and only the most-recent duplicate's id would later resolve when the issue got addressed. - fetch_review_comments URL now includes `{pull_number}` — `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments` is the canonical endpoint; the old form 404s, so comment ids were never stored and watch-mode resolution couldn't run. Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix` flag in the executed command, and that publish_review does not re-post findings whose `github_review_comment_id` is set. * fix(reviewer): default publish cap from 15 to 4 A clean PR with one critical issue padded out by three lower-severity findings is fine; fifteen is review spam. The agent can override per call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
return _stdout_from_result(result)
def _stdout_from_result(result: object) -> str:
"""Best-effort extraction of stdout from a sandbox execute() result.
Different sandbox providers return different shapes; this normalizes them.
"""
if isinstance(result, str):
return result
if isinstance(result, dict):
for key in ("stdout", "output", "text"):
value = result.get(key)
if isinstance(value, str):
return value
stdout = getattr(result, "stdout", None)
if isinstance(stdout, str):
return stdout
text = getattr(result, "text", None)
if isinstance(text, str):
return text
return ""