open-swe/agent/tools/update_finding.py

81 lines
2.8 KiB
Python
Raw Normal View History

feat: implement reviewer findings, publish_review, and watch mode (#1253) * feat: implement reviewer findings, publish_review, and watch mode Build out the reviewer agent end-to-end against the design in REVIEWER_DESIGN.md: - Findings as first-class state on the reviewer thread metadata (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line ranges, suggestion text for ```suggestion blocks, github_review_comment_id for cross-run reconciliation, diff_hunk for UI rendering. Thread-level metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a future frontend can list reviewer threads via the langgraph SDK. - Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff, compute_diff_line_set for in-diff validation, extract_diff_hunk for caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA diffs against the prepped repo. - Tools: `add_finding` (validates against the diff line set so out-of-diff ranges fail at creation, not at GitHub-publish), `update_finding`, `list_findings`, `publish_review`. The reviewer agent's tool list is swapped from `[]` (direct shell `gh api` calls) to these four. - Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`): one POST /reviews call with body + inline comments + ```suggestion blocks, per-comment IDs stored back on findings, GraphQL `resolveReviewThread` fired for findings transitioning open->resolved on a re-review. - Reviewer graph: deterministic clone-or-fetch + checkout in the factory before the agent's first model call (warm- and cold-path symmetric); computed diff and in-diff line set passed via runnable config; system prompt rewritten for the single-evolving-findings model, severity ladder, in-diff-only discipline, and watch-mode reconciliation flow. - Watch mode in webapp.py: `push` event + `pull_request` closed/reopened added to supported events. New `process_github_push_event` resolves the open PR for the pushed branch, gates on the reviewer thread's `watch` flag, builds a re-review configurable, and triggers a run on the same canonical thread. `process_github_pr_close` toggles watch on closed/reopened. `set_reviewer_thread_metadata` is called on first review to install `kind=reviewer` + PR identity + watch=True. - Eval harness: target.py now extracts `add_finding` calls (mapped to the legacy {file, line, body, severity} shape the judge expects) and passes the right configurable so the prep step has base/head SHAs. - Tests: new unit suites for findings helpers, diff parsing, finding tools, publish rendering + GraphQL resolve, and watch-mode webhook handlers (push triggers re-review only when watching, idempotent on unchanged head SHA, PR close disables watch). Updated existing reviewer-webhook tests to mock `set_reviewer_thread_metadata`. - REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context into REVIEWER_DESIGN.md. * fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL Address PR #1253 review findings: - compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag (`option no-prefix takes no value` — every prep run was failing silently and the agent saw an empty diff). - compute_diff_in_sandbox grew a `merge_base` flag. First-review path now uses three-dot `base...head` (the merge-base diff GitHub renders on Files-changed) so we don't pick up changes that landed on the base branch after the PR diverged. Re-review delta keeps two-dot `last_reviewed_sha..head` since that's exactly the new commits. - publish_review skips findings that already carry `github_review_comment_id`. Without this, watched re-reviews re-posted every previously surfaced finding, and only the most-recent duplicate's id would later resolve when the issue got addressed. - fetch_review_comments URL now includes `{pull_number}` — `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments` is the canonical endpoint; the old form 404s, so comment ids were never stored and watch-mode resolution couldn't run. Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix` flag in the executed command, and that publish_review does not re-post findings whose `github_review_comment_id` is set. * fix(reviewer): default publish cap from 15 to 4 A clean PR with one critical issue padded out by three lower-severity findings is fine; fifteen is review spam. The agent can override per call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
"""Tool: ``update_finding``. Mutate an existing finding by id."""
from __future__ import annotations
import asyncio
from typing import Any
from langgraph.config import get_config
from ..reviewer_findings import (
get_thread_id_from_runtime,
update_finding_fields,
)
def update_finding(
finding_id: str,
status: str | None = None,
severity: str | None = None,
description: str | None = None,
suggestion: str | None = None,
note: str | None = None,
) -> dict[str, Any]:
"""Update fields on an existing finding.
Use this on a re-review run to mark an existing finding as resolved or
dismissed, or to revise its severity/description/suggestion if the new
commits changed the situation.
Args:
finding_id: The id returned by ``add_finding`` (or shown in the
``Existing findings`` block of the re-review user message).
status: New status (``open``, ``resolved``, ``dismissed``).
Use ``resolved`` when the new commits address the issue.
severity: New severity, if reassessing.
description: New description body, if revising.
suggestion: New replacement text. Pass an empty string to clear it.
note: Optional free-form note explaining the change. Persisted on the
finding under ``last_update_note``.
Returns:
Dictionary with ``success`` and (on success) the updated ``finding``.
"""
if status is not None and status not in {"open", "resolved", "dismissed"}:
return {"success": False, "error": f"Invalid status: {status}"}
if severity is not None and severity not in {
"informational",
"low",
"medium",
"high",
"critical",
}:
return {"success": False, "error": f"Invalid severity: {severity}"}
updates: dict[str, Any] = {}
if status is not None:
updates["status"] = status
if severity is not None:
updates["severity"] = severity
if description is not None:
updates["description"] = description
if suggestion is not None:
updates["suggestion"] = suggestion or None
if note is not None:
updates["last_update_note"] = note
config = get_config()
configurable = config.get("configurable", {}) if isinstance(config, dict) else {}
head_sha = configurable.get("head_sha", "") if isinstance(configurable, dict) else ""
if status == "open" and isinstance(head_sha, str) and head_sha:
updates["last_confirmed_sha"] = head_sha
if not updates:
return {"success": False, "error": "No fields provided to update"}
thread_id = get_thread_id_from_runtime()
updated = asyncio.run(update_finding_fields(thread_id, finding_id, updates))
if updated is None:
return {"success": False, "error": f"No finding found with id {finding_id}"}
return {"success": True, "finding": updated}