open-swe/agent/utils/prompt_data.py
Adam Moussa 0f0f616cd4
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214)
* feat(reviewer): explicit-request verdicts + shell verdict guard

Mention-triggered reviews that explicitly ask for a verdict now submit a
real APPROVE/REQUEST_CHANGES through publish_review; auto-reviews stay
advisory (COMMENT). Authorization is enforced in code: publish_review
honors a verdict only when the dispatching webhook set verdict_requested,
which only the explicit-mention path does.

- request_pr_review gains instructions (forwarded verbatim into an escaped
  requester_instructions data block) and request_verdict
- self-review guard downgrades verdicts on Open SWE-authored PRs; stale
  APPROVEs are best-effort dismissed when later findings land
- new PullRequestVerdictGuardMiddleware blocks gh pr review
  --approve/-a/--request-changes/-r, gh api, and curl verdict fallbacks on
  both the coding-agent and reviewer graphs
- shared escape helper moved to agent/utils/prompt_data.py

* fix(reviewer): harden verdict path against security-review findings

Adversarial security review (detector fan-out + proof-or-kill verifier)
of the verdict feature surfaced several verdict-integrity gaps; resolve
the confirmed ones:

- head-drift (high): a mid-run push moves the resolved head, so an APPROVE
  could anchor to an unreviewed commit. Downgrade any verdict to a comment
  when the resolved head differs from the reviewed head (verdict_ignored
  reason head_moved); the push's own re-review submits a fresh verdict.
- self-review fail-open: downgrade to comment when the PR author cannot be
  confirmed (author_unknown), and compare bot logins case-insensitively.
- verdict_submitted now reflects GitHub's returned review state, not just
  the event we asked for, so a coerced APPROVE isn't reported as submitted.
- an authorized verdict whose findings all anchor outside the diff now
  posts as a bodied review with zero inline comments instead of failing.
- add finding_reply to the shared data-block escape tag superset.
2026-07-20 15:28:00 -04:00

32 lines
1.1 KiB
Python

"""Escaping for untrusted text embedded in XML-wrapped prompt data blocks."""
from __future__ import annotations
import re
# Closing tags of every XML wrapper used for untrusted data blocks across the
# reviewer prompt and webhook-built run prompts. XML tolerates whitespace
# around the tag name (e.g. `</body >`, `</ body\n>`), so a literal
# `.replace()` of the canonical spelling alone is insufficient — we match each
# end tag whitespace-tolerantly and rewrite it to an inert, human-readable
# form. A shared superset is safe: escaping a closing tag that a given block
# doesn't use only neutralizes attacker-controlled text.
DATA_BLOCK_WRAPPER_TAGS = (
"pr_review_threads",
"thread",
"comment",
"body",
"pr_overview",
"title",
"finding_reply",
"requester_instructions",
)
_CLOSING_TAG_RE = re.compile(
r"</\s*(" + "|".join(DATA_BLOCK_WRAPPER_TAGS) + r")\s*>",
re.IGNORECASE,
)
def escape_for_data_block(text: str) -> str:
"""Neutralize closing tags so an attacker-controlled body can't break out."""
return _CLOSING_TAG_RE.sub(lambda m: f"</{m.group(1).lower()}_>", text)