open-swe/agent/middleware
Adam Moussa 0f0f616cd4
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214)
* feat(reviewer): explicit-request verdicts + shell verdict guard

Mention-triggered reviews that explicitly ask for a verdict now submit a
real APPROVE/REQUEST_CHANGES through publish_review; auto-reviews stay
advisory (COMMENT). Authorization is enforced in code: publish_review
honors a verdict only when the dispatching webhook set verdict_requested,
which only the explicit-mention path does.

- request_pr_review gains instructions (forwarded verbatim into an escaped
  requester_instructions data block) and request_verdict
- self-review guard downgrades verdicts on Open SWE-authored PRs; stale
  APPROVEs are best-effort dismissed when later findings land
- new PullRequestVerdictGuardMiddleware blocks gh pr review
  --approve/-a/--request-changes/-r, gh api, and curl verdict fallbacks on
  both the coding-agent and reviewer graphs
- shared escape helper moved to agent/utils/prompt_data.py

* fix(reviewer): harden verdict path against security-review findings

Adversarial security review (detector fan-out + proof-or-kill verifier)
of the verdict feature surfaced several verdict-integrity gaps; resolve
the confirmed ones:

- head-drift (high): a mid-run push moves the resolved head, so an APPROVE
  could anchor to an unreviewed commit. Downgrade any verdict to a comment
  when the resolved head differs from the reviewed head (verdict_ignored
  reason head_moved); the push's own re-review submits a fresh verdict.
- self-review fail-open: downgrade to comment when the PR author cannot be
  confirmed (author_unknown), and compare bot logins case-insensitively.
- verdict_submitted now reflects GitHub's returned review state, not just
  the event we asked for, so a coerced APPROVE isn't reported as submitted.
- an authorized verdict whose findings all anchor outside the diff now
  posts as a bodied review with zero inline comments instead of failing.
- add finding_reply to the shared data-block escape tag superset.
2026-07-20 15:28:00 -04:00
..
__init__.py feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214) 2026-07-20 15:28:00 -04:00
check_message_queue.py chore: sync upstream/main, defer #1621 modular webhooks (#81) 2026-06-30 16:45:19 -04:00
ensure_no_empty_msg.py chore: sync upstream/main, defer #1621 modular webhooks (#81) 2026-06-30 16:45:19 -04:00
exclude_tools.py feat: add reviewer graph + eval target wiring (#1241) 2026-05-06 10:15:58 -07:00
model_fallback.py feat: port model-fallback resilience from upstream (#1694, #1695) (#161) 2026-07-09 18:28:21 -04:00
notify_step_limit.py fix: notify users via Slack when agent hits model call step limit (#1204) 2026-05-01 14:24:25 -07:00
plan_mode.py chore: sync upstream/main, defer #1621 modular webhooks (#81) 2026-06-30 16:45:19 -04:00
pr_creation_guard.py feat: surface attributed PR creation failures (#180) 2026-07-13 14:45:19 -04:00
pr_verdict_guard.py feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214) 2026-07-20 15:28:00 -04:00
refresh_github_proxy.py fix: refresh sandbox GitHub proxy token before mid-run expiry (#1496) 2026-06-11 10:59:21 -07:00
refresh_slack_status.py feat: add Linear issue search tool (#1748) 2026-07-17 16:22:38 -04:00
repair_orphaned_tool_calls.py fix: repair orphaned tool calls before model calls (#1604) 2026-06-24 12:58:10 -07:00
sandbox_circuit_breaker.py fix: recover from mid-run sandbox death (#1274) 2026-05-08 12:55:36 -07:00
sanitize_fireworks_messages.py feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) 2026-07-09 14:44:15 -04:00
sanitize_openai_responses.py feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155) 2026-07-09 14:44:15 -04:00
sanitize_thinking_blocks.py feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62) 2026-06-29 15:57:19 -04:00
sanitize_tool_inputs.py fix: coerce malformed integer strings in read_file offset/limit params (#1216) 2026-05-01 14:29:48 -07:00
settle_review_check.py refactor: consolidate reviewer modules into agent/review/ 2026-07-17 13:52:03 -04:00
subdir_agents.py feat: re-land scoped AGENTS auto-load (#1684) + platform issue reporting tool (#1685) (#129) 2026-07-08 18:52:01 -04:00
task_retry.py feat: port durable dispatch hardening and startup latency improvements (#160) 2026-07-09 17:11:25 -04:00
timeout_wrapup.py feat: port durable dispatch hardening and startup latency improvements (#160) 2026-07-09 17:11:25 -04:00
tool_artifact.py feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475) 2026-06-11 09:54:35 -07:00
tool_error_handler.py fix: keep sandbox backend stable across recovery (#1294)w 2026-05-11 16:03:38 -07:00
workflow_push_guard.py fix(open-swe): port core GitHub-App scope fallback (#1701), workflows:write kept out of standing scope (#181) 2026-07-13 14:24:37 -04:00