Commit graph

13 commits

Author SHA1 Message Date
259329a971 fix(reviewer): preserve explicit verdict authority 2026-08-01 20:21:16 -04:00
98cde35812 refactor(reviewer): align verdict prompts with authorization 2026-08-01 20:21:16 -04:00
4fea5f21c5 fix(reviewer): enforce blocking review checks 2026-08-01 20:21:16 -04:00
39c768aa3f feat(reviewer): add verdict-aware review checks 2026-08-01 20:21:16 -04:00
9a2535a9ec fix(reviewer): clarify recorded verdict outcomes 2026-08-01 20:21:16 -04:00
404b854544 fix(reviewer): handle verdict publication edge cases 2026-08-01 20:21:16 -04:00
b2b5957233 fix(reviewer): harden verdict enforcement 2026-08-01 20:21:16 -04:00
1b8f8d480d feat(reviewer): add dispatch verdict authorization 2026-08-01 20:21:16 -04:00
Adam Moussa
c99bd78179
feat: clean-review auto-approve for unsolicited verdicts [BLOCKED — security] (#217)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): clean-review auto-approve for unsolicited publish_review verdicts

An unsolicited publish_review(verdict="approve") — a run dispatched
without verdict_requested — is now honored when the review has zero open
findings, so clean auto-reviews land a real APPROVE. With open findings
it downgrades to a comment review (verdict_ignored_reason=
"approve_with_open_findings"). request_changes stays explicit-request-
only; the self-review, head-moved, and author-unknown downgrades and the
shell verdict guard are unchanged. The reviewer base prompt now instructs
the clean-approve call on auto-reviews.

* chore(security): record accepted-risk suppressions for clean-review auto-approve

Two confirmed-HIGH findings from /sh-security-review on the clean-review
auto-approve change are accepted and deferred (Adam, 2026-07-21), tracked
in #218. Machine-recorded per the mandatory-security-review policy; the
revisit trigger is promotion from dev to main/prod.
2026-07-21 20:18:25 -04:00
Adam Moussa
0f0f616cd4
feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): explicit-request verdicts + shell verdict guard

Mention-triggered reviews that explicitly ask for a verdict now submit a
real APPROVE/REQUEST_CHANGES through publish_review; auto-reviews stay
advisory (COMMENT). Authorization is enforced in code: publish_review
honors a verdict only when the dispatching webhook set verdict_requested,
which only the explicit-mention path does.

- request_pr_review gains instructions (forwarded verbatim into an escaped
  requester_instructions data block) and request_verdict
- self-review guard downgrades verdicts on Open SWE-authored PRs; stale
  APPROVEs are best-effort dismissed when later findings land
- new PullRequestVerdictGuardMiddleware blocks gh pr review
  --approve/-a/--request-changes/-r, gh api, and curl verdict fallbacks on
  both the coding-agent and reviewer graphs
- shared escape helper moved to agent/utils/prompt_data.py

* fix(reviewer): harden verdict path against security-review findings

Adversarial security review (detector fan-out + proof-or-kill verifier)
of the verdict feature surfaced several verdict-integrity gaps; resolve
the confirmed ones:

- head-drift (high): a mid-run push moves the resolved head, so an APPROVE
  could anchor to an unreviewed commit. Downgrade any verdict to a comment
  when the resolved head differs from the reviewed head (verdict_ignored
  reason head_moved); the push's own re-review submits a fresh verdict.
- self-review fail-open: downgrade to comment when the PR author cannot be
  confirmed (author_unknown), and compare bot logins case-insensitively.
- verdict_submitted now reflects GitHub's returned review state, not just
  the event we asked for, so a coerced APPROVE isn't reported as submitted.
- an authorized verdict whose findings all anchor outside the diff now
  posts as a bodied review with zero inline comments instead of failing.
- add finding_reply to the shared data-block escape tag superset.
2026-07-20 15:28:00 -04:00
8d8d5bbbbf
fix: bind cached GitHub tokens to users (#1736)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 1ea0e600dcc234fa5a333c6f4b80b90e2e6679d3)

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-17 17:08:09 -04:00
ead210927d
fix: defensive copy in get_reviewer_agent and get_chat_agent [closes #1584] (#1742)
* fix: defensive copy in get_reviewer_agent and get_chat_agent [closes #1584]

Factory functions were mutating the caller's RunnableConfig in-place via
config['recursion_limit'] = DEFAULT_RECURSION_LIMIT. Add copy.deepcopy(config)
at the top of each factory and switch the recursion_limit write to setdefault
so a caller-supplied ceiling is respected.

* fix: preserve runtime config object identities

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Fleet Agent <fleet-agent@langchain.dev>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit c34e04f44da7ec7638cdafc71e0bb40e77070efd)

Co-authored-by: Jacob Albert <122248719+jacobalbert3@users.noreply.github.com>
2026-07-17 16:22:38 -04:00
ae1f883b4c
refactor: move tests into tests/<domain>/ layout
Applies the plan's C5 step: git mv every test per the domain-reorg
move-map (movemap-m50.txt) into tests/{agent,analyzer,auth,dashboard,
github,middleware,models,reviewer,sandbox,slack,tools,webhooks}/, plus
the 13 fork-only placements from the scoping report §2c (Atlassian
webhook tests -> tests/webhooks/, test_atlassian_connect.py and
test_auth_error_leak.py -> tests/auth/, jira/confluence util tests ->
tests/tools/, test_repo_binding_isolation.py -> tests/sandbox/,
bot-identity/autofix tests -> tests/github/).

Path-only move: the only content edits are parents[1] -> parents[2]
fixes in test_e2b_integration.py and test_daytona_integration.py,
required because their __file__-relative ROOT path gained one more
directory level in the move.

Monkeypatch retargets for these files were already completed in C4;
none remained outstanding here.
2026-07-17 14:42:45 -04:00