Commit graph

40 commits

Author SHA1 Message Date
4fea5f21c5 fix(reviewer): enforce blocking review checks 2026-08-01 20:21:16 -04:00
39c768aa3f feat(reviewer): add verdict-aware review checks 2026-08-01 20:21:16 -04:00
9a2535a9ec fix(reviewer): clarify recorded verdict outcomes 2026-08-01 20:21:16 -04:00
404b854544 fix(reviewer): handle verdict publication edge cases 2026-08-01 20:21:16 -04:00
b2b5957233 fix(reviewer): harden verdict enforcement 2026-08-01 20:21:16 -04:00
1b8f8d480d feat(reviewer): add dispatch verdict authorization 2026-08-01 20:21:16 -04:00
Adam Moussa
c99bd78179
feat: clean-review auto-approve for unsolicited verdicts [BLOCKED — security] (#217)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): clean-review auto-approve for unsolicited publish_review verdicts

An unsolicited publish_review(verdict="approve") — a run dispatched
without verdict_requested — is now honored when the review has zero open
findings, so clean auto-reviews land a real APPROVE. With open findings
it downgrades to a comment review (verdict_ignored_reason=
"approve_with_open_findings"). request_changes stays explicit-request-
only; the self-review, head-moved, and author-unknown downgrades and the
shell verdict guard are unchanged. The reviewer base prompt now instructs
the clean-approve call on auto-reviews.

* chore(security): record accepted-risk suppressions for clean-review auto-approve

Two confirmed-HIGH findings from /sh-security-review on the clean-review
auto-approve change are accepted and deferred (Adam, 2026-07-21), tracked
in #218. Machine-recorded per the mandatory-security-review policy; the
revisit trigger is promotion from dev to main/prod.
2026-07-21 20:18:25 -04:00
Adam Moussa
0f0f616cd4
feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): explicit-request verdicts + shell verdict guard

Mention-triggered reviews that explicitly ask for a verdict now submit a
real APPROVE/REQUEST_CHANGES through publish_review; auto-reviews stay
advisory (COMMENT). Authorization is enforced in code: publish_review
honors a verdict only when the dispatching webhook set verdict_requested,
which only the explicit-mention path does.

- request_pr_review gains instructions (forwarded verbatim into an escaped
  requester_instructions data block) and request_verdict
- self-review guard downgrades verdicts on Open SWE-authored PRs; stale
  APPROVEs are best-effort dismissed when later findings land
- new PullRequestVerdictGuardMiddleware blocks gh pr review
  --approve/-a/--request-changes/-r, gh api, and curl verdict fallbacks on
  both the coding-agent and reviewer graphs
- shared escape helper moved to agent/utils/prompt_data.py

* fix(reviewer): harden verdict path against security-review findings

Adversarial security review (detector fan-out + proof-or-kill verifier)
of the verdict feature surfaced several verdict-integrity gaps; resolve
the confirmed ones:

- head-drift (high): a mid-run push moves the resolved head, so an APPROVE
  could anchor to an unreviewed commit. Downgrade any verdict to a comment
  when the resolved head differs from the reviewed head (verdict_ignored
  reason head_moved); the push's own re-review submits a fresh verdict.
- self-review fail-open: downgrade to comment when the PR author cannot be
  confirmed (author_unknown), and compare bot logins case-insensitively.
- verdict_submitted now reflects GitHub's returned review state, not just
  the event we asked for, so a coerced APPROVE isn't reported as submitted.
- an authorized verdict whose findings all anchor outside the diff now
  posts as a bodied review with zero inline comments instead of failing.
- add finding_reply to the shared data-block escape tag superset.
2026-07-20 15:28:00 -04:00
62d9945df4
refactor: consolidate reviewer modules into agent/review/
Part of the domain-reorg adoption (build plan step C2): fork content,
upstream layout. Nine 1:1 module moves (reviewer_diff/eval_store/
findings/groups/publish/reconcile/trace_context + review_style_
collector/guidance) into agent/review/, with internal relative
imports re-wired to the new package depth. agent/review/__init__.py
mirrors upstream's thin re-export shim (one of the 21 verified "A"
structural adds).

Rewrote the 38 grep hits across importer files (agent/{analyzer,
ci_autofix,reviewer,webapp}.py, agent/dashboard/*, agent/middleware/
settle_review_check.py, agent/tools/*, agent/utils/github_feedback.py,
agent/webhooks/github.py, evals/reviewer/*, and the reviewer test
suite) to point at agent.review.*; 4 of the 38 hits were name
collisions (list_reviewer_findings, reviewer_outcomes,
_reviewer_thread_id, reviewer_thread_id — not the moved modules) and
were left untouched. tests/test_github_checks.py's module-alias
import (`from agent import reviewer_publish`) follows upstream's own
`from agent.review import publish as reviewer_publish` pattern so
downstream `reviewer_publish.*` call sites needed no changes.
agent/reviewer.py and agent/webapp.py stay in place per the hard
rule (fork content, import-only rewire) and are not part of this
package.

Gates: ruff check + ruff format --check, pytest --co -q (1637
collected), full unit suite (1637 passed), and the reviewer/findings
suite in isolation (pytest -k "review or finding", 421 passed).
2026-07-17 13:52:03 -04:00
Adam Moussa
032d3889e4
fix: align reviewer eval with published findings (upstream #1713) (#201)
* fix: align reviewer eval with published findings (#1713)

* fix: make reviewer eval reflect published findings

Serialize and deduplicate finding persistence, align review calibration around the final six-finding publication, and make judge matching order-independent and auditable.

* fix: honor reviewer eval limits

Forward configured caps into publication snapshots and keep recall-at-cap bounded for diagnostic all-findings runs.

(cherry picked from commit 71e3b8183882bcc42e318f3f220c291617ebcb67)
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* Empty commit to trigger CI

---------

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-17 12:10:54 -04:00
Adam Moussa
1f060f2a1d
chore: sync upstream/main, defer #1621 modular webhooks (#81)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
* chore: bake sfw binary into sandbox image (#1611)

sfw only ships a launcher that fetches its real binary at first run and does a
daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in
the sandbox (restricted egress; the proxy injects the GitHub App installation
token, which lacks access to that repo), so `sfw yarn install` errors with
"could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at
build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline.

* feat: editable plan mode + fix review-plan banner overlap (#1610)

* feat: editable plan mode + fix review-plan banner overlap

Lets the thread owner edit the plan markdown by hand from the plan-review
page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id}
endpoint that re-publishes the plan and mirrors it into the sandbox
plan.md, so approve hands the edited plan to the agent as the source of
truth. Also fixes the collapsed git-panel's floating expand button
covering the "Review plan ->" banner by reserving space for it.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: abort plan approval when the published plan read fails

get_plan_content() swallowed store errors and returned None, so a
transient failure during approve would still mark the plan approved and
dispatch the generic fallback text — silently dropping an owner's edited
plan. Read the plan strictly (raise_on_error=True) so approval aborts
instead, matching the comment read.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show message timestamps (#1609)

* feat: show message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: suppress fallback message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: stable message + tool-call hover timestamps

Stamp a stable client-side arrival time per message and tool call (keyed
by id, persisted to localStorage). Messages render the timestamp inline;
tool rows reveal a dim timestamp chip on hover. Real backend created_at
still takes precedence when present.

* fix: hide client-stamped message timestamps

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add PR trace resolution (#1612)

* feat: add PR trace resolution

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: inject reviewer trace context as JSON

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on PR trace resolution

Use the documented LangSmith metadata filter syntax
(and(eq(metadata_key,...), eq(metadata_value,...))) instead of
has(metadata, '{...}'), which does not match runs — _list_thread_runs
was silently returning nothing. Bound full-text searches to a 90-day
window so they don't hit LangSmith's large-window rate limit.

Also folds in the best-effort branch->head-sha resolver (dropping the
weighted scoring/threshold + repo/file evidence + GitHub hydration),
sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint.

The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session
were removed; resolution now runs deterministically from the trusted run
config with no model-controlled pr_url or thread_id.

* fix: scope branch trace search to the repo

Branch names like fix-tests aren't unique across repos (or older PRs) in
a shared tracing project, so an unscoped branch hit could resolve to an
unrelated thread and write its runs into the reviewer sandbox. Require
the repo slug to co-occur with the branch in matched runs; the full head
SHA stays unscoped since it is globally unique. Addresses open-swe review
on PR #1612.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: include plan links in PR descriptions (#1613)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: gate workflow pushes with approval (#1614)

* feat: gate workflow pushes with approval

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve proxy refresh test compatibility

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: bind workflow approvals to pushed ref

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: recover thread work as patch (#1615)

* feat: recover thread work as patch

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: search sandbox cwd for recovery patches

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: omit plan link in PR description when no plan exists (#1618)

Plan links in PR descriptions were always built from the thread id, so
runs that never produced a plan linked to an empty plan-review page.
Now the plan content store is consulted first; the link is only added
when a plan with non-empty markdown actually exists. A transient store
failure degrades gracefully (no link) rather than blocking PR creation.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add filter & grouping menu to agents threads sidebar (#1617)

Add a Cursor-style control to the agents sidebar that groups (None/Date/
Status/Project), filters (ownership, status, source, pull request, model,
repo, include-resolved), and compacts the threads list. All client-side over
already-fetched sidebar threads; preferences persist in localStorage.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: update langsmith sdk to 0.9.3 (#1616)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: clickable shared PR header in git panel and reviews (#1620)

* feat: clickable shared PR header in git panel and reviews

Replace the standalone "View PR" button in the agent git panel with a
clickable PR title, matching the reviews view. Extract a shared PrHeader
component reused by both the git panel and the review main body.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: drop PrHeader wrapper, use shared component directly

The review-side PrHeader was just a thin adapter mapping detail -> the
shared component's props. Inline it at the call site and use the shared
PrHeader directly so there's a single component.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: durable interrupt dispatch + completion webhook (#1621)

* wip(rebuild): core reliability spine

- remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring)
- dispatch core: agent/dispatch.py with multitask_strategy=interrupt +
  durability=sync + completion webhook; reroute all webhook + plan triggers;
  drop the racy in-process lock + is_thread_active busy-check
- completion webhook: agent/completion.py + /webhooks/run-complete loopback
  route for failure/timeout replies (idempotent)

Co-authored-by: open-swe[bot]

* feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning

Parallel batch on top of the reliability spine:
- async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the
  http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden
  the IP check to 'not is_global' (+ IPv4-mapped unwrap)
- reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list
  -> cancel_many), wired into the scheduler graph via task='reconcile'
- shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare
  httpx.AsyncClient() across utils/dashboard/webapp/middleware
- run budget: MODEL_CALL_RECURSION_LIMIT 5000->250
- fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8)
- drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls)
- confirm tool-result eviction + summarization auto-wired via backend
- slim system prompt ~8% (full harness-profile rewrite deferred)

Co-authored-by: open-swe[bot]

* feat(rebuild): harness-profile prompt + split webhooks out of webapp

- prompt.py: own the system prompt via a registered harness profile
  (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that
  share it stay safe), registered across all 4 providers; per-thread values
  stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k
  tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped
  ALL-CAPS markers.
- webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into
  agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the
  routes + tests; moved handlers reach shared helpers via the webapp namespace
  to preserve the test suite's monkeypatch targets.

Full suite: 1168 passing, lint clean.

Co-authored-by: open-swe[bot]

* Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks

Reverts the 250 cap from the run-budget change — long-running tasks legitimately
need many model calls. The notify_step_limit_reached safety net still fires if a
run does hit the cap, so runs end with a signal either way.

Co-authored-by: open-swe[bot]

* fix: address PR review (auth, SSRF, interrupted status, redirect headers)

- completion.py: drop `interrupted` from failure statuses — with
  multitask_strategy=interrupt a follow-up ends the prior run as interrupted,
  which is healthy, not a failure to report. [open-swe]
- /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when
  RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest.
  [corridor-security]
- SSRF: extract the URL validator to agent/utils/url_safety.py and apply it
  before server-side image fetches in multimodal.fetch_image_block.
  [corridor-security]
- http_request: preserve caller headers/extensions across redirect hops instead
  of dropping them on the first hop. [open-swe]

Co-authored-by: open-swe[bot]

* chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo)

Co-authored-by: open-swe[bot]

* fix: fail closed on run-complete webhook auth when secret unset

Corridor follow-up: verify_run_complete_token returns False (not True) when
RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never
unauthenticated. Logs a startup warning when the secret is absent, and dispatch
skips registering the webhook when there's no secret (no rejected callbacks).

Co-authored-by: open-swe[bot]

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: restore forced tool call to prevent premature run stops (#1622)

Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task.

Shipping to test whether it fixes runs that stop halfway through.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619)

Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1.
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1)

---
updated-dependencies:
- dependency-name: langgraph-checkpoint
  dependency-version: 4.1.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>

* fix: post reviewer resolution notes verbatim (#1624)

* fix: post reviewer resolution notes verbatim

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: stabilize dashboard follow-up e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve dashboard attribution in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: make e2e attribution marker durable

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: only echo found e2e attribution

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: check live dashboard attribution in e2e

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625)

Installs hung when prefixed with sfw inside the sandbox (trace 019f0608
stalled on a pending `sfw npm install` execute, never returned). Strip the
Socket Firewall guidance from the agent and reviewer prompts so installs run
through the project's package manager directly. sfw stays in the Docker image;
nothing invokes it now.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: make plan view mobile friendly (#1636)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: fall back to vision model for image threads (#1626)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: surface Slack thread errors (#1627)

* fix: surface Slack thread errors

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: don't set failure_reply_posted on Slack preprocessing errors

The preprocessing error handler was setting failure_reply_posted=True,
the same idempotency flag handle_run_completion checks to suppress
duplicate run-failure replies. Since preprocessing failures happen
before any run exists but the flag persists on the thread, a subsequent
run failure on the same thread would be silently ignored.

The preprocessing handler already posts its own Slack reply, so the
run-completion idempotency flag should not be set here.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: avoid recapping Slack replies (#1629)

* chore: avoid recapping Slack replies

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: simplify Slack reply prompt wording

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: update Slack trace reply on web handoff (#1630)

* fix: update Slack trace reply on web handoff

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: trigger web handoff on dashboard starts

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: format web handoff as contextual fragment

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve trace_message_ts when overwriting Slack run mapping

When store_slack_run_mapping is called without trace_message_ts (e.g. on
follow-up Slack mentions), it was unconditionally overwriting the
thread-level mapping and clobbering the timestamp captured from the
initial trace reply. After that, _notify_slack_web_handoff could not find
the original message, so a subsequent move to Web silently skipped the
Slack trace update.

Now, when trace_message_ts is not passed, the existing thread mapping is
read first and its trace_message_ts is preserved.

* style: ruff format

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643)

* fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures

shiki lazy-imports a grammar per language and these libs only live inside
lazy route components, so Vite's startup scanner never sees them. They get
discovered on first thread navigation, triggering a dep re-optimize +
force-reload that aborts the in-flight route-chunk import, surfacing as
"Failed to fetch dynamically imported module: .../$threadId.tsx".

Pre-bundle them (and the github themes + common code-block languages) via
optimizeDeps.include so the optimize happens once at startup. Dev-only;
production bundles are unaffected.

* fix: pre-bundle canonical shiki docker/make langs instead of aliases

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: show queued dashboard follow-ups (#1631)

* feat: show queued dashboard follow-ups

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: de-dupe queued follow-ups while streaming

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* feat: notify Slack on plan approval (#1632)

* feat: notify Slack on plan approval

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: post Slack approval notice after successful dispatch

Move the _maybe_post_plan_approved_to_slack call until after
_dispatch_followup succeeds so the Slack thread is not told
implementation is beginning before the LangGraph run is created.

Addresses PR review comment.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>

* feat: include Slack channel context in prompts (#1633)

Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>

* chore: keep plan guidance high-level (#1634)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: publish plans from sandbox files (#1635)

* feat: publish plans from sandbox files

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: avoid fixed plan filenames

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: virtualize local sandbox file paths

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve plan_file_path across set_plan_status

set_plan_status was rewriting the content record with only markdown
and status, dropping plan_file_path. After a reject, the owner's
dashboard edit would mirror to a different file than the agent's
original, and the next save_plan could republish the stale file.
Preserve plan_file_path when updating status.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: return to thread after plan approval (#1637)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add Slack breakout thread tool (#1638)

* feat: add Slack breakout thread tool

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: make fake LLM scripts declarative

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: exclude slack_start_new_thread from plan mode

The breakout tool can dispatch a fresh agent run that starts outside the
current plan-mode state, bypassing the approval flow. Add it to
PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating
tools while planning.

---------

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: require bun for ui agent work (#1639)

Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: request actions read for sandbox logs (#1642)

* fix: request actions read for sandbox logs

Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: restore actions:read scope after workflow push

After an approved workflow push, the guard was restoring the proxy with
BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read
scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which
includes actions: read) and fall back to BASE if the install hasn't
granted Actions read — mirroring the pattern in _create_sandbox_with_proxy.

Addresses review comment on PR #1642.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: widen split review diffs (#1647)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: install missing deps before verification (#1646)

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: switch ui to pnpm (#1645)

* chore: require pnpm for ui agent work

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: switch ui to pnpm

Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* ci: use corepack for ui pnpm e2e build

Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add Sonnet 5 to model picker (#1651)

* chore: update Sonnet examples to Sonnet 5

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: add Sonnet 5 to model picker

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* Remove dead breakout-thread e2e scenario after dropping the tool

The merge resolution deferred upstream's Slack breakout-thread tool
(slack_start_new_thread, #1638) since it depends on the #1621 dispatch
module, but the e2e harness still scripted it. Removing the tool name
from fake_llm.py's _tool_step call left a malformed scenario, crashing
the langgraph-dev web server at import (TypeError: _tool_step() missing
'call_id') and failing Playwright E2E.

Drop the "breakout" script scenario, its _is_breakout_request helper +
ScriptRule, and the corresponding full_flow.spec.ts test.

* Revert upstream pnpm switch; keep bun for the UI build

The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/
global-setup.ts and ui/package.json, but our fork builds the UI with
bun (vercel.json + the E2E workflow's setup-bun). That left the
Playwright globalSetup running `corepack pnpm install --frozen-lockfile`
with no pnpm-lock.yaml, failing E2E at UI build time.

Revert global-setup.ts and ui/package.json to the dev (bun) baseline,
drop the merge-added ui/pnpm-lock.yaml, and remove the re-added
ui/AGENTS.md (our fork had deleted it).

* Align plan-review e2e + UI with the HEAD (pre-#1635) backend

The merge left a split plan vertical: the backend save_plan/plan_api are
HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/
#1637 per #80), but the plan UI and e2e harness were upstream's. The
fake_llm scenario called save_plan(plan_file_path=...) — upstream's
file-based #1635 contract — while HEAD save_plan takes plan_markdown,
so the plan never saved and PlanReview never rendered (E2E failure on
the plan-review locator).

Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts /
$threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the
whole plan flow (save -> render -> approve -> implement) is consistent
with the HEAD backend.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com>
Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
2026-06-30 16:45:19 -04:00
Johannes du Plessis
055b83e723
feat: surface sub-threshold findings in review summary with web app link (#1571)
Instead of silently swallowing findings below the severity threshold,
the review summary now mentions them with a count and links to the
web app where they can be viewed. For example, if 2 low-severity
findings are filtered out, the PR comment says "No issues found" and
"2 additional findings can be viewed in the web app." with the
existing [Open in Web] link.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-18 11:15:42 -07:00
Johannes du Plessis
889b2663d9
fix: point View trace link at the correct tracing project (#1521)
* fix: resolve trace URL project id by tracing project name

Graphs were split into separate LangSmith tracing projects
(open-swe-agent, open-swe-review) but the "View trace" link still used
a single fixed project-id env var pointing at the old combined project.
Resolve the project id from the tracing project name so agent and
reviewer links point at their respective projects, falling back to the
env var when resolution is unavailable.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* docs: document per-graph tracing projects for trace links

View trace links now resolve project IDs from the open-swe-agent /
open-swe-review project names. Document this so fresh deployments create
the right projects instead of relying on a single project ID.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-13 09:37:15 -07:00
Johannes du Plessis
bd5d5b24d4
fix: point reviewer "Open in Web" link to the review page (#1519)
The top-level review comment's "Open in Web" link pointed at the
agent thread (/agents/{thread_id}). Point it at the dashboard review
detail page (/agents/reviews/{owner}/{repo}/{number}) instead.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 14:15:58 -07:00
Johannes du Plessis
258b5b4034
feat: disable out-of-diff reviewer findings (#1516)
* feat: disable out-of-diff reviewer findings

PR reviews were surfacing findings about code outside the PR's changed
lines, which read as random/off-topic noise. Reject out-of-diff findings
at add_finding and stop surfacing them in publish_review so only findings
anchored to changed lines reach the PR.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: align reviewer bar with out-of-diff rejection

The filing bar still permitted proven regressions in files absent from
the diff, but add_finding now rejects those, so a concrete regression
could be silently dropped after a failed tool call. Update the bar and
the "Do NOT file" list so the prompt only directs in-diff findings.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-12 10:19:30 -07:00
Johannes du Plessis
e8770e1004
feat: report Open SWE Review as a PR check run (#1484)
* feat: report Open SWE Review as a PR check run

Auto-review dispatch now creates an in-progress 'Open SWE Review' check run
on the PR head SHA; publish_review completes it (neutral with findings,
success when clean). An after-agent hook fails the check if the run dies
before publishing. Requires the GitHub App's Checks: Read & write permission;
all calls are best-effort so a missing permission never breaks reviews.

* fix: address review feedback on check-run settling

Keep review_check_run_id when the completion PATCH fails so a later
publish or the after-agent hook can retry instead of hanging the check;
count out-of-diff findings toward the check conclusion.

* fix: retry failed check completion with the real publish conclusion

A transient PATCH failure after a successful publish previously left the
check id for the after-agent hook, which settled it as 'failure'. Persist
the intended result as review_check_pending_result and have the hook
prefer it over the generic failure fallback.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:53:48 -07:00
Johannes du Plessis
8461979b0d
fix: honest publish_review reporting + structured thread-not-found errors (#1481)
* fix: honest publish_review reporting + structured thread-not-found errors

- Document skipped_empty_re_review and dry_run in the publish_review
  docstring and add a closing-summary contract to the reviewer prompt so
  the agent never claims a review was published when review_id is null.
- Raise ReviewerThreadMissingError from replace_findings on SDK
  NotFoundError; add_finding/update_finding/publish_review return a
  structured do-not-retry result instead of raising, so the agent reports
  the blocker after one failure instead of retrying 10-30 times.

* fix: translate thread 404s across all reviewer tool boundaries

get_thread_metadata now raises ReviewerThreadMissingError instead of
swallowing a missing thread as {} (which produced misleading 'No finding
found' results), set_reviewer_thread_metadata translates the SDK 404 the
same way, and every reviewer tool entrypoint (add/update/list findings,
publish_review incl. eval dry-run, resolve/reply thread) returns the
structured do-not-retry result.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:44:13 -07:00
Johannes du Plessis
7cd882bb67
fix: make review publish idempotent (partial-failure recovery) (#1477)
Publish had partial-failure windows that double-posted summaries or
corrupted findings state. This hardens the recovery paths:

- open_swe_review_exists is now tri-state (True/False/None). On a
  pagination/API failure it returns None ("unknown") instead of False,
  and the empty-summary dedup keys off the durable last_reviewed_sha
  before consulting GitHub, so a transient failure never double-posts a
  "no issues found" summary.
- Comment-id backfill matches strictly on the embedded open-swe marker;
  the colliding (path, line, body) fallback is gone, so similar findings
  no longer share a comment id and break resolve-on-fix.
- Review-id and comment-id stamping collapse into one guarded
  read-modify-write (re-reads latest before writing), removing the
  half-stamped intermediate states the prior multi-write flow left open.
- New mutate_findings primitive centralizes read-modify-write so finding
  updates operate on the freshest persisted list and skip no-op writes.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 10:17:31 -07:00
Ramon Nogueira
d2268871eb
feat: surface dashboard UI link on PR reviews [INF-0000] (#1440)
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]

Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.

The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.

* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
2026-06-07 05:17:09 +00:00
Johannes du Plessis
f64ab2bcd7
feat: surface out-of-diff findings in a collapsed dropdown (#1427)
* fix: stop reviewer retrying out-of-diff findings

add_finding rejects findings anchored outside the PR diff, but the agent
retried the same finding 2-3x with adjacent line ranges before giving up,
burning a model turn each. Add a reviewer-prompt recovery block telling the
agent the rejection is authoritative (drop or re-anchor to a + line, don't
retry adjacent), and enrich the rejection payload with nearby in-diff line
ranges for the file so a single re-anchor needs no guessing.

* feat: surface out-of-diff findings in a collapsed dropdown

Instead of rejecting findings anchored outside the PR diff, accept them
(marked in_diff=false) and surface them in a collapsed <details> section of
the review summary, Devin-style. Inline comments stay reserved for in-diff
findings; out-of-diff are severity-gated and capped the same way.

Re-review normally suppresses the empty summary, but now makes an exception
when there are new out-of-diff findings to surface. Surfaced out-of-diff
findings carry a github_review_id so they aren't reposted on later pushes.

Supersedes the earlier 'drop/re-anchor out-of-diff' prompt guidance.

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-05 10:41:48 -07:00
Johannes du Plessis
1d4f1aed33
fix: reviewer publishes against stale head_sha on mid-run re-review (#1393)
* fix: resolve reviewer head_sha from thread metadata, not frozen run config

A push that lands while a reviewer run is in flight is delivered as a
queued message into that run. The run's configurable is frozen at
creation, so its head_sha still names the commit the run was created for
— not the commit just pushed. publish_review then anchored the GitHub
review to the stale commit and regressed last_reviewed_sha to it, and
add_finding/update_finding stamped findings with the stale SHA.

Persist the current head in thread metadata at every reviewer dispatch
(both the ready-for-review and push paths, before they branch to create
a run or queue a message), and add resolve_review_head_sha() which
prefers the metadata head over the run config. Wire it into
publish_review (review commit_id + last_reviewed_sha), add_finding
(first_seen_sha) and update_finding (last_confirmed_sha). Falls back to
the run config when metadata carries no head (first review, eval, tests).

* fix: persist head_sha in manual review dispatch (trigger_pr_review_from_ref)

resolve_review_head_sha prefers metadata[head_sha] over the run config,
and the push/ready dispatchers write it — but trigger_pr_review_from_ref
(Slack/GitHub @open-swe review, request_pr_review tool) created a run
with a freshly-fetched config head while leaving metadata's head stale
from a prior dispatch. A manual re-review at a newer commit would then
resolve to the old head and publish/advance findings against it.

Persist head_sha in that dispatch's metadata write too, so every
run-creating reviewer dispatch keeps metadata in sync with the head its
run targets. Caught by the Open SWE reviewer on this PR.
2026-06-03 11:38:56 -07:00
Johannes du Plessis
18f8ca56fb
fix: dedup empty reviewer summary by PR state, not stale re_review flag (#1391)
A push that lands while a reviewer run is in flight is delivered as a
queued message into the still-running first-review run, whose
configurable still has re_review=False. The empty-review guard in
publish_review only skipped the 'No issues found' summary when
is_re_review was True, so the queued reconcile published a second,
duplicate top-level 'No issues found' review.

Key the empty-review skip off actual PR state instead: add
open_swe_review_exists(), which detects the marker render_review_body
embeds in every Open SWE review body, and skip the summary when a prior
Open SWE review already exists (regardless of the re_review flag). Fails
open on API error so a genuine first review is never suppressed.
2026-06-03 10:30:13 -07:00
open-swe[bot]
a361ee8f2e
feat: let reviewer set comment titles (#1356)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-28 16:04:38 -07:00
open-swe[bot]
d8d3794649
fix: include reviewer trace links (#1351)
* fix: include reviewer trace links

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* feat: move reviewer trace-link toggle to dashboard

Replace the OPEN_SWE_REVIEW_TRACE_LINK_ENABLED env var with a team-level
'Trace Links' toggle in the Open SWE Review dashboard tab. The toggle is
read per-publish via get_team_review_trace_links_enabled(); the per-run
review_trace_link_enabled config override still forces it off.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-28 13:48:14 -07:00
Johannes du Plessis
fee209601b
feat: structured reviewer comments + auto resolution comments (#1352)
* feat: structured reviewer comments + auto resolution comments

Restructure inline review comment bodies (severity emoji, bold title from
the first line, line reference, feedback footer) without duplicating the
first description line, and post an automatic resolution/dismissal comment
to the GitHub thread when a finding is resolved or dismissed.

- Centralize render_resolution_comment in reviewer_publish; fix a crash when
  last_reconciliation_note is None and drop the misleading generic fallback.
- Post the resolution comment in resolve_finding_thread (the normal
  update_finding path), not only in publish_review, so it actually fires on
  re-review. Dedupe via github_posted_resolution_comment_ids.
- Add pytest coverage for rendering and the resolution-comment flow.

* fix: post resolution comment to every closed thread in resolve_finding_thread

Per-thread iteration (matching _resolve_threads_for_resolved_findings) so
duplicate threads after the first also receive the resolved/dismissed
explanation before being closed.
2026-05-28 12:44:00 -07:00
Johannes du Plessis
197d339df4
fix: Reconcile reviewer findings with PR threads (#1346)
* feat: reconcile reviewer findings with PR threads

* fix: harden reviewer finding reply handling

* fix: queue reviewer finding reply body

Ensure review-comment replies that arrive during an active reviewer run include the sanitized reply body in the queued reassessment prompt.

* fix: apply reviewer reply formatting

Apply the repository formatter so the reviewer reply handling fix passes CI format checks.
2026-05-27 17:26:08 -07:00
Johannes du Plessis
0d4d1c5a3b
fix: dedupe reviewer comments from PR state (#1341)
* fix: dedupe reviewer comments from PR state

Use GitHub review-thread markers to repair reviewer publication state before posting or resolving findings, so re-reviews do not duplicate comments and resolved findings close all matching PR threads.

* fix: require all duplicate reviewer threads resolved

Avoid treating a marker-backed finding as resolved when only one duplicate thread is outdated while another matching thread remains open.
2026-05-27 10:52:30 -07:00
Johannes du Plessis
fa147ce256
fix: fetch PR diff via GitHub API to re-enable add_finding validation (#1339)
* reviewer: fetch PR diff via GitHub API to re-enable add_finding validation

The previous hotfix in reviewer.py set diff_line_set=None because the
sandbox-based diff prep was sometimes producing empty diffs. That made
every bad anchor a publish-time 422 instead of a creation-time
rejection — the agent burned tokens producing unanchorable findings,
and we had to add a publish-time retry safety net (#1338) to clean up.

Fetch the PR's unified diff via the GitHub REST API at reviewer
startup and populate diff_text + diff_line_set so add_finding can
reject bad anchors immediately. The API path is reliable and is the
same diff GitHub validates against when posting inline review
comments. If the fetch fails, fall back to the previous behavior
(validation disabled, publish-time retry handles it).

Also extract the PR-diff fetch into reviewer_diff.fetch_pr_diff so
both reviewer.py and publish_review.py share one implementation
instead of two copies.

* reviewer: make diff_line_set validation side-aware

compute_diff_line_set previously returned only new-side line numbers,
so re-enabling add_finding's validation would wrongly reject findings
with side=LEFT (deleted-line bugs whose only anchor is an old-side
line). Return {file: {"RIGHT": {new_lines}, "LEFT": {old_lines}}}
instead, and have is_range_in_diff select the matching side from the
finding's recorded side. add_finding and publish_review's retry
filter both pass the finding's side through.
2026-05-27 17:03:33 +00:00
langsmith-engine[bot]
58b1d52fee
fix: publish_review HTTP 422 "Path/Line could not be resolved" — agent retries with identical args instead of dropping unresolvable findings (#1338)
* publish_review: drop unresolvable findings and retry once on GitHub 422

GitHub returns 422 with 'Path could not be resolved' or 'Line could not be
resolved' when an inline comment anchors to a file/line not in the PR diff.
Previously the agent retried publish_review with byte-identical args
multiple times before draining to skipped_empty_re_review=true, silently
losing findings.

- reviewer_publish.post_pull_request_review: parse 422 body and tag with
  _error_kind='unresolved_anchor' plus _raw_errors so callers can act.
- tools/publish_review._publish_review_async: when that signal fires,
  cross-check each finding's range against the run config's diff_line_set,
  drop the bad ones, and re-POST once with only the valid findings. Return
  unresolvable_findings + hint so the agent calls update_finding instead of
  retrying the same payload.
- reviewer.py: one-line prompt addendum telling the agent that
  unresolvable_findings means update_finding, not retry.
- tests: cover 422 tagging (path + line), the drop-and-retry success path,
  the retry-still-fails path, and the don't-blind-retry path when no
  diff_line_set is available.

* publish_review: fetch PR diff on demand for 422 retry filter

Reviewer runs clear configurable['diff_line_set'] before the agent
starts, so the unresolved-anchor retry path had no diff data to filter
against — in the reachable production case it dropped nothing and
returned success=False with empty unresolvable_findings, losing the
otherwise-valid comments.

Fall back to fetching the PR's unified diff via the GitHub REST API
and recomputing the line set on the fly when no cached set is
available. The cached set is still preferred when present.

---------

Co-authored-by: issues-agent <issues-agent@langchain.dev>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-26 18:37:56 -07:00
Johannes du Plessis
5702a9d452
feat: reconcile reviewer comment lifecycle (#1332)
* feat: reconcile reviewer comment lifecycle

Track GitHub review threads for reviewer findings so re-reviews can resolve or reply to existing comments, and collect thumbs feedback on new review comments in LangSmith.

* fix: clarify reviewer comment lifecycle

* chore: apply reviewer formatting

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-26 16:24:34 -07:00
Johannes du Plessis
b2a0ac3b79
feat: auto-review PRs on opened / ready-for-review (#1325)
* feat: auto-review PRs on opened / ready-for-review

Trigger Open SWE Review on `pull_request` actions `opened` and
`ready_for_review` against the canonical reviewer thread (no need to
request open-swe[bot] as a reviewer). `converted_to_draft` now also
flips watch=False on the existing reviewer thread.

Draft PRs are gated by a tri-state user setting on the profile:
inherit team default, always on, or always off. The team-wide
`review_draft_prs` setting is the org-wide default; each user can
override it in My Settings.

External contributors with no Open SWE profile fall back to the team
default.

* fix: PR review comments — auth source + draft-aware watch toggle

- `process_github_pr_ready` now dispatches with `source="github"` so the
  auth resolver finds the bot token persisted on the thread. The previous
  `source="github_auto"` fell through to the email-based path in non
  bot-token-only deployments and failed with a missing-user-email error.

- `converted_to_draft` no longer unconditionally clears `watch`. When the
  PR author's effective `review_draft_prs` setting is on, watch stays on
  so subsequent pushes still trigger re-reviews while the PR is in draft.

* feat(reviewer): skip "no issues found" comment on empty re-reviews

A re-review run with no new findings to surface no longer posts another
"Open SWE Review: No issues found" comment on the PR. The "no issues"
summary now only appears on the first review of a PR — matching Devin's
behavior, where subsequent reviews are silent unless there's something
new to flag.

Resolved-thread reconciliation and ``last_reviewed_sha`` persistence
still happen on the skipped path, so findings the user just fixed still
get their GitHub threads marked resolved, and the next push event sees
an up-to-date dedup SHA.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-22 13:36:14 -07:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
Johannes du Plessis
834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00
langsmith-forge[bot]
74f5df5df8
fix: publish_review tool returns generic "Failed to POST PR review" without GitHub API status/body, agent retries with no signal (#1299)
* fix(reviewer): surface HTTP status and body excerpt for non-dict GitHub PR review responses

When post_pull_request_review received a non-dict body, it returned
None and publish_review surfaced a generic 'Failed to POST PR review'
string with no signal for the agent to adapt — leading to blind
retries with permuted cap/severity_threshold args.

Now the non-dict-body path mirrors the existing HTTPStatusError /
HTTPError paths: it returns {'_error': 'HTTP <status>: non-dict
response body: <excerpt>'} so the user-facing tool can include the
underlying detail. The bare-None branch in publish_review.py is kept
as a defensive guard with a clearer message.

* ci: apply ruff format to reviewer_publish.py

Collapse the multi-line return dict into a single line so it matches the
output of `ruff format`, unblocking the Agent lint / format-check CI jobs.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: LangSmith Issues Agent <issues-agent@langsmith.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-12 13:55:39 -07:00
open-swe[bot]
85343fab63
feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322] (#1280)
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]

Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* webapp: forward installation-token expiry to reviewer cache writes

The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:57:01 +00:00
Johannes du Plessis
3f319b7d8a
reviewer: surface GitHub error details when publish_review fails (#1282)
post_pull_request_review used to swallow the HTTPError and return None,
so the tool result was just "Failed to POST PR review" with no status
or body — the actual cause (e.g. 422 invalid inline comment, 404 app
not installed) only lived in logs. Capture status + body and propagate
into the tool result so the agent (and traces) can see why.
2026-05-08 15:40:05 -07:00
Johannes du Plessis
892347041f
feat(open-swe): post review summary to Slack on first review (#1258)
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.

The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
2026-05-07 16:11:36 -07:00
Johannes du Plessis
88a55057ea
fix(open-swe): host-format the review summary, drop agent prose (#1257)
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.

Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
2026-05-07 15:41:40 -07:00
Johannes du Plessis
03d9f23645
fix(open-swe): always post a review summary, even with no findings (#1256)
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable

Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.

Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
  (with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
  already-reviewed head_sha)

* fix(reviewer): always post a summary review, even with no findings

The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.

- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
  findings list is empty, vs. `**No issues at or above \`<sev>\`
  severity.**` with hidden count when only sub-threshold findings
  exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
  body is meaningful; calls out specifically not to skip on a clean PR.
2026-05-07 22:16:20 +00:00
Johannes du Plessis
378b95266e
feat: implement reviewer findings, publish_review, and watch mode (#1253)
* feat: implement reviewer findings, publish_review, and watch mode

Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:

- Findings as first-class state on the reviewer thread metadata
  (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
  ranges, suggestion text for ```suggestion blocks, github_review_comment_id
  for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
  metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
  future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
  compute_diff_line_set for in-diff validation, extract_diff_hunk for
  caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
  diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
  ranges fail at creation, not at GitHub-publish), `update_finding`,
  `list_findings`, `publish_review`. The reviewer agent's tool list is
  swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
  one POST /reviews call with body + inline comments + ```suggestion blocks,
  per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
  fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
  before the agent's first model call (warm- and cold-path symmetric);
  computed diff and in-diff line set passed via runnable config; system
  prompt rewritten for the single-evolving-findings model, severity ladder,
  in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
  added to supported events. New `process_github_push_event` resolves the
  open PR for the pushed branch, gates on the reviewer thread's `watch`
  flag, builds a re-review configurable, and triggers a run on the same
  canonical thread. `process_github_pr_close` toggles watch on
  closed/reopened. `set_reviewer_thread_metadata` is called on first
  review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
  legacy {file, line, body, severity} shape the judge expects) and passes
  the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
  publish rendering + GraphQL resolve, and watch-mode webhook handlers
  (push triggers re-review only when watching, idempotent on unchanged
  head SHA, PR close disables watch). Updated existing reviewer-webhook
  tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
  into REVIEWER_DESIGN.md.

* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL

Address PR #1253 review findings:

- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
  (`option no-prefix takes no value` — every prep run was failing
  silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
  now uses three-dot `base...head` (the merge-base diff GitHub renders
  on Files-changed) so we don't pick up changes that landed on the base
  branch after the PR diverged. Re-review delta keeps two-dot
  `last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
  `github_review_comment_id`. Without this, watched re-reviews
  re-posted every previously surfaced finding, and only the most-recent
  duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
  `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
  is the canonical endpoint; the old form 404s, so comment ids were
  never stored and watch-mode resolution couldn't run.

Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.

* fix(reviewer): default publish cap from 15 to 4

A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00