* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt
Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.
Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
claims without a concrete attacker/interleaving/scale, style preferences
the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
vendored / pure-rename hunks)
Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.
* trim prompt
* subagent prompting
* confidence ratings
* added medium
* enforce confidence threshold
* .
* reviewer: precision-tuned prompt + drop confidence gate
Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.
Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.
Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.
* benchmax
* adding google provider
* slight steering
* tuning
* more tuning
* fix
* cleanup
* reducing overfitting
* Add per-repo review style profiles and inject them into the reviewer.
Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.
Co-authored-by: Cursor <cursoragent@cursor.com>
* Fix review style job errors leaking exception details to clients.
Return generic dashboard messages while logging full stack traces server-side.
Co-authored-by: Cursor <cursoragent@cursor.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat: tighten reviewer eval workflow
Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.
* chore: move reviewer eval settings to config
Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.
* feat: allow reviewer eval model overrides
Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.
* fix: use adaptive thinking for Opus 4.7
Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.
* refactor: use latest Anthropic effort API
Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.
* revert prompting
Avoid treating regex-parsed Slack text as authoritative routing or allowlist input so repository mentions remain model context and access is bounded by installation permissions.
PR #1303 was merged before this URL change landed on the branch. Updates
the rewrite destination from the open-swe-test deployment to the prod
open-swe-v3 deployment so openswe.vercel.app proxies to the right
backend.
Browsers (Safari, Brave, Firefox, increasingly Chrome) refuse to set or
send SameSite=None cookies cross-site. With the frontend on
openswe.vercel.app and the API on *.langgraph.app, the osw_session
cookie from `/auth/callback` was being silently dropped, so every
subsequent /me call returned 401.
Adding a Vercel rewrite makes the API same-origin from the browser's
point of view: the request goes to openswe.vercel.app/dashboard/api/...,
Vercel proxies it to the LangSmith deployment, and the Set-Cookie comes
back attributed to openswe.vercel.app — a first-party cookie that all
browsers honour.
Pair with the matching deployment-side config changes:
- GitHub App callback URL → https://openswe.vercel.app/dashboard/api/auth/callback
- DASHBOARD_API_BASE_URL → https://openswe.vercel.app (LangSmith env)
- VITE_DASHBOARD_API_BASE_URL → unset / empty (Vercel env)
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints
Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token
CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.
* feat: apply dashboard profile model/effort overrides in get_agent
Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.
* feat: ui/ TanStack Start dashboard for profile config
Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:
- /login — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin — Admin-only: list users and edit other profiles
API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.
Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.
* feat: searchable Combobox for default repo picker
Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.
* fix: address review comments + wire default_repo and Anthropic thinking
Security/correctness fixes from PR review:
* Open redirect: validate `redirect_to` in `/auth/login` against
`DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
into the state JWT. Anything off-allowlist falls back to the dashboard
base URL. (PR #1302 r3250054386)
* Login CSRF: bind the OAuth `state` to the requesting browser. At
`/auth/login` we generate a fresh nonce, set it as a short-lived
HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
we require the cookie nonce to hash-match the state JWT's nonce_hash
(constant-time compare). (PR #1302 r3250054395)
* RMW race in profile vs token writes: split storage into two
namespaces — `["profiles"]` for user-editable settings and
`["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
only writes its own namespace so an in-flight profile save can no
longer clobber a fresh token from a concurrent re-login (and vice
versa). (PR #1302 r3250054393)
* /repos pagination: follow `Link: rel="next"` for both
`/user/installations` and per-installation `/repositories` with
per_page=100, capped at 1000 items. (PR #1302 r3250054401)
Feature wires:
* default_repo: applied as a fallback in `get_slack_repo_config` (after
explicit-repo / thread metadata, before the env defaults) and in the
Linear webhook (after comment-body extraction, before team mapping).
Both paths resolve the triggering user's GitHub login via
GITHUB_USER_EMAIL_MAP and read the profile's default_repo.
* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
anthropic. OpenAI path still ignores "max" since the Literal doesn't
accept it.
* fix(reviewer): surface HTTP status and body excerpt for non-dict GitHub PR review responses
When post_pull_request_review received a non-dict body, it returned
None and publish_review surfaced a generic 'Failed to POST PR review'
string with no signal for the agent to adapt — leading to blind
retries with permuted cap/severity_threshold args.
Now the non-dict-body path mirrors the existing HTTPStatusError /
HTTPError paths: it returns {'_error': 'HTTP <status>: non-dict
response body: <excerpt>'} so the user-facing tool can include the
underlying detail. The bare-None branch in publish_review.py is kept
as a defensive guard with a clearer message.
* ci: apply ruff format to reviewer_publish.py
Collapse the multi-line return dict into a single line so it matches the
output of `ruff format`, unblocking the Agent lint / format-check CI jobs.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: LangSmith Issues Agent <issues-agent@langsmith.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
The SLACK_ASSISTANTS_API_ENABLED env flag and its gating function
are removed. set_slack_assistant_status now always proceeds when a
bot token and channel/thread are provided.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Drops the add_slack_reaction call in process_slack_mention; the Slack
assistant status indicator and the agent's first reply still signal
acknowledgement. Linear and GitHub eyes reactions are unchanged.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Adds a scheduled workflow that force-pushes main onto a prod branch at
08:00 UTC (midnight PST / 1am PDT). The LangGraph deployment will track
prod instead of main, so routine merges no longer trigger redeploys
that cancel in-flight 30-minute coding runs. workflow_dispatch is
retained for hotfixes that need to ship immediately.
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]
Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* webapp: forward installation-token expiry to reviewer cache writes
The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: add 10 more tips to the trace-reply rotation
The TRACE_REPLY_TIPS pool only had 9 tips, so users mostly saw the same
ones. Added 10 more grounded in actual features (review command, image
attachments, GitHub-issue triggers, persistent sandboxes, OAuth fallback,
etc.) so the rotation surfaces more of what open-swe can actually do.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* swap out 3 tips per review feedback
Replaced the suggestion-blocks, sandbox-auto-recovery, and OAuth-fallback
tips with three more practically useful ones: cross-posted Slack message
resolution, the web_search tool, and the Linear ticket-management tools.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
post_pull_request_review used to swallow the HTTPError and return None,
so the tool result was just "Failed to POST PR review" with no status
or body — the actual cause (e.g. 422 invalid inline comment, 404 app
not installed) only lived in logs. Capture status + body and propagate
into the tool result so the agent (and traces) can see why.
When the primary model raises a transient provider error (5xx, 429,
connection/timeout) the request is retried once against a fallback
model from the other provider. Anthropic primaries fall back to
OpenAI and vice versa. Also bumps the SDK max_retries from the
default 2 to 6 so quick blips stay on the primary and keep prompt
caching warm.
Triggered by 529 OverloadedError traces that ended runs silently
with no Slack/Linear/PR reply.
* fix: harden http_request SSRF guard against DNS rebinding [closes AB-2321]
Pin DNS resolution per request hop so urllib3's connection-time lookup
cannot rebind to a private IP after _is_url_safe validated a public one.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: scope DNS pin to urllib3 with reference-counted install
Address review feedback: the previous version permanently overwrote the
process-global socket.getaddrinfo on first use. Now the patch targets
urllib3.util.connection.create_connection (much narrower blast radius),
and is installed/uninstalled via reference count so no global mutation
persists once no http_request calls are in flight.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: forward timeout and socket_options through pinned create_connection
urllib3 calls create_connection with timeout positional and socket_options
as a keyword. The previous wrapper only read kwargs, silently dropping the
caller's connect timeout (so a slow validated IP could hang) and TCP
options like TCP_NODELAY. Accept both positionally and forward them to
the underlying socket.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Appending "Please react with 👍 or 👎..." to every Slack completion
message ended up being repetitive noise. Drop the prompt-side instruction
and add the same ask as one of the rotating tips on the trace reply, so
users still see it occasionally without it cluttering each final summary.
Now that assistant.threads.setStatus carries the "is thinking…"
indicator, the per-run greeting phrase + auto-unfurl on the LangSmith
link are just visual noise stacking on top of it. The trace reply now
posts only `<url|View trace>` + a tip, with link unfurling off so the
smith.langchain.com card no longer appears.
Adds an `unfurl_links`/`unfurl_media` knob to post_slack_thread_reply_with_ts
(default-on to preserve behaviour for every other caller).
Threat model T6: Fernet token-encryption key had no rotation path. Rotating
TOKEN_ENCRYPTION_KEY immediately invalidated every github_token_encrypted
value in thread metadata, forcing re-auth or bot-token re-resolution.
Switch to cryptography.fernet.MultiFernet and parse TOKEN_ENCRYPTION_KEY as a
comma- or newline-separated ordered list (most-recent-first). New writes
encrypt under the first key; reads try every key. Deployers can prepend a new
key, let active threads roll over, then drop the old key.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: add Slack reaction feedback to LangSmith
Record Slack reaction feedback against explicitly mapped LangGraph runs so user ratings are idempotent and tied to the message they reacted to.
* Address review feedback on Slack reaction → LangSmith feedback
- langsmith.py: drop lru_cache on _build_langsmith_feedback_clients so
rotated keys / late env hydration are picked up; dedupe by (key, url)
tuple instead of key alone so the same key pointing at different
endpoints (cloud + self-hosted) builds both clients.
- langsmith.py: treat LangSmithNotFoundError on delete_feedback as
success — out-of-order or redelivered reaction_removed events would
otherwise loop forever on Slack's retry policy.
- slack_feedback.py: include channel_id in _feedback_key so the same
message_ts in two channels can't collide on the same feedback id.
- slack_feedback.py: treat conflicting +/- reactions from one user as
ambiguous (clear feedback) instead of averaging to a misleading 0.5.
- slack_feedback.py + slack.py + webapp.py: gate reaction handling to
the user who triggered the run (stored in the slack_run_map mapping
alongside run_id). Prevents bystanders in shared channels from
polluting eval feedback.
* feat: add ALLOWED_GITHUB_REPOS env var for owner/repo-level webhook filtering
Previously only org-level filtering was supported via ALLOWED_GITHUB_ORGS.
This adds ALLOWED_GITHUB_REPOS for finer-grained control, allowing specific
owner/repo pairs to be allowlisted independently of org membership.
* fix: update test patches for _is_repo_org_allowed rename
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: tell agent its source code lives at langchain-ai/open-swe
* fix: scope self-reference to only when user talks to the agent about itself
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: recover from mid-run sandbox death
Recreate dead sandboxes during tool execution and stop repeated unrecoverable timeout loops with a user-facing notification.
* fix: count repeated sandbox recreations
Treat consecutive sandbox recreations as an unrecovered failure streak so outages cannot loop until the model-call limit.
Adds a webhook-level check so only members of $PUBLIC_REPO_ORG_GATE
(e.g. langchain-ai) can trigger Open SWE via mentions or review
requests on public repositories. Private repos remain governed by the
existing org/repo allowlists. Internal bots bypass the gate.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
The gh-cli migration removed agent/middleware/open_pr.py and
agent/tools/commit_and_open_pr.py — the only callers of
add_user_coauthor_trailer / add_pr_collaboration_note. Since then the
agent has been driving commits and PRs entirely via gh, with no
attribution back to the Slack/Linear/GitHub user who triggered the run.
Resolve the triggering user's identity in get_agent (reusing the
existing authorship helpers) and inject a Collaborative Attribution
section into the system prompt with the exact trailer and PR-body note
to use. The section is only rendered when an identity is resolvable, so
runs without a known triggering user are unchanged.
* feat: add optional Slack Assistants API typing status indicator
Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.
* fix(slack): drop redundant clear, add status heartbeat across model calls
- Slack auto-clears the typing indicator on bot post; remove the explicit
assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
that refreshes it on every model tick so it stays visible across long
agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
way out per Slack docs); no scope or app-config change required.
* feat(slack): contextual status text + rotating loading_messages
- set_slack_assistant_status now accepts an optional loading_messages list
(capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
assistant message's tool calls (e.g. "searching the codebase…" after
grep, "running commands…" after execute), falling back to the default
"is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
contextual status on each refresh.
* fix slack assistant status lifecycle
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat(open-swe): always anchor findings to a single line
GitHub renders multi-line review comment ranges as walls of context above
the comment, which buries the point. Match Devin Review's behavior and
always collapse `end_line` to `start_line` so each finding is anchored to
the single most relevant line.
Removes the (now unused) `MAX_FINDING_RANGE_LINES` cap and
`clip_finding_range` helper.
* fix(open-swe): collapse end_line before diff-range check
Anchoring a finding to a 25-line range (e.g. an entire function) makes
GitHub render the whole block as context above the comment, burying the
review text. Cap ranges at 10 lines and collapse to the start line when
exceeded; small ranges (≤10 lines) still render as a multi-line anchor.
Updated the reviewer prompt to anchor tightly rather than range-select
the whole function.
* feat(open-swe): cap reviewer suggestions at 4 lines
Long suggestion blocks read as the reviewer rewriting the code rather
than flagging an issue, which clutters PR comments. Steer the reviewer
toward description-only findings for non-trivial fixes, and enforce the
cap in `add_finding` / `update_finding` so suggestions over 4 lines are
dropped (the finding itself still publishes).
* fix(open-swe): don't clobber prior suggestion on over-cap update
`update_finding` was setting `suggestion=None` whenever clip_suggestion
dropped the input, which silently wiped any existing suggestion on the
finding. Distinguish the three cases: empty string clears, valid value
sets, over-cap value is rejected without touching the stored field.
The host-side prep (clone + git diff in the sandbox) was producing empty
diffs for some PRs. The agent then silently published a "no issues
found" review (e.g. langchainplus#24069, #24458). Fix is to rip out the
prep entirely and instruct the agent to fetch the diff itself with
\`gh pr diff <num> --repo <owner>/<repo>\` (or the compare API on
re-review). \`add_finding\` line-range validation is skipped when no
\`diff_line_set\` is provided in config — we trust the agent's anchors.
We can revisit a host-side diff path once we understand why the prep
was failing; this hotfix unblocks reviews on private PRs in the
meantime.
When `_ensure_repo_checked_out` or `compute_diff_in_sandbox` fail, the
reviewer used to log+continue and hand the agent an empty diff. The
agent would then emit a misleading "no issues found" review and the
underlying error stayed buried in server logs.
Now the helpers raise on non-zero exit codes (with the sandbox output)
and the prep block re-raises, so the failure shows up in the LangSmith
run trace and we can debug the real issue.
* feat(open-swe): trigger reviewer agent from `@open-swe review` PR comment
Mirrors the Slack `@open-swe review` flow on GitHub: a comment containing
`@open-swe review` (optionally followed by a PR URL) on a PR triggers the
reviewer agent. Without a URL it reviews the commenting PR; with a URL it
targets that PR. Works for `issue_comment`, `pull_request_review_comment`,
and `pull_request_review` events, gated by the existing reviewer repo
allowlist and reusing `trigger_pr_review_from_ref`.
* fix(open-swe): require URL after `@open-swe review`, don't swallow trailing text
The previous regex matched any non-whitespace token after `review`, including
across newlines. Comments like `@open-swe review\nthanks!` parsed as
`(True, "thanks!")`, which then failed PR-URL parsing and was silently
dropped — the user got no review and the comment never reached the regular
PR-comment handler.
Restrict the optional URL token to `https?://\S+` so non-URL trailing text
falls through to `process_github_pr_comment` instead of being eaten by the
review-command branch. Adds regression tests for the multiline and
trailing-word cases.
When a Slack user kicks off a PR review with `@open-swe review <pr-url>`,
the reviewer agent now posts a one-line summary back to the Slack thread
when it finishes — either "No issues found" or "found N potential
issue(s)" with a link to the GitHub review.
The reviewer agent has no Slack tools by design, so the summary is sent
host-side from `publish_review` after the GitHub review POST succeeds.
The Slack channel/thread_ts is persisted on reviewer thread metadata at
trigger time and read back on publish. Re-reviews triggered by push
events stay silent in Slack to avoid noise on the original thread.
The reviewer agent was writing a 1–2 sentence "top-level take" as the
review body, which produced noisy paragraph-style summaries on PRs
("Reviewed the PR. The new ALLOWED_GITHUB_REPOS allowlist…"). Devin's
review comment is just a one-liner (`✅ No Issues Found` or
`**Devin Review** found N potential issue.`), and was preferred in the
internal A/B vs Graphite.
Drop the `summary` parameter from `publish_review`; render a fixed,
host-formatted body in `render_review_body` instead. Update the reviewer
prompt to forbid prose summaries.
* fix(reviewer): log every push/close early-return so 'silent ignore' is debuggable
Pushes to PRs that haven't had a first review fall through the watch
handler because the reviewer thread doesn't have kind=reviewer set.
Without log lines on the early-return paths, this scenario was
indistinguishable from 'webhook reached the handler at all' in the
hosted log stream.
Now every early-return logs at info or debug:
- info when a real PR exists but the reviewer thread isn't set up
(with a hint pointing at the trigger paths the user can use)
- info when the repo isn't in the reviewer allowlist
- debug for benign skips (non-branch refs, branch deletions,
already-reviewed head_sha)
* fix(reviewer): always post a summary review, even with no findings
The publish_review tool gated POSTing on `inline_comments or summary`,
so when the agent called publish_review() with no args on a clean PR
the result returned `success: true` but no GitHub review was posted —
the user got silence instead of a "no issues found" comment.
- Drop the gate so publish_review always POSTs.
- Friendlier no-findings render: `**No issues found.**` when the
findings list is empty, vs. `**No issues at or above \`<sev>\`
severity.**` with hidden count when only sub-threshold findings
exist. Agent summary renders below.
- Prompt now requires the agent to always pass a `summary` so the
body is meaningful; calls out specifically not to skip on a clean PR.
* feat: implement reviewer findings, publish_review, and watch mode
Build out the reviewer agent end-to-end against the design in
REVIEWER_DESIGN.md:
- Findings as first-class state on the reviewer thread metadata
(`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line
ranges, suggestion text for ```suggestion blocks, github_review_comment_id
for cross-run reconciliation, diff_hunk for UI rendering. Thread-level
metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a
future frontend can list reviewer threads via the langgraph SDK.
- Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff,
compute_diff_line_set for in-diff validation, extract_diff_hunk for
caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA
diffs against the prepped repo.
- Tools: `add_finding` (validates against the diff line set so out-of-diff
ranges fail at creation, not at GitHub-publish), `update_finding`,
`list_findings`, `publish_review`. The reviewer agent's tool list is
swapped from `[]` (direct shell `gh api` calls) to these four.
- Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`):
one POST /reviews call with body + inline comments + ```suggestion blocks,
per-comment IDs stored back on findings, GraphQL `resolveReviewThread`
fired for findings transitioning open->resolved on a re-review.
- Reviewer graph: deterministic clone-or-fetch + checkout in the factory
before the agent's first model call (warm- and cold-path symmetric);
computed diff and in-diff line set passed via runnable config; system
prompt rewritten for the single-evolving-findings model, severity ladder,
in-diff-only discipline, and watch-mode reconciliation flow.
- Watch mode in webapp.py: `push` event + `pull_request` closed/reopened
added to supported events. New `process_github_push_event` resolves the
open PR for the pushed branch, gates on the reviewer thread's `watch`
flag, builds a re-review configurable, and triggers a run on the same
canonical thread. `process_github_pr_close` toggles watch on
closed/reopened. `set_reviewer_thread_metadata` is called on first
review to install `kind=reviewer` + PR identity + watch=True.
- Eval harness: target.py now extracts `add_finding` calls (mapped to the
legacy {file, line, body, severity} shape the judge expects) and passes
the right configurable so the prep step has base/head SHAs.
- Tests: new unit suites for findings helpers, diff parsing, finding tools,
publish rendering + GraphQL resolve, and watch-mode webhook handlers
(push triggers re-review only when watching, idempotent on unchanged
head SHA, PR close disables watch). Updated existing reviewer-webhook
tests to mock `set_reviewer_thread_metadata`.
- REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context
into REVIEWER_DESIGN.md.
* fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL
Address PR #1253 review findings:
- compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag
(`option no-prefix takes no value` — every prep run was failing
silently and the agent saw an empty diff).
- compute_diff_in_sandbox grew a `merge_base` flag. First-review path
now uses three-dot `base...head` (the merge-base diff GitHub renders
on Files-changed) so we don't pick up changes that landed on the base
branch after the PR diverged. Re-review delta keeps two-dot
`last_reviewed_sha..head` since that's exactly the new commits.
- publish_review skips findings that already carry
`github_review_comment_id`. Without this, watched re-reviews
re-posted every previously surfaced finding, and only the most-recent
duplicate's id would later resolve when the issue got addressed.
- fetch_review_comments URL now includes `{pull_number}` —
`/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments`
is the canonical endpoint; the old form 404s, so comment ids were
never stored and watch-mode resolution couldn't run.
Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix`
flag in the executed command, and that publish_review does not re-post
findings whose `github_review_comment_id` is set.
* fix(reviewer): default publish cap from 15 to 4
A clean PR with one critical issue padded out by three lower-severity
findings is fine; fifteen is review spam. The agent can override per
call when a PR genuinely warrants more.
* feat: only post Slack 'Working on it!' on first thread mention
* feat: randomize Slack trace reply phrase
Pick from a small list of friendly phrases instead of always saying
'Working on it!' so the bot feels less robotic. Explicit messages (e.g.
'Taking a look...' from PR review path) are unaffected.
* adjust phrases
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
The Slack mention path already routes mid-run messages through the
store-based queue + check_message_queue_before_model middleware (the
same path Linear uses), so multitask_strategy="enqueue" on runs.create
was a leftover no-op — that line is only reached when is_thread_active
returned False. The busy-path payload was also hardcoding image_urls=[]
which silently dropped any images attached to mid-run Slack mentions;
forward the resolved image_urls so the middleware can rebuild image
blocks like Linear does.