A push that lands while a reviewer run is in flight is delivered as a
queued message into the still-running first-review run, whose
configurable still has re_review=False. The empty-review guard in
publish_review only skipped the 'No issues found' summary when
is_re_review was True, so the queued reconcile published a second,
duplicate top-level 'No issues found' review.
Key the empty-review skip off actual PR state instead: add
open_swe_review_exists(), which detects the marker render_review_body
embeds in every Open SWE review body, and skip the summary when a prior
Open SWE review already exists (regardless of the re_review flag). Fails
open on API error so a genuine first review is never suppressed.
* fix: scope public reviewer tokens
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: simplify reviewer token wiring; fix push re-scope + red test
- Remove the redundant _check_or_recreate_sandbox_for_proxy /
_refresh_github_proxy_or_recreate_for_proxy wrappers and call the
underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
payload lacked repo privacy/id but PR metadata reveals a public repo, so
reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
and add coverage for _reviewer_token_for_repo + push-event scoping.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Replace the floating absolute-positioned prompt bar (gradient overlay +
128px bottomInset) with a plain flex column: MessageView as flex-1 with a
shrink-0 prompt bar beneath it, matching open-swe-app's ChatView.
Also remove the redundant always-on 'Files Changed' panel — MessageView
already renders an inline TurnChangedFilesCard per agent turn, and the
duplicate was being clipped by the floating prompt bar. This eliminates
the large whitespace gap and the broken files-changed box.
* fix: deliver Slack account-link prompt as a visible threaded reply
Blocked Slack users got no prompt at all. Prod logs show chat.postEphemeral
returns ok, but ephemeral messages are silently dropped in Slack's assistant
threads (where Open SWE runs), so the user sees nothing. Post the prompt as a
normal threaded reply instead — the same channel the agent uses to reply.
* fix: deliver Slack auth-failure prompt as a visible threaded reply
leave_failure_comment() tried an ephemeral message first and only fell back
to a thread reply on failure. Ephemeral messages succeed (ok) but are dropped
in Slack's assistant threads, so the fallback never fired and the user saw no
auth-failure prompt. Post the visible threaded reply directly, matching the
account-link prompt fix.
* fix: prompt blocked Slack users with a generic, token-free dashboard link
Addresses the review findings that posting the per-user account-link token /
auth URL in a visible thread lets any channel member bind their GitHub account
to the triggering user's Slack identity.
Drop the per-user signed link entirely. Both the account-link prompt
(_post_account_link_prompt) and the runtime auth-failure prompt
(leave_failure_comment) now post a plain dashboard settings link
(build_settings_url) as a visible threaded reply. The user signs in with GitHub
from their own session and connects Slack via verified OIDC on the settings
page — no secret in the thread, nothing to hijack, and no DM machinery.
* feat: nudge first-time users to connect Slack from the dashboard home
Show a Connect Slack banner on the agents landing page whenever Slack OAuth is
enabled and the user hasn't linked Slack yet. A first-time user (no Slack
mapping) sees it immediately after signing in; it disappears once connected.
* feat: prompt first-time users to connect Slack via a dialog
Replace the inline Connect Slack card on the agents home with a modal dialog
(Base UI). It opens automatically once the mapping query resolves to
"not connected" and closes itself once Slack is linked; "Maybe later" dismisses
it for the session. No new dependency — uses the design system's Base UI.
* copy: frame Slack connect as resolving the user's GitHub account
Drop 'act/reply on your behalf' wording across the connect-Slack dialog, the
Slack thread prompts (blocked + auth-failure), and the settings description.
Connecting Slack lets Open SWE resolve the user's GitHub account when they tag
it in Slack.
Replace the double-attribution PR body footer (_Opened collaboratively by
{user} and open-swe._) with a single Cursor-style footer linking to the
project. Since PRs are now opened as the triggering user, the user no longer
needs to be named in the footer. The commit Co-authored-by trailer is kept.
Legacy footers are migrated on PR updates.
* fix: clearer Slack prompt for users who must sign in with GitHub
The Slack gate labeled any mapped-but-tokenless user as having an
"expired or revoked" authorization, keyed off whether a *mapping* row
existed. Most blocked users are team members in the legacy hardcoded
mapping who have simply never completed a dashboard GitHub login, so the
message was misleading and steered them toward reconnecting Slack (which
never creates a token).
Base the prompt on whether an oauth_tokens *record* exists:
- no record -> ask the user to sign in with GitHub and connect Slack
- record but unusable -> ask the user to sign in with GitHub again
Add has_access_token_record() to distinguish the two cases.
* fix: guard token-record check so a store failure still prompts sign-in
The has_access_token_record() lookup ran outside the defensive handling
around token resolution. If the store read fails it would raise before
posting the sign-in prompt and clearing the Slack assistant status.
Wrap it like get_valid_access_token: on failure, default to the
sign-in/connect prompt.
Add an open_pull_request tool that creates a new PR via the GitHub REST
API using the triggering user's OAuth token (resolved by login from the
dashboard store), so the PR creator is the user rather than open-swe[bot].
Falls back to the GitHub App installation token for GitHub-triggered runs,
unmapped users, and bot-token-only deployments.
The user token never enters the sandbox: clone/push/comments still go
through the bot proxy via gh. The agent is steered to use the tool only
for OPENING a new PR; updates (body edits, mark ready) and pasted/existing
PRs continue to use gh pr edit. Existing-PR (422) returns the open PR's
URL so re-runs don't create duplicates.
* docs: document dashboard + refresh local dev setup in INSTALLATION
- add the web dashboard (ui/) and its FastAPI backend API to the setup flow
- document dashboard-login OAuth (GITHUB_APP_CLIENT_ID/SECRET, second callback
URL) as distinct from the LangSmith-brokered agent-runtime OAuth
- add new env vars: DASHBOARD_API_BASE_URL/BASE_URL/JWT_SECRET/ALLOWED_ORIGINS,
CONFIGURED_ADMINS, X_SERVICE_AUTH_JWT_SECRET, LANGGRAPH_URL, SANDBOX_TYPE,
REVIEWER_OUTCOMES_DATASET, PUBLIC_REPO_ORG_GATE, SLACK_CLIENT_ID/SECRET/TEAM_ID,
SLACK_REPO_OWNER/NAME
- add "Sign in with Slack" OIDC account-linking setup
- rewrite local dev: make dev serves 3 graphs + FastAPI on :2024; new step for
running the UI (bun) on :3000, incl. the required DASHBOARD_ALLOWED_ORIGINS
CORS setting and cookie wiring
- correct production langgraph.json to three graphs; document Vercel UI deploy
- add dashboard verify + troubleshooting sections
- README: dashboard feature bullet + updated install blurb
* docs: clarify dashboard API setup
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: open Slack-triggered PRs as the triggering user
Route the Slack per-user GitHub token through the dashboard OAuth store
(the backend the self-service link prompt populates) and block runs that
lack a valid user token, prompting the user to (re-)link. Per-user OAuth
now wins over bot-token-only mode for mapped Slack/dashboard users.
Flip commit/PR authorship across all sources: the triggering user is the
commit author (via repo-local git identity using their resolvable GitHub
noreply email) and open-swe[bot] is the Co-authored-by collaborator.
* fix: address PR review — shell-escape commit identity, fix token cache impersonation
- Shell-escape the triggering user's name/email with shlex.quote before
embedding them in the repo-setup `git config` command, so a name like
O'Connor (or a crafted one) can't break or inject into the command.
- Stop consulting the shared thread-metadata token cache in
_resolve_dashboard_user_token. Slack thread ids are shared across the
conversation, so a cached token from a prior triggering user could be
returned for the current github_login. Always resolve by login from the
dashboard OAuth store instead.
* feat: dashboard self-service user mapping + UI cleanup
- Add session-scoped GET/PUT /dashboard/api/my-mapping so users can set their
own work email / Slack member ID (keyed by their GitHub login, source=self).
- Slack account-link prompt now redirects to Profile Settings after auth.
- Rename "My Settings" -> "Profile Settings" and "Cloud Agents" -> "Open SWE
Agent"; remove the Integrations tab/section (folded out, low value for now)
and redirect /integrations to Profile Settings.
- Add a "User mapping" section to Profile Settings (work email used by Slack
and Linear, optional Slack member ID).
- Make dashboard auth cookies scheme-aware: Secure;SameSite=None over HTTPS,
non-Secure;SameSite=Lax over http://localhost so local login works.
* feat: self-service Slack account linking via Sign in with Slack (OIDC)
Replace the spoofable manual work-email/Slack-ID form with a verified
"Sign in with Slack" flow so a logged-in GitHub user can only ever link
their own Slack identity.
- New agent/dashboard/slack_oauth.py: OIDC authorize URL, code exchange,
userInfo identity parse, optional workspace gate, configured check.
- routes.py: session-gated GET /slack/login and /slack/callback that upsert
the mapping from Slack-verified user_id + email (source=slack_oauth).
Remove the spoofable PUT /my-mapping; expose slack_oauth_enabled on /me.
- UI: drop the editable inputs; add a Connect Slack button + status to the
User mapping section.
Admin-managed mappings are unaffected and still resolve at trigger time.
* fix: include GitHub username in collaboration footer
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: clarify legacy footer replacement
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
When sending a message to a thread from the Agents chat UI, the run used
the OAuth profile email (which can be a personal account that isn't an org
member). For Slack-originated threads resumed from the UI, auth routes
through the email-based path and fails ("Could not find a LangSmith account").
Resolve the run email via the GitHub->email mapping first, falling back to
the profile email only when no mapping exists, matching the github-source
auth path.
* chore: remove legacy mapping + admin profiles section, page user mappings
- Delete the hardcoded github_user_email_map.py and the one-time
POST /admin/user-mappings/import endpoint + Import legacy mapping UI
button (the Store is now the sole source of truth post-import).
- Remove the per-user profiles admin section and its GET/PUT
/admin/profiles endpoints + ProfileForm component; users still manage
their own profile via My Settings.
- Page the user mappings list: /admin/user-mappings now takes
page/page_size and returns {items,total,page,page_size}; the admin UI
shows 20 rows per page with Previous/Next controls.
* Address review: fix stale import-button doc + clamp mappings page on shrink
* Replace hardcoded GitHub-email map with Store-backed user mapping
Move the static GITHUB_USER_EMAIL_MAP to a Store-backed bidirectional
mapping (GitHub login <-> work email <-> optional Slack ID) with an
in-process cache, self-service onboarding, and admin management.
- agent/dashboard/user_mappings.py: Store CRUD + login/email/slack-id
indexes, sync cache readers for hot paths, async fallthrough, and a
bulk_import that preserves existing richer records.
- Migrate all read sites (auth.py, agent_overrides.py, authorship.py,
github_comments.py, webapp.py x2) off the dict.
- Unmapped Slack tags now run on the GitHub App installation token
(use_installation_token_fallback) and get an ephemeral "link your
GitHub account" prompt carrying the Slack id + email via a signed
account-link token threaded through the OAuth state.
- OAuth callback completes a self-service (org-gated) mapping from that
token, falling back to the verified GitHub email.
- Admin CRUD endpoints + one-time legacy import; dashboard UI section.
- Legacy dict retained only as the import payload (no longer read).
Tests: mapping store, account-link round-trip + completion, mapped vs
unmapped Slack flows; existing trust-gate tests updated to prime cache.
* Address review: cold-cache email resolution + stale alias de-indexing
- agent_overrides: add resolve_login_from_email_async that falls through to
the Store on a cold cache; use it at the async repo-resolution call sites
(Slack repo config, Linear comment, owner-metadata) so a mapped user still
resolves to their GitHub login + dashboard default_repo on a fresh worker.
- user_mappings.upsert_mapping: de-index the existing login before re-indexing
so a changed email/Slack id no longer leaves stale aliases resolving to the
login in-process.
- Tests for both fixes; update Slack repo-config test to patch the async resolver.
* Inject PR title and body into reviewer context
The reviewer agent previously received the PR url, number, and SHAs but not
the PR title or description, so it sometimes missed the original intent of
the PR. Fetch the title/body fresh from the GitHub API on every run (never
cached) so edits to the title/description are reflected on re-reviews, and
inject them into the first-review, re-review, and finding-reply contexts as
an untrusted-data block (author-controlled text, guarded against prompt
injection).
* Harden PR overview escaping against whitespace-padded closing tags
* Lock dashboard login to GitHub org members
Add an org-membership gate to the dashboard OAuth callback. After
resolving the GitHub login, enforce_org_login_gate(login) checks the
existing ALLOWED_GITHUB_ORGS allowlist before issuing a session.
- Reuses ALLOWED_GITHUB_ORGS (no new config knob) and
is_user_active_org_member (installation-token check, so no extra
OAuth scope and private memberships are visible).
- Fail-open when unset/blank so existing deployments keep working;
fail-closed on API errors.
- Gate runs before the session cookie/token is persisted.
Adds unit tests and documents the behavior in INSTALLATION.md.
* docs: document Organization Members permission required for org login gate
* fix: reset stale sandbox creation sentinel
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: treat SANDBOX_CREATING as a timestamped cross-process lock
Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.
* feat(analyzer): outcomes dataset + bootstrap/continual split via skills
Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.
- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
(openswe-reviewer-outcomes), keyed deterministically per finding+source.
Emit points wired into update_finding, resolve_finding_thread, and the
GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
continual-learning), served as virtual files via a CompositeBackend /skills/
route + StateBackend (seeded into the run files channel at invoke time, never
written to the sandbox). Mode is set by the launcher; continual runs fall
back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
continual playbook.
Tests for outcome label mapping, skills helper, and cron idempotency.
* fix(analyzer): anchor continual cron runs to a real thread_id
The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.
* refactor(analyzer): move cron lifecycle calls out of the review-styles store
Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.
* refactor: hoist reviewer_outcomes imports to module level
Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: surface all of a user's agent threads in the Agents UI
Previously the Agents UI only listed and allowed opening threads with
source=dashboard. Threads triggered from GitHub, Slack, or Linear were
hidden and could not be opened.
- Persist owner-identifying metadata (source, github_login,
triggering_user_email, source_context) onto the thread for each
webhook-triggered main-agent run via upsert_agent_thread_owner_metadata,
resolving a github_login from the triggering email where possible.
- Relax dashboard thread listing/ownership to surface github/slack/linear
threads owned by the logged-in user (matched by github_login or email),
and preserve the original source + reply-routing context when continuing
such threads from the UI.
- Add a source field to the thread summary and render a source icon
(GitHub/Slack/Linear) on thread items in the sidebar and run cards.
* Normalize triggering email when persisting thread owner metadata
Mixed-case Slack/Linear emails were stored verbatim, but the dashboard
searches with a lowercased value, so those threads never surfaced for
their owner. Normalize at write time to match _thread_owner_email.
* fix: preserve pushed commit history
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* refactor: consolidate no-force-push prompt guidance
Collapse the repeated force-push rule into a single IMPORTANT block,
drop the confusing rebase-onto-origin step, and stop mandating a
no-op trailer-only commit for already-pushed work.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: propagate Slack reply errors
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: classify Slack 429 as rate_limited with retry-after
Slack chat.postMessage rate limiting returns HTTP 429, which raise_for_status
turned into a generic http_error and told the agent to retry immediately,
ignoring Slack's retry window. Special-case 429 (threading Retry-After) and
normalize the ratelimited body code before the generic HTTP path so the
existing rate_limited hint actually fires.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: include reviewer trace links
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: move reviewer trace-link toggle to dashboard
Replace the OPEN_SWE_REVIEW_TRACE_LINK_ENABLED env var with a team-level
'Trace Links' toggle in the Open SWE Review dashboard tab. The toggle is
read per-publish via get_team_review_trace_links_enabled(); the per-run
review_trace_link_enabled config override still forces it off.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: structured reviewer comments + auto resolution comments
Restructure inline review comment bodies (severity emoji, bold title from
the first line, line reference, feedback footer) without duplicating the
first description line, and post an automatic resolution/dismissal comment
to the GitHub thread when a finding is resolved or dismissed.
- Centralize render_resolution_comment in reviewer_publish; fix a crash when
last_reconciliation_note is None and drop the misleading generic fallback.
- Post the resolution comment in resolve_finding_thread (the normal
update_finding path), not only in publish_review, so it actually fires on
re-review. Dedupe via github_posted_resolution_comment_ids.
- Add pytest coverage for rendering and the resolution-comment flow.
* fix: post resolution comment to every closed thread in resolve_finding_thread
Per-thread iteration (matching _resolve_threads_for_resolved_findings) so
duplicate threads after the first also receive the resolved/dismissed
explanation before being closed.
* feat: upgrade default agent + reviewer model to Opus 4.8
Replace Opus 4.7 with Opus 4.8 (claude-opus-4-8) as the supported
Anthropic model surfaced in the profile editor and used by the main
agent and reviewer graphs. Effort levels (low/medium/high/xhigh/max)
and the high default are unchanged, matching the official Opus 4.8
docs. Updates eval config comment and tests accordingly.
* fix: provider-aware fallback for stale stored model ids
Dropping claude-opus-4-7 from the supported set meant persisted
profile/team-settings still holding it failed the SUPPORTED_MODEL_IDS
check and fell through to default_model_pair() — a cross-provider jump
to the OpenAI global default.
Add provider_fallback_pair: when a stored id is no longer supported but
its provider still has a supported model, resolve to that provider's
newest supported model (anthropic:claude-opus-4-7 -> 4.8), preserving
effort when valid. Resolution order is now: valid stored pair ->
same-provider fallback -> global default_model_pair(). Profile overrides
keep deferring to the team default when no model is set or the provider
is unknown.
* feat: reconcile reviewer findings with PR threads
* fix: harden reviewer finding reply handling
* fix: queue reviewer finding reply body
Ensure review-comment replies that arrive during an active reviewer run include the sanitized reply body in the queued reassessment prompt.
* fix: apply reviewer reply formatting
Apply the repository formatter so the reviewer reply handling fix passes CI format checks.
* fix: dedupe reviewer comments from PR state
Use GitHub review-thread markers to repair reviewer publication state before posting or resolving findings, so re-reviews do not duplicate comments and resolved findings close all matching PR threads.
* fix: require all duplicate reviewer threads resolved
Avoid treating a marker-backed finding as resolved when only one duplicate thread is outdated while another matching thread remains open.
* fix: route PR review requests through agent tool
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: align public repo gate test with agent routing
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* fix: allow app-token PR review requests
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* reviewer: fetch PR diff via GitHub API to re-enable add_finding validation
The previous hotfix in reviewer.py set diff_line_set=None because the
sandbox-based diff prep was sometimes producing empty diffs. That made
every bad anchor a publish-time 422 instead of a creation-time
rejection — the agent burned tokens producing unanchorable findings,
and we had to add a publish-time retry safety net (#1338) to clean up.
Fetch the PR's unified diff via the GitHub REST API at reviewer
startup and populate diff_text + diff_line_set so add_finding can
reject bad anchors immediately. The API path is reliable and is the
same diff GitHub validates against when posting inline review
comments. If the fetch fails, fall back to the previous behavior
(validation disabled, publish-time retry handles it).
Also extract the PR-diff fetch into reviewer_diff.fetch_pr_diff so
both reviewer.py and publish_review.py share one implementation
instead of two copies.
* reviewer: make diff_line_set validation side-aware
compute_diff_line_set previously returned only new-side line numbers,
so re-enabling add_finding's validation would wrongly reject findings
with side=LEFT (deleted-line bugs whose only anchor is an old-side
line). Return {file: {"RIGHT": {new_lines}, "LEFT": {old_lines}}}
instead, and have is_range_in_diff select the matching side from the
finding's recorded side. add_finding and publish_review's retry
filter both pass the finding's side through.
* publish_review: drop unresolvable findings and retry once on GitHub 422
GitHub returns 422 with 'Path could not be resolved' or 'Line could not be
resolved' when an inline comment anchors to a file/line not in the PR diff.
Previously the agent retried publish_review with byte-identical args
multiple times before draining to skipped_empty_re_review=true, silently
losing findings.
- reviewer_publish.post_pull_request_review: parse 422 body and tag with
_error_kind='unresolved_anchor' plus _raw_errors so callers can act.
- tools/publish_review._publish_review_async: when that signal fires,
cross-check each finding's range against the run config's diff_line_set,
drop the bad ones, and re-POST once with only the valid findings. Return
unresolvable_findings + hint so the agent calls update_finding instead of
retrying the same payload.
- reviewer.py: one-line prompt addendum telling the agent that
unresolvable_findings means update_finding, not retry.
- tests: cover 422 tagging (path + line), the drop-and-retry success path,
the retry-still-fails path, and the don't-blind-retry path when no
diff_line_set is available.
* publish_review: fetch PR diff on demand for 422 retry filter
Reviewer runs clear configurable['diff_line_set'] before the agent
starts, so the unresolved-anchor retry path had no diff data to filter
against — in the reachable production case it dropped nothing and
returned success=False with empty unresolvable_findings, losing the
otherwise-valid comments.
Fall back to fetching the PR's unified diff via the GitHub REST API
and recomputing the line set on the fly when no cached set is
available. The cached set is still preferred when present.
---------
Co-authored-by: issues-agent <issues-agent@langchain.dev>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix(reviewer): inject existing PR review threads into reviewer context
The reviewer agent was filing the same inline comment on every re-review
because it only saw findings recorded on its own thread metadata — not
the live PR review-thread state on GitHub. When a previous finding was
still open (code unchanged, or a human reply explained it), the agent
rediscovered the same defect on the next push and called `add_finding`
again, producing duplicate comments.
This change fetches the PR's review threads (across all reviewers, with
replies and isResolved status) via GraphQL and renders them into the
first-review and re-review contexts as a "Pre-existing PR review
threads" block. The system prompt now lists overlap with that block as
a hard "Do NOT file" rule, and treats threads addressed by a human
reply as resolved.
This also gives the reviewer comment-awareness on its very first run
on a PR, so it skips findings already raised by another reviewer or
bot.
* fix(reviewer): wrap PR review threads in untrusted-data XML block
Addresses the reviewer comment on this PR
(https://github.com/langchain-ai/open-swe/pull/1331#discussion_r3295497533):
PR review comment bodies are attacker-controlled (anyone who can comment
on the PR can put anything in them), and they were being concatenated
into the reviewer's system prompt with instruction-priority.
Switches the existing-threads section from a Markdown block to an XML
data block:
<pr_review_threads>
<thread location="path:line" status="open">
<comment author="open-swe[bot]">
<body>...</body>
</comment>
<comment author="romain-priour-lc">
<body>We added defaults in the template</body>
</comment>
</thread>
</pr_review_threads>
The system prompt now explicitly names the wrapper, tells the agent that
everything inside it is untrusted data from the PR (not instructions),
and that prompt-injection payloads inside bodies must be disregarded.
We keep the bodies so the agent can actually read engineer replies —
that's the whole point of comment-awareness — but they're delimited as
data, not concatenated as prose. Modern frontier models are well-trained
to honor this contract.
Additional defenses:
- Author logins are validated against the GitHub username grammar; any
unexpected value is rendered as "unknown" so the `author` attribute
can't smuggle freeform text.
- Literal closing tags (`</body>`, `</pr_review_threads>`, etc.) in
bodies are neutered so a body can't break out of its wrapper.
- Body length is capped at 4000 chars per comment to bound the prompt.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix(dashboard): stream agent tokens to the Agents UI
The dashboard SSE endpoint was serializing the SDK's `StreamPart`
namedtuple via `str(part)`, so every event arrived at the browser
with `event = "StreamPart(event='values', data=..., id=None)"`. The
frontend's `event.startsWith("messages")` branch never fired and
every event fell through to a query refetch — which is why tool
calls appeared to stream (state refetch) but assistant text did not.
Also opt the run into `messages-tuple` stream mode so the server
actually buffers per-token `AIMessageChunk` events, and forward
`Last-Event-ID` so reconnects resume the in-flight run instead of
restarting from `seq=0`.
* fix(dashboard): drop unsupported stream_mode arg from join_stream
threads.join_stream's stream_mode is a ThreadStreamMode literal
(run_modes/lifecycle/state_update); passing run-level modes is
invalid. The default 'run_modes' already replays the modes the
run was created with, which we set correctly on runs.create.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(reviewer): inline AGENTS.md into reviewer system prompt
Fetches AGENTS.md from the PR's head_sha via the GitHub contents API
during reviewer setup and inlines it as a "Repository conventions"
block in the system prompt. Mirrors how the per-repo review style
prompt is already wired.
The main agent has long had a mandatory step to read AGENTS.md after
cloning, but the reviewer often skips cloning entirely (it can
`gh pr diff` directly), so it never saw the file. Loading it
deterministically means the reviewer judges findings against the
project's own conventions instead of relying on the model to fetch
the file itself.
* fix(reviewer): fetch AGENTS.md from base_sha, not head_sha
The reviewer inlines AGENTS.md into its system prompt. Reading from
head_sha means a PR author can edit AGENTS.md in the same PR being
reviewed and smuggle instructions like "ignore all bugs" / "publish
no findings" into the reviewer's prompt. Switch to base_sha (the
target branch's pre-PR state, which is trusted) and update the prompt
text to reflect that the contents come from the base, not the head.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Bring the agent-facing docs back in sync with the codebase: the reviewer
and review_style_analyzer graphs, the dashboard router and Agents UI,
auto-review on PR opened/ready_for_review, the current middleware order
in get_agent, model/profile/team-default resolution, and the leaner
reviewer middleware stack.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The base reviewer prompt had grown into a curated catalog of defect
patterns and grep recipes tailored to the Martian eval set (Optional.get,
forEach(async), React keys, ERB templates, ORM increment, X-Frame-Options,
etc.). That bias toward one dataset hurts generalization across the
repos the reviewer actually runs on. Trim it down to universal review
discipline (the bar, do-not-file rules, severity rubric, publish
discipline, a generic workflow) and rely on per-repo prompts in memory
to carry repo-specific patterns at runtime.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: auto-review PRs on opened / ready-for-review
Trigger Open SWE Review on `pull_request` actions `opened` and
`ready_for_review` against the canonical reviewer thread (no need to
request open-swe[bot] as a reviewer). `converted_to_draft` now also
flips watch=False on the existing reviewer thread.
Draft PRs are gated by a tri-state user setting on the profile:
inherit team default, always on, or always off. The team-wide
`review_draft_prs` setting is the org-wide default; each user can
override it in My Settings.
External contributors with no Open SWE profile fall back to the team
default.
* fix: PR review comments — auth source + draft-aware watch toggle
- `process_github_pr_ready` now dispatches with `source="github"` so the
auth resolver finds the bot token persisted on the thread. The previous
`source="github_auto"` fell through to the email-based path in non
bot-token-only deployments and failed with a missing-user-email error.
- `converted_to_draft` no longer unconditionally clears `watch`. When the
PR author's effective `review_draft_prs` setting is on, watch stays on
so subsequent pushes still trigger re-reviews while the PR is in draft.
* feat(reviewer): skip "no issues found" comment on empty re-reviews
A re-review run with no new findings to surface no longer posts another
"Open SWE Review: No issues found" comment on the PR. The "no issues"
summary now only appears on the first review of a PR — matching Devin's
behavior, where subsequent reviews are silent unless there's something
new to flag.
Resolved-thread reconciliation and ``last_reviewed_sha`` persistence
still happen on the skipped path, so findings the user just fixed still
get their GitHub threads marked resolved, and the next push event sees
an up-to-date dedup SHA.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>