prepare_review_repo is best-effort, but the system prompt unconditionally
told the agent the repo is checked out at the PR head. On a reused sandbox,
a dirty worktree (or a transient fetch failure) makes 'git checkout <sha>'
fail; prep returns False, the old checkout stays in place, and the reviewer
confidently reads stale code — observed as 'I rechecked the current head'
replies quoting pre-push code.
- checkout with --force so leftover worktree state can't block it, verify
HEAD matches the requested sha, tolerate 'git fetch --all' failures
- when prep fails, the prompt now warns the checkout may be stale and tells
the agent to re-fetch/checkout (or fall back to API file contents) before
trusting local files
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preflight stream proxy before SSE starts, let optimistic thread survive refetch
* fix: seed sidebar thread details as stale so mark-viewed fetch still fires
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: track PR lifecycle state per thread for sidebar
Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: consolidate PR state mapping into shared derive_pr_state helper
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: auto-recover from expired GitHub refresh tokens
When a user's GitHub OAuth refresh token was permanently dead (revoked or
expired), token refresh failed but get_valid_access_token still handed back
the known-stale access token, so dashboard GitHub calls kept 401ing until the
user manually logged out and back in.
Now we distinguish unrecoverable refresh failures (bad_refresh_token /
unauthorized_client) from transient ones: on an unrecoverable failure we drop
the dead stored authorization and return None, so callers prompt a clean
re-login. Transient failures still fall back to the stored token.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't delete fresh re-auth when stale refresh fails
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
- _origin_of: treat invalid ports as invalid origin instead of letting
urlparse ValueError turn CSRF rejections into 500s
- AgentsHome: reset submitting/draft when stream.submit rejects before a
thread id is minted, so the prompt isn't left disabled
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add View PR link in git panel header
Surface a clickable "View PR" link in the agent git panel header so
users can jump straight from a chat thread to its pull request instead
of hunting for the URL in the conversation. Uses the pr.url already
captured in thread metadata.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: use shared buttonVariants for View PR link
Style the View PR link with the shared buttonVariants (outline/sm)
helper instead of hand-rolled classes, matching the existing
anchor-as-button pattern used in login.tsx and AutomationsList.tsx.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: show PR diff in git panel diff tab
Adds /threads/{id}/pr-diff endpoint fetching PR files + contents from GitHub;
git panel prefers it and falls back to the live message-derived diff.
* feat: auto-expand git panel when a PR lands mid-run
Panel still defaults to collapsed and remembers manual toggles; the
auto-expand is ephemeral and not written to localStorage.
* fix: git panel always starts collapsed
Drop localStorage persistence of collapsed state; panel opens only via
manual toggle or when a PR lands mid-run, and re-collapses on thread switch.
* fix: address PR review findings
Require user OAuth token for pr-diff (no app-token fallback) so GitHub
enforces current repo access; render placeholder for binary/oversized files.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: refresh sandbox GitHub proxy token before mid-run expiry
GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve repo-scoped proxy token on mid-run refresh
Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: update passthrough stub for github_proxy_repositories param
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add right-click open trace context menu for threads
Add a context menu to sidebar thread rows exposing an Open trace action
that links to the LangSmith thread trace, surfaced via a new traceUrl
field on the thread summary.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: drop stray yarn artifacts from PR
Remove accidentally committed ui/.yarnrc.yml and ui/.yarn/install-state.gz
that broke immutable installs against the v1 lockfile, and gitignore yarn
artifacts to prevent recurrence.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: reuse thread_id var, drop redundant danger-color fallback
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: keep review check visible on follow-up commits and stop using neutral
The "Open SWE Review" check completed as `neutral` whenever findings were
surfaced, which GitHub renders as a confusing "neutral check" group. Always
complete it as `success` (informational/non-blocking), matching Devin and
Corridor — the finding count stays in the title and findings post as comments.
Also create a fresh check run on the new head SHA in the push re-review path:
GitHub only shows check runs on a PR's current head, so the check vanished
after a follow-up push (and the stale id settled on an outdated commit).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: surface settled review check when push leaves diff unchanged
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: report Open SWE Review as a PR check run
Auto-review dispatch now creates an in-progress 'Open SWE Review' check run
on the PR head SHA; publish_review completes it (neutral with findings,
success when clean). An after-agent hook fails the check if the run dies
before publishing. Requires the GitHub App's Checks: Read & write permission;
all calls are best-effort so a missing permission never breaks reviews.
* fix: address review feedback on check-run settling
Keep review_check_run_id when the completion PATCH fails so a later
publish or the after-agent hook can retry instead of hanging the check;
count out-of-diff findings toward the check conclusion.
* fix: retry failed check completion with the real publish conclusion
A transient PATCH failure after a successful publish previously left the
check id for the after-agent hook, which settled it as 'failure'. Persist
the intended result as review_check_pending_result and have the hook
prefer it over the generic failure fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Fable 5 is currently unusable with our API key. Remove it from the
selectable model list; it can be re-added later.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: server-side Datadog/LangSmith observability tools + team creds
Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on observability tools
Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
read-modify-write race dropping the other provider on concurrent saves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: async email resolution in observability authorization gate
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: honest publish_review reporting + structured thread-not-found errors
- Document skipped_empty_re_review and dry_run in the publish_review
docstring and add a closing-summary contract to the reviewer prompt so
the agent never claims a review was published when review_id is null.
- Raise ReviewerThreadMissingError from replace_findings on SDK
NotFoundError; add_finding/update_finding/publish_review return a
structured do-not-retry result instead of raising, so the agent reports
the blocker after one failure instead of retrying 10-30 times.
* fix: translate thread 404s across all reviewer tool boundaries
get_thread_metadata now raises ReviewerThreadMissingError instead of
swallowing a missing thread as {} (which produced misleading 'No finding
found' results), set_reviewer_thread_metadata translates the SDK 404 the
same way, and every reviewer tool entrypoint (add/update/list findings,
publish_review incl. eval dry-run, resolve/reply thread) returns the
structured do-not-retry result.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Publish had partial-failure windows that double-posted summaries or
corrupted findings state. This hardens the recovery paths:
- open_swe_review_exists is now tri-state (True/False/None). On a
pagination/API failure it returns None ("unknown") instead of False,
and the empty-summary dedup keys off the durable last_reviewed_sha
before consulting GitHub, so a transient failure never double-posts a
"no issues found" summary.
- Comment-id backfill matches strictly on the embedded open-swe marker;
the colliding (path, line, body) fallback is gone, so similar findings
no longer share a comment id and break resolve-on-fix.
- Review-id and comment-id stamping collapse into one guarded
read-modify-write (re-reads latest before writing), removing the
half-stamped intermediate states the prior multi-write flow left open.
- New mutate_findings primitive centralizes read-modify-write so finding
updates operate on the freshest persisted list and skip no-op writes.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: prep reviewer repo at init and load repo skills
Clone + checkout the PR head during reviewer agent init so SkillsMiddleware
can discover the repo's .agents/skills and .claude/skills from disk at its
one-shot scan, and so the LLM no longer narrates the clone mid-run.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: load reviewer skills from trusted base sha and fetch PR pull ref
Address review findings: skills are now extracted from the PR base sha via
git archive into a dir outside the checkout (prevents PR-authored SKILL.md
prompt injection), and repo prep fetches refs/pull/<n>/head with a strict
checkout so fork PRs fail loudly instead of silently reviewing the default
branch.
* fix: drop ref from skill-extraction log to satisfy CodeQL
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The viewed-marker feature (#1441) reintroduced an owner-only gate on
thread reads, regressing #1425 which made all threads readable by any
org user. Reads no longer assert ownership; viewed markers are only
written for the thread owner.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: per-repo custom instructions for the coding agent
Adds per-repository custom instructions for the main coding agent,
mirroring the reviewer's per-repo style prompts. Instructions are stored
in the LangGraph Store, managed via dashboard API + UI (Monaco editor),
and appended to the agent's system prompt for runs targeting that repo.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: wire agent instructions route into generated route tree
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce repo access on instruction routes
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: out-of-process usage-snapshot builder (Phase 1)
Move all usage-tab compute off the run-serving HTTP process (the #1434 bug
class). Read path is now a pure cache read with a typed computing placeholder
on cold miss; a dedicated usage_snapshot graph + global ~10min cron rebuild
every period's snapshot out-of-process, wrapped in asyncio.timeout and gated by
USAGE_SNAPSHOT_CRON_ENABLED. Lifespan makes one fire-and-forget scheduling call,
never a retained/looping task.
* fix: address PR review on usage-snapshot builder
- admin guard on /admin/usage/rebuild (was any logged-in user)
- cap lifespan loopback calls with asyncio.timeout(5) so a startup hang
can't block boot
- memoize cron registration so steady-state requests skip the loopback
check; reap duplicate crons from concurrent-replica races
- thread the computing flag through the usage payload too
* feat: add Claude Fable 5 as a supported model
Add Anthropic's claude-fable-5 (Mythos-class, released June 9 2026) to
the supported model list so it can be selected in the profile editor.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: order Fable 5 after Opus 4.8 so stale-Anthropic fallback is preserved
provider_fallback_pair picks the first same-provider model, so placing
Fable 5 first redirected stale Opus selections to it. Keep Opus 4.8 first.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: prevent thread prefetches from marking threads viewed
Sidebar prefetches now request thread details with mark_viewed=false so
loading /agents no longer clears the unread/finished indicator for
threads the user has not opened.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: always refetch active thread on mount to mark viewed
Prefetches cache details fetched with mark_viewed=false under the same
query key as the active thread. Forcing refetchOnMount="always" ensures
opening a thread issues a mark_viewed=true request, so last_viewed_* is
recorded even when prefetched data is still fresh.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Pull the api-standards skill from the LangSmith Context Hub at reviewer
run start and inject it into the system prompt, gated on the PR adding or
modifying an API surface. Best-effort: failures fall back to no supplement.
Co-authored-by: GowriH-1 <218394553+GowriH-1@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Add conversations.info fetch so a repo:owner/name (or GitHub URL) token
in a Slack channel's topic/purpose pins the channel to a repo, slotting
in just below thread metadata in get_slack_repo_config.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce client-side deadline on sandbox execute
The langsmith SDK's default execute path is now a WebSocket stream with
no client-side read deadline. On a live socket where the dataplane never
emits an exit/error frame, CommandHandle.result blocks forever in an
uncancellable thread, wedging the run (introduced by the langsmith
0.8.3 -> 0.8.8 bump in #1385). The command `timeout` is only enforced
server-side, so it doesn't fire.
TimeoutLangSmithSandbox drives a non-blocking CommandHandle and kills the
command if it overruns its timeout by a grace window
(SANDBOX_EXECUTE_CLIENT_GRACE_SECONDS, default 30), returning a timed-out
tool result instead of hanging. WS connect failures fall back to the base
wait=True path, whose HTTP fallback carries its own request deadline.
* fix: fall back to HTTP when WS execute connect fails
run(wait=False) eagerly opens the WebSocket and reads the "started" frame,
so connect/setup failures (and connect timeouts) raise from the run() call
itself, not from handle.result. The previous structure left run() outside
the try, so those failures bypassed the HTTP fallback and would fail every
sandbox command in any environment where the WS path is unavailable.
Move handle creation inside the fallback handler in both execute and
aexecute, and run it via to_thread in the async path since it now blocks on
connect. Add tests for connect-failure and connect-timeout fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]
Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.
The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.
* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
Adds an admin-managed, org-wide review guidelines field to team settings
that the reviewer injects into every PR review across all repos, alongside
the existing per-repo style prompt and AGENTS.md context. Repo-specific
rules take precedence when they conflict.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: stop reviewer retrying out-of-diff findings
add_finding rejects findings anchored outside the PR diff, but the agent
retried the same finding 2-3x with adjacent line ranges before giving up,
burning a model turn each. Add a reviewer-prompt recovery block telling the
agent the rejection is authoritative (drop or re-anchor to a + line, don't
retry adjacent), and enrich the rejection payload with nearby in-diff line
ranges for the file so a single re-anchor needs no guessing.
* feat: surface out-of-diff findings in a collapsed dropdown
Instead of rejecting findings anchored outside the PR diff, accept them
(marked in_diff=false) and surface them in a collapsed <details> section of
the review summary, Devin-style. Inline comments stay reserved for in-diff
findings; out-of-diff are severity-gated and capped the same way.
Re-review normally suppresses the empty summary, but now makes an exception
when there are new out-of-diff findings to surface. Surfaced out-of-diff
findings carry a github_review_id so they aren't reposted on later pushes.
Supersedes the earlier 'drop/re-anchor out-of-diff' prompt guidance.
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
GitHub returns repository: null when the token can't read the repo (SAML,
expired token, private/deleted). dict.get(k, {}) doesn't coalesce explicit
null, so fetch_pr_review_threads crashed with AttributeError and publish_review
could never post a review. Guard with isinstance checks and return collected
threads on null repository; sweep the same pattern in resolve_review_thread.
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: make web threads viewable by any org user
Reads (get/stream) now allow any logged-in org user to view a thread
whose source is surfaced (dashboard/github/slack/linear/schedule),
instead of requiring ownership. Internal reviewer/analyzer threads stay
hidden via the same source filter. Writes (send/cancel/delete) remain
owner-only.
Adds GET /threads?all=true to list every surfaced thread; the default
list stays per-user.
* feat: make all web threads readable by any org user
Drop the per-thread source/owner gate on reads: any logged-in org user
can now view and stream any thread, including reviewer/analyzer threads.
GET /threads?all=true returns every thread regardless of source. Writes
(send/cancel/delete) stay owner-only. All routes remain behind the
session + org-login gate.
The bot's numeric noreply (215916821+open-swe[bot]@users.noreply.github.com)
introduced in ae946d1a doesn't resolve to a GitHub account Vercel accepts,
breaking preview deploys on PRs in langchainplus. Revert OPEN_SWE_BOT_EMAIL
to open-swe@users.noreply.github.com.
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: add scheduled web agents
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: secure scheduled agent repositories
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: rebuild scheduled agents as Automations tab
Scrap the inline ScheduledAgentsPanel and replace it with a dedicated
Automations tab: sidebar nav entry, list view with stat cards + empty
state, and a full editor (name, Active toggle, repo, scheduled trigger
picker, agent instructions + model).
* fix: clear collapsed-sidebar button on mobile in Automations
* fix: allow clearing automation repo on update
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: stabilize agent defaults and repo selector
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: align default repo selector styling, restore text fallback
Match the repo selector to sibling settings controls (h-7, bg-input/20,
text-xs). Fall back to a text input when the repo list is empty so a
default repo can still be entered.
* fix: make repo selector dropdown more compact
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>