* feat: report Open SWE Review as a PR check run
Auto-review dispatch now creates an in-progress 'Open SWE Review' check run
on the PR head SHA; publish_review completes it (neutral with findings,
success when clean). An after-agent hook fails the check if the run dies
before publishing. Requires the GitHub App's Checks: Read & write permission;
all calls are best-effort so a missing permission never breaks reviews.
* fix: address review feedback on check-run settling
Keep review_check_run_id when the completion PATCH fails so a later
publish or the after-agent hook can retry instead of hanging the check;
count out-of-diff findings toward the check conclusion.
* fix: retry failed check completion with the real publish conclusion
A transient PATCH failure after a successful publish previously left the
check id for the after-agent hook, which settled it as 'failure'. Persist
the intended result as review_check_pending_result and have the hook
prefer it over the generic failure fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
evals/reviewer/{target,run_eval,build_dataset} called load_dotenv() at
import time, so importing them in tests injected the real .env (live
LANGGRAPH_URL, tokens) into the whole pytest process. The slack-context
default-repo tests then reached the real LangGraph store via
get_team_default_repo() and picked up the developer's actual team
default repo, failing in full-suite runs while passing in isolation.
Move load_dotenv() into the CLI entrypoints (all env reads were already
lazy), and patch get_team_default_repo in the two affected tests so they
stay hermetic regardless of environment.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: server-side Datadog/LangSmith observability tools + team creds
Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on observability tools
Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
read-modify-write race dropping the other provider on concurrent saves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: async email resolution in observability authorization gate
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: honest publish_review reporting + structured thread-not-found errors
- Document skipped_empty_re_review and dry_run in the publish_review
docstring and add a closing-summary contract to the reviewer prompt so
the agent never claims a review was published when review_id is null.
- Raise ReviewerThreadMissingError from replace_findings on SDK
NotFoundError; add_finding/update_finding/publish_review return a
structured do-not-retry result instead of raising, so the agent reports
the blocker after one failure instead of retrying 10-30 times.
* fix: translate thread 404s across all reviewer tool boundaries
get_thread_metadata now raises ReviewerThreadMissingError instead of
swallowing a missing thread as {} (which produced misleading 'No finding
found' results), set_reviewer_thread_metadata translates the SDK 404 the
same way, and every reviewer tool entrypoint (add/update/list findings,
publish_review incl. eval dry-run, resolve/reply thread) returns the
structured do-not-retry result.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Publish had partial-failure windows that double-posted summaries or
corrupted findings state. This hardens the recovery paths:
- open_swe_review_exists is now tri-state (True/False/None). On a
pagination/API failure it returns None ("unknown") instead of False,
and the empty-summary dedup keys off the durable last_reviewed_sha
before consulting GitHub, so a transient failure never double-posts a
"no issues found" summary.
- Comment-id backfill matches strictly on the embedded open-swe marker;
the colliding (path, line, body) fallback is gone, so similar findings
no longer share a comment id and break resolve-on-fix.
- Review-id and comment-id stamping collapse into one guarded
read-modify-write (re-reads latest before writing), removing the
half-stamped intermediate states the prior multi-write flow left open.
- New mutate_findings primitive centralizes read-modify-write so finding
updates operate on the freshest persisted list and skip no-op writes.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: prep reviewer repo at init and load repo skills
Clone + checkout the PR head during reviewer agent init so SkillsMiddleware
can discover the repo's .agents/skills and .claude/skills from disk at its
one-shot scan, and so the LLM no longer narrates the clone mid-run.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: load reviewer skills from trusted base sha and fetch PR pull ref
Address review findings: skills are now extracted from the PR base sha via
git archive into a dir outside the checkout (prevents PR-authored SKILL.md
prompt injection), and repo prep fetches refs/pull/<n>/head with a strict
checkout so fork PRs fail loudly instead of silently reviewing the default
branch.
* fix: drop ref from skill-extraction log to satisfy CodeQL
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The viewed-marker feature (#1441) reintroduced an owner-only gate on
thread reads, regressing #1425 which made all threads readable by any
org user. Reads no longer assert ownership; viewed markers are only
written for the thread owner.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: per-repo custom instructions for the coding agent
Adds per-repository custom instructions for the main coding agent,
mirroring the reviewer's per-repo style prompts. Instructions are stored
in the LangGraph Store, managed via dashboard API + UI (Monaco editor),
and appended to the agent's system prompt for runs targeting that repo.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: wire agent instructions route into generated route tree
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce repo access on instruction routes
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: out-of-process usage-snapshot builder (Phase 1)
Move all usage-tab compute off the run-serving HTTP process (the #1434 bug
class). Read path is now a pure cache read with a typed computing placeholder
on cold miss; a dedicated usage_snapshot graph + global ~10min cron rebuild
every period's snapshot out-of-process, wrapped in asyncio.timeout and gated by
USAGE_SNAPSHOT_CRON_ENABLED. Lifespan makes one fire-and-forget scheduling call,
never a retained/looping task.
* fix: address PR review on usage-snapshot builder
- admin guard on /admin/usage/rebuild (was any logged-in user)
- cap lifespan loopback calls with asyncio.timeout(5) so a startup hang
can't block boot
- memoize cron registration so steady-state requests skip the loopback
check; reap duplicate crons from concurrent-replica races
- thread the computing flag through the usage payload too
* fix: prevent thread prefetches from marking threads viewed
Sidebar prefetches now request thread details with mark_viewed=false so
loading /agents no longer clears the unread/finished indicator for
threads the user has not opened.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: always refetch active thread on mount to mark viewed
Prefetches cache details fetched with mark_viewed=false under the same
query key as the active thread. Forcing refetchOnMount="always" ensures
opening a thread issues a mark_viewed=true request, so last_viewed_* is
recorded even when prefetched data is still fresh.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Pull the api-standards skill from the LangSmith Context Hub at reviewer
run start and inject it into the system prompt, gated on the PR adding or
modifying an API surface. Best-effort: failures fall back to no supplement.
Co-authored-by: GowriH-1 <218394553+GowriH-1@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Add conversations.info fetch so a repo:owner/name (or GitHub URL) token
in a Slack channel's topic/purpose pins the channel to a repo, slotting
in just below thread metadata in get_slack_repo_config.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce client-side deadline on sandbox execute
The langsmith SDK's default execute path is now a WebSocket stream with
no client-side read deadline. On a live socket where the dataplane never
emits an exit/error frame, CommandHandle.result blocks forever in an
uncancellable thread, wedging the run (introduced by the langsmith
0.8.3 -> 0.8.8 bump in #1385). The command `timeout` is only enforced
server-side, so it doesn't fire.
TimeoutLangSmithSandbox drives a non-blocking CommandHandle and kills the
command if it overruns its timeout by a grace window
(SANDBOX_EXECUTE_CLIENT_GRACE_SECONDS, default 30), returning a timed-out
tool result instead of hanging. WS connect failures fall back to the base
wait=True path, whose HTTP fallback carries its own request deadline.
* fix: fall back to HTTP when WS execute connect fails
run(wait=False) eagerly opens the WebSocket and reads the "started" frame,
so connect/setup failures (and connect timeouts) raise from the run() call
itself, not from handle.result. The previous structure left run() outside
the try, so those failures bypassed the HTTP fallback and would fail every
sandbox command in any environment where the WS path is unavailable.
Move handle creation inside the fallback handler in both execute and
aexecute, and run it via to_thread in the async path since it now blocks on
connect. Add tests for connect-failure and connect-timeout fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]
Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.
The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.
* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
Adds an admin-managed, org-wide review guidelines field to team settings
that the reviewer injects into every PR review across all repos, alongside
the existing per-repo style prompt and AGENTS.md context. Repo-specific
rules take precedence when they conflict.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: stop reviewer retrying out-of-diff findings
add_finding rejects findings anchored outside the PR diff, but the agent
retried the same finding 2-3x with adjacent line ranges before giving up,
burning a model turn each. Add a reviewer-prompt recovery block telling the
agent the rejection is authoritative (drop or re-anchor to a + line, don't
retry adjacent), and enrich the rejection payload with nearby in-diff line
ranges for the file so a single re-anchor needs no guessing.
* feat: surface out-of-diff findings in a collapsed dropdown
Instead of rejecting findings anchored outside the PR diff, accept them
(marked in_diff=false) and surface them in a collapsed <details> section of
the review summary, Devin-style. Inline comments stay reserved for in-diff
findings; out-of-diff are severity-gated and capped the same way.
Re-review normally suppresses the empty summary, but now makes an exception
when there are new out-of-diff findings to surface. Surfaced out-of-diff
findings carry a github_review_id so they aren't reposted on later pushes.
Supersedes the earlier 'drop/re-anchor out-of-diff' prompt guidance.
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
GitHub returns repository: null when the token can't read the repo (SAML,
expired token, private/deleted). dict.get(k, {}) doesn't coalesce explicit
null, so fetch_pr_review_threads crashed with AttributeError and publish_review
could never post a review. Guard with isinstance checks and return collected
threads on null repository; sweep the same pattern in resolve_review_thread.
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: add scheduled web agents
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: secure scheduled agent repositories
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* feat: rebuild scheduled agents as Automations tab
Scrap the inline ScheduledAgentsPanel and replace it with a dedicated
Automations tab: sidebar nav entry, list view with stat cards + empty
state, and a full editor (name, Active toggle, repo, scheduled trigger
picker, agent instructions + model).
* fix: clear collapsed-sidebar button on mobile in Automations
* fix: allow clearing automation repo on update
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
* fix: tag Slack threads with stored identity so they surface in web
process_slack_mention gated the run on mapped_login (resolved from the stable
Slack user id), but upsert_agent_thread_owner_metadata independently re-resolved
the GitHub login from the Slack profile email. When that email differs from the
user's mapping email (e.g. a personal vs work address), the lookup returned None,
so github_login was never stamped on the thread and the thread never surfaced in
the web Agents UI (which searches by github_login / triggering_user_email).
Resolve the GitHub user from the store via the Slack id, pass that login through
to the owner metadata, and use the mapping's stored work email (falling back to
the Slack profile email for unmapped users) for both the run config and the
thread tagging, so Slack-started threads reliably appear in web.
* fix: stamp github_login on Slack threads so they surface in web
process_slack_mention gated the run on mapped_login (resolved from the stable
Slack user id) but upsert_agent_thread_owner_metadata re-resolved the login from
the Slack profile email; when that email isn't the user's mapping email the
lookup returns None and github_login is never stamped, so the thread is invisible
in the web Agents UI (which searches by github_login / triggering_user_email).
Pass the already-resolved mapped_login through to the owner metadata. The
dashboard match keys on the stable GitHub login, so this is sufficient; the
triggering email stays the live Slack profile value.
* fix: preserve Slack email during account mapping
* fix: require Slack OIDC for email mappings
* chore: format Slack OIDC mapping cleanup
Reviewer/push runs execute in a worker process, but the GitHub token was
cached by the webhook handler in the API server process — a different
process — so the worker's in-process cache was always cold. resolve_github_token
then failed (User not authenticated / Unknown source: github_push) and only the
app-token fallback kept reviews working, noisily.
The reviewer always acts as the GitHub App (open-swe[bot]), so resolve the
installation token directly at run start, scoped to the repo. This also bypasses
org SAML enforcement that blocks user OAuth tokens. Drop the now-dead
cross-process cache writes in the webhook reviewer-dispatch handlers, and stop
leave_failure_comment raising on the github_push source.
create_local_sandbox now mkdir -p's the resolved root dir, so a custom
LOCAL_SANDBOX_ROOT_DIR (or a /tmp path cleared on reboot) no longer fails
sandbox work-dir resolution.
Slack thread replies and Linear comments showed only the bare tool name in the dashboard chat. Map them to dedicated 'slack'/'linear' toolKinds and render the message body in a ReplyCard, so Open-in-Web shows what the agent actually posted.
* fix: make repository optional when starting a dashboard run
The agent infers and clones the target repo from the task itself, so a
default repo is never actually required to run — but the dashboard 400'd
("no default repository configured") when a user had none set.
Treat repo as optional: _resolve_repo_config returns {} instead of
raising, repo metadata/config are only written when a repo is present,
and the "missing repository metadata" gate on follow-up messages is
dropped. UI hides the repo chip when absent.
* feat: add repo picker to the run prompt bar
Adds an optional, searchable repository selector next to the model picker
on the Agents home prompt bar (Cursor-style). It pre-fills the user's saved
default repo and can be cleared to "No repository" for a repo-less run.
Because the picker now resolves the default on the client, the create
endpoint honors the request value verbatim: _resolve_repo_config just parses
what's sent ({} when empty) instead of falling back to the saved default,
so an explicit "No repository" is respected.
* refactor: match Cursor layout for repo placement
Move the repo selector out of the prompt-box footer to a pill row above
the input (folder + caret, dropdown opens downward); the model picker
stays inside the box. In the thread view, show the thread title and repo
in a header at the top of the chat, with the follow-up input pinned to
the bottom as before.
Open SWE Review now runs from automated PR triggers, so drop the old Slack/GitHub review keyword entrypoints and keep PR comments on the regular agent path.
* feat: add Slack Open in Web link
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: skip web link for Slack reviewer runs
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: accept include_dashboard_link kwarg in Slack reviewer test double
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
* fix: resolve reviewer head_sha from thread metadata, not frozen run config
A push that lands while a reviewer run is in flight is delivered as a
queued message into that run. The run's configurable is frozen at
creation, so its head_sha still names the commit the run was created for
— not the commit just pushed. publish_review then anchored the GitHub
review to the stale commit and regressed last_reviewed_sha to it, and
add_finding/update_finding stamped findings with the stale SHA.
Persist the current head in thread metadata at every reviewer dispatch
(both the ready-for-review and push paths, before they branch to create
a run or queue a message), and add resolve_review_head_sha() which
prefers the metadata head over the run config. Wire it into
publish_review (review commit_id + last_reviewed_sha), add_finding
(first_seen_sha) and update_finding (last_confirmed_sha). Falls back to
the run config when metadata carries no head (first review, eval, tests).
* fix: persist head_sha in manual review dispatch (trigger_pr_review_from_ref)
resolve_review_head_sha prefers metadata[head_sha] over the run config,
and the push/ready dispatchers write it — but trigger_pr_review_from_ref
(Slack/GitHub @open-swe review, request_pr_review tool) created a run
with a freshly-fetched config head while leaving metadata's head stale
from a prior dispatch. A manual re-review at a newer commit would then
resolve to the old head and publish/advance findings against it.
Persist head_sha in that dispatch's metadata write too, so every
run-creating reviewer dispatch keeps metadata in sync with the head its
run targets. Caught by the Open SWE reviewer on this PR.
The Co-authored-by trailer and bot git identity used
open-swe@users.noreply.github.com, which resolves to the separate
open-swe *user* account rather than the open-swe[bot] GitHub App.
Switch OPEN_SWE_BOT_EMAIL to the bot's noreply address
(215916821+open-swe[bot]@users.noreply.github.com) so co-author credit
and the fallback author identity point at the bot.
Drive the prompt trailer and sandbox git config from the constant
instead of hardcoding the address.
A push that lands while a reviewer run is in flight is delivered as a
queued message into the still-running first-review run, whose
configurable still has re_review=False. The empty-review guard in
publish_review only skipped the 'No issues found' summary when
is_re_review was True, so the queued reconcile published a second,
duplicate top-level 'No issues found' review.
Key the empty-review skip off actual PR state instead: add
open_swe_review_exists(), which detects the marker render_review_body
embeds in every Open SWE review body, and skip the summary when a prior
Open SWE review already exists (regardless of the re_review flag). Fails
open on API error so a genuine first review is never suppressed.
* fix: scope public reviewer tokens
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: simplify reviewer token wiring; fix push re-scope + red test
- Remove the redundant _check_or_recreate_sandbox_for_proxy /
_refresh_github_proxy_or_recreate_for_proxy wrappers and call the
underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
payload lacked repo privacy/id but PR metadata reveals a public repo, so
reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
and add coverage for _reviewer_token_for_repo + push-event scoping.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: deliver Slack account-link prompt as a visible threaded reply
Blocked Slack users got no prompt at all. Prod logs show chat.postEphemeral
returns ok, but ephemeral messages are silently dropped in Slack's assistant
threads (where Open SWE runs), so the user sees nothing. Post the prompt as a
normal threaded reply instead — the same channel the agent uses to reply.
* fix: deliver Slack auth-failure prompt as a visible threaded reply
leave_failure_comment() tried an ephemeral message first and only fell back
to a thread reply on failure. Ephemeral messages succeed (ok) but are dropped
in Slack's assistant threads, so the fallback never fired and the user saw no
auth-failure prompt. Post the visible threaded reply directly, matching the
account-link prompt fix.
* fix: prompt blocked Slack users with a generic, token-free dashboard link
Addresses the review findings that posting the per-user account-link token /
auth URL in a visible thread lets any channel member bind their GitHub account
to the triggering user's Slack identity.
Drop the per-user signed link entirely. Both the account-link prompt
(_post_account_link_prompt) and the runtime auth-failure prompt
(leave_failure_comment) now post a plain dashboard settings link
(build_settings_url) as a visible threaded reply. The user signs in with GitHub
from their own session and connects Slack via verified OIDC on the settings
page — no secret in the thread, nothing to hijack, and no DM machinery.
* feat: nudge first-time users to connect Slack from the dashboard home
Show a Connect Slack banner on the agents landing page whenever Slack OAuth is
enabled and the user hasn't linked Slack yet. A first-time user (no Slack
mapping) sees it immediately after signing in; it disappears once connected.
* feat: prompt first-time users to connect Slack via a dialog
Replace the inline Connect Slack card on the agents home with a modal dialog
(Base UI). It opens automatically once the mapping query resolves to
"not connected" and closes itself once Slack is linked; "Maybe later" dismisses
it for the session. No new dependency — uses the design system's Base UI.
* copy: frame Slack connect as resolving the user's GitHub account
Drop 'act/reply on your behalf' wording across the connect-Slack dialog, the
Slack thread prompts (blocked + auth-failure), and the settings description.
Connecting Slack lets Open SWE resolve the user's GitHub account when they tag
it in Slack.