* feat: CI auto-fix and PR babysitting for agent PRs
Watch CI failures and review feedback on PRs Open SWE opened, then dispatch
confidence-gated fix runs on the originating agent thread. Adds CI webhook
ingestion (check_run/check_suite/workflow_run/status), a per-PR @open-swe
autofix on|off toggle, auto-response to review comments, and a polling
ci_monitor graph that also flags merge conflicts. Gated by the existing
autofix_mode/trigger_mode settings, the enabled-repos opt-in, base-branch and
human-commit skip rules, dedupe, and a per-PR attempt cap.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review feedback on CI auto-fix
- Security: gate the no-mention review-feedback path on author trust —
require a trusted author_association (OWNER/MEMBER/COLLABORATOR) plus a
GitHub write/maintain/admin permission check before dispatching a
write-capable agent run, preventing privilege escalation from
read/triage/outside reviewers.
- Auth: reuse the originating PR thread's source + login/email when
dispatching fix runs so the GitHub-token resolver authenticates them in
non-bot-token deployments (bespoke github_ci source failed to resolve).
- Docs: document the Commit statuses: Read-only permission required for the
Status webhook event.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: stream reviewer eval logs on a dedicated admin page
Stream the eval subprocess output into a rolling log_tail and persist it
during the run (was only captured at exit), so the live output is visible
while the eval runs. Move the eval runner off the admin page onto its own
/admin/evals page (linked like Review Style Prompts) with a live log viewer.
* chore: drop unrelated SSR-register drift from generated route tree
* fix(ui): pre-bundle workbox-window to stop dev re-optimize reload
The PWA service worker (devOptions.enabled) pulls workbox-window, which
Vite discovers after first render and re-optimizes, forcing a reload that
cancels in-flight code-split route imports (Failed to fetch dynamically
imported module). Pre-bundling it via optimizeDeps.include avoids the
mid-session reload.
* fix(ui): suppress html hydration warning for pre-hydration theme script
The inline theme script sets class="dark"/color-scheme on <html> before
React hydrates, so the prerendered HTML never matches. suppressHydrationWarning
on <html> silences the (expected) one-level attribute mismatch.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: trigger reviewer evals from the admin page
Add an admin-only "Reviewer eval" section + endpoints that launch the
reviewer benchmark as an isolated subprocess against the running
deployment, with live status and the LangSmith experiment link. Route
eval traces to a dedicated open-swe-evals project so they stay out of
the production tracing project.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: reconcile reviewer eval status via heartbeat, not local process
The persisted record is shared across workers but _PROCS is process-local.
The owning worker now refreshes a heartbeat while the subprocess runs, and
status is only reconciled to failed once the heartbeat is stale, so a poll on
a worker without the local handle no longer kills a live run (and a duplicate
start is rejected across workers).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Kimi K2.7 is now the supported version, so drop the older K2.6 entry.
Update the related fireworks provider test to use K2.7 instead.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: resolve trace URL project id by tracing project name
Graphs were split into separate LangSmith tracing projects
(open-swe-agent, open-swe-review) but the "View trace" link still used
a single fixed project-id env var pointing at the old combined project.
Resolve the project id from the tracing project name so agent and
reviewer links point at their respective projects, falling back to the
env var when resolution is unavailable.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* docs: document per-graph tracing projects for trace links
View trace links now resolve project IDs from the open-swe-agent /
open-swe-review project names. Document this so fresh deployments create
the right projects instead of relying on a single project ID.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add Kimi K2.7 model option
Surface Moonshot's newly released Kimi K2.7 in the model picker, routed
through Fireworks following the existing kimi-k2pN convention, with the
same none/low/medium/high reasoning efforts (default high) as K2.6.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: correct Kimi K2.7 Fireworks model id to kimi-k2p7-code
Fireworks now hosts the K2.7 release as
accounts/fireworks/models/kimi-k2p7-code (the `kimi-k2p7` slug 404s).
Point the model option and its test at the live model id.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: drop unsupported "none" effort for Kimi K2.7
Kimi K2.7 Code is a thinking-only model — Moonshot documents that it does
not support non-thinking mode (passing thinking={"type":"disabled"} or
reasoning_effort="none" errors). Removing "none" from the advertised
efforts so the picker can't send an invalid Fireworks request.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The top-level review comment's "Open in Web" link pointed at the
agent thread (/agents/{thread_id}). Point it at the dashboard review
detail page (/agents/reviews/{owner}/{repo}/{number}) instead.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* perf: speed up Reviews list (My PRs + All PRs)
Push the "My PRs" author filter into the threads.search metadata
(pr.author containment) instead of scanning up to 1000 reviewer threads
in Python, and replace the per-repo GitHub access check (an N+1 of
sequential GET /repos calls) with a single per-login accessible-repo set,
cached for 60s. Detail endpoints still re-validate access live.
Frontend: prefetch the inactive tab and adjacent page on hover/focus so
tab switches and pagination are instant.
* fix: don't cache repo access for the reviews list
The /reviews list is an authorization boundary for private PR metadata
(repo/PR titles, branches, authors, finding counts). A cross-request TTL
cache on the accessible-repo set could surface that metadata for up to
60s after a user lost repo access. Resolve the set fresh per request
instead — still a fixed, repo-count-independent burst of GitHub calls
(no per-repo N+1), with no staleness.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: mark threads resolved to hide from sidebar [closes resolve-threads]
Add a resolved flag (thread metadata) so users can clear finished
threads from the sidebar without deleting them. Resolved threads move to
a collapsible "Resolved" group (capped at 20 with a "Show all" link) and
are fully searchable/filterable + paginated on a new /agents/threads
page driven by URL query params. Sending a new message auto-unresolves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: refresh paginated agent thread lists
Invalidate all agent thread list queries when thread lifecycle events change sidebar data, keeping the new paginated sidebar cache in sync.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: render Reviews page diffs with pierre MultiFileDiff
The Reviews detail page used a hand-rolled hunk renderer with no syntax
highlighting. Switch it to the same pierre MultiFileDiff + theming the agent
chat git panel uses.
- review-diff API now returns full original/modified file contents instead of
hunks, via a shared build_pr_diff_files helper extracted from thread_api
- findings render as right-anchored markers (pierre line annotations); focus
highlight uses selectedLines; floating finding card still anchors to the marker
* feat: auto-collapse a review diff card when marked as viewed
* fix: URL-encode file path in Contents API fetch
Filenames containing reserved URL characters (#, ?) were truncated, so those
files rendered as empty/unrenderable. quote(path, safe='/') preserves the path
separators while escaping the rest.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: disable out-of-diff reviewer findings
PR reviews were surfacing findings about code outside the PR's changed
lines, which read as random/off-topic noise. Reject out-of-diff findings
at add_finding and stop surfacing them in publish_review so only findings
anchored to changed lines reach the PR.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: align reviewer bar with out-of-diff rejection
The filing bar still permitted proven regressions in files absent from
the diff, but add_finding now rejects those, so a concrete regression
could be silently dropped after a failed tool call. Update the bar and
the "Do NOT file" list so the prompt only directs in-diff findings.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: link Slack thread/Linear ticket in PRs and append ticket to title
When opening PRs, include any referenced Slack thread or Linear ticket
in the description and append the resolvable ticket number to the title.
Adds a Slack permalink to the webhook prompt so the agent has a
ready-to-link thread URL.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: use [closes <TICKET>] format in PR title suffix
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: deterministically append source refs to private-repo PRs only
Move the Slack/Linear source-reference linking out of the prompt and into
open_pull_request, gated to private repos so private Slack thread URLs and
Linear identifiers are never published to a public PR. Reverts the prompt
instruction and webhook permalink injection in favor of this server-side append.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
prepare_review_repo is best-effort, but the system prompt unconditionally
told the agent the repo is checked out at the PR head. On a reused sandbox,
a dirty worktree (or a transient fetch failure) makes 'git checkout <sha>'
fail; prep returns False, the old checkout stays in place, and the reviewer
confidently reads stale code — observed as 'I rechecked the current head'
replies quoting pre-push code.
- checkout with --force so leftover worktree state can't block it, verify
HEAD matches the requested sha, tolerate 'git fetch --all' failures
- when prep fails, the prompt now warns the checkout may be stale and tells
the agent to re-fetch/checkout (or fall back to API file contents) before
trusting local files
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preflight stream proxy before SSE starts, let optimistic thread survive refetch
* fix: seed sidebar thread details as stale so mark-viewed fetch still fires
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: track PR lifecycle state per thread for sidebar
Persist a PR's draft/open/merged/closed state on the agent thread and keep
it in sync as PR webhooks fire, so the Agents UI can show per-thread PR
status the way Cursor does. open_pull_request now records the initial
draft/open state and pr_title; thread summaries expose diffStats; and the
PR webhook refreshes pr_state on close/reopen/draft/ready transitions.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: consolidate PR state mapping into shared derive_pr_state helper
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: auto-recover from expired GitHub refresh tokens
When a user's GitHub OAuth refresh token was permanently dead (revoked or
expired), token refresh failed but get_valid_access_token still handed back
the known-stale access token, so dashboard GitHub calls kept 401ing until the
user manually logged out and back in.
Now we distinguish unrecoverable refresh failures (bad_refresh_token /
unauthorized_client) from transient ones: on an unrecoverable failure we drop
the dead stored authorization and return None, so callers prompt a clean
re-login. Transient failures still fall back to the stored token.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: don't delete fresh re-auth when stale refresh fails
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
- _origin_of: treat invalid ports as invalid origin instead of letting
urlparse ValueError turn CSRF rejections into 500s
- AgentsHome: reset submitting/draft when stream.submit rejects before a
thread id is minted, so the prompt isn't left disabled
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: refresh sandbox GitHub proxy token before mid-run expiry
GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve repo-scoped proxy token on mid-run refresh
Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: update passthrough stub for github_proxy_repositories param
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add right-click open trace context menu for threads
Add a context menu to sidebar thread rows exposing an Open trace action
that links to the LangSmith thread trace, surfaced via a new traceUrl
field on the thread summary.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: drop stray yarn artifacts from PR
Remove accidentally committed ui/.yarnrc.yml and ui/.yarn/install-state.gz
that broke immutable installs against the v1 lockfile, and gitignore yarn
artifacts to prevent recurrence.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: reuse thread_id var, drop redundant danger-color fallback
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: keep review check visible on follow-up commits and stop using neutral
The "Open SWE Review" check completed as `neutral` whenever findings were
surfaced, which GitHub renders as a confusing "neutral check" group. Always
complete it as `success` (informational/non-blocking), matching Devin and
Corridor — the finding count stays in the title and findings post as comments.
Also create a fresh check run on the new head SHA in the push re-review path:
GitHub only shows check runs on a PR's current head, so the check vanished
after a follow-up push (and the stale id settled on an outdated commit).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: surface settled review check when push leaves diff unchanged
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: report Open SWE Review as a PR check run
Auto-review dispatch now creates an in-progress 'Open SWE Review' check run
on the PR head SHA; publish_review completes it (neutral with findings,
success when clean). An after-agent hook fails the check if the run dies
before publishing. Requires the GitHub App's Checks: Read & write permission;
all calls are best-effort so a missing permission never breaks reviews.
* fix: address review feedback on check-run settling
Keep review_check_run_id when the completion PATCH fails so a later
publish or the after-agent hook can retry instead of hanging the check;
count out-of-diff findings toward the check conclusion.
* fix: retry failed check completion with the real publish conclusion
A transient PATCH failure after a successful publish previously left the
check id for the after-agent hook, which settled it as 'failure'. Persist
the intended result as review_check_pending_result and have the hook
prefer it over the generic failure fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
evals/reviewer/{target,run_eval,build_dataset} called load_dotenv() at
import time, so importing them in tests injected the real .env (live
LANGGRAPH_URL, tokens) into the whole pytest process. The slack-context
default-repo tests then reached the real LangGraph store via
get_team_default_repo() and picked up the developer's actual team
default repo, failing in full-suite runs while passing in isolation.
Move load_dotenv() into the CLI entrypoints (all env reads were already
lazy), and patch get_team_default_repo in the two affected tests so they
stay hermetic regardless of environment.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: server-side Datadog/LangSmith observability tools + team creds
Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on observability tools
Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
read-modify-write race dropping the other provider on concurrent saves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: async email resolution in observability authorization gate
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: honest publish_review reporting + structured thread-not-found errors
- Document skipped_empty_re_review and dry_run in the publish_review
docstring and add a closing-summary contract to the reviewer prompt so
the agent never claims a review was published when review_id is null.
- Raise ReviewerThreadMissingError from replace_findings on SDK
NotFoundError; add_finding/update_finding/publish_review return a
structured do-not-retry result instead of raising, so the agent reports
the blocker after one failure instead of retrying 10-30 times.
* fix: translate thread 404s across all reviewer tool boundaries
get_thread_metadata now raises ReviewerThreadMissingError instead of
swallowing a missing thread as {} (which produced misleading 'No finding
found' results), set_reviewer_thread_metadata translates the SDK 404 the
same way, and every reviewer tool entrypoint (add/update/list findings,
publish_review incl. eval dry-run, resolve/reply thread) returns the
structured do-not-retry result.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Publish had partial-failure windows that double-posted summaries or
corrupted findings state. This hardens the recovery paths:
- open_swe_review_exists is now tri-state (True/False/None). On a
pagination/API failure it returns None ("unknown") instead of False,
and the empty-summary dedup keys off the durable last_reviewed_sha
before consulting GitHub, so a transient failure never double-posts a
"no issues found" summary.
- Comment-id backfill matches strictly on the embedded open-swe marker;
the colliding (path, line, body) fallback is gone, so similar findings
no longer share a comment id and break resolve-on-fix.
- Review-id and comment-id stamping collapse into one guarded
read-modify-write (re-reads latest before writing), removing the
half-stamped intermediate states the prior multi-write flow left open.
- New mutate_findings primitive centralizes read-modify-write so finding
updates operate on the freshest persisted list and skip no-op writes.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: prep reviewer repo at init and load repo skills
Clone + checkout the PR head during reviewer agent init so SkillsMiddleware
can discover the repo's .agents/skills and .claude/skills from disk at its
one-shot scan, and so the LLM no longer narrates the clone mid-run.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: load reviewer skills from trusted base sha and fetch PR pull ref
Address review findings: skills are now extracted from the PR base sha via
git archive into a dir outside the checkout (prevents PR-authored SKILL.md
prompt injection), and repo prep fetches refs/pull/<n>/head with a strict
checkout so fork PRs fail loudly instead of silently reviewing the default
branch.
* fix: drop ref from skill-extraction log to satisfy CodeQL
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
The viewed-marker feature (#1441) reintroduced an owner-only gate on
thread reads, regressing #1425 which made all threads readable by any
org user. Reads no longer assert ownership; viewed markers are only
written for the thread owner.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: per-repo custom instructions for the coding agent
Adds per-repository custom instructions for the main coding agent,
mirroring the reviewer's per-repo style prompts. Instructions are stored
in the LangGraph Store, managed via dashboard API + UI (Monaco editor),
and appended to the agent's system prompt for runs targeting that repo.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* chore: wire agent instructions route into generated route tree
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce repo access on instruction routes
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: out-of-process usage-snapshot builder (Phase 1)
Move all usage-tab compute off the run-serving HTTP process (the #1434 bug
class). Read path is now a pure cache read with a typed computing placeholder
on cold miss; a dedicated usage_snapshot graph + global ~10min cron rebuild
every period's snapshot out-of-process, wrapped in asyncio.timeout and gated by
USAGE_SNAPSHOT_CRON_ENABLED. Lifespan makes one fire-and-forget scheduling call,
never a retained/looping task.
* fix: address PR review on usage-snapshot builder
- admin guard on /admin/usage/rebuild (was any logged-in user)
- cap lifespan loopback calls with asyncio.timeout(5) so a startup hang
can't block boot
- memoize cron registration so steady-state requests skip the loopback
check; reap duplicate crons from concurrent-replica races
- thread the computing flag through the usage payload too
* fix: prevent thread prefetches from marking threads viewed
Sidebar prefetches now request thread details with mark_viewed=false so
loading /agents no longer clears the unread/finished indicator for
threads the user has not opened.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: always refetch active thread on mount to mark viewed
Prefetches cache details fetched with mark_viewed=false under the same
query key as the active thread. Forcing refetchOnMount="always" ensures
opening a thread issues a mark_viewed=true request, so last_viewed_* is
recorded even when prefetched data is still fresh.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Pull the api-standards skill from the LangSmith Context Hub at reviewer
run start and inject it into the system prompt, gated on the PR adding or
modifying an API surface. Best-effort: failures fall back to no supplement.
Co-authored-by: GowriH-1 <218394553+GowriH-1@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Add conversations.info fetch so a repo:owner/name (or GitHub URL) token
in a Slack channel's topic/purpose pins the channel to a repo, slotting
in just below thread metadata in get_slack_repo_config.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce client-side deadline on sandbox execute
The langsmith SDK's default execute path is now a WebSocket stream with
no client-side read deadline. On a live socket where the dataplane never
emits an exit/error frame, CommandHandle.result blocks forever in an
uncancellable thread, wedging the run (introduced by the langsmith
0.8.3 -> 0.8.8 bump in #1385). The command `timeout` is only enforced
server-side, so it doesn't fire.
TimeoutLangSmithSandbox drives a non-blocking CommandHandle and kills the
command if it overruns its timeout by a grace window
(SANDBOX_EXECUTE_CLIENT_GRACE_SECONDS, default 30), returning a timed-out
tool result instead of hanging. WS connect failures fall back to the base
wait=True path, whose HTTP fallback carries its own request deadline.
* fix: fall back to HTTP when WS execute connect fails
run(wait=False) eagerly opens the WebSocket and reads the "started" frame,
so connect/setup failures (and connect timeouts) raise from the run() call
itself, not from handle.result. The previous structure left run() outside
the try, so those failures bypassed the HTTP fallback and would fail every
sandbox command in any environment where the WS path is unavailable.
Move handle creation inside the fallback handler in both execute and
aexecute, and run it via to_thread in the async path since it now blocks on
connect. Add tests for connect-failure and connect-timeout fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000]
Post a transient "review in progress" comment (with an "Open in Web"
dashboard link) when a reviewer run starts, then delete it once the
review lands. The published review body now carries the same
"Open in Web" link, so the link persists on the review itself.
The transient comment's id is tracked in reviewer thread metadata
(status_comment_id) so it can be deleted on completion.
* refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000]
Adds an admin-managed, org-wide review guidelines field to team settings
that the reviewer injects into every PR review across all repos, alongside
the existing per-repo style prompt and AGENTS.md context. Repo-specific
rules take precedence when they conflict.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: stop reviewer retrying out-of-diff findings
add_finding rejects findings anchored outside the PR diff, but the agent
retried the same finding 2-3x with adjacent line ranges before giving up,
burning a model turn each. Add a reviewer-prompt recovery block telling the
agent the rejection is authoritative (drop or re-anchor to a + line, don't
retry adjacent), and enrich the rejection payload with nearby in-diff line
ranges for the file so a single re-anchor needs no guessing.
* feat: surface out-of-diff findings in a collapsed dropdown
Instead of rejecting findings anchored outside the PR diff, accept them
(marked in_diff=false) and surface them in a collapsed <details> section of
the review summary, Devin-style. Inline comments stay reserved for in-diff
findings; out-of-diff are severity-gated and capped the same way.
Re-review normally suppresses the empty summary, but now makes an exception
when there are new out-of-diff findings to surface. Surfaced out-of-diff
findings carry a github_review_id so they aren't reposted on later pushes.
Supersedes the earlier 'drop/re-anchor out-of-diff' prompt guidance.
---------
Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>