Commit graph

69 commits

Author SHA1 Message Date
Johannes du Plessis
29015fadc4
feat: gate workflow pushes with approval (#1614)
* feat: gate workflow pushes with approval

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve proxy refresh test compatibility

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: bind workflow approvals to pushed ref

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-25 17:17:51 -07:00
Johannes du Plessis
39681102d6
fix: repair orphaned tool calls before model calls (#1604)
When a run is cancelled or the sandbox dies mid-tool-call, LangGraph persists
the AIMessage tool_call but never the matching ToolMessage. The next run sends
the provider an orphaned tool_use (Anthropic 400: "tool_use ids were found
without tool_result blocks"), permanently wedging the thread on every retry.

Add RepairOrphanedToolCallsMiddleware, which inserts a synthetic error
ToolMessage immediately after any tool_call lacking a result so the agent can
retry instead of dying. Wired into the agent and reviewer graphs.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-24 12:58:10 -07:00
Johannes du Plessis
3992d3ef5d
feat: repo-scoped dynamic sandbox snapshots (#1595)
* feat: repo-scoped dynamic sandbox snapshots

Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.

Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden repo snapshot builds

Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: document repo snapshot base image config

Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:24:11 -07:00
Johannes du Plessis
860aee48ee
feat: add user-scoped Notion MCP OAuth (#1593)
* feat: add user-scoped Notion MCP OAuth

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: refresh Notion token per tool call

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: normalize Notion wrapper response format

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:07:13 -07:00
Ramon Nogueira
3a0e2b4672
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning

Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce plan-mode read-only at tool layer and disable subagents

Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: harden plan-mode shell guard against wrapped mutations

Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow

- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline

Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.

* test(open-swe): add Playwright E2E for the Slack → PR → web handoff

Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.

- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).

Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.

* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright

The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.

Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.

* test(open-swe): record Playwright trace + video on every E2E run

Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.

* feat(plan-mode): collaborative plan review with BlockNote + Yjs

When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.

- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
  the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
  snapshots; plan content/status store; plan REST API (get/approve/reject,
  owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
  mounted under the agents shell, with a "Review plan" banner in the thread view
  and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
  flow, including cross-user comment sync and owner-only approval.

* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)

- plan-collab WS: authorize per-thread before joining a room (same read gate as
  the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
  opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
  resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
  dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
  unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
  before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.

Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).

* style: ruff format plan_collab.py

* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS

- Slack "Approve & Implement" now verifies the clicking user is the plan
  requester (owner, via the stored triggering_user_id) before implementing —
  matching the dashboard API's owner-only approval. Non-owners are pointed to
  Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
  allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
  the REST require_same_origin CSRF defense.

* fix(plan-mode): enter plan mode only via the model + local mock dev harness

Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.

- enter_plan_mode returns a terminating ToolMessage, fixing the missing
  ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
  remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
  call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
  Alice/Bob mock users, and a GitHub login picker.

* docs(plan-mode): drop stale references to removed profile/team defaults

The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.

* feat(plan-mode): let any reviewer edit the plan, not just comment

Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.

* test(plan-mode): assert plan-mode entry via the tool's success message

plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.

---------

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 12:06:58 -07:00
Johannes du Plessis
450a25d33f
feat: add schedule_thread_wakeup tool for self-polling (#1592)
* feat: add schedule_thread_wakeup tool for self-polling

Add a new agent tool that schedules a one-shot re-trigger of the
current thread after a configurable delay (1–1440 minutes). Uses a
LangGraph cron with end_time to fire exactly once, then retire.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: prevent early thread wakeups

Round scheduled wakeup times up to the next whole minute so cron minute precision cannot fire before the requested delay.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 11:06:01 -07:00
John Kennedy
9949077bc7
fix: allow GitHub logins for dashboard admins (#1582)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-20 09:44:12 -07:00
John Kennedy
9db7eab134
feat: Add Corridor MCP analyzePlan integration (#1572)
* Add Corridor MCP analyzePlan integration

* Update agent/integrations/corridor_mcp.py

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* Add Corridor analysis prompt

* Only include Corridor prompt when tool loads

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-18 14:01:25 -07:00
Johannes du Plessis
60b7f4677a
feat: add user-scoped Currents.dev API key for e2e test investigation (#1566)
* feat: add user-scoped Currents.dev API key for e2e test investigation

Allow each user to configure their own Currents.dev API key on the
Profile Settings page. The key is encrypted at rest in a per-user
LangGraph Store namespace and feeds server-side read-only tools that
query the Currents REST API (runs, instances, projects, test results)
so agent runs can inspect e2e test failures including screenshots and
DOM snapshots.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: add pagination cursors to currents_list_project_runs

Address review feedback: forward starting_after/ending_before cursor
parameters to /projects/{projectId}/runs so the agent can paginate
beyond the first 50 results.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-17 13:59:55 -07:00
Caroline di Vittorio
554d754585
feat: link PR attribution footer to the originating thread (#1539)
The "Made by Open SWE" PR footer linked to the generic dashboard
homepage. Point it at the dashboard thread that generated the PR
(/agents/<thread_id>), falling back to the homepage when no thread
id is available.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 09:50:18 -07:00
Johannes du Plessis
ddd7190864
feat: let the agent stop naturally without forced tool calls (#1535)
Remove the hardcoded "call a tool every turn" instruction from the system
prompt and delete the ensure_no_empty_msg middleware that re-injected no_op /
confirming_completion tool calls. The agent now ends its turn naturally when
the model emits a final message with no tool call, which avoids needlessly
extending trajectories (and token spend) on tasks that are already complete.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 14:54:01 -07:00
Johannes du Plessis
aa98d195a8
feat: route graphs to separate LangSmith tracing projects (#1508)
* feat: route graphs to separate LangSmith tracing projects

* docs: point graph entrypoint references at traced wrappers

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 17:57:16 -07:00
Johannes du Plessis
da20e57bcc
fix: refresh sandbox GitHub proxy token before mid-run expiry (#1496)
* fix: refresh sandbox GitHub proxy token before mid-run expiry

GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve repo-scoped proxy token on mid-run refresh

Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: update passthrough stub for github_proxy_repositories param

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 10:59:21 -07:00
Christian Bromann
abf354bb05
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol

Replace the bespoke SSE + React Query polling path with LangGraph’s
v2 event stream through credentialed dashboard proxies. Run starts go
through stream commands; mid-run follow-ups still queue via /messages.

* fix import path

* fix tests after rebase

* format

* PR feedback

* improved model fallback

* fix image handling

* embrace sdk

* cleanup

* cr

* more cleanup

* fix cors

* harden security

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-11 09:54:35 -07:00
Johannes du Plessis
f539962c73
feat: server-side Datadog/LangSmith observability tools + team creds [closes OPE-54] (#1476)
* feat: server-side Datadog/LangSmith observability tools + team creds

Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on observability tools

Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
  OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
  untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
  read-modify-write race dropping the other provider on concurrent saves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: async email resolution in observability authorization gate

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:07:42 -07:00
Johannes du Plessis
da8f003933
feat: per-repo custom instructions for the coding agent (#1460)
* feat: per-repo custom instructions for the coding agent

Adds per-repository custom instructions for the main coding agent,
mirroring the reviewer's per-repo style prompts. Instructions are stored
in the LangGraph Store, managed via dashboard API + UI (Monaco editor),
and appended to the agent's system prompt for runs targeting that repo.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: wire agent instructions route into generated route tree

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: enforce repo access on instruction routes

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-09 15:46:29 -07:00
Johannes du Plessis
449cb5d1a8
fix: make default repository configurable (#1429)
* fix: make default repository configurable

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: preserve dashboard repo-less runs

* fix: distinguish explicit repo-less dashboard runs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-05 13:48:47 -07:00
Johannes du Plessis
a511add856
feat: add agent usage leaderboard (#1418)
* feat: add agent usage leaderboard

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: bound leaderboard refresh and hide emails

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-04 23:33:31 +00:00
Johannes du Plessis
8b18e7e95d
fix: Reduce graph load and dashboard repo failures (#1412)
* fix: reduce graph load and dashboard repo failures

* fix: surface repo listing timeouts
2026-06-04 13:54:06 -07:00
Brace Sproul
2070a770c2
fix: stop storing GitHub tokens in metadata (#1405)
* fix: stop storing GitHub tokens in metadata

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* fix: bound in-process GitHub token cache with 24h TTL + sweep

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
Co-authored-by: open-swe[bot] <johannes@langchain.dev>
2026-06-04 09:33:51 -07:00
open-swe[bot]
ae946d1aa5 chore: attribute commit co-author to open-swe[bot], not open-swe user
The Co-authored-by trailer and bot git identity used
open-swe@users.noreply.github.com, which resolves to the separate
open-swe *user* account rather than the open-swe[bot] GitHub App.
Switch OPEN_SWE_BOT_EMAIL to the bot's noreply address
(215916821+open-swe[bot]@users.noreply.github.com) so co-author credit
and the fallback author identity point at the bot.

Drive the prompt trailer and sandbox git config from the constant
instead of hardcoding the address.
2026-06-03 11:19:24 -07:00
Johannes du Plessis
0a2e682364
fix: scope public reviewer tokens (#1389)
* fix: scope public reviewer tokens

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* refactor: simplify reviewer token wiring; fix push re-scope + red test

- Remove the redundant _check_or_recreate_sandbox_for_proxy /
  _refresh_github_proxy_or_recreate_for_proxy wrappers and call the
  underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
  payload lacked repo privacy/id but PR metadata reveals a public repo, so
  reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
  and add coverage for _reviewer_token_for_repo + push-event scoping.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-03 09:25:17 -07:00
Johannes du Plessis
04c346176d
feat: open PRs as the triggering user via dedicated tool (#1378)
Add an open_pull_request tool that creates a new PR via the GitHub REST
API using the triggering user's OAuth token (resolved by login from the
dashboard store), so the PR creator is the user rather than open-swe[bot].
Falls back to the GitHub App installation token for GitHub-triggered runs,
unmapped users, and bot-token-only deployments.

The user token never enters the sandbox: clone/push/comments still go
through the bot proxy via gh. The agent is steered to use the tool only
for OPENING a new PR; updates (body edits, mark ready) and pasted/existing
PRs continue to use gh pr edit. Existing-PR (422) returns the open PR's
URL so re-runs don't create duplicates.
2026-06-02 16:01:54 -07:00
Johannes du Plessis
4a55145bb1
feat: outcomes dataset + bootstrap/continual split via skills (#1365)
* fix: reset stale sandbox creation sentinel

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* fix: treat SANDBOX_CREATING as a timestamped cross-process lock

Only reset the sentinel when proven stale (older than the creation
timeout); otherwise wait for the worker that holds the lock so a
concurrent run does not create a duplicate sandbox.

* feat(analyzer): outcomes dataset + bootstrap/continual split via skills

Rename the review_style_analyzer graph to `analyzer` and split it into two
modes, plus capture reviewer finding outcomes for continual learning.

- Outcomes dataset: upsert resolved-by-commit (positive), dismissed (false
  positive), and GitHub/Slack thumbs findings into a single LangSmith dataset
  (openswe-reviewer-outcomes), keyed deterministically per finding+source.
  Emit points wired into update_finding, resolve_finding_thread, and the
  GitHub/Slack reaction handlers.
- Two playbooks delivered as deepagents skills (bootstrap-repo-analysis,
  continual-learning), served as virtual files via a CompositeBackend /skills/
  route + StateBackend (seeded into the run files channel at invoke time, never
  written to the sandbox). Mode is set by the launcher; continual runs fall
  back to the GitHub App installation token.
- Split launcher into start_bootstrap_analysis + start_continual_run; register
  a per-repo nightly continual-learning cron when bootstrap completes.
- New read_finding_outcomes tool feeds confirmed/dismissed findings back to the
  continual playbook.

Tests for outcome label mapping, skills helper, and cron idempotency.

* fix(analyzer): anchor continual cron runs to a real thread_id

The nightly continual-learning cron is threadless, and get_analyzer
early-returns an empty agent when configurable.thread_id is missing — so
every cron-launched run no-op'd before reading outcomes or saving a refined
prompt. Include the repo's deterministic analyzer thread_id in the continual
run configurable so the run executes; the threadless run carries no message
history, so nightly runs don't accumulate context.

* refactor(analyzer): move cron lifecycle calls out of the review-styles store

Drop the inline `analyzer_cron` imports from review_styles.py (added only to
dodge a circular import) by relocating the cron-trigger calls to the layer
above the store: registration to the save_review_style tool (after a prompt is
saved) and removal to the dashboard delete route. review_styles.py is now a
pure store again with top-level imports only.

* refactor: hoist reviewer_outcomes imports to module level

Move the two inline emit_finding_status_outcome imports introduced in this PR
(update_finding, resolve_finding_thread) to top-level imports. reviewer_outcomes
only depends on langsmith, so there is no circular import to avoid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-01 13:25:12 -07:00
open-swe[bot]
8130a188ef
chore: bump deepagents to 0.6.6 (#1359)
* chore: bump deepagents to 0.6.6

Co-authored-by: Mason Daugherty <61371264+mdrxy@users.noreply.github.com>

* chore: remove obsolete deepagents reducer patch

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Mason Daugherty <61371264+mdrxy@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-29 11:08:55 -07:00
open-swe[bot]
65acc4c9d3
fix: sanitize malformed Anthropic thinking blocks (#1357)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-28 16:15:13 -07:00
Johannes du Plessis
aeaa92eb8a
feat: Configure subagent model defaults (#1342) 2026-05-27 12:34:52 -07:00
open-swe[bot]
e347aed851
feat: make PR creation policy opt-in (#1334)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
2026-05-26 13:30:18 -07:00
Johannes du Plessis
e87085139b
feat: add Agents chat UI for cloud threads (#1323)
* feat(ui): add Agents chat UI ported from open-swe-app

Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(dashboard): wire Agents UI to LangGraph thread APIs

Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(dashboard): single agent reply per turn in Agents UI

Use UUID thread IDs LangGraph accepts, skip confirming_completion for
dashboard threads, and merge adapter agent messages so duplicate bubbles
do not render.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): polish Agents UI with floating prompt and layout cleanup

Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar
from open-swe-app, and refine chat layout so messages scroll behind the input.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(agent): patch deepagents reducer for None messages on checkpoint replay

LangGraph thread state could 500 when cancelled runs left messages as None.
Apply the reducer guard before graph import, fall back to metadata in the
dashboard API, and adjust Agents prompt bar layout.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(ui): unify sidebar user menu and clean up Agents UI navigation

Extract SidebarUserMenu so the dashboard and Agents sidebars render the
same profile button, drop the redundant Agents nav row in favor of the
existing Back to Agents link, add the open-swe logo header to the Agents
sidebar, flatten the New Agent button, and cap the home screen run list
to keep the prompt input in view.

* feat(ui): resizable/collapsible sidebar shared across dashboard and Agents

Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars
share a persisted width (default 260px, drag to resize, 200-420 range)
and a collapse toggle that hides the panel and surfaces a floating
reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on-
hover thread delete control in the Agents sidebar.

* feat(ui): instant user message and busy indicator on Agents transition

Stash submitted prompts in sessionStorage, pre-populate the new thread
detail cache, and merge pending prompts into the rendered message list
so the Agents page renders the user bubble plus the existing thinking
spinner immediately instead of flashing a skeleton and "Agent is
starting" while the run boots.

* feat(ui): token-stream agent replies in the Agents thread view

Opt the LangGraph runs into messages-tuple streaming and forward those
events through the existing SSE channel. The frontend now applies
AIMessageChunk deltas directly to the cached thread (cancelling any
in-flight refetch first so optimistic tokens are not clobbered) and
keeps positional pending prompts so the user bubble stays in the right
place while the agent streams its reply.

* fix(dashboard): await threads.join_stream before iterating

threads.join_stream is async def returning an AsyncIterator, so it must
be awaited before async for. The SSE endpoint was raising
TypeError: 'async for' requires an object with __aiter__ method, got
coroutine on every connection.

* fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns

Setting stream_mode=["values","messages-tuple","updates"] on
runs.create forces langchain_anthropic into streaming, and on the
second model call (after tool execution) its serialized thinking
blocks come back malformed, so Anthropic rejects the request with
'messages.1.content.0.thinking.thinking: Field required'. Revert to
the default stream_mode so claude-opus thinking + tool use runs to
completion. The frontend keeps the messages-event handler in place
as a no-op fallback for when streaming is re-enabled.

* feat(agents): per-thread model picker wired through to the run

Add optional model_id/effort to the create-thread and send-message
request bodies, forward them as agent_model_id/agent_effort in the
LangGraph run configurable, and record the resolved choice in thread
metadata so the UI can show the model the run is actually using.
get_agent now picks the per-thread override last (highest priority over
team default + profile override) and falls back gracefully when it is
absent or unsupported.

The frontend prompt bar becomes a controlled component fed by a
shared useModelOptions hook (options + profile -> defaultSelection).
AgentsHome seeds the picker from the user's profile default; the
thread view seeds from the thread's recorded model/effort and lets
each follow-up retarget the run.

* refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar

Drop the absolute-positioned send button, restore the original
px-4 py-3.5 min-h-[106px] flex-col container, and move the model
picker into a mt-auto pt-2 footer row so the placeholder text and
the model selector share the same horizontal padding.

* chore: fix lint/format CI failures

Remove unused imports and reformat two files flagged by ruff.

* fix(tests): stop messages-reducer patch tests from polluting the suite

Restore agent modules after reducer patch tests and import LangSmithSandbox
from agent.server in proxy refresh tests so isinstance checks stay valid.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-22 18:15:59 +00:00
Johannes du Plessis
32ec9b485f
feat: restructure Open SWE Review tab + wire create_prs (#1319)
* feat(dashboard): restructure Open SWE Review tab + wire create_prs

Restructures the dashboard around two related changes the reviewer settings
have been asking for:

- Wire profile.create_prs. Defaults to true (opt-out); when off the system
  prompt gets a `Pull Request Policy Override` section telling the agent
  to push the branch and notify with the branch URL instead of opening a
  PR. Removes the noop Slack Notifications / Allow Artifacts / First Name
  / Last Name controls and their schema fields.
- Repositories opt-in for Open SWE Review. New per-team enabled list
  stored in the LangGraph Store (`["enabled_review_repos"]`). Every
  reviewer webhook chokepoint now goes through `_is_repo_enabled_for_review`
  which AND-combines the existing env allowlist with the dashboard list.
  Default is empty (opt-in) — admins enable repos per-installation from
  the new Repositories page nested under Open SWE Review.
- Open SWE Review tab now mirrors the Cursor "rules" pattern: main page
  shows installation rows + a Rules entry; both drill into nested pages
  (/review/repositories/$owner and /review/styles) with a back link.
- Adds the new logo/favicon assets shipped from sidebar + html head.

Tests pass with a new autouse fixture (`tests/conftest.py`) that defaults
`is_review_repo_enabled` to True for existing allowlist tests.

* fix(dashboard): make main content scroll independently of the sidebar

Outer flex container was min-h-svh, so it grew with main's content and the
whole page scrolled — sidebar moved with it. Pin to h-svh + overflow-hidden
so the sidebar stays put and only <main> scrolls.

* fix(dashboard): make disabled repo toggles obviously disabled

Switch's disabled state used opacity-50 against a muted background, so
the not-admin state looked nearly identical to the off state. Bump to
opacity-40 + grayscale, and wrap each repo toggle in a span carrying a
native hover tooltip explaining why it's disabled.

* fix(switch): handle base-ui's data-disabled state

base-ui's Switch.Root sets data-disabled (not the HTML disabled attribute)
when disabled, so Tailwind's disabled: variant never matches and the
button keeps its cursor-pointer + clickable look. Mirror the styling
under the data-[disabled] variant and add pointer-events-none so the
disabled state is both visible and actually unclickable.

* feat(dashboard): paginate per-installation repository list

20 repos per page with Prev / page X of Y / Next controls at the bottom.
Pager only renders when there are more than 20 repos. Page resets to 0
when navigating between installations.

* feat(dashboard): global default model selectors for Agent + Reviewer

Adds team-wide default model + reasoning effort for both agents in the
Admin tab so operators can switch models without redeploying.

Resolution chain:
  Agent:    hardcoded -> LLM_MODEL_ID env -> team default -> user profile
  Reviewer: hardcoded -> LLM_MODEL_ID env -> team default -> per-call configurable

Team defaults live in team_settings and are validated against the
SUPPORTED_MODELS allowlist + the model's supported reasoning efforts.
'Inherit from env' clears the override and falls back to LLM_MODEL_ID.

* refactor(models): drop LLM_MODEL_ID env in favour of the team default

The team default is now the single source of truth for the runtime model
choice; per-user (agent) and per-call configurable (reviewer) selections
still win on top. When no admin has touched the team default, it surfaces
the hardcoded fallback (DEFAULT_MODEL_ID + its default effort), so the
admin UI's dropdown is always pre-populated with a sensible value.

The Admin UI loses the 'Inherit from env' option since there is no longer
an env layer to inherit from.

* chore(models): set hardcoded fallback to gpt-5.5 medium

Decouple the team-default boot value (gpt-5.5 / medium) from each model's
ProfileForm-suggested default_effort so we can change one without nudging
the other. The Opus xhigh default for new user profiles is unchanged.

* feat(dashboard): trigger-mode copy, Coming Soon badges, logout in My Settings

- Rename trigger mode 'ready_for_review' -> 'once_per_pr' with new
  description copy that matches the screenshot. Legacy stored values
  fall back to 'every_push' on read so the UI never shows an unknown
  selection.
- Add a 'Coming soon' badge + greyed-out + disabled state on the
  controls that don't have runtime consumers yet: Trigger Mode,
  Autofix Mode, Autofix Severity Threshold, and Automatically fix CI
  failures. SettingsRow grew a comingSoon prop to keep this consistent.
- My Settings drops the noop PR Preferences section and adds a Sign
  Out button. preferred_pr_destination is removed from the profile
  schema; old records get the field popped on next write.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-21 09:17:07 -07:00
Johannes du Plessis
82852f9eda
feat: tune reviewer for precision — web/wiki tools + recalibrated prompt (#1312)
* feat: tune reviewer for precision — web/wiki tools + recalibrated prompt

Reviewer agent now has web_search, fetch_url, and http_request alongside the
finding tools, so it can verify library semantics and consult the DeepWiki
auto-generated wiki for public repos (https://deepwiki.com/<owner>/<repo>)
before flagging cross-file or architectural concerns.

Prompt rewritten to push precision over recall:
- explicit severity ladder pushing reviews toward bimodal high/low instead of
  defaulting to medium
- ≤200-char description target (gold set averages ~186 chars; we were at ~436)
- mandatory docs / wiki / code lookup before flagging concurrency, security,
  or perf — the three categories that dominated false positives
- "do not flag" list covering compiler/linter-catchable nits, speculative
  claims without a concrete attacker/interleaving/scale, style preferences
  the codebase doesn't share, and test-quality nits on non-test diffs
- smart file-selection guidance for large PRs (deprioritize generated /
  vendored / pure-rename hunks)

Eval config switched to openai:gpt-5.5 + high reasoning effort for the next
benchmark run.

* trim prompt

* subagent prompting

* confidence ratings

* added medium

* enforce confidence threshold

* .

* reviewer: precision-tuned prompt + drop confidence gate

Rewrites the reviewer system prompt around a defensibility bar (anchor +
failure mode + maintainer wouldn't say "not a bug"), an explicit do-not-file
list (style nits, speculation, scope-policing, same-bug fan-out), and a
checklist of 10 bug archetypes drawn from a per-PR audit of the eval golden
set. The audit showed 145 FPs in the last eval split ~28% speculative, ~26%
style-nit, ~31% real-but-unscored (mostly same-archetype fan-out); the new
prompt targets each class directly.

Confidence is still recorded on every finding for post-hoc calibration but
no longer gates publication — the audit showed the gate was a no-op (agent
self-rated 65% of findings "high" regardless), and the prompt's defensibility
bar is the actual discipline. Drops CONFIDENCE_ORDER, CONFIDENCE_THRESHOLD,
the confidence_threshold kwarg on filter_findings_for_publish, the
confidence_filtered score_mode, and the min_confidence kwarg on the eval
target's _extract_comments — all dead once the gate is gone.

Also removes the "informational" severity tier from the Severity enum,
SEVERITY_ORDER, and all validators / tests / docstrings. It was reserved for
FYI observations the dataset never rewards.

* benchmax

* adding google provider

* slight steering

* tuning

* more tuning

* fix

* cleanup

* reducing overfitting

* Add per-repo review style profiles and inject them into the reviewer.

Dashboard users can analyze historical PR review feedback per repository,
edit the resulting style guide, and have it loaded from LangGraph Store at
reviewer runtime (including Martian eval runs) keyed by owner/name.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix review style job errors leaking exception details to clients.

Return generic dashboard messages while logging full stack traces server-side.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-20 18:35:00 +00:00
Johannes du Plessis
834efbc33c
feat: Adds ability to run evals against deployment (#1311)
* feat: tighten reviewer eval workflow

Require the reviewer to verify and dedupe findings before recording them, and make benchmark runs safe to execute against deployed reviewer graphs without posting GitHub reviews.

* chore: move reviewer eval settings to config

Load reviewer benchmark settings from the default eval config file so deployed eval runs do not require a wide CLI surface.

* feat: allow reviewer eval model overrides

Pass reviewer model and reasoning effort from the eval config into reviewer runs so isolated benchmark deployments can test Opus 4.7 high thinking.

* fix: use adaptive thinking for Opus 4.7

Switch Opus 4.7 model overrides to Anthropic adaptive thinking with effort instead of the deprecated budgeted thinking payload rejected by the API.

* refactor: use latest Anthropic effort API

Remove legacy Anthropic budget-token thinking support and route Anthropic efforts through adaptive thinking plus effort.

* revert prompting
2026-05-18 15:47:13 -07:00
Johannes du Plessis
88856a04fa
feat: open-swe dashboard for per-user profile config (#1302)
* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints

Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token

CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.

* feat: apply dashboard profile model/effort overrides in get_agent

Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.

* feat: ui/ TanStack Start dashboard for profile config

Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:

- /login   — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin   — Admin-only: list users and edit other profiles

API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.

Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.

* feat: searchable Combobox for default repo picker

Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.

* fix: address review comments + wire default_repo and Anthropic thinking

Security/correctness fixes from PR review:

* Open redirect: validate `redirect_to` in `/auth/login` against
  `DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
  into the state JWT. Anything off-allowlist falls back to the dashboard
  base URL. (PR #1302 r3250054386)

* Login CSRF: bind the OAuth `state` to the requesting browser. At
  `/auth/login` we generate a fresh nonce, set it as a short-lived
  HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
  embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
  we require the cookie nonce to hash-match the state JWT's nonce_hash
  (constant-time compare). (PR #1302 r3250054395)

* RMW race in profile vs token writes: split storage into two
  namespaces — `["profiles"]` for user-editable settings and
  `["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
  only writes its own namespace so an in-flight profile save can no
  longer clobber a fresh token from a concurrent re-login (and vice
  versa). (PR #1302 r3250054393)

* /repos pagination: follow `Link: rel="next"` for both
  `/user/installations` and per-installation `/repositories` with
  per_page=100, capped at 1000 items. (PR #1302 r3250054401)

Feature wires:

* default_repo: applied as a fallback in `get_slack_repo_config` (after
  explicit-repo / thread metadata, before the env defaults) and in the
  Linear webhook (after comment-body extraction, before team mapping).
  Both paths resolve the triggering user's GitHub login via
  GITHUB_USER_EMAIL_MAP and read the profile's default_repo.

* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
  kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
  to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
  anthropic. OpenAI path still ignores "max" since the Literal doesn't
  accept it.
2026-05-15 11:23:53 -07:00
Johannes du Plessis
5c7c78406c
fix: keep sandbox backend stable across recovery (#1294)w
Use a per-thread proxy so in-flight tools continue through the latest recreated sandbox instead of holding a stale backend reference.
2026-05-11 16:03:38 -07:00
open-swe[bot]
85343fab63
feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322] (#1280)
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]

Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.

Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>

* webapp: forward installation-token expiry to reviewer cache writes

The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 22:57:01 +00:00
Johannes du Plessis
094b2df939
feat: cross-provider model fallback on transient errors (#1281)
When the primary model raises a transient provider error (5xx, 429,
connection/timeout) the request is retried once against a fallback
model from the other provider. Anthropic primaries fall back to
OpenAI and vice versa. Also bumps the SDK max_retries from the
default 2 to 6 so quick blips stay on the primary and keep prompt
caching warm.

Triggered by 529 OverloadedError traces that ended runs silently
with no Slack/Linear/PR reply.
2026-05-08 15:35:13 -07:00
Johannes du Plessis
743b2b9ba4
fix: recover from mid-run sandbox death (#1274)
* fix: recover from mid-run sandbox death

Recreate dead sandboxes during tool execution and stop repeated unrecoverable timeout loops with a user-facing notification.

* fix: count repeated sandbox recreations

Treat consecutive sandbox recreations as an unrecovered failure streak so outages cannot loop until the model-call limit.
2026-05-08 12:55:36 -07:00
Johannes du Plessis
5d1020e8d7
fix: restore Co-authored-by trailer for triggering user (#1271)
The gh-cli migration removed agent/middleware/open_pr.py and
agent/tools/commit_and_open_pr.py — the only callers of
add_user_coauthor_trailer / add_pr_collaboration_note. Since then the
agent has been driving commits and PRs entirely via gh, with no
attribution back to the Slack/Linear/GitHub user who triggered the run.

Resolve the triggering user's identity in get_agent (reusing the
existing authorship helpers) and inject a Collaborative Attribution
section into the system prompt with the exact trailer and PR-body note
to use. The section is only rendered when an identity is resolvable, so
runs without a known triggering user are unchanged.
2026-05-08 10:42:15 -07:00
open-swe[bot]
96f97710ad
feat: add optional Slack Assistants API typing status indicator (#1269)
* feat: add optional Slack Assistants API typing status indicator

Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.

* fix(slack): drop redundant clear, add status heartbeat across model calls

- Slack auto-clears the typing indicator on bot post; remove the explicit
  assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
  that refreshes it on every model tick so it stays visible across long
  agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
  configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
  way out per Slack docs); no scope or app-config change required.

* feat(slack): contextual status text + rotating loading_messages

- set_slack_assistant_status now accepts an optional loading_messages list
  (capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
  payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
  assistant message's tool calls (e.g. "searching the codebase…" after
  grep, "running commands…" after execute), falling back to the default
  "is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
  contextual status on each refresh.

* fix slack assistant status lifecycle

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-08 10:21:55 -07:00
Johannes du Plessis
88d3659d00
start sandbox before proxy refresh (#1249) 2026-05-07 11:26:25 -07:00
Johannes du Plessis
1319347dd9
feat: route Slack PR review requests (#1245)
* feat: route Slack PR review requests

Add a lightweight Slack review command path that starts the reviewer graph directly and gives the core agent a handoff tool when review requests are misrouted.

* fix: harden Slack PR review routing

* fix: validate Slack PR review URLs

* fix: preserve malformed GitHub review routing
2026-05-06 17:14:43 -07:00
Johannes du Plessis
ace71b0fd0
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring

- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
  alongside the main `agent` graph. Reuses the same sandbox lifecycle,
  GH proxy auth, and middleware primitives from `agent.server`, but with
  a narrower tool set, a reviewer-specific system prompt, no
  commit/push, and the `task` (subagent) tool stripped via
  `_ToolExclusionMiddleware` so review stays in one context.

- New `github_comment` tool: agents call it once per issue with
  `(file, line, body, severity)` and the eval scores those calls
  against golden comments.

- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
  *not* on the reviewer's stack — that middleware exists to enforce the
  main agent's "always finalize via Slack/Linear/PR" contract, which
  the reviewer doesn't have. The main agent's behavior is unchanged.

- `evals/reviewer/target.py`: send PR info as a user message, extract
  every `github_comment` tool call (multiple expected per review) into
  the run output.

- `evals/reviewer/judge.py`: per-example evaluator now returns a list
  of metrics under `{"results": [...]}` so LangSmith averages each
  numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
  the UI. Dropped the broken `aggregate_pr` summary evaluator that
  reached for an attribute that doesn't exist on `RunTree`.

- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
  `client.list_examples(limit=N)` since `aevaluate` doesn't accept
  `max_examples`.

- Makefile: `dev` and `run` targets now use `uv run` so they work
  without an activated venv.

* resolve comments

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
Johannes du Plessis
7d4d7e5833
fix: re-apply git identity every run, drive from env vars (#1242)
* fix: re-apply git identity on every agent invocation, source from env

Cached/reconnected sandboxes weren't getting `git config --global` re-applied
(the call was gated on new-sandbox creation), so commits picked up whatever
identity the sandbox happened to have — sometimes an email not associated
with any GitHub account, which Vercel rejects on preview deploys.

Move the git config call out of the new-sandbox branch so it runs on every
get_agent, and source the values from OPEN_SWE_GIT_AUTHOR_NAME /
OPEN_SWE_GIT_AUTHOR_EMAIL env vars (defaulting to the existing bot identity)
so forks can override without code changes.

* cleanup

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-05 14:53:27 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
Brace Sproul
6579b84b74
fix: Set model recursion limit to 5k (#1234) 2026-05-03 14:41:24 -07:00
Brace Sproul
3405d145ac
feat: add edit_pull_request tool for editing PR titles/descriptions (#1063)
* feat: add edit_pull_request tool for editing PR titles and descriptions

* fix: patch auth flow in open PR middleware tests

* fix: support app token for editing PRs

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 18:12:41 -07:00
langsmith-forge[bot]
48965a82a6
fix: coerce malformed integer strings in read_file offset/limit params (#1216)
- Root cause: LLM occasionally generates strings like '1, 80' or '170, "limit": 60'
  for integer fields, causing a Pydantic ValidationError and wasting an LLM turn
- Change: add SanitizeToolInputsMiddleware in agent/middleware/sanitize_tool_inputs.py
  that extracts the leading integer from any string value in offset/limit before
  the call reaches Pydantic validation; registered before ToolErrorMiddleware in server.py
- Verified: 14 unit tests covering all three production trace patterns pass

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:29:48 -07:00
langsmith-forge[bot]
e5bc27a0ad
fix: notify users via Slack when agent hits model call step limit (#1204)
* fix: notify users via Slack when agent hits model call step limit

- Root cause: GraphRecursionError at 1000 steps bypassed all @after_agent
  middleware including open_pr_if_needed, leaving users with no notification
- Change: Added ModelCallLimitMiddleware(run_limit=60) to intercept gracefully
  before the hard recursion limit, and added notify_step_limit_reached
  @after_agent middleware to post a Slack thread reply when the limit fires
- Verified: 107 existing tests pass, no regressions

* fix: harden step-limit Slack notification

Ensure the step-limit notification runs after the PR safety net and cover the new middleware behavior with focused unit tests.

---------

Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 14:24:25 -07:00
langsmith-forge-dev[bot]
4060933ce4
feat: add github CI check run tools for shepherding CI (#1121)
* feat: add github CI check run tools for shepherding CI

Add get_pr_check_runs and rerun_failed_check_runs tools that authenticate
using the GitHub App installation token so the agent can query and retry
CI status on private repos without relying on GH_TOKEN or unauthenticated
http_request calls.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: handle paginated GitHub CI results

* refactor(github_ci): address review feedback

- Rename rerun_failed_check_runs -> rerun_failed_workflow_runs and clarify
  in docstrings that the tool only retries GitHub Actions workflow runs
  (not third-party CI checks surfaced by get_pr_check_runs).
- Skip action_required workflow runs when rerunning; those need manual
  approval, not a rerun.
- Run rerun-failed-jobs requests concurrently via asyncio.gather instead
  of sequentially.
- Fix latent pagination bug in _fetch_paginated_items where caller-supplied
  params could overwrite per_page/page and break the end-of-pagination
  check; reserved keys now always win and the threshold uses a PER_PAGE
  constant.
- Set an explicit 30s httpx timeout so a hung GitHub call cannot stall
  the agent loop.
- Restore alphabetical ordering of tools in agent/tools/__init__.py.
- Add tests for: a 500 surfaced on a later pagination page, and
  action_required runs being filtered out of rerun candidates.

---------

Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-05-01 12:19:04 -07:00
langsmith-forge-dev[bot]
d63780a77f
fix: add get_pr_review_comments tool to fetch PR review comments with auth (#1043)
* feat: add get_pr_review_comments tool for authenticated GitHub API access

The agent was asking users to paste PR review comments because it had no
tool to fetch them with auth. This adds get_pr_review_comments, which uses
the GitHub App installation token to fetch all three comment types (thread
comments, inline review comments, review submissions) from private repos.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* u

* u

---------

Co-authored-by: Forge Agent <agent@forge.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Palash Shah <palash@langchain.dev>
Co-authored-by: Palash Shah <35114859+Palashio@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-30 17:01:13 -07:00