* feat: repo-scoped dynamic sandbox snapshots
Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.
Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden repo snapshot builds
Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: document repo snapshot base image config
Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: refresh sandbox GitHub proxy token before mid-run expiry
GitHub App installation tokens expire after exactly 1 hour. The LangSmith
sandbox proxy was configured once at run start with a snapshot of that
token, so runs longer than ~1h hit 401s on every gh/git call. Record the
proxy token's expiry per thread and add a before-model hook that
re-configures the proxy with a fresh token when it nears expiry.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: preserve repo-scoped proxy token on mid-run refresh
Reviewer runs mint a repository-scoped installation token. Record the
repo scope per thread alongside the expiry so the before-model refresh
re-mints a token with the same scope instead of an installation-wide
token, avoiding privilege expansion on long reviewer runs.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: update passthrough stub for github_proxy_repositories param
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: scope public reviewer tokens
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor: simplify reviewer token wiring; fix push re-scope + red test
- Remove the redundant _check_or_recreate_sandbox_for_proxy /
_refresh_github_proxy_or_recreate_for_proxy wrappers and call the
underlying functions directly (they already default the token to None).
- process_github_push_event: re-scope the GitHub App token when the push
payload lacked repo privacy/id but PR metadata reveals a public repo, so
reviewer.py never proxies a full-installation token for a public PR.
- Clarify the two-token sequence in trigger_pr_review_from_ref.
- Fix pre-existing failing test test_proxy_refresh_failure_recreates_sandbox
and add coverage for _reviewer_token_for_repo + push-event scoping.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat(ui): add Agents chat UI ported from open-swe-app
Introduce a Cursor-style Agents surface separate from the dashboard, with ported chat/diff components and mock thread data until LangGraph APIs land.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(dashboard): wire Agents UI to LangGraph thread APIs
Add dashboard thread list/detail/run/message/stream endpoints with a LangGraph message adapter, dashboard OAuth auth for runs, and TanStack Query hooks replacing mock data.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(dashboard): single agent reply per turn in Agents UI
Use UUID thread IDs LangGraph accepts, skip confirming_completion for
dashboard threads, and merge adapter agent messages so duplicate bubbles
do not render.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(ui): polish Agents UI with floating prompt and layout cleanup
Remove no-op chrome (git panel, headers, sidebar search), port CloudPromptBar
from open-swe-app, and refine chat layout so messages scroll behind the input.
Co-authored-by: Cursor <cursoragent@cursor.com>
* fix(agent): patch deepagents reducer for None messages on checkpoint replay
LangGraph thread state could 500 when cancelled runs left messages as None.
Apply the reducer guard before graph import, fall back to metadata in the
dashboard API, and adjust Agents prompt bar layout.
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat(ui): unify sidebar user menu and clean up Agents UI navigation
Extract SidebarUserMenu so the dashboard and Agents sidebars render the
same profile button, drop the redundant Agents nav row in favor of the
existing Back to Agents link, add the open-swe logo header to the Agents
sidebar, flatten the New Agent button, and cap the home screen run list
to keep the prompt input in view.
* feat(ui): resizable/collapsible sidebar shared across dashboard and Agents
Add a useSidebarLayout hook + SidebarFrame wrapper so both sidebars
share a persisted width (default 260px, drag to resize, 200-420 range)
and a collapse toggle that hides the panel and surfaces a floating
reopen button. Also adds a DELETE /threads/{id} endpoint and an X-on-
hover thread delete control in the Agents sidebar.
* feat(ui): instant user message and busy indicator on Agents transition
Stash submitted prompts in sessionStorage, pre-populate the new thread
detail cache, and merge pending prompts into the rendered message list
so the Agents page renders the user bubble plus the existing thinking
spinner immediately instead of flashing a skeleton and "Agent is
starting" while the run boots.
* feat(ui): token-stream agent replies in the Agents thread view
Opt the LangGraph runs into messages-tuple streaming and forward those
events through the existing SSE channel. The frontend now applies
AIMessageChunk deltas directly to the cached thread (cancelling any
in-flight refetch first so optimistic tokens are not clobbered) and
keeps positional pending prompts so the user bubble stays in the right
place while the agent streams its reply.
* fix(dashboard): await threads.join_stream before iterating
threads.join_stream is async def returning an AsyncIterator, so it must
be awaited before async for. The SSE endpoint was raising
TypeError: 'async for' requires an object with __aiter__ method, got
coroutine on every connection.
* fix(dashboard): drop messages-tuple stream_mode that broke thinking-mode tool turns
Setting stream_mode=["values","messages-tuple","updates"] on
runs.create forces langchain_anthropic into streaming, and on the
second model call (after tool execution) its serialized thinking
blocks come back malformed, so Anthropic rejects the request with
'messages.1.content.0.thinking.thinking: Field required'. Revert to
the default stream_mode so claude-opus thinking + tool use runs to
completion. The frontend keeps the messages-event handler in place
as a no-op fallback for when streaming is re-enabled.
* feat(agents): per-thread model picker wired through to the run
Add optional model_id/effort to the create-thread and send-message
request bodies, forward them as agent_model_id/agent_effort in the
LangGraph run configurable, and record the resolved choice in thread
metadata so the UI can show the model the run is actually using.
get_agent now picks the per-thread override last (highest priority over
team default + profile override) and falls back gracefully when it is
absent or unsupported.
The frontend prompt bar becomes a controlled component fed by a
shared useModelOptions hook (options + profile -> defaultSelection).
AgentsHome seeds the picker from the user's profile default; the
thread view seeds from the thread's recorded model/effort and lets
each follow-up retarget the run.
* refactor(ui): align Agents prompt bar layout with open-swe-app PromptBar
Drop the absolute-positioned send button, restore the original
px-4 py-3.5 min-h-[106px] flex-col container, and move the model
picker into a mt-auto pt-2 footer row so the placeholder text and
the model selector share the same horizontal padding.
* chore: fix lint/format CI failures
Remove unused imports and reformat two files flagged by ruff.
* fix(tests): stop messages-reducer patch tests from polluting the suite
Restore agent modules after reducer patch tests and import LangSmithSandbox
from agent.server in proxy refresh tests so isinstance checks stay valid.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]
Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* webapp: forward installation-token expiry to reviewer cache writes
The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: move github workflows to gh cli
Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.
* docker ignore + snapshot and docker image updates
* updated image and instructions
* removing open_pr if needed after agent call