* feat: optional separate LangSmith key/endpoint for sandboxes
Adds optional SANDBOX_LANGSMITH_API_KEY / SANDBOX_LANGSMITH_ENDPOINT env
overrides so sandboxes can run against a different LangSmith workspace than
the one used for tracing and other API calls. Both fall back to the existing
LANGSMITH_API_KEY / LANGSMITH_ENDPOINT resolution, so default behavior is
unchanged.
Applied to sandbox create/connect/delete, the GitHub proxy config, and repo
snapshot builds.
* feat: name langsmith sandboxes openswe-<b32(thread id)>
New sandboxes get a deterministic, thread-traceable name derived from the
LangGraph thread id (UUID base32-encoded lowercase, no padding), e.g.
openswe-ci2fm6asgrlhqerukz4bencwpa. Falls back to an unset name when no thread
id is present. Reconnect/delete still key off the server-assigned sandbox id.
* fix: pass sandbox base URL (root + /v2/sandboxes) to langsmith SDK clients
The SDK's api_endpoint is the sandbox base, not the API root — its methods
append /boxes, /snapshots, etc. Passing the bare root sent calls to
<root>/boxes instead of <root>/v2/sandboxes/boxes. Add _get_sandbox_api_endpoint
for the SDK clients (async client, provider, snapshot SandboxClient) while the
proxy-config PATCH keeps using the root.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit e826864dce0e56cda7decbc48254b1e13eef07e2)
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
Add SANDBOX_CREATE_EXTRA_JSON so operators can merge extra fields (e.g.
{"_internal_runtime":"v2"}) into the LangSmith sandbox-create request body.
The SDK's create_sandbox builds a fixed payload with no passthrough, so we
wrap the HTTP client's post to inject the fields on the POST /boxes request
only. Malformed JSON fails at startup validation.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 2238303306493ae6fcd0c2d4ab4236adf283a896)
Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev>
* feat: repo-scoped dynamic sandbox snapshots
Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.
Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden repo snapshot builds
Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: document repo snapshot base image config
Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce client-side deadline on sandbox execute
The langsmith SDK's default execute path is now a WebSocket stream with
no client-side read deadline. On a live socket where the dataplane never
emits an exit/error frame, CommandHandle.result blocks forever in an
uncancellable thread, wedging the run (introduced by the langsmith
0.8.3 -> 0.8.8 bump in #1385). The command `timeout` is only enforced
server-side, so it doesn't fire.
TimeoutLangSmithSandbox drives a non-blocking CommandHandle and kills the
command if it overruns its timeout by a grace window
(SANDBOX_EXECUTE_CLIENT_GRACE_SECONDS, default 30), returning a timed-out
tool result instead of hanging. WS connect failures fall back to the base
wait=True path, whose HTTP fallback carries its own request deadline.
* fix: fall back to HTTP when WS execute connect fails
run(wait=False) eagerly opens the WebSocket and reads the "started" frame,
so connect/setup failures (and connect timeouts) raise from the run() call
itself, not from handle.result. The previous structure left run() outside
the try, so those failures bypassed the HTTP fallback and would fail every
sandbox command in any environment where the WS path is unavailable.
Move handle creation inside the fallback handler in both execute and
aexecute, and run it via to_thread in the async path since it now blocks on
connect. Add tests for connect-failure and connect-timeout fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: move github workflows to gh cli
Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.
* docker ignore + snapshot and docker image updates
* updated image and instructions
* removing open_pr if needed after agent call
Previous defaults (4 vCPU / 15 GiB) exceed the maximum sandbox size
the LangSmith API accepts, causing a 400 on every sandbox creation:
sandbox size 4 vCPU / 15360 MiB exceeds the maximum supported size;
must fit within one of: small (1 vCPU / 1792 MiB),
medium (1 vCPU / 3840 MiB), or large (2 vCPU / 7936 MiB)
Default to the "large" cap (2 vCPU / 7936 MiB) so deployments without
DEFAULT_SANDBOX_VCPUS / DEFAULT_SANDBOX_MEM_BYTES env overrides boot
into a working state.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: migrate LangSmith sandbox creation to snapshot API
Replaces the template-based sandbox flow (DEFAULT_SANDBOX_TEMPLATE_NAME /
DEFAULT_SANDBOX_TEMPLATE_IMAGE) with the new snapshot-based flow.
- New required env var DEFAULT_SANDBOX_SNAPSHOT_ID (UUID of a pre-built
LangSmith snapshot; build out-of-band via UI or SandboxClient.create_snapshot)
- Optional DEFAULT_SANDBOX_SNAPSHOT_FS_CAPACITY_BYTES overrides the root FS
size at boot (default 32 GiB)
- Startup-time validation via a FastAPI lifespan hook: the server refuses
to boot with a clear ValueError if SANDBOX_TYPE=langsmith and
DEFAULT_SANDBOX_SNAPSHOT_ID is unset, so failures surface in boot logs
rather than on the first thread
- Reconnect-to-existing-sandbox path unchanged
- Docs (INSTALLATION.md, CUSTOMIZATION.md) updated to describe the new
snapshot workflow
* fix: format create_sandbox_snapshot.py to pass ruff
---------
Co-authored-by: aran-yogesh <yogesh.mahendran@langchain.dev>
* fix: add retry with delay for sandbox proxy config to avoid 500 when proxy isn't ready
* fix: add retry with exponential backoff and connection error handling for proxy config