* fix: enforce client-side deadline on sandbox execute
The langsmith SDK's default execute path is now a WebSocket stream with
no client-side read deadline. On a live socket where the dataplane never
emits an exit/error frame, CommandHandle.result blocks forever in an
uncancellable thread, wedging the run (introduced by the langsmith
0.8.3 -> 0.8.8 bump in #1385). The command `timeout` is only enforced
server-side, so it doesn't fire.
TimeoutLangSmithSandbox drives a non-blocking CommandHandle and kills the
command if it overruns its timeout by a grace window
(SANDBOX_EXECUTE_CLIENT_GRACE_SECONDS, default 30), returning a timed-out
tool result instead of hanging. WS connect failures fall back to the base
wait=True path, whose HTTP fallback carries its own request deadline.
* fix: fall back to HTTP when WS execute connect fails
run(wait=False) eagerly opens the WebSocket and reads the "started" frame,
so connect/setup failures (and connect timeouts) raise from the run() call
itself, not from handle.result. The previous structure left run() outside
the try, so those failures bypassed the HTTP fallback and would fail every
sandbox command in any environment where the WS path is unavailable.
Move handle creation inside the fallback handler in both execute and
aexecute, and run it via to_thread in the async path since it now blocks on
connect. Add tests for connect-failure and connect-timeout fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
create_local_sandbox now mkdir -p's the resolved root dir, so a custom
LOCAL_SANDBOX_ROOT_DIR (or a /tmp path cleared on reboot) no longer fails
sandbox work-dir resolution.
* feat: move github workflows to gh cli
Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.
* docker ignore + snapshot and docker image updates
* updated image and instructions
* removing open_pr if needed after agent call
Previous defaults (4 vCPU / 15 GiB) exceed the maximum sandbox size
the LangSmith API accepts, causing a 400 on every sandbox creation:
sandbox size 4 vCPU / 15360 MiB exceeds the maximum supported size;
must fit within one of: small (1 vCPU / 1792 MiB),
medium (1 vCPU / 3840 MiB), or large (2 vCPU / 7936 MiB)
Default to the "large" cap (2 vCPU / 7936 MiB) so deployments without
DEFAULT_SANDBOX_VCPUS / DEFAULT_SANDBOX_MEM_BYTES env overrides boot
into a working state.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: migrate LangSmith sandbox creation to snapshot API
Replaces the template-based sandbox flow (DEFAULT_SANDBOX_TEMPLATE_NAME /
DEFAULT_SANDBOX_TEMPLATE_IMAGE) with the new snapshot-based flow.
- New required env var DEFAULT_SANDBOX_SNAPSHOT_ID (UUID of a pre-built
LangSmith snapshot; build out-of-band via UI or SandboxClient.create_snapshot)
- Optional DEFAULT_SANDBOX_SNAPSHOT_FS_CAPACITY_BYTES overrides the root FS
size at boot (default 32 GiB)
- Startup-time validation via a FastAPI lifespan hook: the server refuses
to boot with a clear ValueError if SANDBOX_TYPE=langsmith and
DEFAULT_SANDBOX_SNAPSHOT_ID is unset, so failures surface in boot logs
rather than on the first thread
- Reconnect-to-existing-sandbox path unchanged
- Docs (INSTALLATION.md, CUSTOMIZATION.md) updated to describe the new
snapshot workflow
* fix: format create_sandbox_snapshot.py to pass ruff
---------
Co-authored-by: aran-yogesh <yogesh.mahendran@langchain.dev>
* fix: add retry with delay for sandbox proxy config to avoid 500 when proxy isn't ready
* fix: add retry with exponential backoff and connection error handling for proxy config