* feat: repo-scoped dynamic sandbox snapshots
Let admins build a per-repo sandbox image from a custom Dockerfile so runs
targeting that repo boot from a snapshot with its deps pre-baked. Snapshot
selection is purely additive: repos without a `ready` repo-scoped snapshot
always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID.
Backend adds a repo_snapshots store module (Dockerfile + build status keyed by
owner/name), threads the resolved repo through the LangSmith sandbox creation
path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a
throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The
UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor +
build status/logs).
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden repo snapshot builds
Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins
cannot accidentally build a repo snapshot from a bare Python image that lacks
Open SWE's sandbox tools. Allow stale building records to be retried by tracking
build_started_at and treating old or missing timestamps as stale.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: document repo snapshot base image config
Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert
missing base-image configuration into a handled dashboard API error so admins see
a clear configuration message instead of an unhandled template-generation error.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add user-scoped Currents.dev API key for e2e test investigation
Allow each user to configure their own Currents.dev API key on the
Profile Settings page. The key is encrypted at rest in a per-user
LangGraph Store namespace and feeds server-side read-only tools that
query the Currents REST API (runs, instances, projects, test results)
so agent runs can inspect e2e test failures including screenshots and
DOM snapshots.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: add pagination cursors to currents_list_project_runs
Address review feedback: forward starting_after/ending_before cursor
parameters to /projects/{projectId}/runs so the agent can paginate
beyond the first 50 results.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: server-side Datadog/LangSmith observability tools + team creds
Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: address review on observability tools
Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
read-modify-write race dropping the other provider on concurrent saves.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: async email resolution in observability authorization gate
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce client-side deadline on sandbox execute
The langsmith SDK's default execute path is now a WebSocket stream with
no client-side read deadline. On a live socket where the dataplane never
emits an exit/error frame, CommandHandle.result blocks forever in an
uncancellable thread, wedging the run (introduced by the langsmith
0.8.3 -> 0.8.8 bump in #1385). The command `timeout` is only enforced
server-side, so it doesn't fire.
TimeoutLangSmithSandbox drives a non-blocking CommandHandle and kills the
command if it overruns its timeout by a grace window
(SANDBOX_EXECUTE_CLIENT_GRACE_SECONDS, default 30), returning a timed-out
tool result instead of hanging. WS connect failures fall back to the base
wait=True path, whose HTTP fallback carries its own request deadline.
* fix: fall back to HTTP when WS execute connect fails
run(wait=False) eagerly opens the WebSocket and reads the "started" frame,
so connect/setup failures (and connect timeouts) raise from the run() call
itself, not from handle.result. The previous structure left run() outside
the try, so those failures bypassed the HTTP fallback and would fail every
sandbox command in any environment where the WS path is unavailable.
Move handle creation inside the fallback handler in both execute and
aexecute, and run it via to_thread in the async path since it now blocks on
connect. Add tests for connect-failure and connect-timeout fallback.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
create_local_sandbox now mkdir -p's the resolved root dir, so a custom
LOCAL_SANDBOX_ROOT_DIR (or a /tmp path cleared on reboot) no longer fails
sandbox work-dir resolution.
* feat: move github workflows to gh cli
Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.
* docker ignore + snapshot and docker image updates
* updated image and instructions
* removing open_pr if needed after agent call
Previous defaults (4 vCPU / 15 GiB) exceed the maximum sandbox size
the LangSmith API accepts, causing a 400 on every sandbox creation:
sandbox size 4 vCPU / 15360 MiB exceeds the maximum supported size;
must fit within one of: small (1 vCPU / 1792 MiB),
medium (1 vCPU / 3840 MiB), or large (2 vCPU / 7936 MiB)
Default to the "large" cap (2 vCPU / 7936 MiB) so deployments without
DEFAULT_SANDBOX_VCPUS / DEFAULT_SANDBOX_MEM_BYTES env overrides boot
into a working state.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: migrate LangSmith sandbox creation to snapshot API
Replaces the template-based sandbox flow (DEFAULT_SANDBOX_TEMPLATE_NAME /
DEFAULT_SANDBOX_TEMPLATE_IMAGE) with the new snapshot-based flow.
- New required env var DEFAULT_SANDBOX_SNAPSHOT_ID (UUID of a pre-built
LangSmith snapshot; build out-of-band via UI or SandboxClient.create_snapshot)
- Optional DEFAULT_SANDBOX_SNAPSHOT_FS_CAPACITY_BYTES overrides the root FS
size at boot (default 32 GiB)
- Startup-time validation via a FastAPI lifespan hook: the server refuses
to boot with a clear ValueError if SANDBOX_TYPE=langsmith and
DEFAULT_SANDBOX_SNAPSHOT_ID is unset, so failures surface in boot logs
rather than on the first thread
- Reconnect-to-existing-sandbox path unchanged
- Docs (INSTALLATION.md, CUSTOMIZATION.md) updated to describe the new
snapshot workflow
* fix: format create_sandbox_snapshot.py to pass ruff
---------
Co-authored-by: aran-yogesh <yogesh.mahendran@langchain.dev>
* fix: add retry with delay for sandbox proxy config to avoid 500 when proxy isn't ready
* fix: add retry with exponential backoff and connection error handling for proxy config