* feat: dashboard backend — GitHub OAuth, profile CRUD, admin endpoints
Adds agent/dashboard/ FastAPI router mounted at /dashboard/api covering:
- GitHub App OAuth login → JWT cookie session (cross-domain ready)
- profile CRUD against LangGraph Store with model+effort validation
- admin gate via CONFIGURED_ADMINS
- /repos via /user/installations using the user's encrypted OAuth token
CORS allowlist on webapp.py is opt-in via DASHBOARD_ALLOWED_ORIGINS so the
Vercel-hosted frontend can call the LangSmith deployment with credentials.
* feat: apply dashboard profile model/effort overrides in get_agent
Look up the triggering user's GitHub login from config (direct field or
GITHUB_USER_EMAIL_MAP reverse lookup), read their profile from the Store,
and apply default_model + reasoning_effort to make_model when both are
valid. Effort 'max' is captured on the profile but not yet wired through —
the OpenAI Reasoning Literal doesn't accept it.
* feat: ui/ TanStack Start dashboard for profile config
Scaffolded with the shadcn b7CScJIjA preset (TanStack Start template,
base-ui primitives, Tailwind v4). Three routes:
- /login — Sign in with GitHub (links to /dashboard/api/auth/login)
- /profile — Edit default model, reasoning effort, default repo
- /admin — Admin-only: list users and edit other profiles
API client (src/lib/api.ts) uses credentials: include so the osw_session
cookie set by the OAuth callback rides cross-origin. VITE_DASHBOARD_API_BASE_URL
points at the LangSmith deployment.
Effort options re-render when the model changes; 'max' on Opus 4.7 is
captured on the profile but ignored downstream until anthropic reasoning
is wired through make_model.
* feat: searchable Combobox for default repo picker
Replaces the Select with a base-ui Combobox so users can filter by typing,
the popup is wider than the trigger so full owner/repo names are readable,
and the list caps at max-h-80 to stay on screen.
* fix: address review comments + wire default_repo and Anthropic thinking
Security/correctness fixes from PR review:
* Open redirect: validate `redirect_to` in `/auth/login` against
`DASHBOARD_BASE_URL` + `DASHBOARD_ALLOWED_ORIGINS` before signing it
into the state JWT. Anything off-allowlist falls back to the dashboard
base URL. (PR #1302 r3250054386)
* Login CSRF: bind the OAuth `state` to the requesting browser. At
`/auth/login` we generate a fresh nonce, set it as a short-lived
HttpOnly SameSite=Lax cookie scoped to `/dashboard/api/auth`, and
embed `hash_state_nonce(nonce)` in the state JWT. At `/auth/callback`
we require the cookie nonce to hash-match the state JWT's nonce_hash
(constant-time compare). (PR #1302 r3250054395)
* RMW race in profile vs token writes: split storage into two
namespaces — `["profiles"]` for user-editable settings and
`["oauth_tokens"]` for the encrypted GitHub token. Each upsert now
only writes its own namespace so an in-flight profile save can no
longer clobber a fresh token from a concurrent re-login (and vice
versa). (PR #1302 r3250054393)
* /repos pagination: follow `Link: rel="next"` for both
`/user/installations` and per-installation `/repositories` with
per_page=100, capped at 1000 items. (PR #1302 r3250054401)
Feature wires:
* default_repo: applied as a fallback in `get_slack_repo_config` (after
explicit-repo / thread metadata, before the env defaults) and in the
Linear webhook (after comment-body extraction, before team mapping).
Both paths resolve the triggering user's GitHub login via
GITHUB_USER_EMAIL_MAP and read the profile's default_repo.
* Anthropic "thinking" effort: `make_model` now accepts a `thinking`
kwarg; `get_agent` maps profile effort {low,medium,high,xhigh,max}
to budget_tokens {1k,4k,12k,32k,60k} when the chosen model is
anthropic. OpenAI path still ignores "max" since the Literal doesn't
accept it.
The SLACK_ASSISTANTS_API_ENABLED env flag and its gating function
are removed. set_slack_assistant_status now always proceeds when a
bot token and channel/thread are provided.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: TTL and revocation handling for cached GitHub OAuth tokens [closes AB-2322]
Persist github_token_expires_at alongside github_token_encrypted, treat
expired cache entries as missing so we re-resolve before kicking off
runs, and invalidate the cached ciphertext on a downstream 401 so the
next invocation gets a fresh token instead of replaying a revoked one.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* webapp: forward installation-token expiry to reviewer cache writes
The three reviewer-thread persist sites in webapp.py were calling
get_github_app_installation_token() (no expiry) and persist_encrypted_github_token
without expires_at, so cached App tokens were treated as never-expiring even
though they actually expire in ~1 hour.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: add 10 more tips to the trace-reply rotation
The TRACE_REPLY_TIPS pool only had 9 tips, so users mostly saw the same
ones. Added 10 more grounded in actual features (review command, image
attachments, GitHub-issue triggers, persistent sandboxes, OAuth fallback,
etc.) so the rotation surfaces more of what open-swe can actually do.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* swap out 3 tips per review feedback
Replaced the suggestion-blocks, sandbox-auto-recovery, and OAuth-fallback
tips with three more practically useful ones: cross-posted Slack message
resolution, the web_search tool, and the Linear ticket-management tools.
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
When the primary model raises a transient provider error (5xx, 429,
connection/timeout) the request is retried once against a fallback
model from the other provider. Anthropic primaries fall back to
OpenAI and vice versa. Also bumps the SDK max_retries from the
default 2 to 6 so quick blips stay on the primary and keep prompt
caching warm.
Triggered by 529 OverloadedError traces that ended runs silently
with no Slack/Linear/PR reply.
Appending "Please react with 👍 or 👎..." to every Slack completion
message ended up being repetitive noise. Drop the prompt-side instruction
and add the same ask as one of the rotating tips on the trace reply, so
users still see it occasionally without it cluttering each final summary.
Now that assistant.threads.setStatus carries the "is thinking…"
indicator, the per-run greeting phrase + auto-unfurl on the LangSmith
link are just visual noise stacking on top of it. The trace reply now
posts only `<url|View trace>` + a tip, with link unfurling off so the
smith.langchain.com card no longer appears.
Adds an `unfurl_links`/`unfurl_media` knob to post_slack_thread_reply_with_ts
(default-on to preserve behaviour for every other caller).
* feat: add Slack reaction feedback to LangSmith
Record Slack reaction feedback against explicitly mapped LangGraph runs so user ratings are idempotent and tied to the message they reacted to.
* Address review feedback on Slack reaction → LangSmith feedback
- langsmith.py: drop lru_cache on _build_langsmith_feedback_clients so
rotated keys / late env hydration are picked up; dedupe by (key, url)
tuple instead of key alone so the same key pointing at different
endpoints (cloud + self-hosted) builds both clients.
- langsmith.py: treat LangSmithNotFoundError on delete_feedback as
success — out-of-order or redelivered reaction_removed events would
otherwise loop forever on Slack's retry policy.
- slack_feedback.py: include channel_id in _feedback_key so the same
message_ts in two channels can't collide on the same feedback id.
- slack_feedback.py: treat conflicting +/- reactions from one user as
ambiguous (clear feedback) instead of averaging to a misleading 0.5.
- slack_feedback.py + slack.py + webapp.py: gate reaction handling to
the user who triggered the run (stored in the slack_run_map mapping
alongside run_id). Prevents bystanders in shared channels from
polluting eval feedback.
Adds a webhook-level check so only members of $PUBLIC_REPO_ORG_GATE
(e.g. langchain-ai) can trigger Open SWE via mentions or review
requests on public repositories. Private repos remain governed by the
existing org/repo allowlists. Internal bots bypass the gate.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
* feat: add optional Slack Assistants API typing status indicator
Mirrors OpenClaw's pragmatic approach: instead of rebuilding around
assistant_thread_started events, just opt into assistants.threads.setStatus
to show 'is thinking…' while the agent is working, and clear it when
post_slack_thread_reply lands. Gated behind SLACK_ASSISTANTS_API_ENABLED so
it can be toggled without touching code.
* fix(slack): drop redundant clear, add status heartbeat across model calls
- Slack auto-clears the typing indicator on bot post; remove the explicit
assistants.threads.setStatus("") call from post_slack_thread_reply.
- The indicator expires after ~2 minutes; add a before_model middleware
that refreshes it on every model tick so it stays visible across long
agent runs. Reuses the existing slack_thread.{channel_id,thread_ts}
configurable already plumbed for notify_step_limit.
- chat:write is sufficient on the bot token (assistant:write is on the
way out per Slack docs); no scope or app-config change required.
* feat(slack): contextual status text + rotating loading_messages
- set_slack_assistant_status now accepts an optional loading_messages list
(capped at 10 per Slack's API), surfaced via the assistants.threads.setStatus
payload so Slack rotates through them client-side.
- The heartbeat middleware derives a contextual status from the last
assistant message's tool calls (e.g. "searching the codebase…" after
grep, "running commands…" after execute), falling back to the default
"is thinking…" when no tool calls or unknown tool name.
- Adds a curated DEFAULT_LOADING_MESSAGES list passed alongside the
contextual status on each refresh.
* fix slack assistant status lifecycle
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat(open-swe): trigger reviewer agent from `@open-swe review` PR comment
Mirrors the Slack `@open-swe review` flow on GitHub: a comment containing
`@open-swe review` (optionally followed by a PR URL) on a PR triggers the
reviewer agent. Without a URL it reviews the commenting PR; with a URL it
targets that PR. Works for `issue_comment`, `pull_request_review_comment`,
and `pull_request_review` events, gated by the existing reviewer repo
allowlist and reusing `trigger_pr_review_from_ref`.
* fix(open-swe): require URL after `@open-swe review`, don't swallow trailing text
The previous regex matched any non-whitespace token after `review`, including
across newlines. Comments like `@open-swe review\nthanks!` parsed as
`(True, "thanks!")`, which then failed PR-URL parsing and was silently
dropped — the user got no review and the comment never reached the regular
PR-comment handler.
Restrict the optional URL token to `https?://\S+` so non-URL trailing text
falls through to `process_github_pr_comment` instead of being eaten by the
review-command branch. Adds regression tests for the multiline and
trailing-word cases.
* feat: only post Slack 'Working on it!' on first thread mention
* feat: randomize Slack trace reply phrase
Pick from a small list of friendly phrases instead of always saying
'Working on it!' so the bot feels less robotic. Explicit messages (e.g.
'Taking a look...' from PR review path) are unaffected.
* adjust phrases
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* feat: move github workflows to gh cli
Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.
* docker ignore + snapshot and docker image updates
* updated image and instructions
* removing open_pr if needed after agent call
* feat: add edit_pull_request tool for editing PR titles and descriptions
* fix: patch auth flow in open PR middleware tests
* fix: support app token for editing PRs
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: prevent futile retry loop when commit_and_open_pr fails with git/API errors
- Root cause: when git checkout or GitHub PR API fails, the tool returned a generic {"success": false} error with no signal that retrying is futile, causing the agent to loop 9-13+ times until hitting the 1000-step recursion limit
- Change: (1) git_checkout_branch now returns (bool, str) so the actual git error output is surfaced in the tool response; (2) checkout and PR creation failures now include "fatal": true and an explicit "Do not retry" message; (3) prompt.py COMMIT_PR_SECTION adds an explicit instruction to stop on fatal errors
- Verified: 109 unit tests pass, no regressions
* fix: skip PR safety net on fatal commit failures
* style(open_pr): ruff-format fatal retry skip condition
---------
Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
- Root cause: create_github_pr found an existing PR on 422 but never
PATCHed it, so callers like commit_and_open_pr could not update the
PR body (e.g. adding "Closes AB-1159") on subsequent invocations.
- Change: after _find_existing_pr succeeds, call new _update_github_pr
helper which PATCHes /repos/{owner}/{repo}/pulls/{number} with the
requested title and body before returning pr_existing=True.
- Verified: self-evident API call addition; proof in production traces.
Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
- Root cause: httpx.HTTPError handler returned (None, None, False) without
checking if the PR was already created on GitHub before the network error
- Change: added _find_existing_pr fallback in except block in agent/utils/github.py
- Verified: 5 production traces showed false failures where PR existed (pr_existing=True on retry)
Co-authored-by: LangSmith Forge <forge-agent@langsmith.ai>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
* fix: stop agent retrying commit_and_open_pr on 403 permission denied
Detect 403/permission-denied push failures in commit_and_open_pr and
return a PERMANENT_FAILURE message so the LLM stops retrying. Also add
prompt-level guidance to the COMMIT_PR_SECTION reinforcing this. Add
unit tests covering both the 403 and non-403 push failure paths.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix: stop safety net retrying permanent push failures
Skip the after-agent PR fallback when commit_and_open_pr reports a permanent GitHub push authorization failure, while preserving fallback behavior for recoverable failures.
---------
Co-authored-by: Claude Agent <agent@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
Calling get_github_token() without arguments always invoked LangGraph get_config internally,
which broke tests that only patch agent.middleware.open_pr.get_config and failed outside
runnable context.
Extend get_github_token with an optional runnable config mapping; the middleware passes
the config dict already resolved from get_config(). Request GitHub App installation tokens
only after detecting sandbox/repo changes worth publishing.
Fixes failing Agent unit tests in tests/test_open_pr_middleware.py.
* feat: open PRs under user's name and add OpenSWE label
* feat: use user token for PR authorship, add OpenSWE label, and consolidate fallback logic
* linting
* fix: address review nits for PR authorship and labeling
Fix docstring casing, add debug logging for 422 existing-PR search
fallback, tighten test type annotations, and add missing HTTPError
fallback test.
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: default to GPT-5.5 medium reasoning
Use OpenAI GPT-5.5 with medium reasoning as the default model and document the completion-token budget semantics for reasoning models.
* fix: use Responses API reasoning config
Pass GPT-5.5 reasoning settings through LangChain's Responses API parameter instead of the Chat Completions-only reasoning_effort field.
* feat: raise GPT-5.5 output budget
Set the default GPT-5.5 output token budget to the model maximum so long-running coding tasks have more room for reasoning and final responses.
* feat: align recursion limit with Deep Agents
Use Deep Agents' default recursion limit so longer coding runs have room to complete without Open SWE imposing a lower cap.
* chore: remove minimal effort level
* chore: reduce max tokens to 64_000
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: migrate LangSmith sandbox creation to snapshot API
Replaces the template-based sandbox flow (DEFAULT_SANDBOX_TEMPLATE_NAME /
DEFAULT_SANDBOX_TEMPLATE_IMAGE) with the new snapshot-based flow.
- New required env var DEFAULT_SANDBOX_SNAPSHOT_ID (UUID of a pre-built
LangSmith snapshot; build out-of-band via UI or SandboxClient.create_snapshot)
- Optional DEFAULT_SANDBOX_SNAPSHOT_FS_CAPACITY_BYTES overrides the root FS
size at boot (default 32 GiB)
- Startup-time validation via a FastAPI lifespan hook: the server refuses
to boot with a clear ValueError if SANDBOX_TYPE=langsmith and
DEFAULT_SANDBOX_SNAPSHOT_ID is unset, so failures surface in boot logs
rather than on the first thread
- Reconnect-to-existing-sandbox path unchanged
- Docs (INSTALLATION.md, CUSTOMIZATION.md) updated to describe the new
snapshot workflow
* fix: format create_sandbox_snapshot.py to pass ruff
---------
Co-authored-by: aran-yogesh <yogesh.mahendran@langchain.dev>
The agent sees users formatted as @Name(USER_ID) in conversation context and
reproduces that pattern in replies, but Slack requires <@USER_ID> for real
mentions. This adds automatic conversion and updates prompt instructions.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Adds 6 new agent tools backed by Linear's GraphQL API, with a shared
_graphql_request helper to reduce boilerplate. Refactors existing
comment_on_linear_issue to use the same helper.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: extract repo parsing into shared util and add linear comment repo override
Moves the repo extraction regex logic (repo:, repo , GitHub URL) into
agent/utils/repo.py so it can be reused. Updates the Linear webhook
handler to check the comment body for a custom repo first, falling back
to the team/project mapping when none is specified.
* chore: document repo extraction util and default org configuration
* feat: add generic DEFAULT_REPO_OWNER/DEFAULT_REPO_NAME env vars replacing Slack-only defaults
* chore: remove deprecated SLACK_REPO_OWNER/NAME env vars from docs
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: send LangSmith trace URL on run trigger from Slack and Linear
* fix: check LANGSMITH_PROJECT before LANGSMITH_PROJECT_PROD for project name lookup
* fix: pass LangSmith API key explicitly to Client and remove global variable
* Update agent/utils/langsmith.py
Co-authored-by: Brace Sproul <braceasproul@gmail.com>
* refactor: move trace notification helpers out of webapp and use env vars for LangSmith URL base
* docs: add LangSmith tenant and project ID env vars to installation guide
* fix: remove unused comment_on_linear_issue import from webapp
* nit: remove lru_cache, make get_langsmith_trace_url sync, and move langsmith import to top level
* Update INSTALLATION.md
Co-authored-by: Brace Sproul <braceasproul@gmail.com>
* Update INSTALLATION.md
Co-authored-by: Brace Sproul <braceasproul@gmail.com>
* Update INSTALLATION.md
Co-authored-by: Brace Sproul <braceasproul@gmail.com>
---------
Co-authored-by: Brace Sproul <braceasproul@gmail.com>