mirror of
https://github.com/Sea-Haven-Industries/open-swe.git
synced 2026-09-30 10:23:14 +00:00
130 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
db05d5826f
|
feat(open-swe): port upstream clean batch (#1775, #1785, #1789) (#215)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* Fix: Fix Improper privilege management in server.py (#1789) Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com> (cherry picked from commit 3ea29d3f231dd66bd7627769b5659564be4525df) * Fix: Fix Stored XSS in ReplyCard.tsx (#1785) Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com> (cherry picked from commit 31263f832a2ecedf669eee2e27b827a358a6a4d7) * feat: add structured Linear issue filters (#1775) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit b5e529252c3012818289eca1b1cdeb8014721310) * chore(triage): mark #1775 #1785 #1789 landed Move the three clean cherry-picks in this batch from deferred to landed in the upstream-sync ledger and re-render triage.md. --------- Co-authored-by: corridor-security[bot] <203152403+corridor-security[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
0f0f616cd4
|
feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): explicit-request verdicts + shell verdict guard Mention-triggered reviews that explicitly ask for a verdict now submit a real APPROVE/REQUEST_CHANGES through publish_review; auto-reviews stay advisory (COMMENT). Authorization is enforced in code: publish_review honors a verdict only when the dispatching webhook set verdict_requested, which only the explicit-mention path does. - request_pr_review gains instructions (forwarded verbatim into an escaped requester_instructions data block) and request_verdict - self-review guard downgrades verdicts on Open SWE-authored PRs; stale APPROVEs are best-effort dismissed when later findings land - new PullRequestVerdictGuardMiddleware blocks gh pr review --approve/-a/--request-changes/-r, gh api, and curl verdict fallbacks on both the coding-agent and reviewer graphs - shared escape helper moved to agent/utils/prompt_data.py * fix(reviewer): harden verdict path against security-review findings Adversarial security review (detector fan-out + proof-or-kill verifier) of the verdict feature surfaced several verdict-integrity gaps; resolve the confirmed ones: - head-drift (high): a mid-run push moves the resolved head, so an APPROVE could anchor to an unreviewed commit. Downgrade any verdict to a comment when the resolved head differs from the reviewed head (verdict_ignored reason head_moved); the push's own re-review submits a fresh verdict. - self-review fail-open: downgrade to comment when the PR author cannot be confirmed (author_unknown), and compare bot logins case-insensitively. - verdict_submitted now reflects GitHub's returned review state, not just the event we asked for, so a coerced APPROVE isn't reported as submitted. - an authorized verdict whose findings all anchor outside the diff now posts as a bodied review with zero inline comments instead of failing. - add finding_reply to the shared data-block escape tag superset. |
||
|
8d8d5bbbbf
|
fix: bind cached GitHub tokens to users (#1736)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit 1ea0e600dcc234fa5a333c6f4b80b90e2e6679d3) Co-authored-by: Johannes du Plessis <johannes@langchain.dev> |
|||
|
bfb36be948
|
feat: add Linear issue search tool (#1748)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit 79df6b2ff283afdbd36660be888c130ea964fe3f) Co-authored-by: Johannes du Plessis <johannes@langchain.dev> |
|||
|
62d9945df4
|
refactor: consolidate reviewer modules into agent/review/
Part of the domain-reorg adoption (build plan step C2): fork content,
upstream layout. Nine 1:1 module moves (reviewer_diff/eval_store/
findings/groups/publish/reconcile/trace_context + review_style_
collector/guidance) into agent/review/, with internal relative
imports re-wired to the new package depth. agent/review/__init__.py
mirrors upstream's thin re-export shim (one of the 21 verified "A"
structural adds).
Rewrote the 38 grep hits across importer files (agent/{analyzer,
ci_autofix,reviewer,webapp}.py, agent/dashboard/*, agent/middleware/
settle_review_check.py, agent/tools/*, agent/utils/github_feedback.py,
agent/webhooks/github.py, evals/reviewer/*, and the reviewer test
suite) to point at agent.review.*; 4 of the 38 hits were name
collisions (list_reviewer_findings, reviewer_outcomes,
_reviewer_thread_id, reviewer_thread_id — not the moved modules) and
were left untouched. tests/test_github_checks.py's module-alias
import (`from agent import reviewer_publish`) follows upstream's own
`from agent.review import publish as reviewer_publish` pattern so
downstream `reviewer_publish.*` call sites needed no changes.
agent/reviewer.py and agent/webapp.py stay in place per the hard
rule (fork content, import-only rewire) and are not part of this
package.
Gates: ruff check + ruff format --check, pytest --co -q (1637
collected), full unit suite (1637 passed), and the reviewer/findings
suite in isolation (pytest -k "review or finding", 421 passed).
|
|||
|
|
cfb5624663
|
feat(open-swe): add E2B sandbox provider (port of upstream 48217b68, #1489) (#199)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
Additive sandbox provider adapted to this fork's synchronous create_sandbox factory: agent/integrations/e2b.py registered as a lazy-import entry in SANDBOX_FACTORIES, so the e2b/langchain-e2b imports only load when SANDBOX_TYPE=e2b (dark-safe on unset env; unset E2B_API_KEY raises a clean ValueError only when the provider is selected). Supports reconnect-by-id and optional E2B_TEMPLATE. Non-langsmith providers skip the GitHub proxy step, so the GitHub-App token flow is untouched. Upstream: langchain-ai/open-swe 48217b68 (#1489), re-implemented against the fork's sync sandbox lifecycle rather than cherry-picked (upstream ships on the deferred async sandbox.py base). Docs: provider tables/lists in README, CUSTOMIZATION, INSTALLATION; the CUSTOMIZATION registration example now shows the lazy-tuple form the code actually uses. uv.lock refreshed (adds e2b, langchain-e2b, dockerfile-parse). |
||
|
|
1f80500cf3
|
feat: Jira + Confluence integration (tools + triggers) (#182)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat(open-swe): add Jira tool plane (Phase 1)
Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear
tools:
- utils/jira.py: service-account REST client (Basic auth) with get/
create/update issue, comments, list projects, trace comment; issue and
comment bodies normalized to markdown.
- utils/adf.py: minimal ADF <-> markdown conversion (read paths convert
Jira ADF to markdown; agent comments convert prose to ADF).
- tools/jira_{comment,get_issue,get_issue_comments,create_issue,
update_issue,list_projects}.py wired into the tool registry and the
main agent tool list.
- tests/test_jira_utils.py: ADF conversion + mocked-transport util tests.
Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env
returns a clean error, so this is safe to land dark. Trigger plane,
prompt guidance, and config plumbing follow in Phase 2.
* feat(open-swe): add Confluence tool plane (Phase 3)
Curated Confluence Cloud REST toolset for the agent, mirroring the Jira
tools:
- utils/confluence.py: service-account REST client (Basic auth) with
get/create/update page, add comment, CQL search. Page bodies are XHTML
storage format (not ADF), with minimal storage<->text converters;
update_page reads the current version and bumps it, as Confluence
requires.
- tools/confluence_{get_page,create_page,update_page,comment,search}.py
registered in the tool registry.
- tests/test_confluence_utils.py: converter + mocked-transport tests
including the version-bump path.
Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN;
unset env returns a clean error. Activation in the agent tool list lands
with the Phase 2 server.py wiring.
* feat(open-swe): add Jira trigger plane (Phase 2)
Make an @openswe comment on a Jira issue spawn an agent run, mirroring
the Linear trigger plane:
- webhooks/jira.py: process_jira_issue clones process_linear_issue —
deterministic thread id, full-issue fetch, actor accountId->email
attribution feeding resolve_login_from_email_async (PRs open as the
human), multimodal image handling, source="jira" + jira_issue config.
- webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time
X-Automation-Webhook-Token check, fails closed), repo-resolution
cascade, get_repo_config_from_jira_mapping.
- utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder
entry — real project->repo mappings still needed).
- utils/jira.py: get_user_email (accountId -> email) for attribution.
- completion.py: source=="jira" failure-reply branch.
- prompt.py: Jira-triggered notify guidance + Refs:/branch key from
{jira_project_key}-{jira_issue_number}.
- server.py: read jira_issue config + pass jira key to the system
prompt; also activates the Phase 3 Confluence tools in the agent list.
Jira Automation lacks native webhook HMAC signing, so trust is a shared
secret header (decision D2); replay protection is weaker than Linear's
HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the
outstanding gate/hardening before push.
* fix(open-swe): harden Jira webhook trust (sh-security-review)
Resolves findings from the Phase 2 security review (detector fan-out +
proof-or-kill verifier). The unsigned Jira Automation webhook body was
trusted for identity, comment content, repo routing, and issue
existence; a JIRA_WEBHOOK_SECRET holder could forge those fields.
- Corroborate against the real Jira record: the webhook body is now only
a pointer (issue_key + required comment_id). The triggering comment's
author and text are re-fetched server-side via get_comment/fetch_jira_
comment, and identity, the @openswe check, prompt text, and project
key are derived from that authoritative record — never payload author/
body fields. An uncorroborated comment is rejected. (closes the
account-id impersonation, unsigned-body prompt injection, and
fabricated-issue findings)
- Validate issue_key against the Jira key format and percent-encode all
untrusted path segments (_seg) so a crafted key can't traverse to a
different Jira REST endpoint or inject query params. (closes the path-
traversal / query-injection findings)
- Route source=="jira" through the bot-token-default / author_prs_as_
user opt-in path in resolve_github_token, matching Linear, instead of
unconditionally resolving a per-user OAuth token from a payload email.
- Gate attribution on an active user mapping (is_login_mapped) so a
pending/unconfirmed mapping can't drive PR authorship.
Adds regression tests: server-corroboration wins over payload, malformed
issue_key rejected, uncorroborated comment rejected, path-segment
encoding, project-key derivation, active-mapping gate.
Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/
REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body +
timestamp on the Automation payload to close the residual replay gap.
* harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist
Folds the two deployment-hardening items from the Phase 2 security review
into code (all opt-in / default-off, so existing and upstream deployments
are unaffected):
- JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must
carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by
verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear
HMAC+freshness model). Closes the static-token model's replay/forgery
gap when enabled.
- JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's
direct client IP (verify_jira_source_ip). Documented as direct-peer
only; behind a proxy/LB, allowlist Atlassian's ranges at that layer.
- REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail
CLOSED instead of the back-compat allow-all, plus a startup fail-open
warning. Applies to all channels for consistency.
Documents all new vars (and a Jira section) in .env.example. Adds tests
for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and
the fail-closed allowlist.
* feat(open-swe): Confluence Atlassian Connect trigger (Phase 4)
Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian
Connect app. Designed and adversarially verified with the ultracode
workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus
skeptics on the implemented crypto).
- utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's
official test vector), PyJWT HS256 webhook verifier with alg-pinning,
issuer binding, and qsh-verified-last ordering; RS256 signed-install
lifecycle verifier against Atlassian's published keys; installation
store keyed by clientKey with the sharedSecret encrypted at rest
(TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned).
- webhooks/confluence.py: install/uninstall lifecycle + comment handler.
The JWT-signed webhook body is only a pointer; the comment's real
author/text/container are re-fetched server-side via the Basic-auth
service account (Phase-2 corroboration lesson), with active-only login
attribution and the repo allowlist.
- utils/confluence.py: get_comment / get_user_email (path-encoded).
- webapp.py: GET /connect/atlassian-connect.json (served dynamically),
POST /connect/{installed,uninstalled,webhook/comment-created}, the
space->repo resolver, thread-id, and fetch helpers.
- completion.py: source=="confluence" failure-reply branch.
Security: the sh-security-review verify pass confirmed one HIGH — the
symmetric signed-install=false first-install was trust-on-first-use gated
only by the public Confluence hostname (webhook-auth bypass). Fixed by
switching to signed-install=true + RS256 verification of lifecycle
callbacks, which cryptographically authenticates the first install. All
other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall
DoS/corroboration/injection) were defeated; residuals are deployment
config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover
bodies; comment-trigger prompt injection, shared with all sources).
New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN,
CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets
require the durable Postgres LangGraph store in prod.
Outstanding before push: /sh-security-review on the real diff and the
GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory.
* docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs
- prompt.py: Confluence-triggered runs notify via confluence_comment on
the triggering page; add Confluence to the shared-base source list.
- CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian
triggers (Jira Automation shared-secret webhook; Confluence Connect app
with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the
server-side corroboration + encrypted install store.
Phase 5 also verified the trigger surface end-to-end against a running
uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed
without valid auth) and recorded the integration in project memory.
* fix(open-swe): resolve /sh-security-review findings on the Atlassian surface
Formal sh-security-review (detector fan-out + verifier) over the Phase-4
Connect surface (esp. the new RS256 signed-install code, unseen by the
earlier adversarial verify) and the Phase-2 opt-in hardening.
CRITICAL — cross-tenant install (origin validation, CWE-346): signed-
install proves the caller is *an* Atlassian tenant, not *ours*, and the
descriptor is served publicly, so any attacker could install the app on
their own Confluence site and drive agent runs against our allowlisted
repos. The baseUrl body field is attacker-controlled and cannot bind the
tenant; only the signature-verified clientKey (JWT iss) can. Added a
MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in
process_install after signature+iss verification.
HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment
ids are per-instance, so generate_thread_id_from_confluence_comment now
salts the hash with the verified clientKey (plumbed from the webhook JWT
iss) to prevent thread hijack across tenants.
HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page
interpolated page_id into the REST path unencoded (update_page on a
mutating PUT with no params= backstop). Now _seg()-encoded, matching the
rest of the module.
MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no
bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID
guard mirroring the Linear botActor / Jira comment_author_is_bot checks.
LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as
opt-in defense-in-depth (the clientKey allowlist is the real gate).
Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp,
kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption,
constant-time comparisons, and the Phase-2 hardening. New regression
tests for each fix; full suite green (1602).
* harden(open-swe): GPT-4.1 cross-family review follow-ups
Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no
critical/high issues and confirmed the auth boundary is fail-closed and
correct. Two low-cost defense-in-depth items applied:
- Validate the signed-install JWT 'kid' against a strict charset before
the public-key fetch, so a malformed kid fails fast with no network
call (on top of the existing fixed host + percent-encoding).
- Make JWT nbf verification explicit (verify_nbf) on both the RS256
lifecycle and HS256 webhook decodes.
Other suggestions triaged as already-handled (aud cross-app replay is
blocked by the per-tenant iss->secret lookup; documented static-token/IP/
baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra
(Fernet rotation via MultiFernet; rate limiting at the gateway).
* docs(open-swe): document Jira + Confluence in installation & customization guides
- INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret,
service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect
app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note,
CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars +
REQUIRE_REPO_ALLOWLIST.
- CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction
note covers all four sources.
- AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth).
- README.md: invocation section, tools table, and overview line.
|
||
|
|
2fb1122630
|
feat(open-swe): open Linear-triggered PRs as the triggering user (#179)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: open Linear-triggered PRs as the triggering user Add 'linear' to the set of sources that carry a mapped GitHub login (Slack, Linear, dashboard, schedule), so Linear-triggered runs resolve the per-user OAuth token when the author_prs_as_user profile flag is enabled. Previously only Slack and dashboard runs could open PRs as the user; Linear runs always used the bot token. The Linear webhook now resolves the GitHub login from the Linear email via the same user-mapping store Slack uses, and passes it through the run configurable and thread owner metadata. Refs: 5003c953 (upstream #1683) * fix(open-swe): restrict Linear token attribution to comment author only Split actor_email (comment_author only, feeds github_login for token attribution) from user_email (full fallback chain, for display/model). This ensures a PR is never opened as a non-actor (creator/assignee). Add test_resolve_github_token_linear_defaults_to_bot to lock the security-critical default: Linear + mapped login + no opt-in → bot. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com> |
||
|
|
e9da186b5f
|
fix(open-swe): port core GitHub-App scope fallback (#1701), workflows:write kept out of standing scope (#181)
* fix: fall back to core GitHub App scope when optional grants missing (#1701) * fix: fall back to core GitHub App scope when optional grants missing Proxy-token minting requested workflows:write and actions:read in the permission set used for every sandbox. GitHub 422s a token request that asks for a permission the installation hasn't granted, so any install without workflows:write failed to mint a token and every run died in before-agent setup with "GitHub App installation token is unavailable". _resolve_proxy_token now walks a permission ladder (full -> +workflows -> core) and returns the first scope that mints, recording the granted scope so hourly proxy refreshes stay consistent. A missing optional grant now degrades to the install-time core scope instead of failing the run; workflow-file HITL pushes still require workflows:write and fail at push time when it is absent. * refactor: flatten proxy-token ladder loop with continue --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> (cherry picked from commit f53caff1aa24a7b29d851b267aa3bdfe62c1e935) Sea Haven fork deviation: upstream #1701 folds workflows:write into the standing BASE/RUNTIME scope. This fork deliberately keeps workflows:write OUT of the standing permission ladder (RUNTIME = core + actions:read; LADDER = (RUNTIME, CORE)) so the sandbox proxy token cannot push .github/workflows/* during normal operation. workflows:write is minted only transiently by WorkflowPushGuardMiddleware for an approved HITL push and dropped on restore, preserving token scope as a backstop for the workflow- push approval control. Security-reviewed (agentic fan-out + GPT-4.1 cross review); the standing-scope-carries-workflows:write bypass was blocked. * fix(open-swe): harden proxy-token restore and mint error handling Two low-severity follow-ups from the security review of the #1701 port. Restore the recorded baseline scope after a workflow-push elevation instead of a hardcoded RUNTIME. An install granted workflows:write but not actions:read resolves its standing token to core; hardcoding RUNTIME on restore requested the ungranted actions:read, 422'd, and fired a false "SECURITY: failed to downscope" error on every approved workflow push before the core fallback recovered. The guard now captures the run's recorded scope before elevating (via the new get_recorded_proxy_permissions) and restores exactly that, falling back to the guaranteed core scope only when the baseline restore fails. Classify installation-token mint failures. get_github_app_installation_token_ with_expiry now treats HTTP 422 (a permission the installation hasn't granted) as the ladder's expected descend signal and keeps it at debug, while a non-422 failure (network/5xx/timeout) is surfaced at WARNING even when errors are otherwise suppressed — so a transient blip no longer silently downscopes a whole run under a debug-only trace. The reduced-scope warning no longer asserts a missing grant as the sole cause. * chore(triage): mark upstream #1701 landed on this branch Ported via PR #181 as Option A (workflows:write kept out of the standing proxy-token scope). Regenerated triage.md from triage.jsonl. --------- Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> |
||
|
|
0546085672
|
feat: port durable dispatch hardening and startup latency improvements (#160)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port plan-review & workflow-approval UX (#135)
Port six upstream commits onto dev:
-
|
||
|
|
f87847baa4
|
feat: port plan-review & workflow-approval UX (#159)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port plan-review & workflow-approval UX (#135)
Port six upstream commits onto dev:
-
|
||
|
|
c2bc7720cd
|
feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) (#155)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat: port LangSmith LLM Gateway routing from upstream (#1671, #1673, #1674, #1678) Ports four upstream commits that add opt-in LLM call routing through the LangSmith Gateway, preserving fork conventions (Bedrock/Fireworks model IDs, no-agent-attribution, bun toolchain). - #1671 (e9dc6e01): opt-in gateway routing — new gateway.py, team-settings toggle, admin UI section, wired into make_model for all graph entrypoints - #1673 (702ef908): dedicated LANGSMITH_GATEWAY_API_KEY precedence over platform LANGSMITH_API_KEY - #1674 (5f7c2f46): fix Fireworks gateway base URL to /fireworks (bare host, SDK appends /v1/chat/completions) + SanitizeFireworksMessagesMiddleware - #1678 (73b7d1c0): fix OpenAI Responses reasoning replay — SanitizeOpenAIResponsesMiddleware, store/include config for encrypted reasoning content, reasoning_effort coercion for Chat Completions fallback Refs #134 * fix: downgrade gateway not-routed log to debug, add Bedrock UI note, add sanitizer parity - Downgrade logger.warning to logger.debug in gateway_overrides for not-routed providers and missing API key (Bedrock is the default provider in this fork, so these are expected steady states) - Add Bedrock to the LLMGatewaySection route-toggle description so admins know it is not routed through the gateway - Add SanitizeOpenAIResponsesMiddleware to chat.py for parity with server.py and reviewer.py - Restore the Bedrock region comment in model.py that explains the AWS_REGION / AWS_DEFAULT_REGION precedence Refs #138 --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
b3b0274403
|
feat: Re-land deferred upstream features on modular webhooks (#80) (#128)
* fix(webhooks): fall back to vision model for Slack/Linear image threads Re-land upstream #1626 onto the modular webhook structure. When a Slack mention or Linear issue carries images but the resolved model is text-only, fall back to a vision-capable model instead of dropping the images. Re-points default_vision_model_pair at the fork's image-capable models (Opus 4.8 default, else any supports_images model) rather than upstream's openai:/anthropic: provider filter. Refs #80, upstream #1626 * fix(slack): persist trace_message_ts so web-handoff updates the trace reply Re-land upstream #1630 onto the modular structure. The first-mention store_slack_run_mapping call did not pass trace_message_ts, so it was never persisted (nothing to preserve from on first mention) and _notify_slack_web_handoff always skipped the trace-reply update on web handoff. Pass it through and cover it with a test. Refs #80, upstream #1630 * feat(slack): include channel context in Slack prompts Re-land upstream #1633 onto the modular structure. Fetch cached Slack channel metadata once per event (_get_slack_channel_context) and thread it through the docs-plz gate, repo resolution, and process_slack_mention so prompts carry the channel name and a clearly-marked untrusted channel description. Avoids duplicate conversations.info calls. Refs #80, upstream #1633 * feat(tools): add slack_start_new_thread breakout tool Re-land upstream #1638 onto the modular structure. Adds the slack_start_new_thread tool (posts a top-level Slack message and dispatches a fresh agent run for a broken-out task via the durable dispatch_agent_run contract), wires it into the agent tool list and tools/__init__, adds prompt guidance, and excludes it from plan mode so it can't bypass the approval flow. Tool imports only live modules. Refs #80, upstream #1638 * feat(plan): notify Slack on plan approval Re-land upstream #1632 onto the modular structure. When a plan is approved via the dashboard approve endpoint, post a thread reply to the originating Slack thread noting the comment count and approver, after the follow-up run is dispatched. Slack post failures never break approval. Adapted to the fork's approve_plan (no plan_markdown read). Refs #80, upstream #1632 * feat(plan): publish plans from sandbox files Re-land upstream #1635 onto the modular structure, completing the partially-ported change so dev is internally consistent. save_plan now takes a plan_file_path, reads the agent-authored Markdown file from /workspace/plans/ (validating extension/location/UTF-8/size) and publishes it, instead of taking a plan_markdown string. Removes write_file/edit_file from PLAN_MODE_EXCLUDED_TOOLS so the agent can author the plan file, updates enter_plan_mode/reject_plan guidance and the e2e fake LLM. Skips the #1610-only update_plan hunk (not on dev). Refs #80, upstream #1635 * fix(security): SSRF-harden server-side image fetch + stop logging raw image URLs INJ-01 (high): fetch_image_block used follow_redirects=True with no per-hop revalidation and discarded the resolved-IP pin, so an attacker-authored Slack/ Linear image URL could 302-redirect the fetch to an internal host / cloud metadata endpoint (blind SSRF), and DNS-rebinding could bypass the one-shot is_url_safe check. Route image fetches through the same per-hop resolve+pin+ revalidate loop the http_request tool uses, lifted into url_safety as the shared request_with_safe_redirects. Also strip the per-host Slack/Linear bearer token on redirect so it can't be replayed to a redirect target. SC-1 (low): linear.py logged full image URLs (which can carry signed tokens) at DEBUG; multimodal logged them at INFO on every fetch. Log host-only. Sink lived in multimodal.py (unchanged by the feature work) but PR #128 widened its reach by no longer dropping images for text-only models. Fixing on the base branch so #130/#129 inherit it on rebase. Adds fetch_image_block SSRF regression tests (redirect-to-internal blocked; auth stripped on redirect). |
||
|
ea1845d41d
|
fix(security): fence and neutralize untrusted Slack channel description in agent prompt
Hardens SLACK-PI-001 (sh-security-review). Slack channel topic/purpose is editable by ordinary channel members and flowed verbatim into the agent LLM prompt behind only a prose 'untrusted' label — an indirect prompt-injection vector for an agent with network egress and repo write. Now strip leading markdown structural tokens per line (so it can't forge the prompt's real request/section delimiters), cap length, and wrap it in a per-render unguessable sentinel fence (so injected text can't spoof a closing marker to escape the data block). Deliberately diverges from upstream #1633. |
|||
|
|
2b01652754
|
refactor: adopt modular webhook architecture (#1621) + port fork customizations (#85)
* Adopt upstream modular webhook skeleton (#1621) Apply the durable-interrupt-dispatch refactor: split the monolithic webapp.py into a thin routing layer plus per-source handlers in webhooks/{github,slack,linear}.py, and add completion.py, dispatch.py, and reconcile.py. Reconcile fork divergence by keeping the Bedrock/ Fireworks cross-provider fallback, the no-agent-attribution prompt policy, the dashboard-handoff re-export, and the Slack channel-info cache. ci_autofix is restored on the new dispatch model in a later commit. Refs: #80 * Port fork webhook security delta onto modular handlers Re-apply the fork's security customizations that #1621 did not carry: Linear webhook replay protection (freshness window on the signed webhookTimestamp), per-repo token-cache binding threaded through the thread token resolvers, the INTERNAL_BOT_LOGINS self-check in the review-finding-reply path, and a user-mapping cache refresh before email resolution on the issue and PR-comment paths (multi-replica staleness). Existing fork security tests pass unchanged. Refs: #80 * Restore CI auto-fix on the modular dispatch model Bring back ci_autofix.py and the ci_monitor graph that #1621 deleted, re-wiring the fork's security-reviewed PR-babysitting onto the new structure: the CI-event, autofix-toggle, and review-feedback handlers move into webhooks/github.py and the github_webhook router re-gains the check_run/check_suite/workflow_run/status routing plus the autofix command and actionable-review branches. Auto-fix runs now dispatch through dispatch_agent_run (durability + completion webhook) while keeping the deliberate batch-while-busy skip-rule via get_thread_active_status. Restore langgraph.json's ci_monitor entry and the fork autofix tests (dispatch mock + import paths re-pointed). Refs: #80 * Reformat and update docs for the modular webhook split Point CLAUDE.md and deploy/MIGRATION.md at the new webhooks/ modules and the dispatch/completion/reconcile contract, and mark the user-mapping cache-refresh fix as applied on the GitHub handlers. Refs: #80 * Restore reject backstop for autofix dispatch A burst of near-simultaneous CI events for one head SHA can slip past the busy-check before the dedupe SHA is recorded, so dispatch the autofix path with multitask_strategy=reject (dev's prior platform default) to drop duplicate concurrent creates instead of letting them interrupt each other. Also make the completion failure-reply dedup claim-then-post and drop the unreachable interrupted branch. --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> |
||
|
|
1f060f2a1d
|
chore: sync upstream/main, defer #1621 modular webhooks (#81)
* chore: bake sfw binary into sandbox image (#1611) sfw only ships a launcher that fetches its real binary at first run and does a daily update check against api.github.com/repos/SocketDev/sfw-free. Both fail in the sandbox (restricted egress; the proxy injects the GitHub App installation token, which lacks access to that repo), so `sfw yarn install` errors with "could not fetch its binary". Pin sfw 2.0.6, warm + verify the binary cache at build, and set SFW_SKIP_UPDATE_CHECK=1 so runs use the baked binary offline. * feat: editable plan mode + fix review-plan banner overlap (#1610) * feat: editable plan mode + fix review-plan banner overlap Lets the thread owner edit the plan markdown by hand from the plan-review page (Edit -> textarea -> Save) via a new PUT /dashboard/api/plan/{id} endpoint that re-publishes the plan and mirrors it into the sandbox plan.md, so approve hands the edited plan to the agent as the source of truth. Also fixes the collapsed git-panel's floating expand button covering the "Review plan ->" banner by reserving space for it. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: abort plan approval when the published plan read fails get_plan_content() swallowed store errors and returned None, so a transient failure during approve would still mark the plan approved and dispatch the generic fallback text — silently dropping an owner's edited plan. Read the plan strictly (raise_on_error=True) so approval aborts instead, matching the comment read. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show message timestamps (#1609) * feat: show message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: suppress fallback message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: stable message + tool-call hover timestamps Stamp a stable client-side arrival time per message and tool call (keyed by id, persisted to localStorage). Messages render the timestamp inline; tool rows reveal a dim timestamp chip on hover. Real backend created_at still takes precedence when present. * fix: hide client-stamped message timestamps Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add PR trace resolution (#1612) * feat: add PR trace resolution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: inject reviewer trace context as JSON Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review on PR trace resolution Use the documented LangSmith metadata filter syntax (and(eq(metadata_key,...), eq(metadata_value,...))) instead of has(metadata, '{...}'), which does not match runs — _list_thread_runs was silently returning nothing. Bound full-text searches to a 90-day window so they don't hit LangSmith's large-window rate limit. Also folds in the best-effort branch->head-sha resolver (dropping the weighted scoring/threshold + repo/file evidence + GitHub hydration), sandbox JSON injection, and the admin "Resolve trace" dry-run endpoint. The IDOR findings are moot: resolve_pr_to_threads/summarize_agent_session were removed; resolution now runs deterministically from the trusted run config with no model-controlled pr_url or thread_id. * fix: scope branch trace search to the repo Branch names like fix-tests aren't unique across repos (or older PRs) in a shared tracing project, so an unscoped branch hit could resolve to an unrelated thread and write its runs into the reviewer sandbox. Require the repo slug to co-occur with the branch in matched runs; the full head SHA stays unscoped since it is globally unique. Addresses open-swe review on PR #1612. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: include plan links in PR descriptions (#1613) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: gate workflow pushes with approval (#1614) * feat: gate workflow pushes with approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve proxy refresh test compatibility Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: bind workflow approvals to pushed ref Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: recover thread work as patch (#1615) * feat: recover thread work as patch Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: search sandbox cwd for recovery patches Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: omit plan link in PR description when no plan exists (#1618) Plan links in PR descriptions were always built from the thread id, so runs that never produced a plan linked to an empty plan-review page. Now the plan content store is consulted first; the link is only added when a plan with non-empty markdown actually exists. A transient store failure degrades gracefully (no link) rather than blocking PR creation. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add filter & grouping menu to agents threads sidebar (#1617) Add a Cursor-style control to the agents sidebar that groups (None/Date/ Status/Project), filters (ownership, status, source, pull request, model, repo, include-resolved), and compacts the threads list. All client-side over already-fetched sidebar threads; preferences persist in localStorage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: update langsmith sdk to 0.9.3 (#1616) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: clickable shared PR header in git panel and reviews (#1620) * feat: clickable shared PR header in git panel and reviews Replace the standalone "View PR" button in the agent git panel with a clickable PR title, matching the reviews view. Extract a shared PrHeader component reused by both the git panel and the review main body. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: drop PrHeader wrapper, use shared component directly The review-side PrHeader was just a thin adapter mapping detail -> the shared component's props. Inline it at the call site and use the shared PrHeader directly so there's a single component. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: durable interrupt dispatch + completion webhook (#1621) * wip(rebuild): core reliability spine - remove PR-babysitting (ci_autofix + ci_monitor graph + webhook wiring) - dispatch core: agent/dispatch.py with multitask_strategy=interrupt + durability=sync + completion webhook; reroute all webhook + plan triggers; drop the racy in-process lock + is_thread_active busy-check - completion webhook: agent/completion.py + /webhooks/run-complete loopback route for failure/timeout replies (idempotent) Co-authored-by: open-swe[bot] * feat(rebuild): async tools, reconcile, shared http timeouts, assembly tuning Parallel batch on top of the reliability spine: - async-ify all 24 tools (drop asyncio.run; requests->httpx); re-implement the http_request/fetch_url SSRF + DNS-rebinding defense httpx-natively and harden the IP check to 'not is_global' (+ IPv4-mapped unwrap) - reconcile.py: stale pending-run sweep (threads.search -> per-thread runs.list -> cancel_many), wired into the scheduler graph via task='reconcile' - shared DEFAULT_HTTP_TIMEOUT (agent/utils/http.py) on every bare httpx.AsyncClient() across utils/dashboard/webapp/middleware - run budget: MODEL_CALL_RECURSION_LIMIT 5000->250 - fix stale OpenAI->Anthropic fallback id (claude-opus-4-5 -> 4-8) - drop redundant custom repair middleware (deepagents auto-adds PatchToolCalls) - confirm tool-result eviction + summarization auto-wired via backend - slim system prompt ~8% (full harness-profile rewrite deferred) Co-authored-by: open-swe[bot] * feat(rebuild): harness-profile prompt + split webhooks out of webapp - prompt.py: own the system prompt via a registered harness profile (OPEN_SWE_SHARED_BASE, kept neutral so the read-only reviewer/analyzer that share it stay safe), registered across all 4 providers; per-thread values stay in construct_system_prompt. Assembled main-agent prompt ~6.8k -> ~3.1k tokens (~55% smaller); de-duped PR/commit/suite/force-push guidance; dropped ALL-CAPS markers. - webapp.py 3325 -> 1890 LOC: moved 14 per-source handlers into agent/webhooks/{linear,slack,github}.py; webapp re-exports them for the routes + tests; moved handlers reach shared helpers via the webapp namespace to preserve the test suite's monkeypatch targets. Full suite: 1168 passing, lint clean. Co-authored-by: open-swe[bot] * Restore MODEL_CALL_RECURSION_LIMIT to 5000 for long-running tasks Reverts the 250 cap from the run-budget change — long-running tasks legitimately need many model calls. The notify_step_limit_reached safety net still fires if a run does hit the cap, so runs end with a signal either way. Co-authored-by: open-swe[bot] * fix: address PR review (auth, SSRF, interrupted status, redirect headers) - completion.py: drop `interrupted` from failure statuses — with multitask_strategy=interrupt a follow-up ends the prior run as interrupted, which is healthy, not a failure to report. [open-swe] - /webhooks/run-complete: shared-secret auth — dispatch appends ?token= when RUN_COMPLETE_WEBHOOK_SECRET is set; route verifies via hmac.compare_digest. [corridor-security] - SSRF: extract the URL validator to agent/utils/url_safety.py and apply it before server-side image fetches in multimodal.fetch_image_block. [corridor-security] - http_request: preserve caller headers/extensions across redirect hops instead of dropping them on the first hop. [open-swe] Co-authored-by: open-swe[bot] * chore: remove REBUILD_PLAN.md (planning doc, not needed in the repo) Co-authored-by: open-swe[bot] * fix: fail closed on run-complete webhook auth when secret unset Corridor follow-up: verify_run_complete_token returns False (not True) when RUN_COMPLETE_WEBHOOK_SECRET is unset, so the public route is never unauthenticated. Logs a startup warning when the secret is absent, and dispatch skips registering the webhook when there's no secret (no rejected callbacks). Co-authored-by: open-swe[bot] --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: restore forced tool call to prevent premature run stops (#1622) Restore the ensure_no_empty_msg middleware and the always-call-a-tool system-prompt instruction that #1535 removed. When the model emits a message with no tool call (and hasn't already messaged the user or confirmed completion), the middleware re-injects a no_op / confirming_completion tool call so the run continues instead of ending mid-task. Shipping to test whether it fixes runs that stop halfway through. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore(deps): bump langgraph-checkpoint from 4.1.0 to 4.1.1 (#1619) Bumps [langgraph-checkpoint](https://github.com/langchain-ai/langgraph) from 4.1.0 to 4.1.1. - [Release notes](https://github.com/langchain-ai/langgraph/releases) - [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1) --- updated-dependencies: - dependency-name: langgraph-checkpoint dependency-version: 4.1.1 dependency-type: indirect ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> * fix: post reviewer resolution notes verbatim (#1624) * fix: post reviewer resolution notes verbatim Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: stabilize dashboard follow-up e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make e2e attribution marker durable Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: only echo found e2e attribution Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: check live dashboard attribution in e2e Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * hotfix: stop prompting agent/reviewer to wrap installs in sfw (#1625) Installs hung when prefixed with sfw inside the sandbox (trace 019f0608 stalled on a pending `sfw npm install` execute, never returned). Strip the Socket Firewall guidance from the agent and reviewer prompts so installs run through the project's package manager directly. sfw stays in the Docker image; nothing invokes it now. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: make plan view mobile friendly (#1636) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: fall back to vision model for image threads (#1626) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface Slack thread errors (#1627) * fix: surface Slack thread errors Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't set failure_reply_posted on Slack preprocessing errors The preprocessing error handler was setting failure_reply_posted=True, the same idempotency flag handle_run_completion checks to suppress duplicate run-failure replies. Since preprocessing failures happen before any run exists but the flag persists on the thread, a subsequent run failure on the same thread would be silently ignored. The preprocessing handler already posts its own Slack reply, so the run-completion idempotency flag should not be set here. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: avoid recapping Slack replies (#1629) * chore: avoid recapping Slack replies Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: simplify Slack reply prompt wording Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update Slack trace reply on web handoff (#1630) * fix: update Slack trace reply on web handoff Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: trigger web handoff on dashboard starts Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: format web handoff as contextual fragment Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve trace_message_ts when overwriting Slack run mapping When store_slack_run_mapping is called without trace_message_ts (e.g. on follow-up Slack mentions), it was unconditionally overwriting the thread-level mapping and clobbering the timestamp captured from the initial trace reply. After that, _notify_slack_web_handoff could not find the original message, so a subsequent move to Web silently skipped the Slack trace update. Now, when trace_message_ts is not passed, the existing thread mapping is read first and its trace_message_ts is preserved. * style: ruff format --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * fix: pre-bundle shiki/@pierre deps to stop dev dynamic-import failures (#1643) * fix(ui): pre-bundle shiki/@pierre deps to stop dev dynamic-import failures shiki lazy-imports a grammar per language and these libs only live inside lazy route components, so Vite's startup scanner never sees them. They get discovered on first thread navigation, triggering a dep re-optimize + force-reload that aborts the in-flight route-chunk import, surfacing as "Failed to fetch dynamically imported module: .../$threadId.tsx". Pre-bundle them (and the github themes + common code-block languages) via optimizeDeps.include so the optimize happens once at startup. Dev-only; production bundles are unaffected. * fix: pre-bundle canonical shiki docker/make langs instead of aliases --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: show queued dashboard follow-ups (#1631) * feat: show queued dashboard follow-ups Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: de-dupe queued follow-ups while streaming --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: notify Slack on plan approval (#1632) * feat: notify Slack on plan approval Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: post Slack approval notice after successful dispatch Move the _maybe_post_plan_approved_to_slack call until after _dispatch_followup succeeds so the Slack thread is not told implementation is beginning before the LangGraph run is created. Addresses PR review comment. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> * feat: include Slack channel context in prompts (#1633) Add cached Slack channel metadata enrichment for Slack-triggered runs so prompts can include channel names and descriptions without duplicate conversations.info calls.\n\nCo-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> * chore: keep plan guidance high-level (#1634) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: publish plans from sandbox files (#1635) * feat: publish plans from sandbox files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: avoid fixed plan filenames Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: virtualize local sandbox file paths Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve plan_file_path across set_plan_status set_plan_status was rewriting the content record with only markdown and status, dropping plan_file_path. After a reject, the owner's dashboard edit would mirror to a different file than the agent's original, and the next save_plan could republish the stale file. Preserve plan_file_path when updating status. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: return to thread after plan approval (#1637) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Slack breakout thread tool (#1638) * feat: add Slack breakout thread tool Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: make fake LLM scripts declarative Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: exclude slack_start_new_thread from plan mode The breakout tool can dispatch a fresh agent run that starts outside the current plan-mode state, bypassing the approval flow. Add it to PLAN_MODE_EXCLUDED_TOOLS so it's hidden alongside the other mutating tools while planning. --------- Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: require bun for ui agent work (#1639) Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: request actions read for sandbox logs (#1642) * fix: request actions read for sandbox logs Request optional Actions read permission for sandbox proxy tokens, with fallback for installations that have not approved it yet. Update setup docs and prompt guidance for safe GitHub Actions log usage. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: restore actions:read scope after workflow push After an approved workflow push, the guard was restoring the proxy with BASE_RUNTIME_PROXY_TOKEN_PERMISSIONS, which excludes the actions: read scope this PR adds. Restore with RUNTIME_PROXY_TOKEN_PERMISSIONS (which includes actions: read) and fall back to BASE if the install hasn't granted Actions read — mirroring the pattern in _create_sandbox_with_proxy. Addresses review comment on PR #1642. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: widen split review diffs (#1647) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: install missing deps before verification (#1646) Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm (#1645) * chore: require pnpm for ui agent work Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: switch ui to pnpm Replace Bun and Yarn lockfiles with pnpm lockfile and update UI/Vercel commands to use pnpm. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * ci: use corepack for ui pnpm e2e build Run pnpm through Corepack in the E2E global setup so CI can use the pinned package manager without a separate pnpm install step. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add Sonnet 5 to model picker (#1651) * chore: update Sonnet examples to Sonnet 5 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * chore: add Sonnet 5 to model picker Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * Remove dead breakout-thread e2e scenario after dropping the tool The merge resolution deferred upstream's Slack breakout-thread tool (slack_start_new_thread, #1638) since it depends on the #1621 dispatch module, but the e2e harness still scripted it. Removing the tool name from fake_llm.py's _tool_step call left a malformed scenario, crashing the langgraph-dev web server at import (TypeError: _tool_step() missing 'call_id') and failing Playwright E2E. Drop the "breakout" script scenario, its _is_breakout_request helper + ScriptRule, and the corresponding full_flow.spec.ts test. * Revert upstream pnpm switch; keep bun for the UI build The merge auto-adopted upstream's pnpm switch (#1645) in tests/e2e/ global-setup.ts and ui/package.json, but our fork builds the UI with bun (vercel.json + the E2E workflow's setup-bun). That left the Playwright globalSetup running `corepack pnpm install --frozen-lockfile` with no pnpm-lock.yaml, failing E2E at UI build time. Revert global-setup.ts and ui/package.json to the dev (bun) baseline, drop the merge-added ui/pnpm-lock.yaml, and remove the re-added ui/AGENTS.md (our fork had deleted it). * Align plan-review e2e + UI with the HEAD (pre-#1635) backend The merge left a split plan vertical: the backend save_plan/plan_api are HEAD (we deferred the editable-plan/sandbox-publish features #1610/#1635/ #1637 per #80), but the plan UI and e2e harness were upstream's. The fake_llm scenario called save_plan(plan_file_path=...) — upstream's file-based #1635 contract — while HEAD save_plan takes plan_markdown, so the plan never saved and PlanReview never rendered (E2E failure on the plan-review locator). Pass plan_markdown to save_plan, and revert PlanReview.tsx / plan.ts / $threadId_.plan.tsx / plan_review.spec.ts to the dev baseline so the whole plan flow (save -> render -> approve -> implement) is consistent with the HEAD backend. --------- Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: Johannes du Plessis <johannes@langchain.dev> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: Ramon Nogueira <ramon.nogueira@langchain.dev> Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com> Co-authored-by: Caroline di Vittorio <43390382+carolinedivittorio@users.noreply.github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: Ankush Gola <9536492+agola11@users.noreply.github.com> Co-authored-by: Mukil Loganathan <mukil@langchain.dev> |
||
|
|
a30ce2ab40
|
feat: managed LangGraph Cloud + Vercel migration (PR2 — code fixes + docs) (#65)
* fix(dashboard): managed-cloud OAuth hardening + admin user-mapping endpoint Prepare the dashboard backend for the managed LangGraph Cloud + Vercel runtime, where the API is HTTPS and cross-site from the UI. - OAuth redirect_uri (#2): coerce a schemeless DASHBOARD_API_BASE_URL to https:// in _api_base_url() so GitHub stops rejecting login with "redirect_uri not associated with this application". _cookie_security() now treats a schemeless (managed) value as Secure; SameSite=None too, consistent with the coerced scheme. - OAuth state cookie (#3): document that osw_oauth_state is host-only by design (a Domain cookie is unsafe across *.vercel.app, a public suffix), so login must always start on the stable alias to avoid "oauth state mismatch". Operational contract; no behavioral change. - Admin user mappings (#4): add POST /admin/user-mappings so an admin can set the github_login -> work_email link from the dashboard instead of a raw Store write. New "admin" MappingSource provenance value. * fix(webapp): refresh user-mapping cache on GitHub webhook paths On managed LangGraph Cloud the backend runs multiple replicas, so the per-process GitHub<->work-email mapping cache can be stale on the replica handling a webhook (a mapping created on another replica is invisible until refresh). process_github_pr_comment and process_github_issue now refresh the cache from the durable Store before resolving the author's email, matching the existing Slack mention path (process_slack_mention). * perf(webapp): defer deepagents import to speed custom-app cold start The custom FastAPI app (agent.webapp:app, the langgraph.json http.app) pulled deepagents -> langchain_anthropic -> anthropic into its import graph via dashboard.routes, only to build skill/chat seed files. Defer those create_file_data imports into the functions that use them. Removes deepagents/langchain_anthropic/anthropic from app import entirely and roughly halves module-import wall time (~0.6-0.8s -> ~0.35s warm; larger cold-start saving since native anthropic init is skipped). Behavior identical. (reviewer_diff already imports deepagents under TYPE_CHECKING.) * feat(ui): set work_email user mappings from the admin dashboard Add an "Add / update" form to the admin User mappings section and the adminUpsertUserMapping API client method, wiring the new POST /admin/user-mappings endpoint. Admins can now create or update a github_login -> work_email mapping directly instead of waiting for the user to self-connect Slack. * docs: document managed LangGraph Cloud + Vercel deployment - INSTALLATION §10: add the managed production env triad (LANGGRAPH_URL, DASHBOARD_BASE_URL + DASHBOARD_API_BASE_URL with https://, empty VITE_DASHBOARD_API_BASE_URL for same-origin), the stable-alias login and vercel.json stable-deployment-URL requirements, multi-replica cache note, plus redirect_uri-scheme and oauth-state-mismatch troubleshooting. Refresh the langgraph.json snippet to all six graphs. - README: reframe deployment around the managed migration; link the plan. - deploy/MIGRATION.md: import the self-hosted -> managed migration plan. |
||
|
|
a4ed19ba61
|
feat: migrate model providers to Bedrock (Claude) + Fireworks (everything else) (#62)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: switch model providers to AWS Bedrock (Claude) and Fireworks (non-Claude)
Migrate off direct provider APIs: AWS Bedrock for Anthropic/Claude via the
cross-region inference profile us.anthropic.claude-opus-4-8, Fireworks AI for
all non-Claude models. Drop OpenAI (gpt-5.5) and Google (gemini-3.5-flash)
entirely. DEFAULT_MODEL_ID is now Bedrock Claude; all Fireworks models stay
freely selectable for the agent and reviewer graphs and via team/profile
defaults.
- pyproject: add langchain-aws (ChatBedrockConverse + boto3)
- options.py: Bedrock Claude entry + default; remove openai/google entries
- model.py: bedrock_converse provider_model_kwargs (effort -> thinking budget),
region pin in make_model, bedrock<->fireworks fallback pairing, AWS_REGION/
FIREWORKS_API_KEY local-dev validation
- server.py: provider-aware fallback kwargs build
- sanitize_thinking_blocks: also sanitize ChatBedrockConverse thinking blocks
- model_fallback: treat transient botocore ClientError codes as fallback-worthy
- eval_jobs: repoint hardcoded eval model id to Bedrock Claude
- tests: repoint dropped model ids; drop obsolete google test module
* fix(bedrock): use adaptive thinking + output_config.effort for Opus 4.8
The handoff spec wired Bedrock Converse thinking as
{type: enabled, budget_tokens: N}, but Opus 4.7+ rejects that with a
ValidationException: thinking.type "enabled" is not supported; it requires
thinking.type "adaptive" plus output_config.effort. Verified by live invoke
against us.anthropic.claude-opus-4-8 (account 328440206208, us-east-1):
the enabled+budget shape 400s, adaptive+effort returns normally.
Map profile effort to additional_model_request_fields:
{thinking: {type: adaptive, display: summarized},
output_config: {effort: <low|medium|high|xhigh|max>}}
reusing anthropic_thinking_for/anthropic_effort_for. Update the two
subagent-model tests asserting the old shape.
* fix(deploy): seed Bedrock/Fireworks models, not the dropped anthropic:/openai: ids
Model selection is store-driven, so seed_store.sh's team_settings/default seed is
what runs in prod. It still seeded the removed providers, which would fail at runtime
after the migration:
- agent/builder: anthropic:claude-opus-4-8 -> bedrock_converse:us.anthropic.claude-opus-4-8
- reviewer: openai:gpt-5.5 (dropped) -> bedrock_converse:us.anthropic.claude-opus-4-8
(set SEED_REVIEWER_MODEL to a Fireworks model for a cross-family reviewer)
- fetch-config REQUIRED_PROVIDER_KEYS default ANTHROPIC_API_KEY,OPENAI_API_KEY ->
FIREWORKS_API_KEY (Bedrock auths via host IAM role; dropping the old keys would
otherwise fail-fast at boot)
- docs (DEPLOYMENT/ROTATION/put-config) updated to match.
Surfaced by the cross-family review + verified against deploy/.
* fix(bedrock): security-review NITs — region resolution, error sanitization, reasoning-block strip
From /sh-security-review (all confirmed-low):
- model.py: resolve region from AWS_REGION OR AWS_DEFAULT_REGION (matches
validate_local_dev_llm_config) so the validated region is the one actually used.
- model_fallback.py: sanitize Bedrock AccessDenied/ResourceNotFound errors to the
error code only, so the role ARN + account id in the raw botocore message never
reach logs or the user channel (CWE-209).
- sanitize_thinking_blocks.py: also strip empty Bedrock reasoning_content blocks
(Converse emits reasoning_content, not thinking) so the middleware is not a no-op
on Bedrock; + unit tests. (Empty blocks replay fine today; defensive.)
* deploy(bedrock): grant instance-role Bedrock invoke + repoint LLM_MODEL_ID / eval model ids
Deployment-readiness for the Bedrock migration (PR #62):
- instance-role.ts: least-privilege bedrock:InvokeModel[WithResponseStream] on the
us.anthropic.claude-opus-4-8 inference-profile ARN + the foundation-model ARN in
each routed region (us-east-1/2, us-west-2). The model runs in the server process
on the box, so the EC2 instance role is the principal. Simulator-verified (allowed
for opus-4-8, implicitDeny for other models) and synth-verified. Passed the
mandatory GPT-4.1 IAM cross-review (no blockers, least-privilege confirmed).
- config-store.ts: IaC SSM LLM_MODEL_ID anthropic:claude-opus-4-8 ->
bedrock_converse:us.anthropic.claude-opus-4-8. This SSM value overrides
seed_store.sh's default via pick precedence, so the seed-script fix alone was
insufficient — both sources now point at the supported Bedrock id.
- infra/README.md + evals/reviewer/config.toml: repoint stale anthropic:/google_genai:
ids to the Bedrock id (config.toml's model_id was an active, now-broken value).
AWS_REGION is already wired via user-data.sh (IMDS -> boot.env), so no change needed there.
* chore(secrets): drop OPENAI/GOOGLE/GROQ key shells (revoked, providers removed)
Those three providers were dropped in the Bedrock/Fireworks migration and their keys
revoked; the live Secrets Manager objects (open-swe-{dev,prod}/{OPENAI,GOOGLE,GROQ}_API_KEY)
were deleted (7-day recovery). Remove them from the IaC so a future cdk deploy does not
recreate the shells, and from fetch-config's mirror array so boot stops requesting them:
- config-store.ts SECRET_VARS + descriptions (28 -> 25 shells)
- fetch-config.sh SECRET_VARS array (kept in lockstep)
- put-config.sh: drop the put_secret lines; ANTHROPIC_API_KEY re-labelled optional
(eval judge only — Bedrock builder/reviewer auth via the host IAM role).
REQUIRED_PROVIDER_KEYS is not set in SSM, so it uses the FIREWORKS_API_KEY default.
|
||
|
|
8a9974c3c4
|
feat: author Slack/dashboard/schedule commits + PRs as the app by default (#57) (#60)
Some checks failed
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Build & publish app artifacts / Publish + deploy (dev) (push) Has been cancelled
Build & publish app artifacts / Publish + deploy (prod) (push) Has been cancelled
Infra CD / Infra CI (pre-deploy) (push) Has been cancelled
Infra CD / Deploy open-swe-dev (push) Has been cancelled
Infra CD / Deploy open-swe-prod (push) Has been cancelled
* feat: default Slack/dashboard/schedule PRs + commits to the app identity (#57) Slack/dashboard/schedule runs now author PRs and run git/gh operations as the GitHub App seahaven-openswe[bot] by default (matching GitHub-issue runs), so the self-review 422 is impossible by construction rather than guarded in the prompt. A profile flag author_prs_as_user restores per-user attribution. - open_pull_request._resolve_pr_author_token + auth.resolve_github_token: default to the installation token for these sources; per-user only when opted in. - authorship: commit identity -> seahaven-openswe[bot] (numeric noreply; accepted Vercel-resolution risk, documented inline). - self-trigger safety: INTERNAL_BOT_LOGINS + webapp/reviewer_reconcile/reply markers recognize seahaven-openswe[bot] (bot-authored events are now ours). Supersedes the prompt-only guard in #58. * fix: author commits as the app bot in the default path (SH-IDSPLIT-01) Security review found the commit identity was NOT actually unified to the bot: resolve_triggering_user_identity got a 403 from the installation token and fell back to configurable['github_login'], so commits were still authored as the triggering user (commit=user, push+PR=bot — a three-way split that missed the stated goal). Now gate the triggering-user identity resolution on the same default-bot decision as the token: slack/dashboard/schedule default to the app bot identity unless author_prs_as_user is set. * docs(security): record AUTHZ-SLACK-BOT-DEFAULT-001 as an accepted residual (#59) Single-user deployment; bounded by App-on-pilot + ALLOWED_GITHUB_REPOS lock. Revisit (add a per-user gate) before expanding users or the App installation. |
||
|
|
a33aaec495
|
fix: resolve security-review findings (sandbox isolation, IAM list scope, webhook replay, info-leak) (#54)
* fix: enforce a replay window on Linear webhooks (AUTHZ-001) verify_linear_signature accepted any correctly-signed body with no freshness check, so a captured request could be replayed indefinitely. Parse the signed webhookTimestamp (Unix ms) and reject requests outside a 60s window, failing closed when the field is missing or malformed — mirroring the Slack verifier. * fix: stop leaking upstream auth-error bodies into user comments get_github_token_for_user folded the raw upstream response text into the error string that becomes a Slack/Linear comment (AUTH-RESP-LEAK-01). Log the full body server-side only and return a generic "GitHub auth failed (status <code>)". Also document the accepted shared-installation-token blast radius on the bot-token-only path (AUTHZ-003). * fix: bind sandbox and token caches to repo to prevent thread-id collision A PR head-branch name is attacker-controllable and get_thread_id_from_branch derives a thread_id from its first UUID with no repo binding (TID-COLLIDE-01). The in-memory sandbox cache and the per-thread GitHub-token cache were keyed on thread_id alone, and a cached sandbox was reused after only an echo-ping, so a different repo's webhook could bind to another thread's sandbox or token. Without changing the persistent thread-id scheme: - Persist the bound repo (owner/name) in thread metadata on sandbox creation and refuse to reuse a sandbox whose bound repo does not match the current event (SandboxRepoMismatchError); the in-memory proxy also carries the binding. - Bind the GitHub-token cache entries to their repo and evict on a cross-repo read so a colliding thread_id cannot be served another repo's token. - Thread repo through the reviewer and the webhook token resolvers. * fix: scope s3:ListBucket to the releases/ prefix (F-1/IAC-04) The instance role and the GitHub deploy app role granted s3:ListBucket on the whole assets bucket. Every caller (deploy.sh, the publish/rollback scripts) only ever lists under releases/, so add a StringLike s3:prefix=releases/* condition. GetBucketLocation has no s3:prefix in its request context, so it moves to its own unconditioned statement. Also document the accepted F-2 cross-env existence-oracle residual on BatchGetSecretValue. * chore: suppress test-fixture credential false positive; document AUTHZ-002 Add a machine-level suppression for the fake Datadog key in the test_team_credentials encryption-roundtrip fixture (CWE-798, not a real credential). Clarify that the within-org thread-write path is intentional by design (AUTHZ-002) — comment only, no behavior change. * fix: casefold repo-binding keys to avoid spurious cross-repo mismatch GitHub owner/name are case-insensitive. Casefold the owner/name key on both the write (binding) and read (compare) sides — repo_cache_key and the metadata bound_repo read — so Org/Repo and org/repo resolve to one repo and a legitimate same-repo run cannot raise a spurious SandboxRepoMismatchError (Gap 2). * fix: stop leaking upstream auth body in unexpected-result branch The 2xx-but-missing-token/url branch echoed the parsed upstream response body into the user-facing error. Return a generic message and log response_data server-side only, mirroring the existing HTTPStatusError fix (Gap 4). * fix: fail closed for unbound-legacy sandboxes and catch repo mismatch Gap 1: a thread with a persisted sandbox_id but no in-memory cache and no recorded bound_repo (a pre-binding legacy thread, post-deploy) previously reconnected-and-served the sandbox to the current repo, then rebound it. Now fail closed: drop the stale id and recreate a fresh sandbox bound to this repo, logging a reconnect-with-missing-binding event. A sandbox is never served to a repo unless its binding is known and matches; new threads bind on first run unchanged. Gap 3: catch SandboxRepoMismatchError at the agent and reviewer run entrypoints, log it for alarming, and surface a clean sanitized error instead of letting an opaque deep-stack exception crash-loop the worker. * chore: suppress test-fixture credential false positive in token-TTL tests Add a machine-level suppression for the fake "ghp_secret" GitHub token used by the cached-token TTL/revocation unit tests (CWE-798). Not a real credential and not a valid PAT; scoped to the unit test only. |
||
|
|
f3db9f02e3
|
Adopt Sea Haven agent conventions, no attribution (#30)
Codify the box-only #4 customizations into Git so the AWS deployment (which deploys from this repo) actually applies them — previously only the retired sh-openswe box had them. - prompt.py: branch names feature|bug|hotfix/<kebab> (optional <KEY->); imperative PR titles with no conventional-commit type: prefix; PR body Summary/Validation/Tests/Notes; handbook commit format. Rewrite the collaboration template from an attribution MANDATE to a PROHIBITION — no Co-authored-by bot trailer, no "Made by [Open SWE]" footer, no agent/AI notes on any artifact. - github_comments.py: add @seahaven-openswe (the deployed App slug) to the mention triggers. - authorship.py: remove the now-unused attribution helpers (build_pr_attribution_footer, add_bot_coauthor_trailer, add_pr_collaboration_note, PR_ATTRIBUTION_*). Keep OPEN_SWE_BOT_* — server.py still uses them for the sandbox git identity. - Flip the attribution unit tests to assert the no-attribution behavior; drop tests for the removed helpers. Commits stay authored as the triggering user for now — flipping authorship to the bot account depends on the Vercel preview-deploy constraint and is deferred to #11. Refs: #4 #11 Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi |
||
|
|
4c7d45273f
|
perf: cut review-chat time-to-first-token (#1598)
* perf: cut review-chat time-to-first-token The sandbox-less PR review chat paid several blocking network round-trips before the first token on every message. Cache GitHub App installation tokens in-process (per scope, until ~10m before expiry, above the proxy's 5m refresh window) so the chat graph factory and proxy stop re-minting one each turn. Also drop the duplicate thread-metadata read in the commands proxy and replace the heavy per-message get_review staleness check with a single lightweight PR head-SHA lookup. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: keep review chat alive when reseed fails Address review: a moved PR head now triggers _build_pr_context (and thus get_review). For an existing chat, fall back to the last seeded context on HTTPException instead of failing the command; fresh chats still surface the error since they have no prior context. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
3992d3ef5d
|
feat: repo-scoped dynamic sandbox snapshots (#1595)
* feat: repo-scoped dynamic sandbox snapshots Let admins build a per-repo sandbox image from a custom Dockerfile so runs targeting that repo boot from a snapshot with its deps pre-baked. Snapshot selection is purely additive: repos without a `ready` repo-scoped snapshot always fall back to the configured DEFAULT_SANDBOX_SNAPSHOT_ID. Backend adds a repo_snapshots store module (Dockerfile + build status keyed by owner/name), threads the resolved repo through the LangSmith sandbox creation path, runs builds via SandboxClient.create_snapshot_from_dockerfile in a throwaway builder sandbox, and exposes admin-only CRUD + build endpoints. The UI adds an admin-only Agents-tab page (repo picker + Monaco Dockerfile editor + build status/logs). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden repo snapshot builds Require REPO_SNAPSHOT_BASE_IMAGE for generated Dockerfile templates so admins cannot accidentally build a repo snapshot from a bare Python image that lacks Open SWE's sandbox tools. Allow stale building records to be retried by tracking build_started_at and treating old or missing timestamps as stale. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: document repo snapshot base image config Document REPO_SNAPSHOT_BASE_IMAGE alongside sandbox snapshot setup and convert missing base-image configuration into a handled dashboard API error so admins see a clear configuration message instead of an unhandled template-generation error. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
3a0e2b4672
|
feat: plan mode with model-driven entry and collaborative review (#1580)
* feat: add plan mode for read-only research and planning
Adds a per-run plan_mode flag that puts the agent in a read-only
research phase: a strong prompt section is injected and mutating tools
are stripped via ExcludeToolsMiddleware so the agent proposes a
reviewable implementation plan before any edits. Surfaced in the
dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: enforce plan-mode read-only at tool layer and disable subagents
Addresses PR review: plan mode previously relied on prompt text to keep
the shell read-only and left the task subagent (built with its own
write/PR/Linear tools) unrestricted. Now `task` is excluded so research
cannot be delegated to a mutating subagent, and a new
PlanModeShellGuardMiddleware enforces a read-only command allowlist on
`execute`, blocking writes, git state changes, installs, redirection,
and command substitution regardless of model/prompt-injection compliance.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* fix: harden plan-mode shell guard against wrapped mutations
Block git global options that take values (-C, --git-dir, ...) from being
misread as the subcommand, reject config-injection options (-c,
--config-env, --exec-path), and drop the env command wrapper that could
run arbitrary commands.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow
- enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True})
- Plan mode resolution: per-thread > profile default > team default > False
- PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool
- profile_plan_mode_default and team plan_mode_default settings
- Slack plan on/off/status commands with thread metadata persistence
- slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons
- Interactivity handler: approve triggers implementation run, cancel posts confirmation
- Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
* refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline
Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell
commands during plan mode. Plan mode now relies on the system prompt to
instruct the agent not to run mutating commands; the mutating-tool exclusion
(ExcludeToolsMiddleware) is retained.
* test(open-swe): add Playwright E2E for the Slack → PR → web handoff
Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox.
- full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread.
- dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer).
Wired into Agent CI as a `Playwright E2E` job that runs on pull requests.
* fix(open-swe): serve E2E UI assets via explicit route; pin Playwright
The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead.
Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state.
* test(open-swe): record Playwright trace + video on every E2E run
Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it.
* feat(plan-mode): collaborative plan review with BlockNote + Yjs
When the agent enters plan mode it writes the plan as a markdown file in the
sandbox (save_plan tool), publishes it, and posts a review link to the source
channel. Reviewers open the plan inside the dashboard (under the /agents shell),
read it rendered in a BlockNote editor, and leave inline comments synced live
over Yjs. Only the thread owner can approve; any reviewer can request changes.
On approve/reject the comments are harvested and handed to the agent for the
follow-up run; the agent never sees comments mid-review.
- agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares
the plan-review link.
- dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed
snapshots; plan content/status store; plan REST API (get/approve/reject,
owner-only approve, client-harvested comments); planStatus on thread summaries.
- ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page
mounted under the agents shell, with a "Review plan" banner in the thread view
and a back-link; theme-aware (dark mode) using the dashboard tokens.
- e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR
flow, including cross-user comment sync and owner-only approval.
* fix(plan-mode): address review feedback (authz, overrides, leaks, deps)
- plan-collab WS: authorize per-thread before joining a room (same read gate as
the REST API) — previously any logged-in user could join any thread (IDOR).
- plan-collab: tie the snapshot flusher to active connections (refcount) so each
opened plan no longer leaks a permanent 1.5s task on the shared event loop.
- plan decisions: include thread_id in the follow-up run configurable so the run
resumes the existing thread; set plan_mode explicitly so approve forces it off.
- get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan,
dashboard toggle) now overrides profile/team defaults instead of falling back.
- plan mode tool gating moved to a state-aware PlanModeMiddleware installed
unconditionally, so a mid-run enter_plan_mode restricts the next model turn;
before_agent resets stale plan_mode so a later run isn't forced back into it.
- exclude write-capable http_request from plan mode.
- pin pycrdt / pycrdt-websocket with upper bounds.
Includes the latest base (#1583): E2E UI assets served via explicit route
(fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts).
* style: ruff format plan_collab.py
* fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS
- Slack "Approve & Implement" now verifies the clicking user is the plan
requester (owner, via the stored triggering_user_id) before implementing —
matching the dashboard API's owner-only approval. Non-owners are pointed to
Revise / feedback.
- The plan-collab WebSocket validates the handshake Origin against the dashboard
allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring
the REST require_same_origin CSRF defense.
* fix(plan-mode): enter plan mode only via the model + local mock dev harness
Plan mode is now entered solely when the model calls enter_plan_mode.
Removed the per-user and team plan_mode_default settings (backend + UI)
and the Slack `plan on/off/status` toggle.
- enter_plan_mode returns a terminating ToolMessage, fixing the missing
ToolMessage error that silently dropped plan mode mid-run.
- PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev
remount doesn't destroy and then reuse the collaboration provider.
- e2e plan_review spec asserts plan_mode actually engages.
- LangSmith trace-url resolution is best-effort: bail before any API
call when the tenant is unset, cache failures, log at debug.
- Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM,
Alice/Bob mock users, and a GitHub login picker.
* docs(plan-mode): drop stale references to removed profile/team defaults
The plan_mode middleware docstring and the approve/reject dispatch comment
still described the profile/team plan_mode_default resolution that no longer
exists; reword to match model-driven entry + the per-thread carry.
* feat(plan-mode): let any reviewer edit the plan, not just comment
Drop the owner/commenter split for the plan document: everyone with read
access edits and comments alike (DefaultThreadStoreAuth "editor" for all,
editor always editable until a decision, anyone seeds the empty doc). This
matches the collab WS, which already relays frames to every readable user.
Plan approval stays owner-gated.
* test(plan-mode): assert plan-mode entry via the tool's success message
plan_mode lives only in run state for tool gating; it is not a persisted
thread-state channel, so the previous `values.plan_mode === true` poll
could never pass. Assert instead that enter_plan_mode's success ToolMessage
("Plan mode is active …") lands in the thread — which only happens when the
tool's Command applies cleanly, the exact regression this guards.
---------
Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
|
||
|
|
eac293d8a7
|
fix: serialize Slack run dispatch (#1591)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
2df0eabd35
|
feat: reviewer enforces AGENTS.md/CLAUDE.md repo rules as mandatory pass (#1569)
* feat: reviewer enforces AGENTS.md/CLAUDE.md repo rules as mandatory pass The reviewer already fetched AGENTS.md but treated violations as optional candidate findings. Now the reviewer runs a dedicated compliance pass that checks every changed hunk against each rule in AGENTS.md (or CLAUDE.md as fallback), treating violations as mandatory findings rather than style nits. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: oversized AGENTS.md returns None instead of falling back to CLAUDE.md Only a 404 (file absent) triggers fallback to CLAUDE.md. Oversize, HTTP errors, and unexpected status codes now return None immediately so the reviewer does not enforce stale rules from a secondary file. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
4030001ebf
|
feat: validate LLM API keys on startup (#1438)
* feat: validate LLM API keys on startup * fix: correct relative import for options module * refactor: move imports to top of file * style: fix linting and formatting issues * refactor: scope LLM validation to local dev and rename function --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> Co-authored-by: open-swe[bot] <johannes@langchain.dev> |
||
|
|
98b824bd54
|
feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue [closes OPE-51] (#1567)
* feat: add size caps for PR diff, fetch_url, Slack threads, pagination, message queue Per-source byte/token caps with explicit truncation markers to prevent unbounded payloads from blowing up LLM context/memory. - reviewer_diff.py: cap PR diff at 200K chars with head+tail truncation - fetch_url.py: cap markdownify output at 100K chars - slack.py: cap thread message fetch at 500 messages - github_comments.py: cap _fetch_paginated at 50 pages - thread_ops.py: cap queued messages at 100 (drop oldest) Closes OPE-51 Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: compute diff line set from full diff, keep most recent Slack messages Address PR review comments: 1. Truncated diffs rejected valid findings: fetch_pr_diff now returns the full diff; truncate_diff is called separately in reviewer.py so the line set used for add_finding/publish_review validation is computed from the complete diff, not the truncated prompt text. 2. Slack cap dropped recent thread context: fetch_slack_thread_messages now keeps the most recent SLACK_THREAD_MAX_MESSAGES messages (was keeping the oldest). The tool surfaces a truncation marker in the formatted output so the LLM knows the thread was truncated. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
7dd758f845
|
feat: shared GitHub HTTP helper with retries, rate-limit handling [closes OPE-45] (#1565)
* feat: shared GitHub HTTP helper with retries, rate-limit handling, and sane timeouts Introduces agent/utils/github_http.py — a single place for GitHub API HTTP calls with 30s/10s-connect timeouts (vs httpx's 5s default), exponential backoff with jitter, Retry-After header support, and 429/secondary-rate-limit detection. Migrates the reviewer publish path (reviewer_publish.py, reviewer_diff.py, github_checks.py, github_ci.py) from one-shot httpx.AsyncClient() calls to the shared helper. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't retry transport errors on non-idempotent GitHub writes POST/DELETE/PATCH can create side effects server-side even when the client gets a timeout or connection reset. Only retry transport errors for idempotent methods (GET, HEAD, PUT, DELETE). 429/5xx status codes are still retried for all methods since the server explicitly did not process the request. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: don't retry 502/504 on non-idempotent GitHub writes 502 (bad gateway) and 504 (gateway timeout) are ambiguous — the upstream may have processed the write before the gateway returned an error. Only retry these for idempotent methods. 429 and 503 are still retried for all methods since the server explicitly did not process the request. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
b19804536c
|
feat: handle images sent to non-vision models in Slack, Linear, and web UI (#1560)
* feat: handle images sent to non-vision models in Slack, Linear, and web UI Add vision capability checks across all image input paths. When a user sends images to a text-only model (e.g. GLM 5.2, DeepSeek V4 Pro), the images are now skipped and a warning is injected into the prompt instead of sending unsupported content to the model. - Slack: resolve model at webhook time, skip image fetch + add warning - Linear: same pattern as Slack - Queued message middleware: read resolved model from thread metadata, strip images from queued payloads for text-only models - Web UI: disable submit + show inline warning when images are attached to a non-vision model selection - Shared: resolve_agent_model_id helper + vision_not_supported_warning Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: mock resolve_agent_model_id in Slack mention test The test_process_slack_mention_queues_active_thread_message test was missing a mock for the new resolve_agent_model_id call added to the Slack webhook handler, causing a TypeError when image URLs triggered the model resolution path. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: include vision warning in queued payload for text-only models Update the prompt variable (not just content_blocks) before clearing image_urls so the queued payload also carries the warning text when a Slack/Linear follow-up arrives while the thread is busy. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
554d754585
|
feat: link PR attribution footer to the originating thread (#1539)
The "Made by Open SWE" PR footer linked to the generic dashboard homepage. Point it at the dashboard thread that generated the PR (/agents/<thread_id>), falling back to the homepage when no thread id is available. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
7397ff93ba
|
feat: CI auto-fix and PR babysitting for agent PRs (#1530)
* feat: CI auto-fix and PR babysitting for agent PRs Watch CI failures and review feedback on PRs Open SWE opened, then dispatch confidence-gated fix runs on the originating agent thread. Adds CI webhook ingestion (check_run/check_suite/workflow_run/status), a per-PR @open-swe autofix on|off toggle, auto-response to review comments, and a polling ci_monitor graph that also flags merge conflicts. Gated by the existing autofix_mode/trigger_mode settings, the enabled-repos opt-in, base-branch and human-commit skip rules, dedupe, and a per-PR attempt cap. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: address review feedback on CI auto-fix - Security: gate the no-mention review-feedback path on author trust — require a trusted author_association (OWNER/MEMBER/COLLABORATOR) plus a GitHub write/maintain/admin permission check before dispatching a write-capable agent run, preventing privilege escalation from read/triage/outside reviewers. - Auth: reuse the originating PR thread's source + login/email when dispatching fix runs so the GitHub-token resolver authenticates them in non-bot-token deployments (bespoke github_ci source failed to resolve). - Docs: document the Commit statuses: Read-only permission required for the Status webhook event. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
889b2663d9
|
fix: point View trace link at the correct tracing project (#1521)
* fix: resolve trace URL project id by tracing project name Graphs were split into separate LangSmith tracing projects (open-swe-agent, open-swe-review) but the "View trace" link still used a single fixed project-id env var pointing at the old combined project. Resolve the project id from the tracing project name so agent and reviewer links point at their respective projects, falling back to the env var when resolution is unavailable. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * docs: document per-graph tracing projects for trace links View trace links now resolve project IDs from the open-swe-agent / open-swe-review project names. Document this so fresh deployments create the right projects instead of relying on a single project ID. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
bd5d5b24d4
|
fix: point reviewer "Open in Web" link to the review page (#1519)
The top-level review comment's "Open in Web" link pointed at the
agent thread (/agents/{thread_id}). Point it at the dashboard review
detail page (/agents/reviews/{owner}/{repo}/{number}) instead.
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
|
||
|
|
aa98d195a8
|
feat: route graphs to separate LangSmith tracing projects (#1508)
* feat: route graphs to separate LangSmith tracing projects * docs: point graph entrypoint references at traced wrappers --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
bfcb714c5c
|
feat: PR review page (#1495)
* feat: PR review page * feat: manual review trigger card on admin page * feat: finding focus mode with floating card and dimmed backdrop * fix: dim only info panel, subtle block highlight for focused finding * refactor: move PR reviews into agents page with file-tree sidebar * fix: address review feedback - paginate list_reviews until enough accessible records collected - include rename/metadata-only files in diff API with empty hunks - remount ReviewBody on head_sha change so viewed-files state resets * fix: review tree background + finding card anchoring * fix: remove no-op PR size chip and dead check links * fix: parallel diff fetch, complete reReview type, shared github_headers --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
194e9cd60e
|
feat: link Slack thread/Linear ticket in PRs and append ticket to title (#1504)
* feat: link Slack thread/Linear ticket in PRs and append ticket to title When opening PRs, include any referenced Slack thread or Linear ticket in the description and append the resolvable ticket number to the title. Adds a Slack permalink to the webhook prompt so the agent has a ready-to-link thread URL. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: use [closes <TICKET>] format in PR title suffix Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: deterministically append source refs to private-repo PRs only Move the Slack/Linear source-reference linking out of the prompt and into open_pull_request, gated to private repos so private Slack thread URLs and Linear identifiers are never published to a public PR. Reverts the prompt instruction and webhook permalink injection in favor of this server-side append. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
ce3af2d9b2
|
fix: reviewer can silently review a stale checkout on reused sandboxes (#1503)
prepare_review_repo is best-effort, but the system prompt unconditionally told the agent the repo is checked out at the PR head. On a reused sandbox, a dirty worktree (or a transient fetch failure) makes 'git checkout <sha>' fail; prep returns False, the old checkout stays in place, and the reviewer confidently reads stale code — observed as 'I rechecked the current head' replies quoting pre-push code. - checkout with --force so leftover worktree state can't block it, verify HEAD matches the requested sha, tolerate 'git fetch --all' failures - when prep fails, the prompt now warns the checkout may be stale and tells the agent to re-fetch/checkout (or fall back to API file contents) before trusting local files Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
e0678e8c01
|
feat: track PR lifecycle state per thread for sidebar (#1492)
* feat: track PR lifecycle state per thread for sidebar Persist a PR's draft/open/merged/closed state on the agent thread and keep it in sync as PR webhooks fire, so the Agents UI can show per-thread PR status the way Cursor does. open_pull_request now records the initial draft/open state and pr_title; thread summaries expose diffStats; and the PR webhook refreshes pr_state on close/reopen/draft/ready transitions. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor: consolidate PR state mapping into shared derive_pr_state helper --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
da20e57bcc
|
fix: refresh sandbox GitHub proxy token before mid-run expiry (#1496)
* fix: refresh sandbox GitHub proxy token before mid-run expiry GitHub App installation tokens expire after exactly 1 hour. The LangSmith sandbox proxy was configured once at run start with a snapshot of that token, so runs longer than ~1h hit 401s on every gh/git call. Record the proxy token's expiry per thread and add a before-model hook that re-configures the proxy with a fresh token when it nears expiry. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: preserve repo-scoped proxy token on mid-run refresh Reviewer runs mint a repository-scoped installation token. Record the repo scope per thread alongside the expiry so the before-model refresh re-mints a token with the same scope instead of an installation-wide token, avoiding privilege expansion on long reviewer runs. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: update passthrough stub for github_proxy_repositories param --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
abf354bb05
|
feat(open-swe): stream agent chat via @langchain/react v2 protocol (#1475)
* feat(dashboard): stream agent chat via @langchain/react v2 protocol Replace the bespoke SSE + React Query polling path with LangGraph’s v2 event stream through credentialed dashboard proxies. Run starts go through stream commands; mid-run follow-ups still queue via /messages. * fix import path * fix tests after rebase * format * PR feedback * improved model fallback * fix image handling * embrace sdk * cleanup * cr * more cleanup * fix cors * harden security --------- Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
259ec57186
|
fix: keep review check on follow-up commits, drop neutral conclusion (#1486)
* fix: keep review check visible on follow-up commits and stop using neutral The "Open SWE Review" check completed as `neutral` whenever findings were surfaced, which GitHub renders as a confusing "neutral check" group. Always complete it as `success` (informational/non-blocking), matching Devin and Corridor — the finding count stays in the title and findings post as comments. Also create a fresh check run on the new head SHA in the push re-review path: GitHub only shows check runs on a PR's current head, so the check vanished after a follow-up push (and the stale id settled on an outdated commit). Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: surface settled review check when push leaves diff unchanged --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
e8770e1004
|
feat: report Open SWE Review as a PR check run (#1484)
* feat: report Open SWE Review as a PR check run Auto-review dispatch now creates an in-progress 'Open SWE Review' check run on the PR head SHA; publish_review completes it (neutral with findings, success when clean). An after-agent hook fails the check if the run dies before publishing. Requires the GitHub App's Checks: Read & write permission; all calls are best-effort so a missing permission never breaks reviews. * fix: address review feedback on check-run settling Keep review_check_run_id when the completion PATCH fails so a later publish or the after-agent hook can retry instead of hanging the check; count out-of-diff findings toward the check conclusion. * fix: retry failed check completion with the real publish conclusion A transient PATCH failure after a successful publish previously left the check id for the after-agent hook, which settled it as 'failure'. Persist the intended result as review_check_pending_result and have the hook prefer it over the generic failure fallback. --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
8f596a421d
|
feat: prep reviewer repo at init and load repo skills (#1480)
* feat: prep reviewer repo at init and load repo skills Clone + checkout the PR head during reviewer agent init so SkillsMiddleware can discover the repo's .agents/skills and .claude/skills from disk at its one-shot scan, and so the LLM no longer narrates the clone mid-run. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: load reviewer skills from trusted base sha and fetch PR pull ref Address review findings: skills are now extracted from the PR base sha via git archive into a dir outside the checkout (prevents PR-authored SKILL.md prompt injection), and repo prep fetches refs/pull/<n>/head with a strict checkout so fork PRs fail loudly instead of silently reviewing the default branch. * fix: drop ref from skill-extraction log to satisfy CodeQL --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
85049290a3
|
feat: apply API standards skill in PR reviews for API changes (#1452)
Pull the api-standards skill from the LangSmith Context Hub at reviewer run start and inject it into the system prompt, gated on the PR adding or modifying an API surface. Best-effort: failures fall back to no supplement. Co-authored-by: GowriH-1 <218394553+GowriH-1@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
ab2e0d1da3
|
feat: resolve Slack repo from channel topic/purpose (#1453)
Add conversations.info fetch so a repo:owner/name (or GitHub URL) token in a Slack channel's topic/purpose pins the channel to a repo, slotting in just below thread metadata in get_slack_repo_config. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> |
||
|
|
d2268871eb
|
feat: surface dashboard UI link on PR reviews [INF-0000] (#1440)
* feat(reviewer): surface dashboard UI link on PR reviews [INF-0000] Post a transient "review in progress" comment (with an "Open in Web" dashboard link) when a reviewer run starts, then delete it once the review lands. The published review body now carries the same "Open in Web" link, so the link persists on the review itself. The transient comment's id is tracked in reviewer thread metadata (status_comment_id) so it can be deleted on completion. * refactor(reviewer): inline dashboard URL helper, drop redundant future import [INF-0000] |
||
|
|
3a91fbfabc
|
fix: revert co-author email to open-swe user noreply (#1424)
The bot's numeric noreply (215916821+open-swe[bot]@users.noreply.github.com)
introduced in
|
||
|
|
6895ddcedc
|
feat: add scheduled web agents (#1422)
* feat: add scheduled web agents Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com> * fix: secure scheduled agent repositories Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com> * feat: rebuild scheduled agents as Automations tab Scrap the inline ScheduledAgentsPanel and replace it with a dedicated Automations tab: sidebar nav entry, list view with stat cards + empty state, and a full editor (name, Active toggle, repo, scheduled trigger picker, agent instructions + model). * fix: clear collapsed-sidebar button on mobile in Automations * fix: allow clearing automation repo on update --------- Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com> |
||
|
|
f1baceb8ab
|
feat: add Gemini 3.5 Flash provider (#1420) |