Commit graph

56 commits

Author SHA1 Message Date
Adam Moussa
0f0f616cd4
feat(open-swe): explicit-request reviewer verdicts + shell verdict guard (#214)
Some checks failed
CI / Lint (push) Has been cancelled
CI / Format check (push) Has been cancelled
CI / Typecheck (push) Has been cancelled
CI / Unit tests (push) Has been cancelled
CI / Playwright E2E (push) Has been cancelled
CI / Docker build smoke (push) Has been cancelled
CI / Triage ledger up to date (push) Has been cancelled
CI / ui bun.lock in sync (push) Has been cancelled
* feat(reviewer): explicit-request verdicts + shell verdict guard

Mention-triggered reviews that explicitly ask for a verdict now submit a
real APPROVE/REQUEST_CHANGES through publish_review; auto-reviews stay
advisory (COMMENT). Authorization is enforced in code: publish_review
honors a verdict only when the dispatching webhook set verdict_requested,
which only the explicit-mention path does.

- request_pr_review gains instructions (forwarded verbatim into an escaped
  requester_instructions data block) and request_verdict
- self-review guard downgrades verdicts on Open SWE-authored PRs; stale
  APPROVEs are best-effort dismissed when later findings land
- new PullRequestVerdictGuardMiddleware blocks gh pr review
  --approve/-a/--request-changes/-r, gh api, and curl verdict fallbacks on
  both the coding-agent and reviewer graphs
- shared escape helper moved to agent/utils/prompt_data.py

* fix(reviewer): harden verdict path against security-review findings

Adversarial security review (detector fan-out + proof-or-kill verifier)
of the verdict feature surfaced several verdict-integrity gaps; resolve
the confirmed ones:

- head-drift (high): a mid-run push moves the resolved head, so an APPROVE
  could anchor to an unreviewed commit. Downgrade any verdict to a comment
  when the resolved head differs from the reviewed head (verdict_ignored
  reason head_moved); the push's own re-review submits a fresh verdict.
- self-review fail-open: downgrade to comment when the PR author cannot be
  confirmed (author_unknown), and compare bot logins case-insensitively.
- verdict_submitted now reflects GitHub's returned review state, not just
  the event we asked for, so a coerced APPROVE isn't reported as submitted.
- an authorized verdict whose findings all anchor outside the diff now
  posts as a bodied review with zero inline comments instead of failing.
- add finding_reply to the shared data-block escape tag superset.
2026-07-20 15:28:00 -04:00
bfb36be948
feat: add Linear issue search tool (#1748)
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit 79df6b2ff283afdbd36660be888c130ea964fe3f)

Co-authored-by: Johannes du Plessis <johannes@langchain.dev>
2026-07-17 16:22:38 -04:00
a82da1f907
refactor: move docs/resources/assets to domain layout
Part of the domain-reorg adoption (build plan step C1): fork content,
upstream layout. Moves INSTALLATION.md/CUSTOMIZATION.md under docs/,
static/ under assets/, and default_prompt.md under agent/resources/
(packaged via agent/resources/__init__.py), then switches prompt.py's
loader to importlib.resources with an explicit DEFAULT_PROMPT_PATH
override, matching upstream's hunk. README and CUSTOMIZATION.md links
updated for the new paths; wheel build verified to still ship
agent/resources/default_prompt.md.
2026-07-17 13:45:39 -04:00
Adam Moussa
cfb5624663
feat(open-swe): add E2B sandbox provider (port of upstream 48217b68, #1489) (#199)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
Additive sandbox provider adapted to this fork's synchronous create_sandbox
factory: agent/integrations/e2b.py registered as a lazy-import entry in
SANDBOX_FACTORIES, so the e2b/langchain-e2b imports only load when
SANDBOX_TYPE=e2b (dark-safe on unset env; unset E2B_API_KEY raises a clean
ValueError only when the provider is selected). Supports reconnect-by-id and
optional E2B_TEMPLATE. Non-langsmith providers skip the GitHub proxy step,
so the GitHub-App token flow is untouched.

Upstream: langchain-ai/open-swe 48217b68 (#1489), re-implemented against the
fork's sync sandbox lifecycle rather than cherry-picked (upstream ships on
the deferred async sandbox.py base).

Docs: provider tables/lists in README, CUSTOMIZATION, INSTALLATION; the
CUSTOMIZATION registration example now shows the lazy-tuple form the code
actually uses. uv.lock refreshed (adds e2b, langchain-e2b, dockerfile-parse).
2026-07-16 18:23:42 -04:00
Adam Moussa
1f80500cf3
feat: Jira + Confluence integration (tools + triggers) (#182)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Typecheck (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
* feat(open-swe): add Jira tool plane (Phase 1)

Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear
tools:

- utils/jira.py: service-account REST client (Basic auth) with get/
  create/update issue, comments, list projects, trace comment; issue and
  comment bodies normalized to markdown.
- utils/adf.py: minimal ADF <-> markdown conversion (read paths convert
  Jira ADF to markdown; agent comments convert prose to ADF).
- tools/jira_{comment,get_issue,get_issue_comments,create_issue,
  update_issue,list_projects}.py wired into the tool registry and the
  main agent tool list.
- tests/test_jira_utils.py: ADF conversion + mocked-transport util tests.

Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env
returns a clean error, so this is safe to land dark. Trigger plane,
prompt guidance, and config plumbing follow in Phase 2.

* feat(open-swe): add Confluence tool plane (Phase 3)

Curated Confluence Cloud REST toolset for the agent, mirroring the Jira
tools:

- utils/confluence.py: service-account REST client (Basic auth) with
  get/create/update page, add comment, CQL search. Page bodies are XHTML
  storage format (not ADF), with minimal storage<->text converters;
  update_page reads the current version and bumps it, as Confluence
  requires.
- tools/confluence_{get_page,create_page,update_page,comment,search}.py
  registered in the tool registry.
- tests/test_confluence_utils.py: converter + mocked-transport tests
  including the version-bump path.

Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN;
unset env returns a clean error. Activation in the agent tool list lands
with the Phase 2 server.py wiring.

* feat(open-swe): add Jira trigger plane (Phase 2)

Make an @openswe comment on a Jira issue spawn an agent run, mirroring
the Linear trigger plane:

- webhooks/jira.py: process_jira_issue clones process_linear_issue —
  deterministic thread id, full-issue fetch, actor accountId->email
  attribution feeding resolve_login_from_email_async (PRs open as the
  human), multimodal image handling, source="jira" + jira_issue config.
- webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time
  X-Automation-Webhook-Token check, fails closed), repo-resolution
  cascade, get_repo_config_from_jira_mapping.
- utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder
  entry — real project->repo mappings still needed).
- utils/jira.py: get_user_email (accountId -> email) for attribution.
- completion.py: source=="jira" failure-reply branch.
- prompt.py: Jira-triggered notify guidance + Refs:/branch key from
  {jira_project_key}-{jira_issue_number}.
- server.py: read jira_issue config + pass jira key to the system
  prompt; also activates the Phase 3 Confluence tools in the agent list.

Jira Automation lacks native webhook HMAC signing, so trust is a shared
secret header (decision D2); replay protection is weaker than Linear's
HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the
outstanding gate/hardening before push.

* fix(open-swe): harden Jira webhook trust (sh-security-review)

Resolves findings from the Phase 2 security review (detector fan-out +
proof-or-kill verifier). The unsigned Jira Automation webhook body was
trusted for identity, comment content, repo routing, and issue
existence; a JIRA_WEBHOOK_SECRET holder could forge those fields.

- Corroborate against the real Jira record: the webhook body is now only
  a pointer (issue_key + required comment_id). The triggering comment's
  author and text are re-fetched server-side via get_comment/fetch_jira_
  comment, and identity, the @openswe check, prompt text, and project
  key are derived from that authoritative record — never payload author/
  body fields. An uncorroborated comment is rejected. (closes the
  account-id impersonation, unsigned-body prompt injection, and
  fabricated-issue findings)
- Validate issue_key against the Jira key format and percent-encode all
  untrusted path segments (_seg) so a crafted key can't traverse to a
  different Jira REST endpoint or inject query params. (closes the path-
  traversal / query-injection findings)
- Route source=="jira" through the bot-token-default / author_prs_as_
  user opt-in path in resolve_github_token, matching Linear, instead of
  unconditionally resolving a per-user OAuth token from a payload email.
- Gate attribution on an active user mapping (is_login_mapped) so a
  pending/unconfirmed mapping can't drive PR authorship.

Adds regression tests: server-corroboration wins over payload, malformed
issue_key rejected, uncorroborated comment rejected, path-segment
encoding, project-key derivation, active-mapping gate.

Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/
REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body +
timestamp on the Automation payload to close the residual replay gap.

* harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist

Folds the two deployment-hardening items from the Phase 2 security review
into code (all opt-in / default-off, so existing and upstream deployments
are unaffected):

- JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must
  carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by
  JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by
  verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear
  HMAC+freshness model). Closes the static-token model's replay/forgery
  gap when enabled.
- JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's
  direct client IP (verify_jira_source_ip). Documented as direct-peer
  only; behind a proxy/LB, allowlist Atlassian's ranges at that layer.
- REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail
  CLOSED instead of the back-compat allow-all, plus a startup fail-open
  warning. Applies to all channels for consistency.

Documents all new vars (and a Jira section) in .env.example. Adds tests
for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and
the fail-closed allowlist.

* feat(open-swe): Confluence Atlassian Connect trigger (Phase 4)

Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian
Connect app. Designed and adversarially verified with the ultracode
workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus
skeptics on the implemented crypto).

- utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's
  official test vector), PyJWT HS256 webhook verifier with alg-pinning,
  issuer binding, and qsh-verified-last ordering; RS256 signed-install
  lifecycle verifier against Atlassian's published keys; installation
  store keyed by clientKey with the sharedSecret encrypted at rest
  (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned).
- webhooks/confluence.py: install/uninstall lifecycle + comment handler.
  The JWT-signed webhook body is only a pointer; the comment's real
  author/text/container are re-fetched server-side via the Basic-auth
  service account (Phase-2 corroboration lesson), with active-only login
  attribution and the repo allowlist.
- utils/confluence.py: get_comment / get_user_email (path-encoded).
- webapp.py: GET /connect/atlassian-connect.json (served dynamically),
  POST /connect/{installed,uninstalled,webhook/comment-created}, the
  space->repo resolver, thread-id, and fetch helpers.
- completion.py: source=="confluence" failure-reply branch.

Security: the sh-security-review verify pass confirmed one HIGH — the
symmetric signed-install=false first-install was trust-on-first-use gated
only by the public Confluence hostname (webhook-auth bypass). Fixed by
switching to signed-install=true + RS256 verification of lifecycle
callbacks, which cryptographically authenticates the first install. All
other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall
DoS/corroboration/injection) were defeated; residuals are deployment
config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover
bodies; comment-trigger prompt injection, shared with all sources).

New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN,
CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets
require the durable Postgres LangGraph store in prod.

Outstanding before push: /sh-security-review on the real diff and the
GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory.

* docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs

- prompt.py: Confluence-triggered runs notify via confluence_comment on
  the triggering page; add Confluence to the shared-base source list.
- CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian
  triggers (Jira Automation shared-secret webhook; Confluence Connect app
  with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the
  server-side corroboration + encrypted install store.

Phase 5 also verified the trigger surface end-to-end against a running
uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed
without valid auth) and recorded the integration in project memory.

* fix(open-swe): resolve /sh-security-review findings on the Atlassian surface

Formal sh-security-review (detector fan-out + verifier) over the Phase-4
Connect surface (esp. the new RS256 signed-install code, unseen by the
earlier adversarial verify) and the Phase-2 opt-in hardening.

CRITICAL — cross-tenant install (origin validation, CWE-346): signed-
install proves the caller is *an* Atlassian tenant, not *ours*, and the
descriptor is served publicly, so any attacker could install the app on
their own Confluence site and drive agent runs against our allowlisted
repos. The baseUrl body field is attacker-controlled and cannot bind the
tenant; only the signature-verified clientKey (JWT iss) can. Added a
MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in
process_install after signature+iss verification.

HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment
ids are per-instance, so generate_thread_id_from_confluence_comment now
salts the hash with the verified clientKey (plumbed from the webhook JWT
iss) to prevent thread hijack across tenants.

HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page
interpolated page_id into the REST path unencoded (update_page on a
mutating PUT with no params= backstop). Now _seg()-encoded, matching the
rest of the module.

MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no
bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID
guard mirroring the Linear botActor / Jira comment_author_is_bot checks.

LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as
opt-in defense-in-depth (the clientKey allowlist is the real gate).

Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp,
kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption,
constant-time comparisons, and the Phase-2 hardening. New regression
tests for each fix; full suite green (1602).

* harden(open-swe): GPT-4.1 cross-family review follow-ups

Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no
critical/high issues and confirmed the auth boundary is fail-closed and
correct. Two low-cost defense-in-depth items applied:

- Validate the signed-install JWT 'kid' against a strict charset before
  the public-key fetch, so a malformed kid fails fast with no network
  call (on top of the existing fixed host + percent-encoding).
- Make JWT nbf verification explicit (verify_nbf) on both the RS256
  lifecycle and HS256 webhook decodes.

Other suggestions triaged as already-handled (aud cross-app replay is
blocked by the per-tenant iss->secret lookup; documented static-token/IP/
baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra
(Fernet rotation via MultiFernet; rate limiting at the gateway).

* docs(open-swe): document Jira + Confluence in installation & customization guides

- INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret,
  service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect
  app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note,
  CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars +
  REQUIRE_REPO_ALLOWLIST.
- CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction
  note covers all four sources.
- AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth).
- README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
seahaven-openswe[bot]
0fc466a06b
docs: add per-repo AGENTS.md templates for stack-specific conventions (#126)
Some checks are pending
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
CI / Docker build smoke (push) Waiting to run
CI / Triage ledger up to date (push) Waiting to run
CI / ui bun.lock in sync (push) Waiting to run
Refs: #114

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
2026-07-08 16:29:35 -04:00
Ramon Nogueira
3c6077c418
feat: add Slack reaction tool (#1650)
Co-authored-by: Ramon Nogueira <270434257+ramon-langchain@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
(cherry picked from commit ee224d3e91576f771e93df7bad4523a7a7036324)
2026-07-03 15:38:56 -04:00
seahaven-openswe[bot]
a52ebed77c
chore: align docs, templates, and metadata with Sea Haven handbook [closes SH-93] (#95)
* Align repo docs, templates, and metadata with Sea Haven handbook

- Replace dynamic upstream-fork badges with static License, Python,
  and TypeScript badges; add the GitHub-native CI badge for dev.
- Repoint SECURITY.md contact to adam@seahavenind.com.
- Add .github/CODEOWNERS assigning reviews to @amoussa1229.
- Add PR template using the handbook structure.
- Add issue templates for bug, feature, and task plus a security link.
- Remove emoji from the README per CONTRIBUTING.md style.

Refs: SH-93

* Restore Linear 👀 acknowledgement and update metadata contacts

- README.md: bring back the documented 👀 reaction on Linear issues.
- README.md: refresh the TypeScript badge to 6.0+ to match ui/package.json.
- SECURITY.md + issue template: switch security contact to role alias.

Refs: SH-93

---------

Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-01 14:38:39 -04:00
Adam Moussa
7f60324f0c
chore: decommission self-hosted AWS LangGraph stack (#64)
* chore: decommission self-hosted AWS LangGraph stack

Removes the now-dead self-host IaC and AWS-only CI/CD after destroying the
dev + prod CloudFormation stacks (open-swe-dev, open-swe-prod, open-swe-iam,
and the dev-exclusive CDKToolkit-oswedev bootstrap) in account 328440206208,
us-east-1. The deployment is now managed (LangGraph Cloud + Vercel).

- remove infra/ (CDK app: app + IAM stacks, constructs, aspects, tests)
- remove deploy/ami (Packer AMI build) and deploy/seahaven (boot/config
  scripts, DEPLOYMENT/ROTATION runbooks)
- remove AWS-only workflows: cd-infra, ci-infra, build-artifacts, rollback
- README: rewrite the Deployment section to the managed LangGraph Cloud +
  Vercel view; drop dead links to infra/ and deploy/seahaven

Preserved: the shared default CDKToolkit bootstrap and promote-dev-to-prod.yml.
The RETAIN'd Secrets Manager shells and open-swe-<env>-assets S3 buckets
survive cdk destroy by design (orphaned) and need a separate deliberate cleanup.

* chore: clean up dangling references left by the AWS decommission

Folds in the FIX-level items from the #64 review gates (GPT-4.1 cross-review +
/sh-security-review), none of which were blockers:

- delete orphaned .github/scripts/{package-artifacts,publish-and-deploy,roll-box,
  rollback}.sh — their only callers were the removed AWS deploy workflows
- drop the deleted /infra dir from dependabot.yml npm directories (was producing
  a recurring Dependabot config error)
- remove the stale OSWE-IAC-SECRETS-LIST-01 suppression (referenced the deleted
  infra/lib/constructs/instance-role.ts)
- repoint the README promotion link to promote-to-main.yml (renamed in #63)

The promote-dev-to-prod.yml comment in check-dev-green.sh is intentionally left
to #63, which rewrites that same line.
2026-06-29 19:54:38 -04:00
Adam Moussa
9444fd7677
docs: document live AWS prod deploy for Open SWE (#53)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Prod went live 2026-06-29 on self-hosted AWS EC2 behind the shared
seahaven-com ALB, superseding the on-prem VM model the runbook described.

- Rewrite deploy/seahaven/DEPLOYMENT.md as the canonical end-to-end runbook:
  infra CD (CDK stacks + OIDC roles + prod approval gate), config seeding
  (put-config.sh, the 13 boot-required prod vars, fetch-config fail-fast),
  app artifact deploy (S3 + SSM roll + is-active gate), promotion/rollback,
  live prod facts, and a RETAIN secret-shell troubleshooting entry that
  cross-references infra/README.md.
- Correct retired *.seahavenind.com hosts to *.seahaven.com throughout and
  document the live GitHub/Slack/Linear webhook + OAuth endpoints.
- Add a concise Deployment section to README pointing at the runbook.
- Fix the stale host in the retired on-prem nginx/openswe.conf and mark it
  superseded by the AMI template.
2026-06-29 11:12:44 -04:00
seahaven-openswe[bot]
f79f49c755
Make PR title repo-aware and auto-link issues (#46)
Some checks are pending
Build & publish app artifacts / Publish + deploy (dev) (push) Waiting to run
Build & publish app artifacts / Publish + deploy (prod) (push) Waiting to run
Infra CD / Infra CI (pre-deploy) (push) Waiting to run
Infra CD / Deploy open-swe-dev (push) Blocked by required conditions
Infra CD / Deploy open-swe-prod (push) Blocked by required conditions
CI / Lint (push) Waiting to run
CI / Format check (push) Waiting to run
CI / Unit tests (push) Waiting to run
CI / Playwright E2E (push) Waiting to run
Hardcoding the Sea Haven no-type-prefix PR title made every PR fail
semantic-PR-title gates (this repo's PR Title Lint, upstream open-swe),
forcing manual retitling. Make the title rule detect a conventional-commit
gate and conform, falling back to the imperative style otherwise. Also add
Closes/Refs issue-linking guidance and the default-branch auto-close caveat.

Refs: #41

Co-authored-by: seahaven-openswe[bot] <296972425+seahaven-openswe[bot]@users.noreply.github.com>
2026-06-27 22:56:01 -04:00
Adam Moussa
f3db9f02e3
Adopt Sea Haven agent conventions, no attribution (#30)
Codify the box-only #4 customizations into Git so the AWS deployment
(which deploys from this repo) actually applies them — previously only
the retired sh-openswe box had them.

- prompt.py: branch names feature|bug|hotfix/<kebab> (optional <KEY->);
  imperative PR titles with no conventional-commit type: prefix; PR body
  Summary/Validation/Tests/Notes; handbook commit format. Rewrite the
  collaboration template from an attribution MANDATE to a PROHIBITION —
  no Co-authored-by bot trailer, no "Made by [Open SWE]" footer, no
  agent/AI notes on any artifact.
- github_comments.py: add @seahaven-openswe (the deployed App slug) to
  the mention triggers.
- authorship.py: remove the now-unused attribution helpers
  (build_pr_attribution_footer, add_bot_coauthor_trailer,
  add_pr_collaboration_note, PR_ATTRIBUTION_*). Keep OPEN_SWE_BOT_* —
  server.py still uses them for the sandbox git identity.
- Flip the attribution unit tests to assert the no-attribution behavior;
  drop tests for the removed helpers.

Commits stay authored as the triggering user for now — flipping
authorship to the bot account depends on the Vercel preview-deploy
constraint and is deferred to #11.

Refs: #4 #11

Claude-Session: https://claude.ai/code/session_01DMhLf4G5V8MStJQyAW95hi
2026-06-27 22:08:45 -04:00
John Kennedy
9db7eab134
feat: Add Corridor MCP analyzePlan integration (#1572)
* Add Corridor MCP analyzePlan integration

* Update agent/integrations/corridor_mcp.py

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>

* Add Corridor analysis prompt

* Only include Corridor prompt when tool loads

---------

Co-authored-by: open-swe[bot] <215916821+open-swe[bot]@users.noreply.github.com>
2026-06-18 14:01:25 -07:00
Johannes du Plessis
f539962c73
feat: server-side Datadog/LangSmith observability tools + team creds [closes OPE-54] (#1476)
* feat: server-side Datadog/LangSmith observability tools + team creds

Add team-wide observability credential settings (Datadog DD_SITE/API/APP
keys, LangSmith API key) stored encrypted server-side, with an admin
dashboard section to connect/disconnect each provider. When connected,
get_agent loads read-only observability tools server-side: Datadog via its
hosted MCP server (langchain-mcp-adapters, toolsets=core) and LangSmith
read tools (langsmith_get_trace, langsmith_list_runs). Credentials live in
the LangGraph server process and are never exposed to the sandbox.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: address review on observability tools

Address PR review feedback:
- Authorize observability tools per triggering user (admins + the
  OBSERVABILITY_AUTHORIZED_EMAILS allowlist) so prompt-injected runs from
  untrusted contributors can't reach team Datadog/LangSmith data.
- Use the documented Datadog MCP auth headers DD_API_KEY / DD_APPLICATION_KEY.
- Store each provider's credentials under its own store key to avoid a
  read-modify-write race dropping the other provider on concurrent saves.

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: async email resolution in observability authorization gate

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-10 11:07:42 -07:00
Johannes du Plessis
0a7fae60cd
docs: document dashboard + refresh local dev setup (#1376)
* docs: document dashboard + refresh local dev setup in INSTALLATION

- add the web dashboard (ui/) and its FastAPI backend API to the setup flow
- document dashboard-login OAuth (GITHUB_APP_CLIENT_ID/SECRET, second callback
  URL) as distinct from the LangSmith-brokered agent-runtime OAuth
- add new env vars: DASHBOARD_API_BASE_URL/BASE_URL/JWT_SECRET/ALLOWED_ORIGINS,
  CONFIGURED_ADMINS, X_SERVICE_AUTH_JWT_SECRET, LANGGRAPH_URL, SANDBOX_TYPE,
  REVIEWER_OUTCOMES_DATASET, PUBLIC_REPO_ORG_GATE, SLACK_CLIENT_ID/SECRET/TEAM_ID,
  SLACK_REPO_OWNER/NAME
- add "Sign in with Slack" OIDC account-linking setup
- rewrite local dev: make dev serves 3 graphs + FastAPI on :2024; new step for
  running the UI (bun) on :3000, incl. the required DASHBOARD_ALLOWED_ORIGINS
  CORS setting and cookie wiring
- correct production langgraph.json to three graphs; document Vercel UI deploy
- add dashboard verify + troubleshooting sections
- README: dashboard feature bullet + updated install blurb

* docs: clarify dashboard API setup

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-02 15:55:35 -07:00
Johannes du Plessis
96774f20ae
feat: move github workflows to gh cli (#1238)
* feat: move github workflows to gh cli

Use LangSmith proxy auth to support gh-driven GitHub workflows while removing custom GitHub wrapper tools.

* docker ignore + snapshot and docker image updates

* updated image and instructions

* removing open_pr if needed after agent call
2026-05-04 18:03:53 -07:00
Johannes du Plessis
448be4a466
feat(open-swe): Default to GPT-5.5 medium reasoning (#1224)
* feat: default to GPT-5.5 medium reasoning

Use OpenAI GPT-5.5 with medium reasoning as the default model and document the completion-token budget semantics for reasoning models.

* fix: use Responses API reasoning config

Pass GPT-5.5 reasoning settings through LangChain's Responses API parameter instead of the Chat Completions-only reasoning_effort field.

* feat: raise GPT-5.5 output budget

Set the default GPT-5.5 output token budget to the model maximum so long-running coding tasks have more room for reasoning and final responses.

* feat: align recursion limit with Deep Agents

Use Deep Agents' default recursion limit so longer coding runs have room to complete without Open SWE imposing a lower cap.

* chore: remove minimal effort level

* chore: reduce max tokens to 64_000

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-04-28 15:03:21 -07:00
Aran Yogesh
2039fe6660
feat: stop auto-cloning and let agent manage repo setup [closes OPE-21] (#1159)
* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* feat: authenticate git operations via sandbox proxy instead of credential files

* removing logger.info

* formatting and linting

* fix: resolve lint errors in server.py (imports, unused vars, undefined names)

* feat: use opaque proxy headers for GitHub auth in sandbox

* linting formatting and test changes

* linting

* Delete .claude directory

* Delete tests/evals directory

* fix: address PR review — guard missing tokens, quote shell paths, add proxy auth tests

* fix: restore authorship, branch_name support, and installation token for PR creation

* linitng

* fix: move installation token fetch before commit, clean up dead proxy validation code

* feat: stop auto-cloning and let agent manage repo setup [closes OPE-21]

* feat: stop auto-cloning and let agent manage repo setup [closes OPE-21]

* fix: address review feedback — restore agents_md, add git user config, lint fixes

* fix: drop github_token arg from sandbox creation, use generic create_sandbox factory with langsmith-only proxy config

* fix: use _get_langsmith_api_key() for prod key fallback, warn when API key missing for proxy config

* linting

* linting

* feat: add installation token auth to list_repos GitHub API call

* agents.md update

* linting

* fix: address PR review feedback — shell precedence bug in prompt, remove dead code

* linting

* Apply suggestion from @bracesproul

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* Apply suggestion from @bracesproul

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* fix: address PR review feedback — restore {working_dir} in prompt, remove clone code block

* fix:Extract check_or_recreate_sandbox utility from inline sandbox health check

* fix: address PR review feedback — async list_repos, restore template name, fix prompt colon

* fix: resolve merge conflicts with main, adopt deepagents v0.5.0a4 LangSmithSandbox

* linting

* yogesh/ope-21-stop-auto-cloning

* Update agent/tools/list_repos.py

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* Update agent/prompt.py

Co-authored-by: Brace Sproul <braceasproul@gmail.com>

* feat: address PR review — list_repos uses GitHub API only, PR trigger includes org/repo

* linting

* feat: address PR review feedback — list_repos pagination, simpler return, sandbox health check

* feat: support listing repos for personal user accounts via is_organization flag

---------

Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-04-10 17:04:55 -07:00
Brace Sproul
0d8647c2cd
fix: Revert dummy readme change (#1132) 2026-03-25 13:35:11 -07:00
Brace Sproul
4cfe2fd8a5
chore: dummy readme change to test committing workflow (#1130)
* chore: dummy readme change to test committing workflow

Co-authored-by: Brace Sproul <46789226+bracesproul@users.noreply.github.com>

* cr

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-03-25 13:34:08 -07:00
Mason Daugherty
d524e3ba92
chore: add badges and tagline to readme header (#1069)
* Add badges, tagline, and ecosystem links to README header

* Match README header syntax exactly to langchain-ai/langchain repo style

* Apply suggestion from @mdrxy

Co-authored-by: Mason Daugherty <github@mdrxy.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-03-17 18:25:48 +00:00
Brace Sproul
633d80bf0e
fix: Add back logo to readme (#1068) 2026-03-17 10:41:41 -07:00
Brace Sproul
82e2bcc94b
chore: Update announcement blog post link in README (#1067) 2026-03-17 10:38:51 -07:00
Harrison Chase
5c493ea1c7 cr 2026-03-07 13:24:10 -08:00
aran-yogesh
abd26fbbaf readme updates 2026-03-06 15:21:38 -08:00
aran-yogesh
229cc08713 readme update 2026-03-06 15:06:16 -08:00
Aran Yogesh
86ea40e04f
Update README.md
Co-authored-by: Brace Sproul <braceasproul@gmail.com>
2026-03-06 14:56:12 -08:00
aran-yogesh
75bc943ac5 order update instructions 2026-03-06 14:47:05 -08:00
aran-yogesh
a6d05a86b2 lineaer webhook update instructions 2026-03-06 14:36:26 -08:00
aran-yogesh
1d19c05b7e chore: readme update 2026-03-06 14:23:29 -08:00
Open SWE Bot
f154609351 docs: Add repository deprecation warning to README
This repository is no longer actively maintained and will not receive further updates. Added a prominent warning notice at the top of the README to inform users that the project has been deprecated.
2026-01-26 01:01:16 +00:00
Open SWE Bot
db8adf2be2 feat: Add deprecation warning for open-swe-max labels
- Add deprecation warning in webhook handler when max labels are used
- Update documentation to indicate open-swe-max labels are deprecated
- Recommend users switch to standard open-swe labels with Opus 4.5
- Update README, best-practices.mdx, github.mdx, and prompts.ts
- Max labels still functional but now log deprecation warnings
2026-01-26 00:56:50 +00:00
Lauren Hirata Singh
bfb2ab7bcc
Fix links and improve README content (#859)
Updated docs links in README.md for accuracy and clarity.
2025-11-13 18:39:29 -08:00
Brace Sproul
8af9ae9e76
docs: Add blog and video links to readme (#687) 2025-08-06 10:48:22 -07:00
Brace Sproul
a7f726df00
feat: Opus 4.1 support, switch default open-swe-max to Opus 4.1 (#678) 2025-08-05 10:34:26 -07:00
Brace Sproul
68be0e816d
fix: readme title wording (#668) 2025-08-05 00:38:21 +00:00
Brace Sproul
cc127000ab
fix: info -> note readme (#666) 2025-08-05 00:02:24 +00:00
Brace Sproul
700bbf20f5
chore: add proper jsonwebtoken deps (#659)
* deps: add proper jsonwebtoken deps

* cr
2025-08-04 13:31:59 -07:00
open-swe[bot]
2c774ee517
feat: Add Max Mode Labels with Claude Opus 4 Support (#577)
* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* cr

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: bracesproul <braceasproul@gmail.com>
2025-07-28 16:33:12 -07:00
Brace Sproul
8e5414ce2b
fix: Typo in readme (#566) 2025-07-28 18:27:21 +00:00
Brace Sproul
f61ea79b13
chore: Cleanup readme (#532)
* chore: Cleanup readme

* cr

* cr

* cr

* cr

* cr
2025-07-25 16:46:52 +00:00
Aliyan Ishfaq
f822fda82d
feat: Added LLM fallback (#433)
* prompt changes and eval fix

* prompt deletions reverted

* formatting

* chore: code cleaning

* feat: add LangGraph evaluation support

* fix: py lg check logging fix

* chore: code cleaning

* dataset added

* chore: code cleaning

* chore: code cleaning

* Readme changes

* feat: LLM fallback added

* feat: add llm fallback

* chore: code cleaning

* chore: code cleaning

* chore: code cleaned

* feat: llm fallback in runtime

* fix: build errors

* refactor: FallbackRunnable now extends ConfigurableModel

* Readme update

* feat: llm as a judge script added

* chore: code formatted

* fix: use GA gemini models instead of preview

* fix: type improvements

* feat: added skeleton for tool py-dev-server

* fix: code cleaning & type improvements

* fix: code cleaning

* cr

---------

Co-authored-by: bracesproul <braceasproul@gmail.com>
2025-07-21 15:54:55 -07:00
Brace Sproul
cf81583473
fix: Cleanup readme, small changes to docs (#458)
* fix: Cleanup readme, small changes to docs

* add formatter to docs

* cr

* cr

* cr

* cr

* cr

* cr

* cr
2025-07-21 21:12:55 +00:00
Brace Sproul
abc984409f
refactor: Proxy req utils, rename GITHUB_TOKEN_ENCRYPTION_KEY to SECRETS_ENCRYPTION_KEY (#391)
* refactor: Proxy req utils, rename GITHUB_TOKEN_ENCRYPTION_KEY to SECRETS_ENCRYPTION_KEY

* cr
2025-07-11 10:48:14 -07:00
Brace Sproul
38dc6ebdf4
fix: Improve DX for getting port for agent (#221)
* fix: Improve DX for getting port for agent

* docs
2025-06-18 18:58:15 +00:00
Brace Sproul
1d0a1cce6f
feat: Clone, commit and push as bot, not user (#174) 2025-06-15 13:41:32 -07:00
Brace Sproul
fda6529419
feat(open-swe): encrypt GitHub tokens in proxy route to prevent exposure in LangGraph runs (#160)
* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* cr

* docs
2025-06-15 12:37:31 -07:00
Brace Sproul
bf72e1faf7
feat: GitHub e2e auth (#106)
* feat: GitHub e2e auth

* cr

* reimplement proxy route

* fix github auth

* cr

* cr

* cr

* cr

* cr
2025-06-11 21:00:08 +00:00
Brace Sproul
8afe1e6803
feat: Update chat UI to be Open SWE specific (#76) (#91)
* init repo-selector

* add branch selector UI, extend GH hook w branch info

* add   reconnectOnMount: true resumeStream, remove repo description from selector, fix styling

* styling repo + branch

* udpate getBranch calls to include targetRepo arg

* format

* add search to repo + branch selectors

* add settings icon link to /github for dev convenience

* add default branch selection

* format

* minor cleanup

* init task list UI demo

* individual task component

* tasks real data

* task visual improvement

* persisting tasks, navigate to thread onClick

* task UI

* TaskSidebar, Navigate to tasks

* org tasks by thread

* group by thread

* improve sidebar, task-list, rm github modal on refresh

* collapse sidebar works when viewing a thread

* remove micro-tasks expansion from threads component, change taskId to taskIndex for persistence + simplicity

* init config sidebar ui

* open and close config sidebar

* add ref for getAllTasks + add graph_id metadata

* format

* create threadItem component

* format

* add threadUtils

* centralize + refactor types

* format

* improve thread status, realtime updates

* fix scroll issue

* remove temperature from config-sidebar

* Change model via sidebar

* fix flashing on updates

* init plan UI

* centralize types

* self CR: refactoring, cleanup, remove comments, add todos

* CR: clean comments + logs, fix repo selector bug

* fix: Update example env file in web (#77)

* disable auto scroll on pageload

* reposelector: use default branch not main master

* remove archived tab temp, default branch listed first in branchSelector dropdown

* auto select first repo

* improve status indicator

* refactor status indicator

* simplify status implementation to use Langgraph

* util moved from provider to thread-utils

* linting, improve task completion logic

* feat: Support followup requests (#60)

* feat: Support followup requests

* cr

* cr

* fix

* account for pr already exists

* remove task provider

* fix: Better prompting and context provided for codebase structure awareness (#85)

* fix: Better prompting and context provided for codebase structure awareness

* cr

* linting + skeleton loader

* fix: branch selector, add github api pagination

* small branch update

* update task completion status realtime

* threadItem stream remove flashing/polling

* feat: Create a shared package (#89)

* feat: Create a shared package

* formatting

* fix: drop passing target repo to getBranchName

* drop docs

* use shared package for config fields in config sidebar

* fix: styling and structure for repo/branch selectors

* fix: use shared types/objects for config fields

* default key for configs

* fix: improved styling for threads

* fix renering num tasks

* disable repo branch selector when chat started

* cr

* cr

* bump deps

* fix: rendering plan in interrupt

* cr

* cr

---------

Co-authored-by: Dylan Boudro <121908331+starmorph@users.noreply.github.com>
2025-06-09 00:29:40 +00:00
Brace Sproul
887fc030a4
feat(agent-generated): Add github oauth (#63)
* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* Apply patch

* format and lint

* fix ui

* Apply patch (#68)

Co-authored-by: Harrison Chase <11986836+hwchase17@users.noreply.github.com>

* fix: Github auth

* update readme

* cr

* cr

* cr

* cr

---------

Co-authored-by: Harrison Chase <11986836+hwchase17@users.noreply.github.com>
Co-authored-by: Harrison Chase <hw.chase.17@gmail.com>
2025-06-02 01:30:29 +00:00