open-swe/agent/tools/__init__.py

157 lines
5.9 KiB
Python
Raw Permalink Normal View History

feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
import sys
from types import ModuleType
from typing import TYPE_CHECKING, Any
_TOOL_MODULES = {
"add_finding": ".add_finding",
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"confluence_comment": ".confluence_comment",
"confluence_create_page": ".confluence_create_page",
"confluence_get_page": ".confluence_get_page",
"confluence_search": ".confluence_search",
"confluence_update_page": ".confluence_update_page",
feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
"enter_plan_mode": ".enter_plan_mode",
"fetch_url": ".fetch_url",
"http_request": ".http_request",
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"jira_comment": ".jira_comment",
"jira_create_issue": ".jira_create_issue",
"jira_get_issue": ".jira_get_issue",
"jira_get_issue_comments": ".jira_get_issue_comments",
"jira_list_projects": ".jira_list_projects",
"jira_update_issue": ".jira_update_issue",
feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
"linear_comment": ".linear_comment",
"linear_create_issue": ".linear_create_issue",
"linear_delete_issue": ".linear_delete_issue",
"linear_get_issue": ".linear_get_issue",
"linear_get_issue_comments": ".linear_get_issue_comments",
"linear_list_teams": ".linear_list_teams",
"linear_search_issues": ".linear_search_issues",
feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
"linear_update_issue": ".linear_update_issue",
"list_findings": ".list_findings",
"list_review_findings": ".list_review_findings",
"open_pull_request": ".open_pull_request",
"publish_review": ".publish_review",
"read_repo_file": ".read_repo_file",
"report_platform_issue": ".report_platform_issue",
"request_pr_review": ".request_pr_review",
"reply_to_finding_thread": ".reply_to_finding_thread",
"resolve_finding_thread": ".resolve_finding_thread",
"save_plan": ".save_plan",
"schedule_thread_wakeup": ".schedule_thread_wakeup",
"search_repo_code": ".search_repo_code",
"slack_add_reaction": ".slack_add_reaction",
"slack_read_thread_messages": ".slack_read_thread_messages",
"slack_start_new_thread": ".slack_start_new_thread",
"slack_thread_reply": ".slack_thread_reply",
"update_finding": ".update_finding",
"web_search": ".web_search",
}
__all__ = [
feat: implement reviewer findings, publish_review, and watch mode (#1253) * feat: implement reviewer findings, publish_review, and watch mode Build out the reviewer agent end-to-end against the design in REVIEWER_DESIGN.md: - Findings as first-class state on the reviewer thread metadata (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line ranges, suggestion text for ```suggestion blocks, github_review_comment_id for cross-run reconciliation, diff_hunk for UI rendering. Thread-level metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a future frontend can list reviewer threads via the langgraph SDK. - Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff, compute_diff_line_set for in-diff validation, extract_diff_hunk for caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA diffs against the prepped repo. - Tools: `add_finding` (validates against the diff line set so out-of-diff ranges fail at creation, not at GitHub-publish), `update_finding`, `list_findings`, `publish_review`. The reviewer agent's tool list is swapped from `[]` (direct shell `gh api` calls) to these four. - Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`): one POST /reviews call with body + inline comments + ```suggestion blocks, per-comment IDs stored back on findings, GraphQL `resolveReviewThread` fired for findings transitioning open->resolved on a re-review. - Reviewer graph: deterministic clone-or-fetch + checkout in the factory before the agent's first model call (warm- and cold-path symmetric); computed diff and in-diff line set passed via runnable config; system prompt rewritten for the single-evolving-findings model, severity ladder, in-diff-only discipline, and watch-mode reconciliation flow. - Watch mode in webapp.py: `push` event + `pull_request` closed/reopened added to supported events. New `process_github_push_event` resolves the open PR for the pushed branch, gates on the reviewer thread's `watch` flag, builds a re-review configurable, and triggers a run on the same canonical thread. `process_github_pr_close` toggles watch on closed/reopened. `set_reviewer_thread_metadata` is called on first review to install `kind=reviewer` + PR identity + watch=True. - Eval harness: target.py now extracts `add_finding` calls (mapped to the legacy {file, line, body, severity} shape the judge expects) and passes the right configurable so the prep step has base/head SHAs. - Tests: new unit suites for findings helpers, diff parsing, finding tools, publish rendering + GraphQL resolve, and watch-mode webhook handlers (push triggers re-review only when watching, idempotent on unchanged head SHA, PR close disables watch). Updated existing reviewer-webhook tests to mock `set_reviewer_thread_metadata`. - REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context into REVIEWER_DESIGN.md. * fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL Address PR #1253 review findings: - compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag (`option no-prefix takes no value` — every prep run was failing silently and the agent saw an empty diff). - compute_diff_in_sandbox grew a `merge_base` flag. First-review path now uses three-dot `base...head` (the merge-base diff GitHub renders on Files-changed) so we don't pick up changes that landed on the base branch after the PR diverged. Re-review delta keeps two-dot `last_reviewed_sha..head` since that's exactly the new commits. - publish_review skips findings that already carry `github_review_comment_id`. Without this, watched re-reviews re-posted every previously surfaced finding, and only the most-recent duplicate's id would later resolve when the issue got addressed. - fetch_review_comments URL now includes `{pull_number}` — `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments` is the canonical endpoint; the old form 404s, so comment ids were never stored and watch-mode resolution couldn't run. Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix` flag in the executed command, and that publish_review does not re-post findings whose `github_review_comment_id` is set. * fix(reviewer): default publish cap from 15 to 4 A clean PR with one critical issue padded out by three lower-severity findings is fine; fifteen is review spam. The agent can override per call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
"add_finding",
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"confluence_comment",
"confluence_create_page",
"confluence_get_page",
"confluence_search",
"confluence_update_page",
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
"enter_plan_mode",
"fetch_url",
"http_request",
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
"jira_comment",
"jira_create_issue",
"jira_get_issue",
"jira_get_issue_comments",
"jira_list_projects",
"jira_update_issue",
"linear_comment",
"linear_create_issue",
"linear_delete_issue",
"linear_get_issue",
"linear_get_issue_comments",
"linear_list_teams",
"linear_search_issues",
"linear_update_issue",
feat: implement reviewer findings, publish_review, and watch mode (#1253) * feat: implement reviewer findings, publish_review, and watch mode Build out the reviewer agent end-to-end against the design in REVIEWER_DESIGN.md: - Findings as first-class state on the reviewer thread metadata (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line ranges, suggestion text for ```suggestion blocks, github_review_comment_id for cross-run reconciliation, diff_hunk for UI rendering. Thread-level metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a future frontend can list reviewer threads via the langgraph SDK. - Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff, compute_diff_line_set for in-diff validation, extract_diff_hunk for caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA diffs against the prepped repo. - Tools: `add_finding` (validates against the diff line set so out-of-diff ranges fail at creation, not at GitHub-publish), `update_finding`, `list_findings`, `publish_review`. The reviewer agent's tool list is swapped from `[]` (direct shell `gh api` calls) to these four. - Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`): one POST /reviews call with body + inline comments + ```suggestion blocks, per-comment IDs stored back on findings, GraphQL `resolveReviewThread` fired for findings transitioning open->resolved on a re-review. - Reviewer graph: deterministic clone-or-fetch + checkout in the factory before the agent's first model call (warm- and cold-path symmetric); computed diff and in-diff line set passed via runnable config; system prompt rewritten for the single-evolving-findings model, severity ladder, in-diff-only discipline, and watch-mode reconciliation flow. - Watch mode in webapp.py: `push` event + `pull_request` closed/reopened added to supported events. New `process_github_push_event` resolves the open PR for the pushed branch, gates on the reviewer thread's `watch` flag, builds a re-review configurable, and triggers a run on the same canonical thread. `process_github_pr_close` toggles watch on closed/reopened. `set_reviewer_thread_metadata` is called on first review to install `kind=reviewer` + PR identity + watch=True. - Eval harness: target.py now extracts `add_finding` calls (mapped to the legacy {file, line, body, severity} shape the judge expects) and passes the right configurable so the prep step has base/head SHAs. - Tests: new unit suites for findings helpers, diff parsing, finding tools, publish rendering + GraphQL resolve, and watch-mode webhook handlers (push triggers re-review only when watching, idempotent on unchanged head SHA, PR close disables watch). Updated existing reviewer-webhook tests to mock `set_reviewer_thread_metadata`. - REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context into REVIEWER_DESIGN.md. * fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL Address PR #1253 review findings: - compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag (`option no-prefix takes no value` — every prep run was failing silently and the agent saw an empty diff). - compute_diff_in_sandbox grew a `merge_base` flag. First-review path now uses three-dot `base...head` (the merge-base diff GitHub renders on Files-changed) so we don't pick up changes that landed on the base branch after the PR diverged. Re-review delta keeps two-dot `last_reviewed_sha..head` since that's exactly the new commits. - publish_review skips findings that already carry `github_review_comment_id`. Without this, watched re-reviews re-posted every previously surfaced finding, and only the most-recent duplicate's id would later resolve when the issue got addressed. - fetch_review_comments URL now includes `{pull_number}` — `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments` is the canonical endpoint; the old form 404s, so comment ids were never stored and watch-mode resolution couldn't run. Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix` flag in the executed command, and that publish_review does not re-post findings whose `github_review_comment_id` is set. * fix(reviewer): default publish cap from 15 to 4 A clean PR with one critical issue padded out by three lower-severity findings is fine; fifteen is review spam. The agent can override per call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
"list_findings",
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
"list_review_findings",
"open_pull_request",
feat: implement reviewer findings, publish_review, and watch mode (#1253) * feat: implement reviewer findings, publish_review, and watch mode Build out the reviewer agent end-to-end against the design in REVIEWER_DESIGN.md: - Findings as first-class state on the reviewer thread metadata (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line ranges, suggestion text for ```suggestion blocks, github_review_comment_id for cross-run reconciliation, diff_hunk for UI rendering. Thread-level metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a future frontend can list reviewer threads via the langgraph SDK. - Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff, compute_diff_line_set for in-diff validation, extract_diff_hunk for caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA diffs against the prepped repo. - Tools: `add_finding` (validates against the diff line set so out-of-diff ranges fail at creation, not at GitHub-publish), `update_finding`, `list_findings`, `publish_review`. The reviewer agent's tool list is swapped from `[]` (direct shell `gh api` calls) to these four. - Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`): one POST /reviews call with body + inline comments + ```suggestion blocks, per-comment IDs stored back on findings, GraphQL `resolveReviewThread` fired for findings transitioning open->resolved on a re-review. - Reviewer graph: deterministic clone-or-fetch + checkout in the factory before the agent's first model call (warm- and cold-path symmetric); computed diff and in-diff line set passed via runnable config; system prompt rewritten for the single-evolving-findings model, severity ladder, in-diff-only discipline, and watch-mode reconciliation flow. - Watch mode in webapp.py: `push` event + `pull_request` closed/reopened added to supported events. New `process_github_push_event` resolves the open PR for the pushed branch, gates on the reviewer thread's `watch` flag, builds a re-review configurable, and triggers a run on the same canonical thread. `process_github_pr_close` toggles watch on closed/reopened. `set_reviewer_thread_metadata` is called on first review to install `kind=reviewer` + PR identity + watch=True. - Eval harness: target.py now extracts `add_finding` calls (mapped to the legacy {file, line, body, severity} shape the judge expects) and passes the right configurable so the prep step has base/head SHAs. - Tests: new unit suites for findings helpers, diff parsing, finding tools, publish rendering + GraphQL resolve, and watch-mode webhook handlers (push triggers re-review only when watching, idempotent on unchanged head SHA, PR close disables watch). Updated existing reviewer-webhook tests to mock `set_reviewer_thread_metadata`. - REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context into REVIEWER_DESIGN.md. * fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL Address PR #1253 review findings: - compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag (`option no-prefix takes no value` — every prep run was failing silently and the agent saw an empty diff). - compute_diff_in_sandbox grew a `merge_base` flag. First-review path now uses three-dot `base...head` (the merge-base diff GitHub renders on Files-changed) so we don't pick up changes that landed on the base branch after the PR diverged. Re-review delta keeps two-dot `last_reviewed_sha..head` since that's exactly the new commits. - publish_review skips findings that already carry `github_review_comment_id`. Without this, watched re-reviews re-posted every previously surfaced finding, and only the most-recent duplicate's id would later resolve when the issue got addressed. - fetch_review_comments URL now includes `{pull_number}` — `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments` is the canonical endpoint; the old form 404s, so comment ids were never stored and watch-mode resolution couldn't run. Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix` flag in the executed command, and that publish_review does not re-post findings whose `github_review_comment_id` is set. * fix(reviewer): default publish cap from 15 to 4 A clean PR with one critical issue padded out by three lower-severity findings is fine; fifteen is review spam. The agent can override per call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
"publish_review",
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
"read_repo_file",
"report_platform_issue",
"request_pr_review",
"reply_to_finding_thread",
"resolve_finding_thread",
feat: plan mode with model-driven entry and collaborative review (#1580) * feat: add plan mode for read-only research and planning Adds a per-run plan_mode flag that puts the agent in a read-only research phase: a strong prompt section is injected and mutating tools are stripped via ExcludeToolsMiddleware so the agent proposes a reviewable implementation plan before any edits. Surfaced in the dashboard UI with a Plan toggle (Shift+Tab) wired through the thread API. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: enforce plan-mode read-only at tool layer and disable subagents Addresses PR review: plan mode previously relied on prompt text to keep the shell read-only and left the task subagent (built with its own write/PR/Linear tools) unrestricted. Now `task` is excluded so research cannot be delegated to a mutating subagent, and a new PlanModeShellGuardMiddleware enforces a read-only command allowlist on `execute`, blocking writes, git state changes, installs, redirection, and command substitution regardless of model/prompt-injection compliance. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * fix: harden plan-mode shell guard against wrapped mutations Block git global options that take values (-C, --git-dir, ...) from being misread as the subcommand, reject config-injection options (-c, --config-env, --exec-path), and drop the env command wrapper that could run arbitrary commands. Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * feat: add plan mode with enter_plan_mode tool, profile/team defaults, Slack commands and approval flow - enter_plan_mode tool: agent self-activates plan mode via Command(update={'plan_mode': True}) - Plan mode resolution: per-thread > profile default > team default > False - PLAN_MODE_GUIDANCE_SECTION: always-present prompt section telling agent about the tool - profile_plan_mode_default and team plan_mode_default settings - Slack plan on/off/status commands with thread metadata persistence - slack_thread_reply plan_approval=True renders Approve/Revise/Cancel buttons - Interactivity handler: approve triggers implementation run, cancel posts confirmation - Frontend: plan_mode_default in Profile/ProfileUpdate/TeamSettings types and UI toggles Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * test: add tests for enter_plan_mode tool, profile/team defaults, Slack plan commands, approval blocks Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com> * refactor(plan-mode): drop shell guard, rely on prompt for read-only discipline Remove PlanModeShellGuardMiddleware and its enforcement of read-only shell commands during plan mode. Plan mode now relies on the system prompt to instruct the agent not to run mutating commands; the mutating-tool exclusion (ExcludeToolsMiddleware) is retained. * test(open-swe): add Playwright E2E for the Slack → PR → web handoff Local, secrets-free end-to-end suite that drives the full happy path through mock Slack/GitHub control panels and the real dashboard UI. Only the LLM and external SaaS HTTP boundaries (GitHub/Slack APIs, OAuth token mint) are faked — the real process_slack_mention, get_agent, deepagents loop, tools, middleware, and dashboard authorization all run under `langgraph dev` with a scripted fake chat model and a local temp-dir sandbox. - full_flow: a Slack mention runs the agent, which implements a change in the sandbox, opens a PR against a fake GitHub remote, and replies with the PR link in the same thread. - dashboard: clicking the bot's real "Open in Web" link loads the built ui/ app (served same-origin); the thread owner can continue the conversation, while a different user sees the same thread read-only (no composer). Wired into Agent CI as a `Playwright E2E` job that runs on pull requests. * fix(open-swe): serve E2E UI assets via explicit route; pin Playwright The dashboard E2E served the built ui/ SPA's /assets via app.mount(StaticFiles), but LangGraph's custom-app loader serves APIRoutes and drops sub-app Mounts, so /assets 404'd under `langgraph dev` in CI — the React app never booted and the composer/transcript never rendered. Serve assets via an explicit route instead. Also pin @playwright/test to the latest (1.61.0) for reproducible runs, and make the owner composer assertion tolerant of either hydration state. * test(open-swe): record Playwright trace + video on every E2E run Capture a replayable trace (DOM snapshots, network, console, source) and a screen recording for every test, not just retries, plus a screenshot on failure. The CI job already uploads playwright-report/ and test-results/, so each run now has a downloadable replay; documented how to open it. * feat(plan-mode): collaborative plan review with BlockNote + Yjs When the agent enters plan mode it writes the plan as a markdown file in the sandbox (save_plan tool), publishes it, and posts a review link to the source channel. Reviewers open the plan inside the dashboard (under the /agents shell), read it rendered in a BlockNote editor, and leave inline comments synced live over Yjs. Only the thread owner can approve; any reviewer can request changes. On approve/reject the comments are harvested and handed to the agent for the follow-up run; the agent never sees comments mid-review. - agent: enter_plan_mode persists plan state; new save_plan tool; prompt shares the plan-review link. - dashboard: Yjs WebSocket collab server (pycrdt-websocket) with store-backed snapshots; plan content/status store; plan REST API (get/approve/reject, owner-only approve, client-harvested comments); planStatus on thread summaries. - ui: BlockNote native comments (CommentsExtension + YjsThreadStore) plan page mounted under the agents shell, with a "Review plan" banner in the thread view and a back-link; theme-aware (dark mode) using the dashboard tokens. - e2e: Playwright coverage of the full Slack -> plan -> review -> approve -> PR flow, including cross-user comment sync and owner-only approval. * fix(plan-mode): address review feedback (authz, overrides, leaks, deps) - plan-collab WS: authorize per-thread before joining a room (same read gate as the REST API) — previously any logged-in user could join any thread (IDOR). - plan-collab: tie the snapshot flusher to active connections (refcount) so each opened plan no longer leaks a permanent 1.5s task on the shared event loop. - plan decisions: include thread_id in the follow-up run configurable so the run resumes the existing thread; set plan_mode explicitly so approve forces it off. - get_agent: an explicit per-thread plan_mode (Slack `plan off`, approved plan, dashboard toggle) now overrides profile/team defaults instead of falling back. - plan mode tool gating moved to a state-aware PlanModeMiddleware installed unconditionally, so a mid-run enter_plan_mode restricts the next model turn; before_agent resets stale plan_mode so a later run isn't forced back into it. - exclude write-capable http_request from plan mode. - pin pycrdt / pycrdt-websocket with upper bounds. Includes the latest base (#1583): E2E UI assets served via explicit route (fixes the Playwright CI failure — LangGraph's app loader drops sub-app mounts). * style: ruff format plan_collab.py * fix(plan-mode): owner-gate Slack approval + same-origin check on collab WS - Slack "Approve & Implement" now verifies the clicking user is the plan requester (owner, via the stored triggering_user_id) before implementing — matching the dashboard API's owner-only approval. Non-owners are pointed to Revise / feedback. - The plan-collab WebSocket validates the handshake Origin against the dashboard allowlist before accept() (no-op when unconfigured, e.g. local/dev), mirroring the REST require_same_origin CSRF defense. * fix(plan-mode): enter plan mode only via the model + local mock dev harness Plan mode is now entered solely when the model calls enter_plan_mode. Removed the per-user and team plan_mode_default settings (backend + UI) and the Slack `plan on/off/status` toggle. - enter_plan_mode returns a terminating ToolMessage, fixing the missing ToolMessage error that silently dropped plan mode mid-run. - PlanReview: defer Yjs provider/doc teardown so React StrictMode's dev remount doesn't destroy and then reuse the collaboration provider. - e2e plan_review spec asserts plan_mode actually engages. - LangSmith trace-url resolution is best-effort: bail before any API call when the tenant is unset, cache failures, log at debug. - Add `pnpm run dev:mock`: same-origin Vite HMR harness with a real LLM, Alice/Bob mock users, and a GitHub login picker. * docs(plan-mode): drop stale references to removed profile/team defaults The plan_mode middleware docstring and the approve/reject dispatch comment still described the profile/team plan_mode_default resolution that no longer exists; reword to match model-driven entry + the per-thread carry. * feat(plan-mode): let any reviewer edit the plan, not just comment Drop the owner/commenter split for the plan document: everyone with read access edits and comments alike (DefaultThreadStoreAuth "editor" for all, editor always editable until a decision, anyone seeds the empty doc). This matches the collab WS, which already relays frames to every readable user. Plan approval stays owner-gated. * test(plan-mode): assert plan-mode entry via the tool's success message plan_mode lives only in run state for tool gating; it is not a persisted thread-state channel, so the previous `values.plan_mode === true` poll could never pass. Assert instead that enter_plan_mode's success ToolMessage ("Plan mode is active …") lands in the thread — which only happens when the tool's Command applies cleanly, the exact regression this guards. --------- Co-authored-by: Johannes du Plessis <51395795+johannes117@users.noreply.github.com> Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-23 15:06:58 -04:00
"save_plan",
"schedule_thread_wakeup",
feat: chat with your PR on the review page (#1534) * feat: chat with your PR on the review page Add a sandbox-less `chat` graph that answers questions about a single PR from its diff, the published review findings, and read-only GitHub access. - agent/chat.py: deepagents graph, no sandbox (default StateBackend, file mutation + execute tools excluded). PR context is seeded as virtual files under /pr/; a repo-scoped App token is resolved in-graph. - tools: read_repo_file, search_repo_code, list_review_findings. - dashboard/review_chat_api.py + routes: per-user chat thread, LangGraph stream/commands/state/history proxy pinned to the chat assistant, seeds diff/findings/overview on first run. Gated by repo access. - UI: Chat tab wired to a chat-scoped StreamProvider (replaces Coming Soon). * feat: admin setting for review-chat default model Add a 'Open SWE Review Chat' default to team settings (default_chat_model / default_chat_reasoning_effort). get_team_default_model("chat") inherits the Agent default when unset; the chat graph resolves through it. Admin RolePicker gains an 'Agent default' inherit option that clears the override. * feat: multi-conversation review chat (tabs, new chat, history) Replace the single per-PR chat thread with multiple per-user conversations: - threads minted client-side; first message persists with a title derived from the prompt. - list + delete endpoints; chat panel gets a tab strip (history), new-chat (+), close (x), refresh, an intro greeting, and suggested prompts. - get_review_chat now returns availability only (ids are client-minted). * ui fixes * ui: review-chat history dropdown, full-width AI replies, resizable side panel * fix(review-chat): enforce per-user thread ownership on proxy endpoints; reseed PR context on head change --------- Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-15 17:17:30 -07:00
"search_repo_code",
"slack_add_reaction",
"slack_read_thread_messages",
"slack_start_new_thread",
"slack_thread_reply",
feat: implement reviewer findings, publish_review, and watch mode (#1253) * feat: implement reviewer findings, publish_review, and watch mode Build out the reviewer agent end-to-end against the design in REVIEWER_DESIGN.md: - Findings as first-class state on the reviewer thread metadata (`agent/reviewer_findings.py`): Finding TypedDict with start_line/end_line ranges, suggestion text for ```suggestion blocks, github_review_comment_id for cross-run reconciliation, diff_hunk for UI rendering. Thread-level metadata gets `kind=reviewer`, `pr`, `last_reviewed_sha`, `watch` so a future frontend can list reviewer threads via the langgraph SDK. - Diff utilities (`agent/reviewer_diff.py`): parse_unified_diff, compute_diff_line_set for in-diff validation, extract_diff_hunk for caching the hunk on a Finding, compute_diff_in_sandbox for SHA-to-SHA diffs against the prepped repo. - Tools: `add_finding` (validates against the diff line set so out-of-diff ranges fail at creation, not at GitHub-publish), `update_finding`, `list_findings`, `publish_review`. The reviewer agent's tool list is swapped from `[]` (direct shell `gh api` calls) to these four. - Publish path (`agent/reviewer_publish.py` + `agent/tools/publish_review.py`): one POST /reviews call with body + inline comments + ```suggestion blocks, per-comment IDs stored back on findings, GraphQL `resolveReviewThread` fired for findings transitioning open->resolved on a re-review. - Reviewer graph: deterministic clone-or-fetch + checkout in the factory before the agent's first model call (warm- and cold-path symmetric); computed diff and in-diff line set passed via runnable config; system prompt rewritten for the single-evolving-findings model, severity ladder, in-diff-only discipline, and watch-mode reconciliation flow. - Watch mode in webapp.py: `push` event + `pull_request` closed/reopened added to supported events. New `process_github_push_event` resolves the open PR for the pushed branch, gates on the reviewer thread's `watch` flag, builds a re-review configurable, and triggers a run on the same canonical thread. `process_github_pr_close` toggles watch on closed/reopened. `set_reviewer_thread_metadata` is called on first review to install `kind=reviewer` + PR identity + watch=True. - Eval harness: target.py now extracts `add_finding` calls (mapped to the legacy {file, line, body, severity} shape the judge expects) and passes the right configurable so the prep step has base/head SHAs. - Tests: new unit suites for findings helpers, diff parsing, finding tools, publish rendering + GraphQL resolve, and watch-mode webhook handlers (push triggers re-review only when watching, idempotent on unchanged head SHA, PR close disables watch). Updated existing reviewer-webhook tests to mock `set_reviewer_thread_metadata`. - REVIEWER_EVAL_PLAN.md removed per user request; folded relevant context into REVIEWER_DESIGN.md. * fix(reviewer): correct git diff flags, scope, dedup, and review-comments URL Address PR #1253 review findings: - compute_diff_in_sandbox dropped the invalid `--no-prefix=false` flag (`option no-prefix takes no value` — every prep run was failing silently and the agent saw an empty diff). - compute_diff_in_sandbox grew a `merge_base` flag. First-review path now uses three-dot `base...head` (the merge-base diff GitHub renders on Files-changed) so we don't pick up changes that landed on the base branch after the PR diverged. Re-review delta keeps two-dot `last_reviewed_sha..head` since that's exactly the new commits. - publish_review skips findings that already carry `github_review_comment_id`. Without this, watched re-reviews re-posted every previously surfaced finding, and only the most-recent duplicate's id would later resolve when the issue got addressed. - fetch_review_comments URL now includes `{pull_number}` — `/repos/{owner}/{repo}/pulls/{pr_number}/reviews/{review_id}/comments` is the canonical endpoint; the old form 404s, so comment ids were never stored and watch-mode resolution couldn't run. Three new tests cover: three-dot vs two-dot wiring, no `--no-prefix` flag in the executed command, and that publish_review does not re-post findings whose `github_review_comment_id` is set. * fix(reviewer): default publish cap from 15 to 4 A clean PR with one critical issue padded out by three lower-severity findings is fine; fifteen is review spam. The agent can override per call when a PR genuinely warrants more.
2026-05-07 14:48:43 -07:00
"update_finding",
"web_search",
]
feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
if TYPE_CHECKING:
from .add_finding import add_finding
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
from .confluence_comment import confluence_comment
from .confluence_create_page import confluence_create_page
from .confluence_get_page import confluence_get_page
from .confluence_search import confluence_search
from .confluence_update_page import confluence_update_page
feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
from .enter_plan_mode import enter_plan_mode
from .fetch_url import fetch_url
from .http_request import http_request
feat: Jira + Confluence integration (tools + triggers) (#182) * feat(open-swe): add Jira tool plane (Phase 1) Curated Jira Cloud REST v3 toolset for the agent, mirroring the Linear tools: - utils/jira.py: service-account REST client (Basic auth) with get/ create/update issue, comments, list projects, trace comment; issue and comment bodies normalized to markdown. - utils/adf.py: minimal ADF <-> markdown conversion (read paths convert Jira ADF to markdown; agent comments convert prose to ADF). - tools/jira_{comment,get_issue,get_issue_comments,create_issue, update_issue,list_projects}.py wired into the tool registry and the main agent tool list. - tests/test_jira_utils.py: ADF conversion + mocked-transport util tests. Reads JIRA_BASE_URL / JIRA_SERVICE_EMAIL / JIRA_API_TOKEN; unset env returns a clean error, so this is safe to land dark. Trigger plane, prompt guidance, and config plumbing follow in Phase 2. * feat(open-swe): add Confluence tool plane (Phase 3) Curated Confluence Cloud REST toolset for the agent, mirroring the Jira tools: - utils/confluence.py: service-account REST client (Basic auth) with get/create/update page, add comment, CQL search. Page bodies are XHTML storage format (not ADF), with minimal storage<->text converters; update_page reads the current version and bumps it, as Confluence requires. - tools/confluence_{get_page,create_page,update_page,comment,search}.py registered in the tool registry. - tests/test_confluence_utils.py: converter + mocked-transport tests including the version-bump path. Reads CONFLUENCE_BASE_URL / CONFLUENCE_EMAIL / CONFLUENCE_API_TOKEN; unset env returns a clean error. Activation in the agent tool list lands with the Phase 2 server.py wiring. * feat(open-swe): add Jira trigger plane (Phase 2) Make an @openswe comment on a Jira issue spawn an agent run, mirroring the Linear trigger plane: - webhooks/jira.py: process_jira_issue clones process_linear_issue — deterministic thread id, full-issue fetch, actor accountId->email attribution feeding resolve_login_from_email_async (PRs open as the human), multimodal image handling, source="jira" + jira_issue config. - webapp.py: POST/GET /webhooks/jira, verify_jira_secret (constant-time X-Automation-Webhook-Token check, fails closed), repo-resolution cascade, get_repo_config_from_jira_mapping. - utils/jira_project_repo_map.py: JIRA_PROJECT_TO_REPO (placeholder entry — real project->repo mappings still needed). - utils/jira.py: get_user_email (accountId -> email) for attribution. - completion.py: source=="jira" failure-reply branch. - prompt.py: Jira-triggered notify guidance + Refs:/branch key from {jira_project_key}-{jira_issue_number}. - server.py: read jira_issue config + pass jira key to the system prompt; also activates the Phase 3 Confluence tools in the agent list. Jira Automation lacks native webhook HMAC signing, so trust is a shared secret header (decision D2); replay protection is weaker than Linear's HMAC+timestamp. /sh-security-review + an Atlassian IP allowlist are the outstanding gate/hardening before push. * fix(open-swe): harden Jira webhook trust (sh-security-review) Resolves findings from the Phase 2 security review (detector fan-out + proof-or-kill verifier). The unsigned Jira Automation webhook body was trusted for identity, comment content, repo routing, and issue existence; a JIRA_WEBHOOK_SECRET holder could forge those fields. - Corroborate against the real Jira record: the webhook body is now only a pointer (issue_key + required comment_id). The triggering comment's author and text are re-fetched server-side via get_comment/fetch_jira_ comment, and identity, the @openswe check, prompt text, and project key are derived from that authoritative record — never payload author/ body fields. An uncorroborated comment is rejected. (closes the account-id impersonation, unsigned-body prompt injection, and fabricated-issue findings) - Validate issue_key against the Jira key format and percent-encode all untrusted path segments (_seg) so a crafted key can't traverse to a different Jira REST endpoint or inject query params. (closes the path- traversal / query-injection findings) - Route source=="jira" through the bot-token-default / author_prs_as_ user opt-in path in resolve_github_token, matching Linear, instead of unconditionally resolving a per-user OAuth token from a payload email. - Gate attribution on an active user mapping (is_login_mapped) so a pending/unconfirmed mapping can't drive PR authorship. Adds regression tests: server-corroboration wins over payload, malformed issue_key rejected, uncorroborated comment rejected, path-segment encoding, project-key derivation, active-mapping gate. Remaining (non-blocking, deployment/hardening): set ALLOWED_GITHUB_ORGS/ REPOS so the shared allowlist isn't fail-open; consider HMAC-over-body + timestamp on the Automation payload to close the residual replay gap. * harden(open-swe): opt-in Jira webhook replay/IP + fail-closed allowlist Folds the two deployment-hardening items from the Phase 2 security review into code (all opt-in / default-off, so existing and upstream deployments are unaffected): - JIRA_WEBHOOK_REQUIRE_SIGNATURE: when set, the Automation payload must carry X-Openswe-Signature (hex HMAC-SHA256 of the raw body keyed by JIRA_WEBHOOK_SECRET) plus a fresh timestamp, verified by verify_jira_signature / _jira_timestamp_is_fresh (mirrors the Linear HMAC+freshness model). Closes the static-token model's replay/forgery gap when enabled. - JIRA_WEBHOOK_IP_ALLOWLIST: optional CIDR allowlist on the webhook's direct client IP (verify_jira_source_ip). Documented as direct-peer only; behind a proxy/LB, allowlist Atlassian's ranges at that layer. - REQUIRE_REPO_ALLOWLIST: makes an empty ALLOWED_GITHUB_ORGS/REPOS fail CLOSED instead of the back-compat allow-all, plus a startup fail-open warning. Applies to all channels for consistency. Documents all new vars (and a Jira section) in .env.example. Adds tests for signature on/off + valid/missing/wrong/stale, IP allow/deny/off, and the fail-closed allowlist. * feat(open-swe): Confluence Atlassian Connect trigger (Phase 4) Adds the @openswe-on-a-Confluence-comment trigger via a private Atlassian Connect app. Designed and adversarially verified with the ultracode workflow (3 divergent Opus designs + judge; 3 proof-or-kill Opus skeptics on the implemented crypto). - utils/atlassian_connect.py: hand-rolled qsh (pinned to Atlassian's official test vector), PyJWT HS256 webhook verifier with alg-pinning, issuer binding, and qsh-verified-last ordering; RS256 signed-install lifecycle verifier against Atlassian's published keys; installation store keyed by clientKey with the sharedSecret encrypted at rest (TOKEN_ENCRYPTION_KEY / Fernet). No new dependency (PyJWT already pinned). - webhooks/confluence.py: install/uninstall lifecycle + comment handler. The JWT-signed webhook body is only a pointer; the comment's real author/text/container are re-fetched server-side via the Basic-auth service account (Phase-2 corroboration lesson), with active-only login attribution and the repo allowlist. - utils/confluence.py: get_comment / get_user_email (path-encoded). - webapp.py: GET /connect/atlassian-connect.json (served dynamically), POST /connect/{installed,uninstalled,webhook/comment-created}, the space->repo resolver, thread-id, and fetch helpers. - completion.py: source=="confluence" failure-reply branch. Security: the sh-security-review verify pass confirmed one HIGH — the symmetric signed-install=false first-install was trust-on-first-use gated only by the public Confluence hostname (webhook-auth bypass). Fixed by switching to signed-install=true + RS256 verification of lifecycle callbacks, which cryptographically authenticates the first install. All other attack lenses (forgery/replay/alg-confusion/overwrite/uninstall DoS/corroboration/injection) were defeated; residuals are deployment config (REQUIRE_REPO_ALLOWLIST) or accepted-by-design (qsh cannot cover bodies; comment-trigger prompt injection, shared with all sources). New env (documented in .env.example): CONFLUENCE_BASE_URL/EMAIL/API_TOKEN, CONNECT_BASE_URL, CONNECT_EXPECTED_BASE_URL (optional). Install secrets require the durable Postgres LangGraph store in prod. Outstanding before push: /sh-security-review on the real diff and the GPT-4.1 cross-family review (auth boundary); README/CLAUDE.md + memory. * docs(open-swe): Phase 5 — Confluence prompt guidance + architecture docs - prompt.py: Confluence-triggered runs notify via confluence_comment on the triggering page; add Confluence to the shared-base source list. - CLAUDE.md: document the Jira + Confluence tool planes and the Atlassian triggers (Jira Automation shared-secret webhook; Confluence Connect app with HS256 webhook + qsh and RS256 signed-install lifecycle), plus the server-side corroboration + encrypted install store. Phase 5 also verified the trigger surface end-to-end against a running uvicorn app (descriptor served; /connect/* and /webhooks/jira fail closed without valid auth) and recorded the integration in project memory. * fix(open-swe): resolve /sh-security-review findings on the Atlassian surface Formal sh-security-review (detector fan-out + verifier) over the Phase-4 Connect surface (esp. the new RS256 signed-install code, unseen by the earlier adversarial verify) and the Phase-2 opt-in hardening. CRITICAL — cross-tenant install (origin validation, CWE-346): signed- install proves the caller is *an* Atlassian tenant, not *ours*, and the descriptor is served publicly, so any attacker could install the app on their own Confluence site and drive agent runs against our allowlisted repos. The baseUrl body field is attacker-controlled and cannot bind the tenant; only the signature-verified clientKey (JWT iss) can. Added a MANDATORY, fail-closed CONNECT_EXPECTED_CLIENT_KEYS allowlist checked in process_install after signature+iss verification. HIGH — cross-tenant thread-id collision (CWE-330/863): Confluence comment ids are per-instance, so generate_thread_id_from_confluence_comment now salts the hash with the verified clientKey (plumbed from the webhook JWT iss) to prevent thread hijack across tenants. HIGH/MEDIUM — path/query injection (CWE-22/88): get_page and update_page interpolated page_id into the REST path unencoded (update_page on a mutating PUT with no params= backstop). Now _seg()-encoded, matching the rest of the module. MEDIUM — self-trigger loop (CWE-405): process_confluence_comment had no bot-authorship early-out. Added an optional CONFLUENCE_BOT_ACCOUNT_ID guard mirroring the Linear botActor / Jira comment_author_is_bot checks. LOW — corrected the CONNECT_EXPECTED_BASE_URL comment to document it as opt-in defense-in-depth (the clientKey allowlist is the real gate). Verified clean by the detectors: RS256/HS256 alg-pinning, aud/iss/exp, kid-fetch SSRF (host-pinned + quote-encoded), at-rest secret encryption, constant-time comparisons, and the Phase-2 hardening. New regression tests for each fix; full suite green (1602). * harden(open-swe): GPT-4.1 cross-family review follow-ups Cross-family review (GPT-4.1 via orchestrator cross_reviewer) found no critical/high issues and confirmed the auth boundary is fail-closed and correct. Two low-cost defense-in-depth items applied: - Validate the signed-install JWT 'kid' against a strict charset before the public-key fetch, so a malformed kid fails fast with no network call (on top of the existing fixed host + percent-encoding). - Make JWT nbf verification explicit (verify_nbf) on both the RS256 lifecycle and HS256 webhook decodes. Other suggestions triaged as already-handled (aud cross-app replay is blocked by the per-tenant iss->secret lookup; documented static-token/IP/ baseUrl tradeoffs; qsh pinned to Atlassian's vector) or ops/infra (Fernet rotation via MultiFernet; rate limiting at the gateway). * docs(open-swe): document Jira + Confluence in installation & customization guides - INSTALLATION.md §5: add Jira (Automation-rule webhook + shared secret, service account, JIRA_PROJECT_TO_REPO) and Confluence (Atlassian Connect app install, CONNECT_EXPECTED_CLIENT_KEYS bootstrap, durable-store note, CONFLUENCE_SPACE_TO_REPO) trigger setup; §6: add the new env vars + REQUIRE_REPO_ALLOWLIST. - CUSTOMIZATION.md: jira_*/confluence_* in the tools table; repo-extraction note covers all four sources. - AGENTS.md: match CLAUDE.md (triggers, webhooks, tool list, auth). - README.md: invocation section, tools table, and overview line.
2026-07-13 19:45:54 -04:00
from .jira_comment import jira_comment
from .jira_create_issue import jira_create_issue
from .jira_get_issue import jira_get_issue
from .jira_get_issue_comments import jira_get_issue_comments
from .jira_list_projects import jira_list_projects
from .jira_update_issue import jira_update_issue
feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
from .linear_comment import linear_comment
from .linear_create_issue import linear_create_issue
from .linear_delete_issue import linear_delete_issue
from .linear_get_issue import linear_get_issue
from .linear_get_issue_comments import linear_get_issue_comments
from .linear_list_teams import linear_list_teams
from .linear_search_issues import linear_search_issues
feat: port durable dispatch hardening and startup latency improvements (#160) * feat: port plan-review & workflow-approval UX (#135) Port six upstream commits onto dev: - c03a6be7 (already ported): keep plan guidance high-level - 546042a4: add workflow approval UI with diff preview, approval URLs, web review links, and polling for approval status during active runs - 216cf181: remove workflow token elevation; approved pushes pass through directly without proxy token rewriting - 3dbc0282: preserve plan redirects after login by accepting relative same-origin redirect_to values and rejecting blocked paths - bb104d93: submit plan comments with cmd+enter - 90cb6caa: terse Slack replies, shared content via save_plan outside plan mode (PLAN_STATUS_SHARED), reject shared-content mutations Refs: #135 * feat: port durable dispatch hardening and startup latency improvements Port five upstream PRs onto dev: - #1621 / #1658: durable dispatch with loopback webhook defense, create_durable_run helper, _config_with_prepare_run_id, degradation to None for relative/loopback completion webhook URLs - #1696: run-level completion webhook deduplication (replace claim-then-post with post-then-flag per run_id), DeferredErrorModel for graph-factory resilience, ToolRetryMiddleware for task subagents, TimeoutWrapupMiddleware for all three graphs - #1697: lazy-load __init__.py for agent.middleware, agent.tools, agent.dashboard (PEP 562); defer heavy imports (exa_py in web_search, agent.webapp in request_pr_review, deepagents in sandbox.py); add ttl_cache.py with stale-while-revalidate for tool loaders Refs: #137 * fix: restore login page render and clear CI lint/format The plan-review port removed the authRedirectUrl import from login.tsx but left its call site, crashing the login page at runtime (blank page, no 'Sign in to open-swe'). Pass the relative path straight to loginUrl, matching the plan route and the backend relative-redirect handling. Also drop an unused os import in the guard test and reformat workflow_push_guard.py to satisfy ruff. * fix: restore RepairOrphaned middleware export and repoint model fake to deferred_model boundary * fix: restore RepairOrphanedToolCallsMiddleware, fix E2E model-fake patch, drop dead ttl_cache - Re-add RepairOrphanedToolCallsMiddleware to the lazy middleware __init__ (_MIDDLEWARE_MODULES, __all__, TYPE_CHECKING) so agent.reviewer can import it. - Reroute E2E model patching to deferred_model.make_model so make_model_or_defer (used by all three graph factories) returns the scripted fake instead of building a real model with fake credentials. - Drop unused agent/utils/ttl_cache.py — no agent module imports it. - Fix import ordering in agent/reviewer.py and agent/analyzer.py (ruff I001). - Format tests/test_dispatch.py. * fix: claim-then-post run-level failure dedup; stop permanent suppression --------- Co-authored-by: amoussa1229 <166072409+amoussa1229@users.noreply.github.com> Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-09 17:11:25 -04:00
from .linear_update_issue import linear_update_issue
from .list_findings import list_findings
from .list_review_findings import list_review_findings
from .open_pull_request import open_pull_request
from .publish_review import publish_review
from .read_repo_file import read_repo_file
from .reply_to_finding_thread import reply_to_finding_thread
from .report_platform_issue import report_platform_issue
from .request_pr_review import request_pr_review
from .resolve_finding_thread import resolve_finding_thread
from .save_plan import save_plan
from .schedule_thread_wakeup import schedule_thread_wakeup
from .search_repo_code import search_repo_code
from .slack_add_reaction import slack_add_reaction
from .slack_read_thread_messages import slack_read_thread_messages
from .slack_start_new_thread import slack_start_new_thread
from .slack_thread_reply import slack_thread_reply
from .update_finding import update_finding
from .web_search import web_search
def _load_tool(name: str) -> Any:
module_name = _TOOL_MODULES.get(name)
if module_name is None:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
from importlib import import_module
value = getattr(import_module(module_name, __name__), name)
globals()[name] = value
return value
class _LazyToolsModule(ModuleType):
def __getattribute__(self, name: str) -> Any:
tool_map = ModuleType.__getattribute__(self, "__dict__").get("_TOOL_MODULES", {})
if name in tool_map:
return _load_tool(name)
return ModuleType.__getattribute__(self, name)
sys.modules[__name__].__class__ = _LazyToolsModule